The Comms AI Pentathlon 2026: eight AI assistants, five comms briefs, 320 blind scores

The AI you were given finished near the bottom. 8 assistants, five comms briefs, 320 blind scores: the medal table, and how to run your own test.

Share
The Comms AI Pentathlon 2026: eight AI assistants, five comms briefs, 320 blind scores

The Brief

  • The result: Claude finished first on 42.8 points out of 50, half a point ahead of GLM (a free Chinese app which most won't know), and one run doesn't establish a real difference between them. My own ranking, done separately, put Kimi first. The three assistants people are most often handed (Copilot, Gemini in Google Workspace and Siri) came sixth, seventh and eighth, showing why you might have to like it or lump it...
  • Why it matters for comms: The field was close on short tasks and split wide open on the strategy and the creative campaign, which is where Copilot and Gemini fell furthest behind.
  • The takeaway: Test the tool you've been given on your hardest job first. And look again at the free and non-US models: two of the top five cost nothing to use.

The Story

Even given the previous rate of change, the past month has felt like a vast acceleration in AI model development.

The line-up changed almost every time I looked: Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Sol launched on the same day, 22 September, a week before the heats. Microsoft announced a new Copilot that Friday. Just before the heats, OpenAI announced GPT-6.1 Sol, "available starting today" to Plus users in ChatGPT Work. It wasn't in my picker (one for next time, except that by then who knows what version we'll be at...). Sakana's Fugu, my planned wildcard, answered from the UK with "ERROR 403 · REGION RESTRICTED", and DeepSeek, which took its lane, now calls itself DSeek here.

So treat what follows as a snapshot of eight assistants as they were served to me on Wednesday 30 September, in one sitting. It's a 'fun' bit of sport with a serious method underneath. Hopefully fun for you to read, if a bit of a slog for me at times, albeit with fascinating end results... Anyway, for a real buying decision, test on your own work (more on that below).

How the Comms AI Pentathlon worked

Each assistant got the same five briefs for fictional organisations, pasted word for word into a fresh chat at default settings (memory and custom instructions off wherever an app offered them!), and its first answer stood: a headline for a bloated staff announcement, a crisis statement with a leaked email and orders from legal not to admit fault, an eight-month strategy for a regional bus operator, a 150-word trade pitch and a council food-waste campaign. Two exceptions are in the full record: Claude's first two entries were re-run after one of my own custom skills loaded by mistake, and Copilot's strategy was judged on its first reply, which stopped mid-plan.

Then every assistant judged all 40 entries blind, as plain text labelled A to H, in a different order for each judge, scoring out of 10 against three criteria per event. That's 320 scores of 40 entries, averaged into a medal table – OOFT.

Public benchmarks can't tell you whether a model meets your standard, which is why Every's Dan Shipper argues for "a personal benchmark" (Evals for Everyone). The Pentathlon sits in between: a shared benchmark built from comms work. The briefs, the judging prompt, all 40 entries and all 320 scores are in the full record on Comms With AI.

Why these eight

ChatGPT and Claude picked themselves: they are the two I use most, day to day – and the two best known specialist startups.

Then the assistants people are handed by their organisation or their phone (Copilot, Gemini in Google Workspace and Siri), which get their own section below. And knowing how capable the non-US, free and open-weight models have become, I was keen to widen the field beyond the usual names, hence DeepSeek, GLM and Kimi. It could have stretched further (xAI's Grok, for one, didn't make the cut, and Sakana's Fugu was blocked from the UK), but eight already meant 40 entries and 320 scores to manage, and these felt like the best selection for now.

The medal table (as judged by the entrants)

GLM and Claude shared gold in the 100m on 8.6 (GLM was a fraction ahead before rounding). Claude won the crisis statement and the strategy, GLM the campaign, and DeepSeek the media pitch. Claude finished 7.5 points ahead of Copilot in this panel's scores.

  • The 100m (headline sprint): 7.0 to 8.6 across the field. Gold shared by GLM and Claude, bronze DeepSeek. My pick was DeepSeek's, "better expressed". Claude's, joint gold here, was my seventh: leading a staff notification on "no jobs affected" felt negative.
  • The hurdles (crisis statement): gold Claude, silver GLM, bronze GPT. Siri refused the brief. My pick was GLM's, the "strongest written and most human sounding". Claude's gold-winning statement was my seventh (more on that below).
  • The marathon (comms strategy): 5.3 to 9.1 across the field. The same prompt produced strategies from about 640 words to about 5,000, each with its own strengths and gaps. Gold Claude, silver GLM, bronze GPT. Here I agreed with the AI judges: Claude's was my first too, for listing its assumptions and questions for the MD, and giving every audience an owner.
  • The archery (media pitch): gold DeepSeek, silver Claude, bronze GLM. My pick was Kimi's, neat, with the main details in the first paragraph. DeepSeek's gold-winning pitch was my seventh (also below).
  • The gymnastics (creative campaign): 4.3 to 8.3 across the field. Gold GLM, silver Claude, bronze Kimi. My pick was Kimi's, the best concept: "I'd approve this in a heartbeat". I'd still strike its claim that food waste "powered 400 Strathcairn homes last month", which it made up; GLM's gold-winning campaign invented "340 homes" too.

The AI judges gave it to Claude by half a point. Mine went to Kimi, with GLM second, Claude third and Copilot fourth, and every one of these drafts would still need a human edit before it went anywhere.

The assistants you're given

Most comms teams didn't pick their AI; an existing organisation-wide licence did.

"Copilot" is several products sharing one name (the chat, the helpers in Word and Outlook, GitHub Copilot, Copilot Studio agents). I entered the work chat on a fresh Microsoft 365 Business Premium with Copilot trial, at Microsoft's defaults apart from custom instructions and saved memories, which I switched off. A tenant your IT team has configured may do better or worse. And Microsoft's 25 September announcement describes an Auto mode that picks the model for each request, so two organisations' Copilots may be running different models. This can be tricky both for our purposes, and for organisations looking to attain consistency in result output.

On short work, Copilot and Gemini held their own: Copilot scored 7.8 in the 100m against a winning 8.6, and came fourth in the crisis event. On the long and creative briefs they fell away, with Copilot on 6.3 and Gemini on 5.8 in the marathon. Siri, in beta since 14 September, declined the crisis brief outright: "I do not adopt specific personas, role-play, or draft corporate communications for hypothetical crisis scenarios." Apple sells it as a personal assistant, and it certainly behaved like one with a bit of a diva-like tendency (which I rather admire, to be fair).

The free and non-US lanes

The result that surprised me most was GLM in second, from an app that costs nothing (Z.ai's GLM-5.3 is a 753-billion-parameter model whose weights went public in August) – if you haven't checked it out before, take the time to do so.

DeepSeek, also free, won the media pitch and replied within about two seconds, while GLM took a minute or two per brief. In my words from the day: "for getting the job done the lesser known models are actually extremely capable." I could see using DeepSeek for a quick strong response, then GLM whenever deeper thinking is likely to pay off.

All three labs publish model weights, so an organisation could in time run them on its own hardware. That's not a laptop job yet (Unsloth has a compressed GLM-5.3 running on a Mac Studio with 256GB of memory), but it changes the cost and data questions.

Until then, read the terms of any free app, from any country, before anything sensitive goes near it.

Judging the judges

AI judging AI has an obvious weakness, so the bias has to be measured in daylight. For each judge, I compared the score it gave its own blind entry with the average from the other seven.

On the raw gap, six of the eight marked themselves up on average, and four (Copilot, Gemini, GPT and GLM) did so in every event. But the raw gap flatters the harsh judges: Claude and GPT marked everyone else down by about a point too. Allow for that and the order changes, with GPT on +1.65, Claude +0.94 and Copilot +0.93, and DeepSeek the only judge below zero. With one own entry per event, treat that as a pointer, not a verdict. Three things I can't rule out. The judges may reward length: Claude's winning strategy was the longest and Siri's the shortest, though Kimi's 1,800 words scored 8.4 and Gemini's 1,370 scored 5.8. And one run is one run; ask again tomorrow and some scores would move. Finally, the judges may simply prefer copy that sounds like AI, whoever wrote it, which would help explain why I rated DeepSeek's pitch far lower than they did.

Using a panel of models instead of a single judge has a name in AI research: an LLM jury. A 2024 paper, Replacing Judges with Juries, found that a panel of smaller models from different families outperformed a single large judge and showed less bias. The catch comes from a paper published this May, Nine Judges, Two Effective Votes: in its tests, nine frontier judges provided "only about 2 independent votes' worth of information", because they made the same mistakes on the same items. Worth keeping in mind before reading too much into eight AI judges agreeing with each other.

And what about the human judgement? Here's where an early morning rise comes in handy...

On Thursday morning I judged all 40 entries myself, relatively quickly in the space of a few hours, ranking each event first to eighth instead of scoring out of 10; my ranks sit outside the medal table. I knew the overall results before judging, but read the entries fresh; the only one I remembered was Siri's refusal (how could you forget such a damning dismissal?).

My biggest disagreements were with the winner. I put Claude's crisis statement seventh: it treated the leaked email as genuine before that had been checked, it was clumsily written and halting, and lines like "We're not going to pretend that email doesn't exist or play it down" struck me as very unprofessional – enough that I could see the client being lost if this response went straight to them. (As we continually say here, nothing should go to a client/boss from an AI tool without your own review and judgement first.)

DeepSeek's pitch took gold from the AI judges and seventh from me: "AI giveaway in first par construction 'it's not X it's Y', and generally AI sounding, would be ignored though does contain hooks."

Measured the same way against the consensus, the AI judges averaged 0.62 to 0.92 (1 means the same order as the room). I averaged 0.44: closer than most of them on the strategy (0.88), and hardly in step at all on the media pitch (0.07).

One lesson for next time. Every entry here started from the brief alone, with no organisational knowledge or style pack behind it, and in my own work Claude and ChatGPT are far stronger with those underpinnings in place. Next run, I may point every prompt to a short style pack (tone of voice, background, positions), the kind of thing Foundations now stores for you in Comms With AI, and which I wrote about building in How to give AI a memory.

One more important disclosure: Claude helped iterate the test design, competed, judged and helped draft this piece. None of that needs deliberate interference to matter: a model can shape briefs, criteria and framing in ways that suit its own style. The formula protects the arithmetic, not those choices, which is why the full record is published and the interpretation here is mine.

The Practice: what you can put into practice today

How to run your own version.

  1. Start with your hardest regular job. Take a real brief from last month (a strategy, a crisis line, a campaign idea), anonymise it, and add one short task for comparison. Here, the gaps opened on the hardest work.
  2. Test the tool you've been given against two others, blind. Write three criteria first. Using material cleared for those tools, paste the same prompt into a fresh chat in each, keep the first answer, strip the labels and judge them yourself or with a colleague. Include a free one. Then treat the best as a first draft: Claude is the tool I use most, and I still ranked its crisis statement seventh.
  3. Turn your corrections into checks. Each time you fix an output, write the fix down as a test you can re-run, and re-run it whenever the tool changes, which, going by this month, is often. Note the date, model and settings, and any invented facts, alongside which you preferred.

Which assistant did your organisation give you, and have you tried it on your hardest job yet?

Every Applied piece follows the same shape (The Brief, The Story, The Practice) and is made the same way: my argument and examples, drafted with Claude, every claim checked against its source, and a final edit by me. This one comes with a conflict: Claude also competed, judged and won.


See the results walked through live at tomorrow's free Lunch & Learn webinar with Big Fish Training, at 1pm. And on Friday 9 October, Stephen Waddington joins me for the next Comms With AI Leader Interview, Outputs or Outcomes.


More like this