Gullible engines: what AI says about you, and the free way to check your GEO/AEO visibility

A fabricated award on one web page was repeated as fact by three chatbots. Here is what AEO and GEO actually mean, what the visibility tools do and do not buy you, and a free way to find out what the engines say about your organisation.

Share
Gullible engines: what AI says about you, and the free way to check your GEO/AEO visibility

The Brief

  • Where we are: the argument about AI visibility is still going, and a good deal of it is an argument about what to call it. GEO, AEO, LLM visibility, share of model: the acronyms arrive faster than the explanations, and often from people using them with more confidence than they have earned (with hallmarks of SEO/search engine visibility before it). The effect on most comms teams is a shrug and collective confusion about what on earth it all means.
  • Why it matters for comms: in late May, PR agency Hard Numbers invented an award for its own founder, published the release on its own site and never needed the wire. Within days three chatbots were repeating it as fact in 21 of 30 attempts. The agency's report in June found that in UK fintech, 70% of the sources shaping AI answers to buyer questions were companies recommending themselves. That was almost four months ago – and none of it has stopped being true.
  • The takeaway: you can find out what the engines say about you in 20 minutes, for nothing, with a spreadsheet. The paid tools buy you scale and a trend line. They do not buy you better questions, or the judgement about what to change. And the sooner you go and look for yourself, the less mysterious the whole subject feels.

The Story

Darryl Sparey was, for one halcyon stretch of this summer, the best-dressed PR practitioner in the UK. He would like it on record that this was never true.

Darryl runs Hard Numbers, and in late May they published a press release on their own website announcing he had won the title at the UK PR Honours 2026 – an award that does not exist. (Though should it? Answers on a postcard...)

It was written as a wire release because they expected to need the wire. They never did. "Seeding it on our site was enough," he told me from a car on the way to a prospective client, "to influence ChatGPT, Perplexity, Claude and Google's AI Overviews." Within days the three big chatbots were naming him 21 times out of 30.

The engines have since come round to his way of thinking, slowly. When I asked them again in September, ChatGPT and Perplexity still gave his name but added that the award was invented, and Claude and Gemini declined to name anyone at all.

There is a gilet in the release, since that is his thing (and who can blame him?). However, the joke was the hook for something less funny. Hard Numbers had been compiling what Darryl calls "a quite boring, vanilla" report on which fintech firms had the best AI visibility, and kept tripping over the methods used to get it: "top ten payments companies in Europe" pages with the author at number one and Stripe mentioned in passing, competitors review-bombed on G2 and Trustpilot. That was where the fintech figure came from, and the whole of it is set out in the Gullible Engine Optimisation report.

The all-important terms, plainly

No wonder people are confused. There are a great many acronyms flying about, employed with more confidence than most of them have earned, so here they are in plain words.

  • SEO, search engine optimisation: getting a page to rank. (This term has been around since the early days of search engines, and inspires the similar pair that follow next.)
  • AEO, answer engine optimisation: being the answer a machine gives.
  • GEO, generative engine optimisation: shaping what a generative model writes, and which sources it cites.
  • LLM visibility, or share of model: the measurement. How often you are named when the questions you care about are asked.

The terms overlap almost entirely, and the naming is a land-grab by whoever wants to sell you the dashboard.

Darryl's agency settled on GEO because Andreessen Horowitz published a piece on SEO versus GEO in spring 2025 and "that week I had three phone calls from current, former and prospective clients saying, what is GEO and can you help me with it? So we followed the money."

The line that stuck with me from our conversation is blunter: "Stop arguing about acronyms. Start by measuring."

One distinction does matter, and it is about the engines rather than the labels: when a chatbot searches the web before answering, its answer can change within days, and when it answers from training alone, what you publish this week may not register for months. Both behaviours sit behind the same chat box, so there is no single answer to "what does AI say about us", and the several you get move at different speeds.

Whose job is this

Of the twenty or thirty GEO audits Hard Numbers have run, Darryl says most have been for mature businesses in regulated sectors, not the startups I had assumed, and that misrepresentation there touches the licence to operate: "the ones with reputations, the ones that get talked about without having to generate the coverage."

As for who inside an organisation should pick this up, Kerry Parkin of The Remarkables, quoted in Helen Dunne's Corporate Affairs Unpacked piece that covers the experiment, argues for comms, because an SEO or digital lead cannot credibly convene legal, HR, finance and public affairs.

Darryl has recently been rethinking that question and is writing it up himself, so I will point you to his feed rather than pre-empt him! What he was happy to say on the record is the lesson from last time: the technical part of SEO was always small; most of it was writing content people and machines find useful and earning citations from other sites, which is comms work by any description. The industry left it to others anyway, for want of the vocabulary, and paid for it in budget and in journalists' patience with link-builders calling themselves PRs. None of it takes much technical knowledge, he says: what schema markup is, what robots.txt and llms.txt are, what query fan-out and Retrieval-Augmented Generation (RAG) are. Enough for PR not to lose this fight the way it lost the SEO one.

The tools, and what they buy you

It can feel as though a respected agency or vendor launches another visibility tool every week, which is its own contribution to the confusion. Most monitoring vendors now have something, from Meltwater's GenAI Lens to Muck Rack's AI Visibility Badges to PR Newswire's AEO and GEO report, and so do a fair number of agencies. Then there are the pure plays: Profound, Otterly, Peec and a dozen more, from about $29 a month to enterprise pricing.

(Current prices and engine coverage are in the Comms With AI tools directory rather than here, because they change monthly and an article does not. Be warned that most of the comparison articles ranking these tools are, with a pleasing circularity, written by the tools.)

What the paid ones buy is scale, cadence and reporting your board will read, which is worth the fee if you have many products in many markets or a client who wants a dashboard. What they do not buy is a good set of questions, or the judgement about what to change when the answers come back. That part is free, and it is the part comms people are good at.

The free method

Darryl's version: a prompt library of 25 to 100 questions, run through the platforms by tool, plug-in or hand, then analysed. Which of your pages are cited and which never are? Which of the coverage you worked for shows up, and which sources you have never been in keep appearing?

His example is a large organisation whose financial wellbeing page is never cited in answers about financial wellbeing, while a handful of other sources are cited again and again: a content brief and a pitch list in one finding. Mine is the same shape, scaled to a firm with one or two offers and no budget line. The short version takes twenty minutes.

Prompts. Six questions a buyer would ask, in a buyer's words: "Who can help a Scottish charity adopt AI responsibly?" rather than "AI consultancy Scotland". Include two phrasings of the most important one, and one that names a competitor, because that reveals who the engine thinks the set is. Score any prompt containing your own name separately from the rest. A prompt that names you is not evidence that anyone found you.

Engines. Two to begin with: one that searches the web before answering and one that does not, since that is the distinction that decides how fast your work can show up at all. Add Perplexity, Copilot and Google's AI Mode when you want the fuller picture, and ChatGPT twice, search on and off, when you want to see the gap between the two modes. Logged out where you can, and where you cannot, turn off the stored profile as well as the chat memory: they are two different settings and only one of them is called memory. Running this on my own company last month, two of three Claude answers still referred to my old employer and my own company with memory off. A signed-in engine reflects you back at yourself, and the number you write down will be wrong.

Runs and log. Six prompts in two engines, once each, is twelve answers and about twenty minutes. That's the version I ran, and it is enough to tell you whether you show up at all. Three runs per prompt per engine is the version that survives an argument about variance, and at six and two that is thirty-six answers and most of a morning. Do the small one first. One row per run: date, engine, model or mode, whether it searched, your account settings, the prompt, whether you were named and where, the sources cited, and every other organisation named.

Score. Share of voice is runs named over total, overall and per engine. More useful is the list of the five sources cited most often, because that is where the engines are getting their answer. Record three things separately: whether you were named, whether your site was cited, and whether what the answer said about you was actually true. The third is the one no dashboard scores, and this piece opens with an engine naming somebody enthusiastically and wrongly. Mention rate, the share of the answers you sampled that name you, is the headline figure, but it is not a quality score: Kevin Indig's analysis of Semrush data across 1,094 US ChatGPT categories found the most-cited domain was the most-mentioned brand only one time in five.

Act. In rough order of impact: a plain, definition-first page per core offer, with an FAQ answering the questions in your prompt list; the same one-line description of who and where you are everywhere you appear, from your site to LinkedIn to Companies House; then being written about, listed or interviewed by somebody other than yourself, which is what the engines treat as corroboration. That last part used to be called PR.

Re-run. Same protocol, a week or two later. Fresh pages can change what a searching engine says once they have been found, and will not touch an engine answering from training. A movement on its own does not tell you what caused it, which is why the log records the settings as well as the answer.

What I would not do

Hard Numbers proved a fake award works in order to expose it. Self-ranking listicles, invented accolades, comparison pages that put you first, negative content about competitors: the award demonstrably works, Hard Numbers found the rest of it in use, and all of it is the GEO equivalent of the link farms Google spent a decade dismantling.

Darryl worked in SEO through those years and expects the same whack-a-mole here. "You can tell when you're getting bad advice because it focuses on one specific outlet," he said. People built strategies on Reddit's citation share, and in August it fell by roughly 86 per cent in about a week by Promptwatch's count. Promptwatch calls its own figure provisional and cannot rule out a collection problem, and whether Reddit blocked the crawler or OpenAI changed how it searches is still being argued over.

That nobody can say for certain is rather the point. Darryl's alternative is unglamorous and right: "Build an outlet-neutral approach to creating as consistent and maximal a surface area of the internet about your business as you possibly can." Do the honest work, everywhere, and keep doing it.

The Practice: what you can put into practice today

  1. Write six buyer questions and ask them, logged out, in two engines. Twenty minutes. Record who gets named and which sources are cited. If your organisation is not named, note who is: that list is your competitor set as the engines see it, and it may not match yours.
  2. Pick one topic you have coverage on and check whether any of it is cited. If it is not, write down the three sources that are. That is your next pitch list, and if someone else owns the web page, your reason to go and talk to them. It works in both directions: last week an agency working for Grammarly, by its own account, sent us ten edits to our tools directory entry, most of them fair corrections, and not one adding a caveat.
  3. Spend an hour on Darryl's five terms. Schema markup, robots.txt, llms.txt, query fan-out, retrieval-augmented generation. Enough to hold your own in the meeting where someone tells you this is too technical for comms.

I have spent the past month running all of this on my own consultancy, and the results are not flattering (eek). That write-up follows later, once I have a month of measurements rather than a week (do subscribe to ensure you read!), because the numbers move on their own and a week is not long enough to tell you why.

In the meantime: which prompt would you run, and who do you think the engine would name instead of you?

Every Applied piece follows the same shape (The Brief, The Story, The Practice) and is made the same way: my argument and examples, drafted with Claude, every claim checked against its source, and a final edit by me. This one also had a second review from ChatGPT, and Darryl saw and approved his quotes before publication.

Coming up: the Comms With AI Leader Interview with Stephen Waddington, Outputs or Outcomes, online on Friday 9 October, 12:00 to 12:45 BST. Darryl's report, Gullible Engine Optimisation, and his primer on GEO for PR people, free with a short sign-up form, are both worth your time.

More like this