You Can't Buy Your Way Into ChatGPT. That's the Good News.

August 2026 · By Adam Valine

The old SEO rewarded budget. The new one rewards being worth citing — and the only way to know if you are is to measure like an engineer, not a marketer.

The number every marketing leader is quoting into their board deck this quarter is fifty-eight percent. That is the click-through-rate reduction Ahrefs measured on top-ranking pages after Google's AI Overviews rolled out — published in March 2026, roughly double the 34.5 percent decline the same study documented in April 2025. It is also, coincidentally or not, the same number that appears in Penske Media's 2026 antitrust filing against Google as the click-decline publishers have absorbed since the feature's launch. AI Overviews now appear on 13.14 percent of Google queries, more than double the 6.49 percent from January 2025. Global publisher Google traffic dropped by roughly a third across 2025. Enterprise SEO teams that spent a decade optimizing for the ten-blue-link era are watching that era get replaced in front of them.

Two reactions are dominant. The first is to hire a GEO agency — the same consultancies that sold SEO retainers now selling "Generative Engine Optimization" or "LLM Optimization" retainers, at higher price points, on the theory that this is SEO 2.0 and the same purchasing pattern applies. The second is to chase tactics: implement Jeremy Howard's llms.txt spec, buy schema markup software, cram FAQ-style content onto every page in the hope that generative engines will pick it up. Both reactions are the same reaction — treat this as a marketing problem and buy the same shape of solution the last version of the problem accepted. Both reactions fail for the same reason. And the failure mode reveals what a working practice actually looks like.


You cannot buy your way into ChatGPT

The reason the old SEO playbook was purchasable is that its ranking mechanism — Google's — was, in the end, a function of things you could pay for. You could pay for backlinks, either explicitly or through outreach agencies that trafficked in link-placement inventory. You could pay for content mills to spin up hundreds of pages targeting long-tail keywords. You could pay for technical SEO audits that identified schema deficiencies. You could pay for a domain with existing authority and 301-redirect your way onto its coattails. None of it was elegant, and much of it was against Google's stated guidance, but the correlation between spend and ranking was strong enough that budget was a competitive moat.

The mechanism generative engines use to select what to include in an answer does not respond to those levers. When a model like Claude or GPT-5 or Gemini answers a question, the information it draws on comes from two sources: what it was trained on, and what it retrieved live from the web via a search or fetch tool. Neither is buyable in the way SEO was buyable. The training corpus was assembled from what was worth including in the training corpus — a filtering step run by the model provider, not by the brand being cited. The live retrieval step ranks pages using a mix of traditional web-search signals and generative-specific factors that the Princeton GEO paper by Aggarwal, Murahari, and Narasimhan (KDD 2024) was the first to formally characterize — factors like citation density, quotation-worthiness, statistical claims, and the presence of authoritative-source references within the page itself. There is no billing address on either mechanism. There is no agency that can, for a monthly fee, get you into the training corpus of a model that was frozen before the contract was signed.

What actually determines whether you are cited is less like a purchasing decision and more like a reputational one. Are you talked about in the trade press the model was trained on? Do independent third parties reference your work? Do your published pages contain the kind of specific, quantitative, attributed claims that a generative model's retrieval layer weights? Are your engineering blog posts, your published post-mortems, your executive white papers, your research reports — the artifacts most enterprises treat as marketing overhead — dense enough with primary information that a summarizing model finds them useful? This is not something an outside agency can create for you on a six-month statement of work. It is the accumulated output of a research organization that has been producing high-quality primary content for years.

The good news, and it is genuine good news, is that this same constraint applies to your competitors. If they could not buy their way in either, then the market for AI-answer visibility is not going to be dominated by whoever spent the most on the new agency category. It will be dominated by whoever built the reputation to be citable — and reputation, unlike ad budget, is durable across quarters.


The tactics floor is real. It is also just a floor.

The second reaction — chase tactics — is not wrong on its face, and there are a small number of tactical interventions that measurably help. The Princeton paper found that adding statistics, quotes, and authoritative citations to source pages increased AI-citation rates by up to 40 percent across the test set. Schema.org markup, particularly Article and FAQ schema, remains one of the clearest signals a generative engine uses to structure a page's claims. Server-side rendering matters again — AI crawlers frequently do not execute JavaScript, so client-rendered content is invisible to them. Making sure your robots.txt and CDN configuration do not block AI user-agents is a real oversight that some enterprises have already made.

What is not measurably worth the effort, for most companies, is the llms.txt file. Adoption sits somewhere between 2.13 percent of top sites (Web Almanac 2025) and 5–10 percent as of Q1 2026, and no major model provider — OpenAI, Google, Anthropic, Meta, or Mistral — has publicly committed to reading or acting on the file in production systems. There is modest measurable uplift in Anthropic and Perplexity citations for sites with sprawling navigation that benefit from explicit curation, which means the file makes sense for documentation-heavy properties and does not obviously make sense for most enterprise sites. Publishing an llms.txt is not wrong. Treating it as a strategic lever is.

The deeper problem with the tactics-only response is that the ceiling is low. Schema markup, factual density, and technical hygiene get you into the pool of pages a generative engine considers using. They do not determine which pages get chosen. When every competitor has clean schema, dense statistics, and inline citations — and this is the state the market is trending toward over the next two years — the differentiator collapses back to reputational signals: entity clarity, named-author track records, cross-site corroboration, third-party references. Those are the moat. The tactics are the moat's rebar.


If you cannot buy it or game it, how do you know if you are winning?

This is the question that trips people up, and it is the question the two dominant reactions both attempt to answer without acknowledging. Hiring an agency answers it by outsourcing the question: the agency will show you a dashboard. Chasing tactics answers it by pretending the tactics themselves are proof of progress: we implemented schema, therefore we are winning. Neither is measurement. Both are theater.

The uncomfortable answer is that if your company wants to know how visible it is in AI-generated answers, someone has to actually measure that visibility — repeatedly, against a defined set of queries that matter to the business, across the specific model surfaces where customers are asking questions. This is what the third-party tool category is trying to sell: Semrush's AI Search Visibility Checker, Ahrefs' Brand Radar, and roughly a dozen newer entrants like Otterly, Peec, Profound, and BrandRank.ai will run a set of prompts across ChatGPT, Perplexity, Claude, and Google's AI Overviews on your behalf and report which answers included your brand. These tools are useful. They are not, for a serious enterprise program, a substitute for the discipline itself.

The discipline itself is one the enterprise probably already has, sitting inside its AI product organization: the eval framework. We wrote about this earlier this year — every serious enterprise LLM application either has, or ought to have, a golden dataset of queries that exercise the failure modes the product is likely to encounter, a repeatable measurement pipeline that scores responses against that dataset, and an owner whose job is to run the pipeline before every meaningful change and publish the numbers. Almost every enterprise AI team has this artifact in some form. Almost no marketing organization inside the same enterprise has thought to ask them for it, or to build the equivalent for the company's own presence in AI answers.

What this looks like in practice is not exotic. It is:

  • A curated set of, say, thirty to two hundred prompts that a real customer might ask a generative engine about your category — "Who are the leading firms doing X?", "What is the standard approach to Y?", "How does company A compare to company B on Z?" — chosen by the people inside the company who actually know what customer intent looks like.
  • A scheduled job, run weekly or nightly, that submits each prompt to each model surface via API and records the response.
  • A scoring pipeline that measures, for each response: whether your brand is mentioned, whether the mention is accurate, whether it is cited with a link back to a page you own, and how the response ranks you against competitors.
  • A dashboard and an alerting layer so that when your inclusion rate on a category-defining prompt drops from 80 percent to 40 percent, someone in the company sees it within a day rather than a quarter.
  • A named owner whose job is to run this and to be the point of accountability when the numbers move — the same role structure the AI product team already has for its own evals.

None of these components is technically hard. The Semrush and Ahrefs offerings solve the "submit prompts, record responses" step; the scoring and alerting can be assembled from off-the-shelf pipeline tools in a matter of weeks. What is hard is the organizational move: recognizing that AI-answer visibility is not a marketing metric the CMO owns from a vendor dashboard, but an engineering-grade measurement problem with a named owner, a versioned dataset, and a change-management discipline. The teams that make this move will be able to detect when a Google model update erases their category presence, when a competitor's new white paper starts displacing them in Perplexity, when a training-cutoff refresh pulls their most-cited product page out of Claude's answer set. The teams that do not will find out from a quarterly board report, six months late.


Both angles are the same lesson

The failure of the "hire an agency" reflex and the success of the "point your evals at your own visibility" pattern are not two separate points. They are the same point in two frames. In both cases the underlying shift is that AI-driven distribution reduces the value of purchased visibility and increases the value of produced visibility. Purchased visibility — the ad units, the agency retainers, the link networks — buys you a slot in a marketplace that is being restructured. Produced visibility — the research your engineering team publishes, the case studies your practice writes up, the eval framework you point at your own presence — accrues to the entity doing the producing.

The companies that are going to win the next five years of AI-answer visibility are the companies that have already learned this lesson somewhere else inside the building. They are the ones with a working eval discipline on their AI products, an engineering blog that produces primary technical content, an executive team willing to attach their names to specific claims in specific pieces, a research organization that measures its own outputs. They will not need to hire a GEO agency, because their production and measurement discipline is already the thing a GEO agency would be nominally selling. They will underspend the market on tactics and overspend the market on producing.

For everyone else, the honest short answer is that there is no purchasable shortcut. The uncomfortable long answer is that the discipline you need is one your AI product team already has — and the first move worth making is to walk down the hall and ask them how their eval pipeline works.


References

Adam Valine leads Veil Consulting's AI Strategy & Roadmap practice. Currently AI Strategy Lead at Solutions Plus Consulting; twenty years of prior product and engineering leadership across SaaS and platform companies.

Continue reading

All notes