Core / Pillar 26 min read Published Updated

How to Choose an AEO Agency (2026 Guide)

I have sat on both sides of this hire. Here is the brief, the questions, and the scoring sheet I now use before I sign.


On this page
Line drawing of a checklist clipboard next to three agency building icons

Key takeaways Read this if nothing else

  1. 01

    I write the query list, entity list, and measurement rules before I take a single pitch.

  2. 02

    I buy named citation logs I can re-run myself, not a dashboard I cannot export.

  3. 03

    I treat a short paid diagnostic as the working interview, then score process, proof, and contract terms on the same sheet.

  4. 04

    I keep the work in-house when the prompt set is narrow and someone on staff already owns the source pages.

Why this hire is not an SEO RFP

<p>I used to send the same RFP I used for SEO: keyword lists, backlink asks, a rankings dashboard. That packet does not describe this work. Answer engines cite sources inside generated answers. I buy named citations, not a position-three URL.</p>

<p>I now write a different brief, ask different questions, and score process before any case slide. I have sat on both sides of this table. What follows is the sheet I keep on my desk for how to choose aeo agency partners. I am not scoring this hire on the same criteria I used for a classic SEO retainer.</p>

What changed for how to choose aeo agency buyers

<p>Pew Research Center surveyed 5,119 U.S. adults from February 17 through 23, 2026. In that window, 49% of U.S. adults reported using AI chatbots, up from 33% in 2024. Twenty-four percent used them daily; 12% several times a day and 4% almost constantly. Among chatbot users, 42% used them to search for information. Thirty-eight percent of employed U.S. adults used chatbots for work tasks. Sixty-three percent of U.S. adults under 50 reported using AI chatbots, against about four-in-ten ages 50 to 64, and ChatGPT was the most commonly used chatbot across the age groups Pew examined.</p>

<p>That is why I now buy citations inside answers rather than another SEO RFP. I keep the adoption numbers in our guide to aeo statistics; I do not let them rewrite my scoring sheet. Adoption explains why the hire exists. It does not tell me how to choose aeo agency partners. I then return to my own buying process.</p>

Google’s published bar for generative features

<p>Before I hear a methodology slide, I check whether the pitch already knows the public bar. In 2026 Google published guidance on optimizing for generative AI features. That Google guidance on generative AI features emphasizes creating valuable, unique, non-commodity content. It also names local, shopping, image, and video content as relevant formats, not only long-form articles.</p>

<p>I treat that as the floor, not a differentiator. I do not need a restatement of the blog post. I need to see which of my pages would meet that bar, in which formats, and which queries those pages are meant to support. Unique source pages and format-specific assets are the test I run against the deck. I also ask how they would treat image, video, local, and shopping pages if my category actually uses those formats. I keep the URL in the brief pack so the standard is not up for debate on the call.</p>

What I now ask before a pitch

<p>I will not sit through a deck until three facts are on the table in writing. First, which engines are in scope. I name ChatGPT, Perplexity, Google AI Overviews, Gemini, Copilot, Grok, and Claude when they matter to my buyers; the agency names which of those they will actually work. Second, which query clusters we will treat as the job. I freeze ten to thirty prompts, written as full questions, not a keyword dump. Third, who owns measurement. I will not rent a dashboard I cannot export. The log has to be mine at the end of the term.</p>

<p>Those three answers tell me whether we are buying the same thing. If engines, prompts, and measurement ownership are still marked as later scope after kickoff, I stop the call. I want those three facts in an email, not in speaker notes. I use that freeze as the gate, not a case study from another category.</p>

How to do AEO - Answer Engine Optimization video thumbnail

Video: How to do AEO - Answer Engine Optimization · The Next New Thing

What hiring an answer engine optimization agency actually covers

<p>The retained work is not SEO plus a chatbot slide. When I am hiring an answer engine optimization agency, I buy three operational things: named citations inside answers, entity consistency across the pages and profiles models already see, and source pages that can be quoted. If a proposal cannot describe those three in the first two pages, it is still an SEO retainer with new vocabulary. I use this section as shared language for the rest of the guide on how to choose aeo agency partners, so later scoring is about delivery, not definitions.</p>

Citation, not rank, is the unit of work

<p>I do not buy a rankings dashboard. I buy a named citation: my brand, product, or URL appearing inside an engine's answer to a prompt I care about. Rank is a position on a results list. Citation is a mention, a quote, or a source chip inside generated text. Those are different units, and I will not let a pitch collapse them.</p>

<p>A monthly report that shows a visibility score without the prompt, the engine, the date, and the quoted passage is not a citation log. I want the passage, the URL the model pointed at, and whether we were the primary source or a passing name. I also ask whether the citation was accurate: a wrong fact with my name attached is not a win. If the agency cannot show those rows from a past engagement, I treat the work as classic rank with a new label.</p>

How engines surface brands

<p>Models do not rank my homepage the way a classic crawler does. They retrieve, compress, and name entities they already trust for a given prompt. For that selection step I use a closer look at how do llms choose brands; on a retainer I expect the agency to work the parts I can actually change.</p>

<p>That means a clean entity: same legal name, product names, and category labels on my site, Wikidata or equivalent, and the third-party pages models already quote. It means source pages that answer the prompt in a quotable block, not a long essay with the answer buried. It means a plan for pages I do not control, reviews, docs, standards, journalist explainers, because those often get cited instead of me. I ask for a list of target entities and source URLs, owned and unowned, in the first month. If that list is missing, we are not doing this work yet.</p>

Where AEO sits beside SEO and PR

<p>I will not pay two teams for the same page. SEO still owns crawl, indexation, internal links, and the technical health of the URLs we want cited. Content still owns the editorial calendar and brand voice. Comms still owns journalist relationships and the embargo calendar. AEO owns the prompt map, the quotable blocks, the entity consistency, and the citation log. Those lanes stay separate on paper.</p>

<p>The handoff is a shared URL list and a shared prompt list. If SEO rewrites a category page, AEO specifies the questions that page must answer in the first screen. If PR lands a trade interview, AEO specifies the entity names and facts a model can reuse. I put that split in the SOW and ask any bundled content-plus-PR-plus-AEO offer to price the AEO slice alone.</p>

The brief I write before I shortlist

<p>I write a one-pager before hiring an answer engine optimization agency. It freezes the prompt clusters, the brand entities, the languages, and the measurement rules. Every pitch then answers the same brief. I used to let agencies invent the query set in the proposal; I then spent the first month arguing about what we were even measuring. If I cannot fill the one-pager, I am not ready to shortlist. That document later attaches to the SOW so the query list cannot drift.</p>

Queries, entities, and markets I lock first

<p>I freeze three lists before outreach. Prompt clusters: I group the questions my buyers actually ask, usually ten to thirty prompts in two or three themes, written as full sentences a person would type into ChatGPT or Perplexity. I do not hand over a keyword spreadsheet. Brand entities: legal name, product names, common misspellings, and the category label I want the model to use. Markets: language and country pairs, not global as a slogan. A German prompt is not a translated English prompt; I say which languages are in scope and which are out.</p>

<p>I put those lists in the one-pager and I date them. They cannot replace my set in the proposal and call it strategy. When I later score how to choose aeo agency finalists, I score them against this freeze, not against a query list invented to make a case study look tidy. I also name competitor entities for comparison.</p>

Measurement you will own, not rent

<p>I require an independent visibility log I can re-run without the agency. Engine, date, prompt, quoted passage, source URL. That row is the unit. I built AI Rank Checker after I spent a retainer reading screenshots I could not verify; what I saw is that a log I do not control is not measurement, it is a slide. I now write the same requirement into the brief I use for how to choose aeo agency partners, whether or not anyone uses that tool.</p>

<p>The agency may run their own tracker. I still re-run a prompt set myself. For how I pick a tracker, I keep notes in how to choose ai visibility tool in 2026. The brief says I own the raw exports at the end of the term. If that sentence is missing from the proposal, I add it before I sign. I will not accept a platform seat that disappears when the contract ends. Measurement is a deliverable, not a login.</p>

Constraints I put on the table

<p>I disclose the brakes before anyone prices the work. Legal reviews every public page and every claim about regulated products. Brand voice is locked: we have a tone guide, and I will not approve a Q&A block that sounds like a different company. In-house subject-matter experts have a five-business-day SLA; if they miss it, the calendar slips, and I will not treat that slip as the agency's miss. Some categories cannot pitch certain comparisons. Some pages cannot be rewritten because they are contractual or compliance-owned.</p>

<p>I put those limits in the one-pager. An agency that needs weekly legal turns I cannot give will fail the first month, and I would rather they decline than discover it after kickoff. Constraints are not a later alignment meeting. They are part of how I decide whether this hire can even run. I also name the single approver on my side.</p>

Process questions I will not skip

<p>I treat the live call as a process interview, not a slide tour. I already locked queries, entities, and measurement in the brief. What I still need is process: mapping, staffing, and markets. If an agency cannot walk me through those without a generic pod slide, I do not advance them. The questions below are the ones I will not skip when I decide how to choose aeo agency partners, even when the chemistry is good. I want names, examples, and a map I can audit.</p>

How they map prompts to pages

<p>I ask them to pick one prompt cluster from my brief and show the path. I want to see whether that cluster becomes a source page I already own, a new Q&A block on an existing URL, or a third-party mention they would pitch. An answer I accept sounds like this: for “best [category] software for [job]”, they would add a comparison table and a sourced FAQ to our product page, then ask two analysts we already know to cite the same numbers.</p>

<p>An answer I do not accept is “we optimize content for AI.” I also ask who decides when a page is the wrong unit and a mention on a trade site is the right one. If they cannot name that decision on the call, I treat the map as incomplete and I keep looking. I want the output named: URL, block type, or outlet. That map has to survive my own prompt re-run the next month intact.</p>

Staffing I expect when hiring an answer engine optimization agency

<p>When I am hiring an answer engine optimization agency, I do not buy a pod. I buy named people. On the call I ask for the researcher who will map my prompt clusters, the writer who will draft the source copy, the person who fact-checks claims against our legal pack, and the person who will sit on a call with our counsel if a sentence needs to change. I want those names in the proposal, not a role matrix.</p>

<p>If the researcher and the writer are the same person, I want that said out loud so I can judge capacity. I also ask who covers holiday weeks and who I email when a citation drops on a Friday. A staffing answer I accept includes hours per week on my account. An answer I do not accept is a slide with three icons labeled strategy, content, and reporting. I write those names into the kickoff doc so the account does not quietly rotate after month one.</p>

Multilingual and multi-market work

<p>The bar I use is work I ran for ERKE, a sportswear brand, across four languages. I did not translate an English FAQ and ship it. I trained source pages in each language against the prompts people actually typed in that market, then checked whether the engines cited the local page or fell back to English.</p>

<p>An agency that claims they can run more than one market has to describe that loop: who writes in the local language, who checks entity names and regulated claims, and how they log citations per language rather than rolling everything into one English dashboard. If they propose machine translation plus a native “review,” I ask to see the review checklist. If the checklist is missing, I keep the second market in-house until the first language proves out. I also ask which engines they will log in each market, because Copilot and Gemini do not always surface the same local sources ChatGPT does on the same prompt.</p>

Proof, reporting, and checks I still run

<p>I read reports the way I read a lab notebook. A citation either appeared, with a date and a URL, or it did not. I do not score decks. I score rows when I decide how to choose aeo agency partners. Before I sign, I ask to see a sample monthly log, I walk through one before-and-after claim in detail, and I keep a prompt set I re-run myself so the vendor’s spreadsheet is never the only evidence in the room. That is the whole proof stack I still run after we sign.</p>

What a citation log should contain

<p>The minimum row I will accept is five fields: engine, date, prompt, quoted passage, and source URL. Engine means ChatGPT, Perplexity, Google AI Overviews, Gemini, Copilot, Grok, or Claude, named, not “AI search.” Date is the day we observed the answer, not the day the report was compiled. Prompt is the exact string, including the market and language. Quoted passage is the sentence the model used, copied, not paraphrased by the account manager. Source URL is the page the engine attributed, even if it is not ours.</p>

<p>I also want a column for whether our brand was named, and a column for the competing brand if we were omitted. A log without those fields is a narrative. I cannot audit a narrative. I export that log to a sheet I own. Screenshots help, but they are not a substitute for the row. If the agency’s tool cannot produce those five fields, I still require the fields, typed by hand if needed.</p>

How I read before-and-after claims

<p>I ran a Q&A rewrite for a sustainability consultancy. The source page already ranked for a handful of classic queries. What it did not do was get named inside chatbot answers about carbon accounting for mid-market manufacturers. I rewrote the FAQ into short, sourced answers with the firm’s method named in the first sentence of each block. I then re-ran the same prompt set for four weeks.</p>

<p>The change I measured was this: the consultancy’s method started appearing as a named citation in answers that had previously named a larger rival or no firm at all. When someone else sends me a case study, I ask for the prompt list, the dates, the engines, and the exact sentences that changed. A percentage without those four things is not a claim I can use. I also ask whether the old page stayed live, because a before-and-after that swaps URLs is a different test.</p>

Independent checks I still run

<p>I keep a frozen prompt set of the clusters in my brief. Once a month I run those prompts myself in the engines we named in the contract. I record the same five fields I demand in the vendor log. If their row and my row disagree, we discuss the discrepancy before we discuss “progress.” I change nothing in the prompt set during a quarter unless the business adds a product or a market. That is the point: the check has to be boring enough to repeat.</p>

<p>I also run one off-brief prompt, the kind a sales prospect would type, so we do not optimize only for the list we invented. If an agency objects to me re-running their work, I treat that as a process answer, not a personality clash. I screenshot the answer on the day I run it and store it next to the log row, so a later model change does not erase the observation.</p>

Commercial terms I put in writing

<p>I put commercial terms in the same one-pager as the brief, before anyone sends a SOW. I want the shape of the fee, what the fee includes on paper, who owns the copy and the logs, and how we stop. I write those as observable clauses, not as trust. If a term is not in writing, I treat it as not agreed. The three subsections below are the clauses I will not leave to a kickoff conversation. That habit has saved me more arguments than any slide about partnership.</p>

Fees I expect when hiring an answer engine optimization agency

<p>When hiring an answer engine optimization agency, I compare three shapes by what is written on the paper, not by a monthly number. A project lists a prompt map, a set number of source pages or Q&A blocks, a round of third-party outreach, and a citation baseline. A retainer lists research hours per month, pages or blocks in scope, outreach volume, and a monthly log. A hybrid is a paid diagnostic plus a retainer that only starts after I accept the diagnostic.</p>

<p>I ask each shape to name what happens if we need extra engines or an extra language. I also ask whether reporting hours sit inside the fee or get billed as a separate line. If those inclusions are not on the SOW, I add them before I sign. I do not treat a lower monthly figure as the better offer if the paper omits research hours or the log. The comparison I use for how to choose aeo agency fee shapes is scope, then price, in that order.</p>

Who owns content, data, and prompts

<p>I require written ownership at the end of the term. Copy we paid to have written is ours, including drafts that did not ship. Prompt libraries built against my brief are ours. Raw citation exports, the rows, not a dashboard login, are ours. I want those three assignments in the contract, not in an email after the last invoice.</p>

<p>If the agency uses a platform I cannot export from, I still require a monthly CSV of the log. I also require a clause that they will not reuse my unpublished claims or my prompt list for another client in the same category during the term. I am not asking for their methods. I am asking for the artifacts my fee produced. If ownership is missing from the SOW, I pause the hire until it is added. I have signed without that clause once. I do not repeat it.</p>

Exit, overlap with in-house SEO, and handoff

<p>I write a notice period in days, not “reasonable notice.” I name the in-house SEO owner in the contract and I list the work they already do, technical crawls, classic rankings, on-site templates, so that scope is listed once. On the last day I want: the prompt library, the source-page map, every citation export, unpublished drafts, login credentials I paid for, and a two-page note on what is in flight.</p>

<p>I also want a two-week overlap where the agency answers questions from the in-house owner. If the handoff is not on paper, I plan to reconstruct the program from screenshots. I have done that. I would rather not. I schedule the overlap before the notice period ends, not after. I also confirm that in-house SEO keeps publishing on the same source URLs so we do not create two maps. I write those two rules into the notice clause before we start.</p>

Why I buy a diagnostic before a retainer

<p>I buy a two-week paid diagnostic before I sign a retainer. A pitch deck shows how they talk. A diagnostic shows how they work on my queries, my entities, and my source pages.</p>

<p>I treat it as a working interview: they get a bounded fee, I get artifacts I can re-run, and neither of us is locked into a six-month shape that does not fit. If the work is thin, I stop. If it is specific, I have a real basis for how to choose aeo agency retainers.</p>

The two-week scope I actually buy

<p>The two-week scope I actually buy is four artifacts, not a strategy slide. First I want a prompt map: the clusters we locked in the brief, the engines I named, and the source page or third-party mention they would attach to each cluster. Second, a source-gap list: URLs that already exist, URLs that would need a rewrite, and claims that have no citable page at all.</p>

<p>Third, sample rewrites on two or three of those gaps so I can see how they handle voice, facts, and question-and-answer structure on my material. Fourth, a citation baseline I can re-run myself: engine, date, prompt, quoted passage, and source URL for the starting state.</p>

<p>I need a file I can open in two weeks and still understand. The baseline has to be a file I keep, not a screenshot inside their deck. If any of those four is missing, the diagnostic is incomplete for how to choose aeo agency work, and I do not treat a verbal walkthrough as a substitute.</p>

A diagnostic before hiring an answer engine optimization agency

<p>I pass a diagnostic when the files match the brief I sent. The prompt map uses my clusters, not a generic industry list. The rewrites keep my brand voice and cite facts I can check. The baseline is a spreadsheet I can re-run on the same prompts without asking them to log in for me.</p>

<p>I also ask who produced each artifact. If the person on the call cannot name the researcher, the writer, and the fact-checker, I pause. I pause when the work arrives late, when sample pages ignore my constraints, or when the only proof is a slide with unnamed engines.</p>

<p>Polish does not move the decision. Named owners, dated rows, and source URLs do. I decide from the files, not from how the room felt. If two shops both pass, I take the one whose diagnostic I could hand to my in-house SEO without a translator. That is the bar I use before hiring an answer engine optimization agency on a retainer.</p>

How I score a shortlist

<p>I score finalists on a sheet I fill in the same week as the last diagnostic. Weights are written before the first call so a charming deck does not rewrite the criteria I use for how to choose aeo agency finalists. I write those weights down.</p>

<p>I am not ranking companies as good or bad. I am measuring process transparency, independent measurement, category fit, and contract terms against the brief I already locked. When the scores are close, I keep the work in-house instead of forcing a hire.</p>

Criteria I use for how to choose aeo agency finalists

<p>I weight four criteria. Process transparency is 30 percent: can they show how a prompt cluster becomes a source page, a Q-and-A block, or a third-party mention, with named people on research, writing, fact-check, and legal. Independent measurement is 25 percent: I need a citation log I own, not a portal I rent, with engine, date, prompt, quoted passage, and source URL.</p>

<p>Category fit is 25 percent: have they shipped work in my market, language, and sales cycle, or am I paying them to learn it on my retainer. Contract terms are 20 percent: ownership of copy, prompt libraries, and raw exports; notice; and a last-day file set my in-house team can open.</p>

<p>I score each from one to five against what they delivered in the diagnostic, not against their homepage. I score only what I can point to in a file. I use the same grid every time I decide how to choose aeo agency finalists. It is not a statement about anyone's worth.</p>

When I keep the work in-house instead

<p>I keep the work in-house when three conditions hold at once. The query set is narrow: a handful of prompt clusters, one or two languages, and a category where I already know the competing entities. The source pages are already strong: they answer the questions the engines quote, they carry unique facts, and they are not commodity roundups.</p>

<p>And someone on staff already owns the loop: research, rewrite, fact-check, and the monthly re-run of the citation log. In that case a specialist retainer mainly adds coordination. I still run a diagnostic if I am unsure whether the pages are as strong as I think.</p>

<p>If the gap list is short and the sample rewrites are work my editor can finish, I do not sign. I would rather spend the hours on the pages than on a vendor call. The hire exists to cover work I cannot staff. It does not exist because the market now has a new label. Those three conditions are the whole test.</p>

Last-call questions that decide the hire

<p>The last call is not a recap of the deck. I already have the diagnostic files and the score sheet in front of me.</p>

<p>I use the hour to hear how they talk about my category and my sales cycle, and to lock the answers I need in writing: who owns the work, what ships in the first 90 days, and who I call when a citation disappears. If those answers stay vague, I do not send paperwork. Charm does not close this hire. Named people and dated outputs do.</p>

Fit with category and sales cycle

<p>I ask for two live examples in my category, not adjacent ones. I want the prompt, the engine, the cited URL, and the date. Then I ask how that citation would show up in a sales conversation: would a buyer have seen the brand named in an answer, and would a sales engineer be able to point to the same source page we cite on a call.</p>

<p>A dashboard row that never reaches a human conversation is not the outcome I am buying. If they cannot walk me through two examples without a prepared slide, I treat category fit as unproven.</p>

<p>I also ask which part of our sales cycle they think the work supports: awareness, shortlist, or proof after a first meeting, because that changes which prompts I fund first. I write their answer into the brief so the first 90 days are not a tour of every engine at once. I need that answer on the call, not in a later email.</p>

Answers that close the decision

<p>I need four replies in the room, then in the statement of work. Who owns the work by name: researcher, writer, fact-checker, and the person who speaks to my legal team. Timeline: when the prompt map is frozen, when the first source pages ship, and when the first citation log lands.</p>

<p>First-90-day output: how many clusters, how many pages or rewrites, how much outreach if any, and which engines are in the log. And who I call when a citation disappears: a named person, a response window, and whether they re-run the prompt or only wait for the next monthly report.</p>

<p>If any of those four is a promise to confirm after kickoff, I do not sign. I also confirm the handoff file set one last time so an exit does not depend on goodwill. I write those names into the contract. I want the call tree on paper. Those answers close the decision for me on how to choose aeo agency partners. Everything else was already on the score sheet.</p>

Frequently asked

An SEO retainer usually buys crawl health, rankings, and classic-result traffic. When I hire for AEO, I buy citation inside generated answers, not a blue-link slot. Work shifts to entity clarity, unique source material, and multi-engine coverage. Google still stresses valuable, unique, non-commodity content for generative AI features, that is the overlap.

Before I shortlist anyone, I write the queries I need cited, the engines that actually matter to my buyers, and a one-page definition of success, branded mention, linked citation, or both. I also list what in-house SEO already owns so the agency is not bidding on work we already do. Without that brief, every pitch sounds the same.

I keep a retainer only when in-house SEO cannot own multi-engine citation work. Classic ranking skill does not automatically produce ChatGPT or Perplexity citations. Pew found 49% of U.S. adults used AI chatbots in February 2026, and 42% of chatbot users used them to search for information, that demand sits outside a typical SEO ticket queue.

I pay for a diagnostic that returns a query list, current citation baselines on each engine, and the pages that already look citable versus commodity. It should include how they will measure mentions versus linked citations, plus a 30-day work plan. If the deliverable is only a slide deck of opportunities, I do not convert it into a retainer.

I require coverage where my buyers actually ask: ChatGPT, Google AI Overviews, Perplexity, Gemini, Copilot, Grok, and Claude. Pew found ChatGPT was the most commonly used chatbot across the age groups it examined, so I never treat it as optional. I also ask how they handle Google’s generative features for local, shopping, image, and video content.

I rerun the agency’s prompt set myself, on a logged-out session, and save dated screenshots of whether my brand is cited, mentioned, or absent. I also sample adjacent phrasings, not just their demo queries. When I built AI Rank Checker, I used the same method: independent prompts, repeated over days, compared against the slides they sent.