Core / Pillar 28 min read Published Updated

How Do LLMs Choose Which Brands to Recommend? (2026 Guide)

I spend my weeks logging which brands ChatGPT, Perplexity, Gemini, and AI Overviews name, and why the same prompt rarely returns a stable shortlist. This is the retrieval-plus-training map I actually work from.


On this page
Line drawing of a practitioner comparing training-memory brand names against retrieval-ranked brand names

Key takeaways Read this if nothing else

  1. 01

    I treat every brand pick as two stacks, training memory plus retrieval ranking, not as a single brand score.

  2. 02

    Marketplace-visibility signals such as search interest and brand conversation associate with prominence; they are not a switch I can flip in isolation.

  3. 03

    A well-known brand can dominate when specs match, but a sub-0.1-star quality edge can erase that baseline, so I invest in checkable proof.

  4. 04

    Wording can move outcomes, so I keep authority language verifiable and I log every recommendation as a run, not as a permanent ranking.

How I Watch an LLM Pick a Brand

<p>I log brand names the way I used to log SERP positions: same prompt, same engines, a timestamp, and a screenshot. I do this because a recommendation is not a rank, and the same query can name a different shortlist on Tuesday than it did on Monday. The rest of this article uses one vocabulary for that event. I need a hard definition, a frozen prompt set, and a clean contrast with classic rank checking, the same gap I covered in ai visibility vs seo in 2026.</p>

What I count as a recommendation

<p>I count a recommendation only when the model names a brand as something to buy, use, hire, or consider in the answer body. A cited URL means the generator retrieved a page and chose to footnote it. A soft mention is a brand that appears in passing, a parenthetical, a quote, a table cell, without being offered as an option. I keep three buckets in every log row when I record how do llms choose brands: named recommendation, cited URL, soft mention.</p>

<p>If ChatGPT writes “consider Brand A or Brand B for this job,” both count. If it footnotes a roundup on Brand C’s domain, that is a citation. If it says “many teams use tools in this category” and never names one, that run has zero recommendations. I do not count logos, shopping modules I cannot capture as text, or a brand that appears only inside a quoted snippet. The spoken shortlist is the line I use from here.</p>

The prompt set I reuse every week

<p>I freeze three prompt families and re-run them every week on the same engines. Shopping prompts look like: “What [product] should I buy for [use case]?” plus a budget and one constraint. Comparison prompts look like: “Brand A vs Brand B vs Brand C for [job].” Category prompts look like: “Best [category] for [audience] this year.” I keep the wording identical across ChatGPT, Perplexity, Gemini, Google AI Overviews, Copilot, Grok, and Claude. I do not add “be unbiased” or “list ten options” unless that clause is the test.</p>

<p>I record the model label when the UI shows it, and I note whether browsing or shopping mode was on. I reuse the same seed brands in comparison prompts so I can see who enters or leaves the shortlist. I built AI Rank Checker to log those runs faster; the prompt text itself stays frozen. Side tests sit outside the weekly core.</p>

Why this is not a classic rank check

<p>A classic rank check snapshots the ten blue links for a keyword. I still run those. They do not tell me who an answer engine will name. A SERP position is a URL competing for a query. An LLM recommendation is a brand the generator chose to speak after retrieval and after whatever priors sit in the weights. One URL can rank and never get named. A brand can get named with no citation attached. Positions also tend to hold more steadily week to week than the spoken shortlist I log. I treat the two as different instruments, not as two views of the same scoreboard.</p>

<p>Mixing them is how I used to overclaim a win: a page ranked, so I assumed the model would name us. The logs did not support that. When I need to know how do llms choose brands in a live answer, I re-run the frozen prompt set.</p>

How to Get AI to Recommend Your Brand in 16 minutes | AEO Playbook video thumbnail

Video: How to Get AI to Recommend Your Brand in 16 minutes | AEO Playbook · Grace Leung

Training Memory Versus Live Retrieval

<p>I used to talk about “getting cited by ChatGPT” as if it were one pipe. It is two. Some brand names live in the weights from pretraining and post-training. Some names only appear because a live retrieval pass fetched pages at answer time. Mixing those stacks is how I misread a named brand as a crawl win. Separating them is how I read how do llms choose brands on a given day. I keep them separate now. That split is also how I think about more on aeo vs geo.</p>

What stays in the weights

<p>Brand priors that do not depend on this week's crawl sit in the weights. Pretraining saw the public web, books, forums, and whatever else the lab mixed in. Post-training, instruction tuning, preference data, safety layers, then reshaped which names the model is willing to speak. I cannot inspect those corpora. I can only watch the residue: a household name that appears even when I disable browsing, or a niche brand that never appears until a retrieved page carries it.</p>

<p>When ChatGPT names a brand with browsing off, I treat that as a weight prior, not as proof my sitemap was fetched today. When the same prompt with browsing on adds a different third name, that third name is a retrieval candidate. I do not pretend I know the cutoff date of the weights. I only log the difference between memory-only and retrieval-on runs, and I treat that gap as the prior I can see in how do llms choose brands.</p>

What gets fetched at answer time

<p>At answer time, engines that browse or ground will fetch. The retrieved set is the shortlist the generator then ranks into prose. If a brand never enters that set, I do not expect it to be recommended, no matter how often I said the name in a blog post last year. What I actually see fetched: category pages, comparison roundups, retailer listings, docs, and the odd Reddit thread. Shopping and comparison prompts pull more product-schema pages; category prompts pull more listicles.</p>

<p>A 2026 arXiv paper on retrieval and ranking inside LLM-driven systems makes the same split I use in logs: visibility in generation is tied to that retrieval-and-ranking behavior, not to the brand name alone. I still confirm it per run: who was cited, who was named, who never appeared. If the citation list is empty, I treat the names as weight-driven until a later run proves otherwise.</p>

Why marketers mix the two up

<p>I used to treat a named brand as proof of a crawl win. A client would ship a new comparison page on Monday. ChatGPT would name them on Wednesday. I would write “the page got retrieved.” Then I would re-run with browsing off and the name was still there. The page had not been the cause. The weights already knew the brand.</p>

<p>The mix-up is emotional as much as technical. Retrieval work is something I can ship this quarter. Training memory is a slow public record. It is easier to celebrate the page I published than the five years of third-party mentions I cannot control. I still ship crawlable pages. I just stop reading a spoken name as evidence that this week's fetch succeeded. The test is the browsing-off run, the citation list, and whether the name survives when I change the live index. That is the check I run now.</p>

How ChatGPT Recommends Brands in Practice

<p>I do not have ChatGPT’s ranking code. I have logs. This section is how ChatGPT recommends brands in the runs I actually keep: frozen prompts, weekday retests, and a split between names in the prose and URLs in the footnotes. I also keep a clean crawl surface so retrieval has something to fetch. I wrote more on what is llms.txt when I started shipping that file, but the file is not the recommendation.</p>

Same prompt, different days

<p>I freeze the prompt and still watch the shortlist drift. Same account, same wording, different days. ChatGPT will keep a core of two or three names and rotate the fourth. Sometimes a brand drops for a week and returns. I do not treat a one-run swap as a ranking event. What I log as drift: a name that was in the spoken shortlist on Monday and gone on Thursday, with no site change on our side. What I do not log as drift: a paraphrase of the same three brands.</p>

<p>Temperature, routing, and whichever retrieval index was hot that hour all sit outside my control. I re-run three times in a sitting when a claim depends on the shortlist. If two of three runs agree, I note a pattern. If they disagree, I wait for the next weekly pass. Stability is a property of the log, not of a single screenshot.</p>

Cited pages versus named brands

<p>A brand ChatGPT speaks and a page it footnotes are different events. I can get a named recommendation with zero citations. I can get three footnotes to review sites and a spoken shortlist that never includes the brands those reviews cover. I record both. When the spoken name and the cited host match, I note an aligned pick. When they diverge, I do not collapse them. Both columns stay.</p>

<p>A footnote to a review host is not a recommendation of that publisher as a product, and a footnote to a brand’s compare page is not automatic proof the brand was recommended, only that the page was retrieved. I also see named brands whose official site never appears in the citation list. Those names are coming from weights, from another retrieved page that talks about them, or from both. I do not guess which. I log the split. That split is the row I trust when I write down how do llms choose brands on that run.</p>

Shopping and comparison prompts

<p>Frame changes the shortlist. A shopping prompt (“what should I buy for X”) tends to return a short, decisive set, often three names, sometimes with a budget pick. A comparison prompt that already contains Brand A vs Brand B vs Brand C usually speaks those three, then sometimes adds a fourth “also consider.” A category prompt (“best Y for Z”) behaves like a roundup: more names, more hedging, more citations.</p>

<p>I keep the product specs in the prompt when I have them, size, price band, constraint, because ChatGPT fills gaps with famous defaults when I leave the spec blank. That is not a moral claim. It is what the logs show. If I ask for a comparison without naming anyone, the model supplies the set. If I name three challengers, the household brand still often appears as the extra slot. I use that extra slot as a signal of prior strength, not as a crawl report.</p>

Retrieval and Ranking Inside Answer Engines

<p>When I log an answer, I treat two steps as separate. Retrieval decides which brands even enter the candidate set. Ranking inside the generator then orders those candidates as the prose is written.</p>

<p>A 2026 study on LLM retrieval puts it the same way I see it in the field: visibility in generation is not just about the brand itself; it is tied to retrieval and ranking behavior inside LLM-driven systems. That is the map I actually use when a brand is named or skipped.</p>

Retrieval is the shortlist

<p>I used to treat a named brand as an awareness win. My logs disagree. If the retrieved set never includes a brand, the generator has nothing to name. On shopping and comparison prompts, the engines I test pull a handful of pages or entity records, then write from that pile. A brand that exists only on an uncrawlable PDF does not show up.</p>

<p>That is why I treat retrieval as the shortlist. I cannot talk a model into recommending a brand it did not fetch. When I ask how do llms choose brands on a category prompt, the first filter is membership: did a product page, comparison table, or third-party roundup enter the set at answer time? If not, ranking work is wasted.</p>

<p>A spoken name with no retrieved page is training memory, not a crawl win. Those two cases need different work. I look at citations, quoted specs, and whether the same URL appears across engines before I call it retrieval.</p>

Ranking inside the generator

<p>Once a brand is in the retrieved set, the generator still decides order, mention length, and whether to recommend or just list. That second pass is ranking inside generation. I do not see a public scoreboard. I see the written answer: who is named first, who gets a qualifier, who is omitted even though a cited page mentioned them.</p>

<p>On comparison prompts I watch a longer retrieved set collapse into a shorter spoken list. Specs that are easy to quote, battery life, price band, warranty years, tend to survive. Vague claims do not. The generator orders what retrieval already handed it.</p>

<p>This is also where how chatgpt recommends brands can look like a popularity contest even when the retrieved set is mixed. A well-documented challenger can sit in the footnotes while a household name takes the spoken slot. I log both. Ranking is who the model elects to write first from the pile it was given. If I want a different order next week, I work on those pages, not on a slogan.</p>

Crawler hints I actually ship

<p>I cannot control the generator's taste. I can control what a crawler is able to fetch cleanly. The surfaces I actually maintain are boring on purpose: a crawlable product page with specs in HTML, not a screenshot; a category page that states the same entity name I use on the product; a comparison table with units; an organization schema block that matches the legal name; a sitemap that lists those URLs.</p>

<p>I ship robots.txt that allows the answer-engine crawlers I care about. I keep important specs out of infinite-scroll widgets and out of PDFs that return 403. I use one canonical URL per product so retrieval does not split the same SKU across near-duplicates.</p>

<p>None of this guarantees a recommendation. It only gives retrieval something unambiguous to pick up. When a brand I work on fails to enter the shortlist, I check fetchability first: status code, indexable HTML, consistent name, visible specs. If those are broken, I do not blame ranking when I ask how do llms choose brands.</p>

Marketplace Visibility Signals I Keep Logging

<p>I also log marketplace-visibility signals because they travel with brand prominence in retrieval. I do not treat them as a formula I can punch in. A 2026 arXiv paper on brand prominence reports that prominence in LLM retrieval and ranking is associated with broader marketplace-visibility signals, especially search interest and online brand conversation.</p>

<p>That matches the pattern in my weekly runs: household names enter the shortlist more often. Association is not a switch. I keep both signals in the log anyway.</p>

Search interest as a prior

<p>Search interest is the prior I notice most often. Brands that people already query in classic search show up in my LLM shortlists more often than quiet specialists with better spec sheets. I do not have a coefficient. I have a co-occurrence: when search interest is high in a category I track, those names survive retrieval even when my prompt never mentioned them.</p>

<p>The 2026 retrieval study treats search interest as one of the marketplace-visibility signals associated with brand prominence in LLM retrieval and ranking. I read that as a prior, not a ranking factor I can buy. A spike I manufacture with a two-week campaign does not rewrite the model's sense of who belongs in a category.</p>

<p>I check whether the brand I am helping even appears in category-query conversation. If search interest is near zero, I expect the live fetch to prefer roundups that already name the incumbents. I still ship crawlable pages. I do not pretend a new URL outruns a decade of query volume in one week when I explain how do llms choose brands in a quiet category.</p>

Online brand conversation

<p>The other signal I keep in the same notebook is public brand conversation: reviews, forum threads, news mentions, and comparison pieces that name the product in the category. I am not scoring sentiment with a secret model. I am checking whether the brand is spoken about in places retrieval can fetch.</p>

<p>The same 2026 paper groups online brand conversation with search interest as an associated prominence signal. In my runs, a brand that is already named in roundups and review hubs is easier for retrieval to justify. A silent brand with a perfect product page still needs a third-party page to co-occur with the category.</p>

<p>I do not seed invented reviews. I also do not treat a viral week as proof the weights moved. Conversation is a visibility companion. When it is absent, the shortlist I log is almost always the names everyone already repeats. I log the URLs, not the vibe. Those URLs are what a live fetch can actually pick up in how chatgpt recommends brands.</p>

What I do not treat as proof

<p>This is the line I draw so I do not overclaim. Search interest and online conversation travel with prominence. They are not a lever I flip on Tuesday and collect a recommendation on Wednesday. The study reports association. My logs show co-occurrence. Neither is a causal recipe.</p>

<p>I used to treat a PR burst as a shortlist mint. I stopped. After a launch week I re-run the same prompt set. Sometimes the brand is named. Often it is not, and the retrieved URLs are still the incumbent roundups. That is the prior still sitting in the shortlist.</p>

<p>What I will claim: if a brand has no crawlable record and no public conversation, I do not expect it to be recommended. What I will not claim: that raising search interest by some amount guarantees a named slot. Prominence signals explain why household names keep appearing. They do not let me write how do llms choose brands as a checklist with a search-volume cell.</p>

Established-Brand Bias and Tiny Quality Flips

<p>Retrieval explains who can be named. It does not explain why two products with the same specs still produce a lopsided shortlist. I keep a separate note for that: established-brand bias, and the tiny quality flips that can erase it.</p>

<p>A 2026 arXiv experiment on brand recommendations is the cleanest write-up I have of both. I read it against my own comparison-prompt logs, not as a ranking of any company's worth. The two findings belong in one notebook.</p>

How do LLMs choose brands with identical specs

<p>When product specifications are identical, the experiment found that well-known brands can receive 100% recommendation rates. That is a baseline bias toward established names, not a quality judgment I can verify from the spec sheet. I see a milder version of the same pattern in my frozen comparison prompts: when I strip adjectives and leave only matching attributes, the household name still takes the spoken slot.</p>

<p>I do not treat 100% as a number I have reproduced across every engine I log. I treat it as the paper's finding under its protocol. What I can say from my runs is directional: identical specs do not produce a coin flip. The known brand is named first, often named alone.</p>

<p>So when someone asks me how do llms choose brands with matching feature lists, my field answer is: they lean on the prior. Fame is doing work the spec table is not. If I need a challenger in that shortlist, matching specs is the floor, not the differentiator.</p>

The 0.1-star swing

<p>The same 2026 experiment found that this dominance can disappear with less than a +0.1-star rating advantage for a competitor. I keep that number taped next to the 100% finding because they belong together. Established-brand bias is strong when everything else is equal. It is not immortal.</p>

<p>I have not run a lab where I edit a rating by a tenth of a star and watch an engine flip. I have seen comparison prompts where a challenger has a slightly better public rating and clearer specs, and the famous name is no longer alone. That is consistent with the paper, not a replica of it.</p>

<p>The practical read: a tiny, checkable quality edge can change who gets recommended. I look for edges I can put on the page, warranty length, measured battery figures, a current public rating, not a claim of being "just as good." The 0.1-star swing is the paper's reminder that the prior can be overridden by a small, visible difference when I watch how do llms choose brands.</p>

What that means for challengers

<p>If I work a brand that is not a household name, identical specs are not enough. The paper's 100% finding is the wall. The 0.1-star finding is the door. My job is quality and proof retrieval can fetch: a real rating edge, a better spec in HTML, a third-party page that states the difference without adjectives I invented.</p>

<p>I tell teams to stop asking how chatgpt recommends brands as if a tagline will substitute for that proof. It will not, in my logs. A challenger that matches the incumbent and then adds one checkable advantage, longer warranty, higher published rating, a measured figure, is the pattern I trust against fame bias.</p>

<p>I keep the association limit in view. Search interest and conversation still travel with the incumbents. A quality flip on the page does not erase a decade of prominence overnight. I retest. I look for a repeated named slot, not a single lucky run. Challenger work is slow: crawlable proof, then another pass of the same prompts.</p>

Wording, Authority Cues, and Persuasion Signals

<p>I used to treat copy as decoration once a page was crawlable. That was a miss on my side. After the identical-spec runs, I started watching how a product paragraph is phrased, not just whether the brand name appears. Authority-style language can change who the model names. I still refuse to ship claims I cannot document. This section is the line I draw between a checkable cue and an invented one when I test how do llms choose brands.</p>

Authority language the study tested

<p>I keep the 2026 experiment beside my prompt log because marketers over-read the wording result. The wording arm of that 2026 study tested authority-style marketing language, including clinical-evidence claims constructed for the experiment, and reported that those phrases can alter who gets named. I cite that as a sensitivity finding, not a playbook. In my own runs I do not inject invented studies. I watch live pages that already say "clinically proven" or "the leading choice," versus pages that state a measurable spec and a dated test method. When a sandbox shortlist flips after I change only the claim sentence, I treat that as the generator reacting to persuasion cues in retrieved text. I do not treat it as a durable ranking rule. ChatGPT, Perplexity, and Gemini do not share one copy filter. I log the exact sentence, the engine, and whether the brand was named, cited, or both.</p>

Persuasion signals I can verify

<p>The cues I actually ship are boring on purpose. On a product page I keep a spec table with units, a warranty length as a number, and the name of any certification that resolves to a real organization URL. If I mention a lab test, I name the lab and the test date. If I mention reviews, the count and star figure have to match a marketplace page I can fetch that week. Category pages get a comparison table with sources, not adjectives. Author bylines match a public profile. I built AI Rank Checker so I could log whether those pages were named or cited after a copy change; a dated test method often travels with a citation even when the brand is not spoken first. I treat that as an association in my notes, not a causal proof. Every cue I add is something a retriever can fetch and a human can check. Uncheckable superlatives stay off the page even when I am studying how chatgpt recommends brands.</p>

Where I stop even if the model reacts

<p>The 2026 paper's constructed-claim result is the line I will not cross. If a model names a brand more often after I paste an invented clinical sentence onto a staging URL, that is a sensitivity I record, then I revert the sentence. I do not ship invented studies, invented awards, or invented clinician quotes, even when a generator appears to reward them. I also skip "as recommended by ChatGPT" badges on product pages; they are not a retrieval surface I can defend. When someone asks me to "sound more authoritative," I translate that into a dated method, a named lab, or a third-party page that already exists. If none of those exist, the page stays quieter. Model sensitivity is not a brief. I leave a checkable record, not a wording trick I cannot reproduce on the next Tuesday run of the same prompt set.</p>

What Marketers Can Actually Influence

<p>I split every request into two piles: what I can change on crawlable surfaces this quarter, and what only a slow public record can feed into future weights. Mixing those piles is how I used to waste months. Retrieval levers are pages, specs, and entity strings I can still edit. Training levers are third-party mentions that persist into a later snapshot. I still run other tactics. I no longer treat them as the reason a brand got named. This is the finite list I actually work.</p>

Levers on the retrieval side

<p>On the retrieval side I still control a short list this quarter. I keep one canonical brand string and reuse it in the title, the H1, the Organization markup, and the about page. Product and category URLs stay fetchable without a login, and the sitemap I serve lists those URLs. Specs live in a table with units, not in a paragraph of adjectives. For comparison prompts I maintain a comparison page that uses the same entity names the prompt uses. I update the visible date only when the facts changed. I remove duplicate brand spellings that used to split the entity. None of that guarantees a recommendation. It only puts a clean candidate into the retrieved set the generator can rank. A brand that never enters that set cannot be named. I re-fetch my own pages the same week I re-run the prompts, so I know what the live pass could have seen when I ask how do llms choose brands.</p>

Levers on the training side

<p>I cannot edit the weights. I can only leave a public trail that a later training or post-training pass might see. That trail is slow. I look for third-party pages that already discuss the category and can name the brand with the same string I use on-site: industry roundups, retailer listings, standards bodies, and news items with a date and a publisher. I do not treat a single guest post as a training win. I treat a consistent, checkable mention that stays up for months as a prior that might survive into the next snapshot. Public entity pages, when they exist and are accurate, are part of that record; I do not invent them. Review corpora and long-lived comparison articles matter more than a social burst I cannot find six weeks later. A household name already has that trail. A challenger has to earn it in public, in language a future crawl can quote, without inventing the proof.</p>

Levers I stopped overrating

<p>I still write meta descriptions, pitch journalists, and post on the channels the brand already uses. I stopped treating any of those as the reason a model named the brand. A homepage hero rewrite does not show up in my logs as a stable shortlist change. Schema that does not match visible text does not either. A one-week social burst is gone by the next re-run. A roundup I cannot fetch a month later is not a training prior I can point to. I also stopped building pages whose only job is to say an answer engine named the brand; they confuse the log. When a named-brand flip happens, I look first at the retrieved set, a spec or rating change, and whether a third-party page started using the canonical string. If none of those moved, I mark the flip as unexplained and wait for the next weekly pass.</p>

How I Retest Brand Recommendations

<p>A one-run swap is not a finding. I keep a structured log so claims stay tied to runs. I re-run a frozen prompt set weekly, and I only talk about a shift after it repeats on the same engine. That stops me from claiming a model now prefers a brand after a Tuesday session that reversed on Thursday. Below: the fields I record, the events that trigger an extra pass, and how I read noise.</p>

The log I keep

<p>Every row in my log is one prompt on one engine on one day. I store the date, the engine name, and the model label if the UI shows one. I paste the exact frozen prompt; I do not rewrite it mid-week. I record named brands in the order they appear, then cited URLs separately, then a flag for named, cited, or both. I note the comparison frame if the prompt was a shopping or versus query. If I changed a page that week, that URL and the change type go in the same row. I keep a copy of the answer text, not just my summary of it. I do not score sentiment or invent a share-of-voice figure from five runs. The log exists so that when I say a brand dropped, I can point at the rows. Soft mentions that never name the brand stay in a notes field and do not count as recommendations when I score how do llms choose brands.</p>

When I re-run the set

<p>I re-run the frozen set every week on ChatGPT, Perplexity, Gemini, and Google AI Overviews. I add Copilot, Grok, and Claude when I already track that category. An extra pass happens when I ship a page change that could enter retrieval: a spec table, a rating figure, a comparison URL, or a brand-string fix. I also re-run when the UI shows a different model label than last week, because that is a stack change I did not make. I do not re-run because someone "has a feeling," and I do not expand the prompt set mid-flight to chase a friendlier answer. If a brand publishes a material product change, I add one extra shopping prompt that names the new spec and I leave the original prompt untouched. I timestamp every extra pass so it does not get mixed into the weekly baseline. Cadence is how I separate a session blip from a pattern I am willing to write down about how chatgpt recommends brands.</p>

Reading a shift without overclaiming

<p>If brand A is named on Monday and brand B on Wednesday, I do not publish a switch. I wait until the same frozen prompt, on the same engine, shows the new shortlist across three weekly runs, or until I can pair the swap with a retrieval change I can fetch: a new cited URL, a moved rating, a spec table that went live. Two-run flips stay in the log as noise. A Perplexity citation is not a ChatGPT naming, and those rows are not a single story of how chatgpt recommends brands. A Google AI Overview link is not a spoken recommendation in Gemini. When I write up a shift, I state the prompt, the engine, the run dates, and the named-versus-cited flag. I do not convert that into a market-share number. A one-run swap is not a theory of how do llms choose brands. It is a row, and it stays a row until it repeats.</p>

Frequently asked

When two products look identical, I have seen models default to the more established name. A 2026 arXiv experiment found well-known brands can receive 100% recommendation rates when specifications match. That same work showed the dominance can disappear with less than a +0.1-star rating advantage, so a tiny quality signal can flip the pick in how do llms choose brands.

I test them as different retrieval stacks, not as one engine. ChatGPT often answers from conversation context and whatever browsing it has enabled. Perplexity leans on cited web snippets in the same turn. Google AI Overviews stays closer to ranking pages already in the search index. I never treat a brand win on one as a win on the others.

In my own prompt logs, the brands that get named are the ones retrieval can find and rank first. A 2026 arXiv study tied that prominence to marketplace-visibility signals, especially search interest and online brand conversation. I also watch ratings, review volume, and how clearly a page states specs, because those change what the model can quote.

Yes. I have watched smaller names surface when a retrieval pass finds a clearer spec sheet or a better rating. A 2026 arXiv experiment showed well-known brands can hit 100% recommendation rates when specs are identical, yet that pattern can disappear with less than a +0.1-star advantage for a competitor. Fame is the default, not a lock.

I re-run the same branded prompts weekly on ChatGPT, then again after any model-label change I can see in the UI. Retrieval drifts even when the prompt does not. I keep a fixed prompt set, log which brands appear, and compare against Perplexity and AI Overviews in the same sitting so I am not mixing dates.

No. Training data sets a prior, but live retrieval can overwrite it. A 2026 arXiv study found visibility in answer generation is tied to retrieval and ranking inside LLM-driven systems, not only to the brand itself. When I change what the web says about ratings or specs, I often see the named brand change on the next test.