Skip to article
ByDefault
How it worksFAQPricingBlog
Sign inGet started
All articles
AI SearchSeptember 28, 202612 min read

How ChatGPT Web Search Works (2026 breakdown w/ API Example)

Josh

Co-founder

On this page
When does ChatGPT search the web?How can you watch ChatGPT search in the API?What does ChatGPT search for?Where do ChatGPT's search queries go?How does ChatGPT pick which pages to cite?How do you get cited by ChatGPT search?

In ChatGPT web search, the model first decides whether a prompt needs the web at all. If it does, it writes its own search queries instead of searching for the prompt directly, fetches many pages in parallel from search providers and OpenAI's own crawler, and cites only a few in the answer.

We ran this live through the OpenAI API, using the same web search tool. One question about brand tracking tools turned into 11 queries, which returned 82 distinct pages, and only 6 of them got cited. We'll use that real output to show what ChatGPT searches for and why most pages it reads never get a link.

Here's how web search works on a high level:

When does ChatGPT search the web?

ChatGPT decides on its own whether to search for each prompt. It searches the web automatically when a question would benefit from current information, and it answers everything else from what the model learned in training.

We saw the same thing in the API. We sent four prompts to gpt-5.6-sol (the newest GPT-5 model in the API) with web search turned on, and let the model decide when to search:

PromptSearched?Search callsQueries sent
What is the boiling point of water at sea level?No00
Write a haiku about autumn.No00
What did OpenAI announce this week?Yes13
Best CRM for a 10-person startup in 2026?Yes28

The model answered the two timeless prompts from memory. It searched the web for the news question and the buying question, and the CRM question alone turned into eight queries.

For brands, the decision to search in the first place is what gets you into citations. If ChatGPT doesn't search, it can't cite your page, and the answer only reflects what the model learned in training.

You can also force a search in the ChatGPT app by picking Search from the tools menu, or by typing / in the message box and selecting Search. Web search works on every plan (including Free), and even for people who aren't signed in.

How can you watch ChatGPT search in the API?

The Responses API runs the same kind of search loop as ChatGPT, but it shows you every step. With the web_search tool turned on and the sources list included, the response shows each query the model sent, every URL it got back, and the few URLs it cited in the answer.

Here's the script we ran. It asks one buying question and prints what happened:

from openai import OpenAI

client = OpenAI(max_retries=0)
response = client.responses.create(
    model="gpt-6-sol",
    input="What are the best tools to track how often ChatGPT mentions my brand?",
    tools=[{"type": "web_search"}],
    tool_choice="auto",
    include=["web_search_call.action.sources"],
    reasoning={"effort": "low"},
    max_output_tokens=3000,
)
source_urls, cited_urls = set(), set()
calls = 0
for item in response.output:
    if item.type == "web_search_call":
        calls += 1
        action = item.action
        print(f"Action: {action.type}")
        for field in ("query", "queries"):
            value = getattr(action, field, None)
            if value is not None:
                print(f"{field}: {value}")
        sources = getattr(action, "sources", None) or []
        print(f"Sources: {len(sources)}")
        for source in sources:
            print(source.url)
            source_urls.add(source.url)
print("Answer:")
print(response.output_text[:600])
for item in response.output:
    if item.type == "message":
        for part in item.content:
            for citation in getattr(part, "annotations", []):
                if citation.type == "url_citation":
                    print(f"Citation: {citation.title} | {citation.url}")
                    cited_urls.add(citation.url)
print(f"Usage: {response.usage}")
print(f"Summary: search calls={calls}, distinct sources={len(source_urls)}, "
      f"distinct citations={len(cited_urls)}, overlap={len(cited_urls & source_urls)}")

It printed this (we took some urls out to keep this shorter):

Action: search
queries: ['best AI visibility monitoring tools track brand mentions ChatGPT 2026', 'ChatGPT brand mention tracking tools AI search visibility official', 'GEO brand monitoring ChatGPT Perplexity Gemini tools']
Sources: 46
https://mentionsflow.com/10-best-ai-brand-visibility-monitoring-tools-compared-2026/
https://launchit.fast/blog/best-chatgpt-brand-mention-tracking-tools-2026
https://beamtrace.com/blog/best-chatgpt-visibility-tracker
[... 43 more]
Action: search
queries: ['site:otterly.ai AI search monitoring ChatGPT brand mentions pricing', 'site:peec.ai AI search analytics brand visibility ChatGPT pricing', 'site:profound.ai answer engine insights ChatGPT brand visibility', 'site:semrush.com AI Visibility Toolkit ChatGPT brand mentions']
Sources: 16
[... 16 URLs, 12 of them on semrush.com]
Action: search
queries: ['site:ahrefs.com brand radar AI visibility ChatGPT official', 'site:profound.ai platform answer engine insights official', 'site:peec.ai platform AI search analytics official', 'site:otterly.ai pricing AI search monitoring official']
Sources: 21
[... 21 URLs]
Citation: How often does OtterlyAI check AI search engines? | https://help.otterly.ai/monitoring-interval?utm_source=openai
Citation: Free AI Visibility Tool: Check Brand Visibility in AI Search | https://www.semrush.com/free-tools/ai-search-visibility-checker/?utm_source=openai
Citation: What is Brand Radar, and how to use it? | Help Center - Ahrefs | https://help.ahrefs.com/en/articles/11064852-what-is-brand-radar-and-how-to-use-it?utm_source=openai
Citation: The 8 Best LLM Monitoring Tools for Brand Visibility in 2026 | https://www.semrush.com/blog/llm-monitoring-tools/?utm_source=openai
Citation: The 8 Best LLM Monitoring Tools for Brand Visibility in 2026 | https://www.semrush.com/blog/llm-monitoring-tools/?utm_source=openai
Citation: Best ChatGPT Brand Mention Tracking Tools (2026 Compared) | https://launchit.fast/blog/best-chatgpt-brand-mention-tracking-tools-2026?utm_source=openai
Citation: Ahrefs FAQ | Frequently asked questions | https://ahrefs.com/faq?utm_source=openai
Usage: ResponseUsage(input_tokens=29280, ..., output_tokens=1125, ...)
Summary: search calls=3, distinct sources=82, distinct citations=6, overlap=6

The response has search calls, sources, and citations:

  • Search calls. Each one has the queries the model sent. A reasoning model like this one can search, read the results, and search again, so one answer can include several calls. Reasoning models can also open a page or search inside a page, but our runs only used plain searches.
  • Sources. This is the full list of URLs the model looked at. It's usually much longer than the citation list.
  • Citations. These are attached to the answer text. Each one has a URL, a title, and the start and end position of the text it supports. Every cited link came back tagged with ?utm_source=openai, so clicks from these links show up in your analytics with openai as the source.

The API bills search per call. Web search costs $10 per 1,000 calls, plus the search content tokens at the model's normal rates. Our run made 3 calls, so the search cost was 3 / 1,000 × $10 = $0.03, plus 29,280 input tokens (mostly page content).

Here's the web search pricing in detail (September 2026):

The Vercel AI SDK exposes the same tool natively, so you can read each search from a TypeScript app too. You add openai.tools.webSearch() to the tools, and every tool result carries the search action with the full list of queries the model sent:

import { generateText } from 'ai';
import { openai } from '@ai-sdk/openai'


const result = await generateText({
  model: openai('gpt-6-sol'),
  prompt: 'What are the best tools to track how often ChatGPT mentions my brand?',
  tools: { web_search: openai.tools.webSearch() },
  toolChoice: 'auto',
  maxRetries: 0,
  maxOutputTokens: 3000,
  providerOptions: { openai: { reasoningEffort: 'low' } },
})

for (const call of result.toolCalls) {
  console.log('Tool call:', call.toolName, 'input:', JSON.stringify(call.input));
}
for (const tool of result.toolResults) {
  if (!tool.dynamic && tool.toolName === 'web_search') {
    console.log('Tool result:', tool.toolName, 'action:', JSON.stringify(tool.output.action));
  }
}
console.log('result.sources count:', result.sources.length);
for (const source of result.sources.filter(s => s.sourceType === 'url').slice(0, 3)) {
  console.log(source.url);
}

It printed three tool results, and this was one of them:

Tool result: web_search action: {"type":"search","query":"site:semrush.com AI Visibility Toolkit official ChatGPT mentions","queries":["site:semrush.com AI Visibility Toolkit official ChatGPT mentions","site:ahrefs.com brand radar official AI ChatGPT mentions","site:scrunchai.com AI visibility official ChatGPT brand monitoring","site:profound.com answer engine insights official"]}

The queries sit in each tool result's output.action.queries, while the tool call input stays empty. result.sources only lists the cited pages. The full retrieved list is on each tool result's output.sources.

The API is close to ChatGPT, but not the same. For example, the ChatGPT app changes your search based on where you are and what it remembers about you, which the API doesn't do by default:

  • It adds your rough location, based on your IP address, to the query, so "restaurants near me" becomes "top restaurants San Francisco".
  • With memory on, it can add saved facts about you, such as a vegan diet, to the rewritten query.
  • It picks its own model and mode, and it also sends some queries to partners like Shopify for shopping results.

In the API, we set the location ourselves and pick the model and reasoning effort. But the fan-out, sources and citations work the same way.

What does ChatGPT search for?

Instead of searching for your exact prompt, ChatGPT rewrites the prompt into one or more targeted queries. It reads the results, and then often sends narrower follow-up queries. This is called query fan-out.

In OpenAI's own example, a researcher asks about cancer drugs that target CCR8. ChatGPT first searches for "CCR8 immunotherapy drug development 2025". Based on those results, it then searches for "CHS-114 conference 2025", even though the researcher never typed that drug name.

When we asked our demo question, the model sent 11 queries over three rounds, and they included details the user never typed:

  • A year. The first query ended in "2026", even though the prompt had no date.
  • Brand names. Five AI visibility and SEO tools showed up in queries: Otterly, Peec, Profound, Semrush and Ahrefs. The repeat runs added two more tools in the same space, Promptwatch and Scrunch.
  • site: lookups. 8 of the 11 queries were limited to one vendor's own domain, often with the word "official". The model first picked a shortlist, then checked each brand's own site.

Here is how the 11 searches played out:

The fan-out also changes every run. When we ran the same prompt three more times, we got 11, 12 and 11 queries, with different brands each time. The other test prompts worked the same way. All eight CRM queries were site: lookups on HubSpot, Attio, Pipedrive and others, each with "2026" added.

Larger studies found the same fan-out across thousands of prompts:

  • In 15,000 prompts, 89.6% triggered two or more follow-up searches, and 95% of the fan-out queries had zero traditional search volume.
  • In a smaller network-traffic study, 21 of 27 first queries named brands the user never mentioned. "Best AI note taking app" became a query that already listed Granola, Notion AI, Otter and four more.
  • Fan-out has grown fast. After the ChatGPT 5.6 update, the average number of fan-out queries per prompt went from 2.17 to 7.61, and site: showed up in 64% of queries.

Here is the fan-out before and after ChatGPT 5.6:

ChatGPT often decides which brands to consider before it even searches, and then checks those brands' own pages. So a keyword list built from what people type into Google misses most of these queries.

Where do ChatGPT's search queries go?

ChatGPT sends its queries to a mix of outside search providers and OpenAI's own data. OpenAI names Microsoft and Shopify as search partners, and it also uses content that news partners provide directly, like the Associated Press, Reuters and the Financial Times.

Early on, Bing was the main source. When researchers checked 500+ citations in February 2025, 87% matched Bing's top results for the same question, and only 56% matched Google. But that study is old now, and a lot has changed since.

Newer research suggests ChatGPT also uses Google results and its own index:

  • Google results. Independent tests suggest ChatGPT also pulls from Google's index, possibly through third-party scrapers.
  • OpenAI's own index. Researchers watching ChatGPT's network traffic found search results labeled with their source between May and July 2026. Three labels were services that scrape Google. The fourth, Labrador, looks like OpenAI's own index, and only about 1.5% of its results showed up in Bing's top 20 for the same searches.

OpenAI hasn't confirmed Labrador. The claim comes from Peec AI, RESONEO and others, and this post sums it up:

No matter which index answers the query, one OpenAI crawler decides whether your site can show up at all. OpenAI runs three bots, and you can allow or block each one separately in robots.txt:

BotWhat it doesIf you block it
OAI-SearchBotCrawls pages for ChatGPT searchYour pages stop showing up in ChatGPT search answers, except as plain navigation links
GPTBotCrawls pages for model trainingOpenAI stops using your content for training. Search is not affected
ChatGPT-UserFetches a page live when a user's chat asks for itrobots.txt rules may not apply, because a person started the request. It has no effect on search eligibility

If you block OAI-SearchBot to stay out of training, your site also drops out of ChatGPT search. ChatGPT can only show your site in search if OAI-SearchBot is allowed and OpenAI's published crawler IP addresses get through your host or CDN (content delivery network). Changes to robots.txt take about 24 hours to apply.

How does ChatGPT pick which pages to cite?

ChatGPT reads many more pages than it cites. In our demo, the model retrieved 82 distinct URLs and cited 6 of them, so only 6 / 82 = 7% of what it read made it into the answer.

The same gap shows up at scale. Across 548,534 retrieved pages, only 15% were ever cited, and the other 85% were read and dropped. A smaller study of 3,554 retrieved pages found just 110 cited, about 3.1%.

Compared to the pages ChatGPT only read, the cited pages had these in common:

  • The brand was named in the fan-out. When ChatGPT named a brand in its own queries, it cited that brand 68.9% of the time. Pages it only fetched got cited 2.1% of the time.
  • They rank well in Google. 55.8% of cited pages ranked in Google's top 20, and position 1 pages were cited 3.5 times more often than pages outside the top 20. But 32.9% of cited pages only showed up in the fan-out results, not in results for the original prompt.
  • They answer the exact sub-question. In our run, two cited pages were help-center articles: one on how often Otterly checks AI engines, one explaining Ahrefs Brand Radar. The other four were a free tool page, two "best tools" roundups (one cited twice), and a FAQ.

Here's the citation rate for brands named in the fan-out versus pages that were only fetched:

The type of question also moves the odds. Product discovery queries had an 18.3% citation rate, how-to queries 16.9%, and validation searches 11.3%.

At the domain level, community and reference sites take a large share. Reddit got 16.8% of ChatGPT citations in US queries in September 2026, and Wikipedia got 7.0%. These shares swing a lot: Reddit went from close to 60% of responses to about 10% within a few weeks in 2025.

Here are the most-cited domains in September 2026:

How do you get cited by ChatGPT search?

To get cited, you need to get through every step. ChatGPT has to be able to crawl your site, name you in its queries, find you with its site: lookups, and then pick your page over the other pages it read.

  • OAI-SearchBot access. If OAI-SearchBot can't crawl a page, ChatGPT can't cite it. GPTBot is a separate bot, so blocking it only keeps your content out of training.
  • Getting named on the pages ChatGPT reads first. Brands in the fan-out get cited 68.9% of the time versus 2.1% for pages that were only fetched. In our run, the first round of searches came back full of "best tools" roundups and Reddit threads. Being listed on those pages is how a brand gets into the next round of queries.
  • Answering the site: lookups on your own domain. Once ChatGPT has a shortlist, it searches each brand's site for pricing, features and "official" pages. It helps to have those facts on your own domain, in pages a crawler can read. In our run, help-center pages that answered one exact question got cited.
  • Making your official domain obvious. ChatGPT builds site: queries from what it thinks your domain is. In one test it kept searching site:census.com for the startup Census, whose real site is getcensus.com. Census.com is a parked domain.

Ranking in Google helps, but Google Search Console doesn't show the fan-out queries ChatGPT ran, or which pages it cited instead of yours. One SEO described the problem in this post:

The fan-out also changes on every run. Our same prompt produced 11, 12 and 11 queries with a different set of brands each time, so one manual check in the ChatGPT app shows one random sample. To see the pattern, you need the same prompts run on a schedule, with the queries and citations saved each time.

That is what ByDefault does. It runs each tracked prompt in the model's native environment and records who the model cites and the exact searches it ran to decide. You can then see which fan-out queries name your competitors but not you, and which of your pages get fetched but never cited.

← All articlesFollow via RSS ↗

Keep reading

AI SearchSeptember 14, 2026

The Best AI Visibility Tools for Tracking Brand Mentions in ChatGPT, Perplexity, and Claude (2026)

Compare eight AI visibility tools by model coverage, pricing, and how they track brand mentions and citations in ChatGPT, Perplexity, and Claude.

Josh9 min read

Turn AI search into your most profitable sales channel.

Check AI visibilityBook a call
ByDefault

ByDefault tracks who ChatGPT and Claude recommend in your category, and shows you exactly what to publish to change the answer.

Quick Links

  • Pricing
  • Blog
  • FAQ
  • Sign in

Company

  • Twitter / X
  • Privacy Policy
  • Terms of Service
  • Imprint