Search changed and most businesses have not noticed. A growing share of buying research now happens inside ChatGPT, Perplexity, Claude, and Google's AI Overviews — not as ten blue links, but as a single synthesized answer with two or three sources cited underneath it. If your website is not one of those sources, you are not losing a ranking. You are not in the conversation at all.
That is a different problem from traditional SEO, and it needs a different kind of check. A page can rank reasonably well in classic search and still be functionally invisible to an AI engine — because AI models do not "rank" pages the way Google does. They read a page, decide whether it contains a clear, citable answer, and either use it or skip it in a fraction of a second. Most business websites were never built with that test in mind.
TL;DR: AI search engines cite pages that state facts plainly, structure content clearly, and carry real trust signals — most business websites do none of these well. You can check exactly where your site fails this test, free, in about 30 seconds, with Revora. Once you know what is broken, get an instant estimate for what fixing it would actually cost — no call needed for either.
What "AI search readiness" actually means
AI search readiness — sometimes called AEO (answer engine optimization) or GEO (generative engine optimization) — is the measure of whether an AI model can read your website, understand what you do, and confidently cite you as a source in an answer it generates for someone else.
This is not the same test as traditional SEO, even though the two overlap. Classic SEO asks: does this page rank for a keyword? AI readiness asks a narrower, harder question: if a model is assembling an answer right now, is there a clean, verifiable fact on this page it can lift and attribute? Vague marketing copy, buried specifics, and pages that require five clicks to find a real answer all fail this test — even if they rank fine in Google.
The businesses that show up in AI answers today are not always the biggest brands. They are frequently the ones whose content states things plainly: clear pricing, specific process steps, direct answers to common questions, and structured data that tells a crawler exactly what it is looking at.
Why this matters more than it did two years ago
For most of the web's history, being invisible to a new discovery channel was a slow-burn problem — you had time to catch up. AI search is moving faster, for one structural reason: the models are trained and updated on a rolling basis, and their citation behaviour rewards content that is already well-structured now. Waiting to fix this later means waiting through cycles where competitors who are already readable get cited, get clicked, and get remembered by users who never open a second tab.
There is also a compounding effect specific to service businesses and SaaS products. A buyer researching "best way to automate invoicing" or "how much does a custom app cost" is increasingly asking an AI assistant directly instead of scrolling search results. If your site cannot answer that plainly, the assistant answers with a competitor's numbers instead of yours — and the buyer never sees your homepage at all.
The six things that determine whether AI engines can read and trust you
Most sites fail this test for a small, repeatable set of reasons. Six dimensions consistently separate pages that get cited from pages that get skipped.
Clarity. Can a reader — or a model — understand what you do and who it is for within the first few seconds on the page? Clever taglines and abstract positioning score badly here. Plain statements score well.
Conversion structure. Is there an obvious next step, or does the page just describe things and stop? Pages with a clear path (a form, a tool, a defined offer) give both users and AI systems a reason to treat the page as a destination, not a dead end.
Trust signals. Real names, real results, verifiable claims, and a coherent "who is actually behind this" story. Generic stock-photo trust badges do very little; specific, checkable detail does a lot.
Search foundations. The basics still matter — correct titles, meta descriptions, header structure, internal linking, and a sitemap that actually reflects what is live. A surprising number of otherwise good sites fail here simply because pages were added faster than the technical foundation was maintained.
AI-search readiness specifically. This is the newer layer: structured data (schema markup) that tells a crawler exactly what a page is, direct question-and-answer content that mirrors how people actually ask AI assistants things, and crawlability for AI-specific bots — many sites unknowingly restrict or slow down the crawlers that matter most for this channel.
Speed. A slow page gets abandoned by human visitors and deprioritised by automated crawlers alike. Speed is one of the oldest ranking signals and one of the least often actually fixed.
How to check your own site's AI visibility, free
We built Revora because we could not find a tool that scored a website against this newer, revenue-shaped version of the question. Most audit tools still dump a long list of technical warnings ranked by severity to an engineer, not by what it actually costs the business.
Revora works differently. Paste a URL, and in about 30 seconds it scores the site across all six dimensions above — clarity, conversion, trust, SEO, AI-search readiness, and speed — and returns a single Revenue Aura Score. More importantly, it ranks every issue it finds not by technical severity but by estimated revenue impact, and explains each one in plain language: what is leaking, what it is probably costing you, and the exact fix.
You can also run a head-to-head comparison against a competitor's site, and turn on Autopilot to get a free weekly re-scan with an email only when your score actually moves. There is no signup required to see your top leaks, and no credit card anywhere in the flow. If you want to see where your own site stands before reading further, that link takes about as long as this paragraph did.
Once you know what's broken, the next question is cost
A revenue audit is only useful if it leads somewhere. Knowing that your site is missing structured data, has a confusing pricing page, or loads slowly on mobile is progress — but most business owners immediately hit the same wall next: what would it actually cost to fix, and is it worth doing now or later?
That is the exact gap Estimate Your Project is built to close. Instead of booking a call to get a vague range three days later, you answer a short set of questions about what you are building — a website rebuild, a web app, a mobile app, or an automation project — and get an instant cost estimate, an ROI calculator, and a recommended tech stack. No call needed, and it takes a similar 30 seconds to Revora's scan.
Used together, the two tools answer the two questions that actually block most decisions: is something broken, and what would fixing it cost? Most agencies make you talk to a salesperson to get either answer. We built these so you do not have to.
The vision behind why we built them this way
Most "free website audit" tools exist purely as a lead-generation trap — a rough score, then a hard wall demanding your phone number before you see anything useful. We built Revora the opposite way on purpose: the top leaks are visible with no signup, because a tool that hides its own value behind a form is not actually trying to help you, it is trying to capture you.
The same logic shaped Estimate Your Project. Pricing opacity is one of the most common complaints about software and web development — vague ranges, "it depends" answers, and quotes that arrive only after a 45-minute discovery call. An instant, honest starting estimate respects that your time has value before you have committed to anything.
Underneath both tools is the same belief: businesses make better decisions when they have real information immediately, not a sales process disguised as one. If either tool tells you what you need to know and you never talk to us again, that is a completely fine outcome. If it turns into a project, you'll have a much clearer idea of what you actually need built and roughly what it should cost before that conversation even starts.
What to do with your results
If Revora surfaces a low AI-search-readiness score, the fix is rarely a full rebuild. It is usually a specific, scoped set of changes: adding schema markup, restructuring a few key pages around direct questions and answers, tightening titles and meta descriptions, and fixing crawlability issues that are blocking AI bots without you realising it. Scoping that work properly before quoting it is what keeps a fix like this fast and contained instead of turning into an open-ended project.
If the leaks are more structural — a confusing user journey, a slow backend, or a product experience that cannot support where the business is headed — that is a bigger conversation, and understanding what custom software actually costs is a useful next read before you scope anything.
Either way, the starting point is the same: find out where you actually stand before deciding what to do about it.
Check your own site's revenue leaks: run the free scan at mymindstudio.ai/revora — or if you already know what needs fixing, get an instant number at mymindstudio.ai/estimateyourproject.
Which AI crawlers to allow, and which to block
"AI bots" are not one thing. Some crawlers exist to retrieve and cite your pages inside an assistant's answers; others exist to absorb your content into model training with no link back; a third group fetches a page in real time because a human asked about you. Blocking them as a bundle is the most common way a site makes itself invisible on purpose. Read the table by band, decide per band rather than per bot, then check both your robots.txt and your CDN — the two can disagree. Every quoted claim below comes from the crawler operator's own published documentation, checked in August 2026; re-check before you change anything, because these pages and defaults are revised often.
| Crawler (exact user-agent string) | What it actually does | Honors robots.txt? (per the operator's own docs) | What blocking it costs you | What allowing it costs you |
|---|---|---|---|---|
| Retrieval / citation crawlers — block these and you are invisible by definition | ||||
OAI-SearchBot |
Surfaces websites in ChatGPT's search features. | Yes | OpenAI: sites opted out "will not be shown in ChatGPT search answers, though can still appear as navigational links." | Little — OpenAI documents this as a separate decision from GPTBot training. |
PerplexityBot |
Perplexity: "designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models." | Yes | You are not surfaced or linked in Perplexity answers. | Little — Perplexity's docs explicitly exclude this crawler from foundation-model training. |
Claude-SearchBot |
Anthropic: "navigates the web to improve search result quality" — the indexing layer behind Claude's search. | Yes | Claude's search layer cannot index you. | Content read for search indexing only, per Anthropic's bot documentation. |
Googlebot |
Google: "AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control" — so Googlebot is the AI Overviews and AI Mode control. | Yes | In robots.txt, classic Google Search and AI Overviews go together — there is no AI-only user-agent to disallow. | Your content can be summarized in an AI Overview with no click. Use the controls, not a block: Search Console's Search generative AI control excludes a site from AI Overviews, AI Mode and Discover generative features and "isn't used as a ranking or inclusion signal affecting other parts of Search" — but Google says it is still rolling out "to a subset of website owners", so confirm your property has it. Per-page: nosnippet, max-snippet, data-nosnippet, noindex. |
bingbot |
Bing's index also grounds Microsoft Copilot; Microsoft's own crawler documentation lists no separate AI-only user-agent. | Yes | Bing search and Copilot at the same time. | Copilot may use full page content unless you add NOARCHIVE (excluded from chat answers and not linked) or NOCACHE (URL, title and snippet only) — both still appear in Bing search results. |
| Training crawlers — blocking these costs you zero citations | ||||
GPTBot |
OpenAI: crawls "content that may be used in training our generative AI foundation models." | Yes | Nothing in ChatGPT search — OpenAI treats GPTBot and OAI-SearchBot as separate access decisions. | Your content may enter foundation-model training with no citation, link or referral back. |
ClaudeBot |
Anthropic: "collecting web content that could potentially contribute to their training." | Yes | No effect on Claude-SearchBot or Claude-User. | Same trade: training input, no attribution. |
Google-Extended |
Google: a "standalone product token" controlling whether crawled content "may be used for training future generations of Gemini models" — covering Gemini Apps, the Vertex AI API and grounding in both. | Yes | Nothing in Search — Google states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." | Content usable for Gemini training and grounding. |
CCBot (full agent CCBot/2.0) |
Common Crawl's open corpus, widely reused as third-party training data. | Yes, plus Common Crawl's separate opt-out registry | Nothing search-facing. | Your pages sit in a public, redistributable corpus you cannot recall later. |
| Agent fetchers — a human asked, so robots.txt is weak or ignored here | ||||
ChatGPT-User |
Visits a page when a ChatGPT or Custom GPT user's question needs it. | No — OpenAI: "Because these actions are initiated by a user, robots.txt rules may not apply." | Little, because the directive may not be honored in the first place. | Real-time fetches on user demand — usually the traffic you actually want. |
Perplexity-User |
Fetches a page when a Perplexity user asks. | No — Perplexity: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules." | Nothing reliable — robots.txt is not the enforcement layer here. | Cloudflare de-listed Perplexity as a verified bot in August 2025 over undeclared stealth crawling, so network-level rules, not robots.txt, are the only real control. |
Claude-User |
Anthropic: "When individuals ask questions to Claude, it may access websites using a Claude-User agent." | Yes — Anthropic states its bots honor robots.txt, unlike the other two agent fetchers. | Claude cannot open your page when a user asks about you by name. | Real-time fetches on user demand. |
| Not a bot, but check it first: your CDN / WAF (e.g. Cloudflare bot controls) | Where blocking is actually enforced — and where it can happen without you touching robots.txt. | n/a | From 15 September 2026, new Cloudflare domains get Training and Agent blocked by default on ad-serving pages while Search stays allowed — and Cloudflare warns that multi-purpose crawlers (Googlebot, Applebot, BingBot) "will be blocked by customers who have selected to block Training." Owners can opt out of the new defaults in Security settings. | Verify here before concluding your robots.txt is the problem. |
Who each choice is wrong for: blanket-blocking every "AI" user-agent is wrong for anyone who wants to be cited — it takes out OAI-SearchBot, PerplexityBot and Claude-SearchBot along with the training crawlers, and the retrieval loss is immediate while the training "win" is invisible. Blanket-allowing is wrong for publishers whose archive is the product, since Common Crawl in particular is redistributable and cannot be recalled. Disallowing Googlebot or bingbot to escape AI answers is wrong for almost everyone — you lose the underlying search index with it, and Search Console's generative AI control (where you have it) plus the per-page snippet directives are the proportionate tools. And treating robots.txt as the whole answer is wrong for any site behind a CDN with bot controls: check the edge configuration first, because that is where a block usually lives.
Frequently Asked Questions
What is AI search readiness?
AI search readiness is how well a website can be read, understood, and cited by AI-powered answer engines like ChatGPT, Perplexity, Claude, and Google's AI Overviews. It depends on factors like structured data (schema markup), clear and directly stated facts, question-and-answer style content, and whether AI crawlers can access the site at all. A site can rank well in traditional Google search and still be effectively invisible to AI answer engines if these factors are missing.
How is AI search readiness different from traditional SEO?
Traditional SEO optimises for ranking position in a list of links. AI search readiness optimises for being selected and cited as the source inside a single generated answer. The overlap is real — site speed, clear structure, and authority signals matter to both — but AI engines specifically reward content that states facts plainly and can be lifted as a direct answer, which many SEO-optimised pages, full of qualifiers and marketing language, do not do well.
How can I check if my website is visible to AI search engines?
Revora runs a free 30-second scan that scores your website across six dimensions, including AI-search readiness specifically, and returns a ranked list of the biggest issues by estimated revenue impact — not just a technical checklist. There is no signup required to see your top leaks. You can also run a head-to-head comparison against a competitor's site to see exactly where you are ahead or behind.
What are the most common reasons a website is invisible to AI search?
The most common causes are missing or incomplete structured data (schema markup), content that is too vague or marketing-heavy to extract a direct answer from, unintentionally blocking or slowing down AI-specific crawlers, and pages that bury specific facts (pricing, process, results) several clicks deep instead of stating them clearly near the top. Most of these are fixable without a full site rebuild once they are identified.
How much does it cost to fix AI search readiness issues?
It depends entirely on what is found. Adding schema markup and restructuring a handful of key pages around clearer, more direct content is typically a scoped, contained project — not a rebuild. More structural issues, like a confusing site architecture or a slow backend, cost more to address properly. The fastest way to get a real number is to run the free scan first, then use the project cost estimator to get an instant, honest range based on what actually needs fixing.