Field note · Being found
Your firm may be invisible to ChatGPT. Here is how to check in ten minutes.
People have started asking assistants the questions they used to type into Google, including which adviser to use. Whether you appear in that answer is decided by things most firms have never checked.
The short answer
Run four checks: can a crawler fetch your pages without JavaScript, does each service have its own URL, are your key facts in text rather than images, and does any page answer a real client question in its first 60 words. Crawlability is the strongest single citation factor — ahead of every content signal — and it is the one most firm sites fail.
A client used to search "mortgage broker Edinburgh" and see ten links. Increasingly they ask an assistant, and get a paragraph naming two or three firms.
You do not rank in that paragraph. You are either cited or you are not, and the mechanics are different enough from classic SEO to be worth ten minutes.
Check 1 — does your site work without JavaScript?
Open your service page, disable JavaScript, reload. If the page is blank, a crawler that does not execute scripts sees the same blank.
Plenty of modern adviser sites — anything built as a single-page app, or a builder relying on client-side rendering — fail here. It is the cheapest and most consequential thing on this list.
Check 2 — does each service have its own URL?
If your mortgage, protection and pension pages are tabs on one page, or sections reached by a #fragment, they are one document to a machine.
An assistant answering "who does equity release in Edinburgh" is looking for a page about equity release. It cannot cite a tab.
Check 3 — are your facts text, or pictures of text?
Fee structures in a PDF. Credentials in a designed image. The team's qualifications inside a carousel. All invisible.
Everything you would want quoted needs to exist as selectable text on a page that can be fetched.
Check 4 — do you answer the question in the first 60 words?
Assistants extract answers, not brand impressions. A page that opens with "Welcome to our practice, established in 2004" has not answered anything by word sixty.
The structural advice from the citation research is unglamorous and specific: put the answer at the top of each section, keep paragraphs to two or three sentences, use question-shaped headings, and name the source of every statistic.
| Check | How | If it fails |
|---|---|---|
| Fetchable without JS | Disable JavaScript, reload | Highest priority. Talk to whoever built it |
| One URL per service | Click each service, watch the address bar | Split the tabs into real pages |
| Facts as text | Try to select the sentence | Move it out of the image or PDF |
| Answer in 60 words | Read the first paragraph aloud | Rewrite the opening, not the whole page |
| Freshness | Find the last-updated date | Cited pages skew ~25% fresher than organic top-10 |
What about schema markup?
Worth doing, and second-order. FAQPage carries the most weight for question-and-answer content, then Article and Organization. But schema on a page a crawler cannot fetch is decoration on a locked door.
Do the four checks first. They are free, they take ten minutes, and most firms fail at least one.
One caveat worth stating
These factor weightings come from third-party analysis of citation patterns, not from OpenAI, Google or Perplexity disclosing how their systems choose sources. Nobody outside those companies knows for certain.
The four checks above are still worth doing, because they are also just — a website that works.
Questions people actually ask
How do I know if my website appears in ChatGPT answers?
Start with whether it can be read at all: check the page renders without JavaScript, that each service has its own URL, that key facts are selectable text rather than images or PDFs, and that the opening 60 words answer a real question. URL accessibility is the strongest single citation factor.
What matters most for being cited by AI search?
In Cyrus Shepard’s May 2026 synthesis of 54 studies, URL accessibility scored 9.5/10 — ahead of search rank (9.4), fan-out rank (9.3), preview controls (9.2) and query-answer match (9.2). Being fetchable outranks every content signal. The scores are correlations, weighted by how often independent studies agree.
Does schema markup help with AI citations?
It helps, but it is second-order. FAQPage carries the most weight for question-and-answer content, followed by Article and Organization. Schema on a page a crawler cannot fetch achieves nothing.
Sources
- Cyrus Shepard, AI Citation Ranking Factors, Zyppy Signal, 7 May 2026 — 23 factors scored from 54 experiments, patents and case studies, weighted by repeatability, strength of evidence and platform documentation: URL accessibility 9.5/10, search rank 9.4, fan-out rank 9.3, preview controls 9.2, query-answer match 9.2. The author states the scores reflect correlation, not causation. source ↗ — third-party synthesis of 54 sources, method stated; not a platform disclosure re-checked every 6 months
- Ryan Law, Do AI assistants prefer to cite fresh content?, Ahrefs, 28 July 2025 — 16.975 million cited URLs across ChatGPT, Perplexity, Gemini, Copilot, AI Overviews and organic Google: cited content averaged 1,064 days since publication against 1,432 for the organic top-10, 25.7% fresher. source ↗ — vendor dataset, very large, method stated re-checked every 6 months
- BuzzStream, 50 AI statistics digital PR and SEO should know, updated 30 June 2026 — across 4 million citations tracked with Xofu, blog and content pages made up 53.46% of citations; syndicated press releases 0.04%. source ↗ — vendor dataset (a digital-PR firm), large; sells the conclusion re-checked every 6 months
Checked 11 September 2026. Next scheduled check 10 March 2027. Numbers that move — leaderboards, live indices — are re-checked every 30 days; annual datasets and rules in force every six months; dated research once a year. If something here has gone stale before we got to it, tell us and we will correct it and say what changed.
Want the four checks run across your whole site?
We fetch every page the way a crawler does, check server-side rendering and bot rules, and return a per-check result — including “we cannot tell”, where we cannot.
Talk to us about this