From Invisible to Recommended by AI: How I Got ChatGPT to Recommend Me in 7 Weeks

Screen-print style illustration of a speaker in profile facing a crowd, one figure standing in an amber spotlight, symbolizing being the one an AI recommends
Key takeaways: I rebuilt my own site until ChatGPT started recommending me, and ChatGPT's mention rate on my query set went from 3.7% to 17.4% (40 of 230 answers, seven weeks after the first measurement). Then I tested whether my method actually explains that, on 52,821 real AI citations.
  • Do make the page the best-matching answer to the exact question. Semantic match between query and passage was worth 3.27x the citation odds per standard deviation, more than every authority, formatting and freshness variable in the model combined.
  • Do aim at how people actually ask, not at the tidy questions a language model invents. Real queries run a median of 11 words, nearly half start lowercase, and 12% are bare keyword fragments a model never produces. Digital Twins calibrated against real behaviour supply both the what and the how.
  • Do put the answer in the first 60 words. An opening that answers the query in the asker's own terms doubles the citation odds per standard deviation (2.02x), the strongest formatting lever in the model. Keep FAQPage schema as cheap hygiene: the markup itself showed no reliable effect once answer-first content was measured directly. Don't run a backlink campaign for AI visibility: both link variables were statistically null.
  • Do run this as a weekly loop rather than a one off project. The same agentic framework of skills runs every week, keeps what it learned about your audience, and keeps you findable on ChatGPT for the queries people actually type.

In spring 2026 my personal site was getting roughly 50 visitors a month. Nothing was broken. It was simply invisible, both to Google and to the assistants people now ask instead of Google. A typical AI answer names about 2.8 brands, and I was not one of them for a single query that mattered to my business.

Since then I have run a weekly optimisation loop on my own site, and in parallel a pre-registered quasi-experiment on 52,821 AI citations to check whether my tactics are the ones that actually move citations. This piece is both halves: what I did, and what the data says about why it worked.

The one thing that decides whether you get cited

Semantic match between the question and a passage on your page. In the pre-registered model on 52,821 citations it is the strongest variable by a distance, worth 3.27 times the citation odds per standard deviation (95% interval 2.68 to 4.01), more than every authority, formatting and freshness variable in the model put together. Everything else in this article is a detail next to it.

Three step screen print: a Digital Twin query and content merge in phase into a golden record, an LLM query robot and mismatched content stamp a cracked record, and a jukebox presents the golden record to listeners while the cracked one lands in the bin
Content in phase with how real people ask gets taken in and loved. Generated noise at the wrong frequency is just as loud and goes straight to the pile.

Without the jargon: the assistant cuts your page into passages of roughly fifty words, turns the question and every passage into a numeric representation of meaning, and cites the one whose meaning sits closest to the question. Matching the meaning is the job, and using the asker's own words is how you get the meaning to line up.

Where the real question sits, and what lands near it far in meaning further away close in meaning the distance that decides how a real person asks* what does an ai keynote in germany cost the passage answers back in phase off frequency flat, no signal the query a language model imagines What is the average cost of hiring an artificial intelligence keynote speaker in Germany? 15 to 25 words, perfect grammar not what people type optimise for this and you optimise for the wrong spot on the map answers in the asker's own terms An AI keynote in Germany costs between four and eight thousand euros, depending on length and travel. gets cited talks around it in marketer language Our speaking services inspire audiences and unlock the potential of your organisation. invisible closer in meaning, more likely to be cited One map, one real question at its centre. Everything else is placed by how near its meaning sits.
Illustrative placement rather than measured coordinates: the rings stand for the meaning space the model works in, and each item sits at the distance its wording earns. The question broadcasts its meaning from the centre, the passage written in the asker's own terms answers back in phase, and the imagined query broadcasts on a frequency nobody uses. The practical read: you can only aim at the real centre if you know how people actually ask, which is what twin simulated queries, calibrated against real behaviour, give you. What counts as how a real person asks is measured below, in the query properties section.

Why you cannot match a query you cannot predict

Because the centre of that map is a real question, you have to know how real people phrase it, and almost nobody checks. Aim a passage at the tidy version an assistant would have invented instead and the semantic match you were chasing never happens.

I can check that against something most people cannot. neuroflash, where I am CIO, sees how people really phrase things, both what they ask and how they type it. Set the profile of real queries against what a language model invents when asked for search queries and only the length of the query comes close.

What a language model invents, against what people type real people generated by a language model Words per query: where the two come closest Real people's queries* median 11 where they overlap A model produces 15 to 25 words curve shape illustrative, range documented Note: in 2023 real queries ran shorter, a median of 7 words. They have drifted longer and sloppier since. 0 10 20 30 40 50 60 Question type: the dimension with no overlap at all Real queries* Generated by a language model 0% not measured per style Amber marks the gap that never closes: 12.3% of real queries are bare keyword fragments, and a language model produces none of them. More than 95% of generated queries carry a question mark, so the rest of that bar sits in question shaped styles. natural question 67.9% keyword fragment 12.3% brief command 10.8% full briefing 3.8% conversational 3.0% detailed command 2.2% How the query is written, real queries against generated Ends with a question mark 41.9% 95% or more Starts in lowercase 45.6% under 5% 0 25% 50% 75% 100% Length overlaps a lot. Question type does not overlap at all, and that is the dimension a language model cannot fake.
Green is what real people type, violet is what a language model produces when it is asked for search queries. The green curve is not raw data: it is interpolated from measured values of real query lengths, with the median marked where it was measured. For the violet curve the shape is illustrative and only the range is documented, 15 to 25 words. The shaded band marks where the two overlap, and length is the dimension where they do: a model can land a plausible query length by accident. Question type is where it cannot. That mix is measured on real queries; on the generated side the only documented cell is bare keyword fragments at 0%, never produced, so the rest of that bar is drawn as not measured rather than filled with a guess.
* Queries were sampled from ~20 million real human queries on neuroflash.com.

So the query set has to be simulated against real behaviour rather than imagined. That is what the Digital Twins do in the loop below: step 1 simulates the audience, step 2 has them produce search behaviour calibrated to the real distribution above, which is how you get the centre of the map instead of a guess at it. Get those two wrong and every later step optimises a page for a question nobody sends.

How the twins build the query set, in two steps

Both halves are needed. The twins supply the what, the thing a given persona would go looking for. The calibration pass supplies the how, the shape that intent takes when somebody types it into a chat window.

From a persona to a query somebody would really type
Step 1. The twins know what matters to their persona
illustrative example twins
Avatar of Alina, an example Digital Twin persona

Alina, 25

advertising and marketing, Berlin

Cares whether a speaker will land with a team in their twenties.

search intent

Find a keynote speaker whose AI material connects with a young audience.

Avatar of Nadine, an example Digital Twin persona

Nadine, 37

HR assistant, Stuttgart

Organising a leadership day, so budget and logistics decide.

search intent

Find out what an AI keynote costs and how travel and timing work for a one day event.

Avatar of Korbinian, an example Digital Twin persona

Korbinian, 63

distribution manager, Rosenheim

Wants to know whether the spend pays off at his company's size.

search intent

Work out whether an AI keynote is worth it for a mid sized company.

calibrated against real user patterns
Step 2. The humanizing pass aligns each query with real human queries, in what and how
the closer to how people actually ask, the more often people find you

ai speaker who gets gen z audience

7 words, lowercase

ai keynote cost leadership event

5 words, bare search keyword

is an ai keynote worth it for a mid sized company

11 words, lowercase

Calibration target, the real queries measured above: median 11 words, 46% starting in lowercase, a question mark on fewer than half.

The what in step 1, the how in step 2. Avatars and personas come from the Digital Twin system, grounded in real survey respondents; the intents and typed queries here are illustrative examples of the process, not measured data.

The starting point: 50 visitors a month and zero AI mentions

Before the rebuild, my site was a WordPress install that loaded 468 KB of homepage HTML, 60 scripts and 9 stylesheets, scored 70 on PageSpeed, and appeared in about 16 Google impressions per day. On my frozen query set, ChatGPT named me in 3.7% of answers. For commercial queries such as "best AI keynote speaker" I was absent entirely, while other names were recommended by default.

I applied my own method to myself first, for a simple reason: I sell this work, and I would rather not sell something I have not survived personally.

The weekly loop I run on my own site

The loop has seven steps and takes one working session a week. Simulate the audience, generate the queries they really type, measure across engines, read the gaps, write content the twins approved, deploy, then measure again seven days later.

One pass of the loop, seven steps, one working session skill in the package neuroflash connector external tool neuroflash Digital Twins: the synthetic audience neuroflash Digital Twins neuroflash mcp neuroflash Digital Twins: the synthetic audience neuroflash Digital Twins neuroflash mcp query skill audit skill gap brief GEO article writer instagram, linkedin, youtube website editor audit skill 1 Simulate theaudience 2 Generate humanqueries 3 Measure acrossengines 4 Read the gaps,branded apart 5 Write twintested content 6 Publishcontent 7 Re-measureseven days on Google Analytics 4: visitors and conversions Google Search Console: impressions and clicks Peec.ai: AI visibility tracking Rankscale: AI visibility and citations Microsoft Clarity: session recordings and friction WordPress: publishing target LinkedIn: publishing target YouTube: publishing target Instagram: publishing target X: publishing target TikTok: publishing target Google Analytics 4: visitors and conversions Google Search Console: impressions and clicks Peec.ai: AI visibility tracking Rankscale: AI visibility and citations Microsoft Clarity: session recordings and friction step 7 becomes step 1 the following week
The loop in one picture. Steps 1 and 2 build the query set, steps 3 and 4 measure, steps 5 and 6 change the site, and step 7 re-runs the same set seven days later, which is the only way to tell a real gain from engine churn. Amber tags are the packaged skills, cyan the neuroflash connector, the marks below are tools and publish targets. Hover any mark for the tool and its role.

The seven steps, and the skill behind each one

Almost none of it is done by hand. Every step runs on a named skill inside an agent framework I built and run on my own machine, and the framework remembers between weeks: the audience, the frozen query set, what has already been published and what the twins rejected all carry forward. That memory is why week ten costs less effort than week one instead of more.

  1. Simulate the audience with Digital Twins. Personas invented by a language model are close to a coin flip. Twins grounded in real survey respondents answer like the people who actually book keynotes, and the audience definition is written down once and reused every week rather than rebuilt from scratch.
  2. Generate queries that sound human. Twins produce search behaviour along the journey, calibrated against the real query distribution shown above rather than against a keyword tool. The query skill keeps that set frozen, which is what makes one week's numbers comparable with the next.
  3. Measure with an AI visibility tracking provider (Peec.ai, Rankscale and others). The query set runs through ChatGPT, Gemini, Copilot, Perplexity and Claude, returning share of voice, citation rate and position per persona and journey stage. The audit skill runs that scan the same way every time and scores it against the competitors I care about, so the baseline is taken before any content changes and stays comparable.
  4. Read the gaps, branded and unbranded kept apart. The interesting cell is not "position 7", it is "does not appear at all". Mixing branded queries into the average flatters the score and hides the gap. The same skill hands back the absent queries as next week's writing brief.
  5. Write twin tested content. The GEO article writer skill drafts answer-first against the query it is aimed at, and before publishing the twins rate title, hook and opening while a fact checking pass runs against every number and source link.
  6. Publish content. WordPress out, static HTML in. PageSpeed went from 70 to 97 and homepage HTML from 468 KB to 14 KB. The website editor skill edits, previews, deploys and verifies with a rollback path, so a gap read on Monday morning can be answered by a live page the same day.
  7. Re-measure and repeat. Weekly on the changed pages, monthly on the whole query set. The gap analysis produced a backlog of 44 review passed articles, published at a maximum of two per week, and the video repurposing and LinkedIn skills take each one into clips and posts, which the engines pick up in days rather than the month a new page needs.

What the loop changed in six months

Four numbers describe the result. PageSpeed 70 to 97. Google impressions from about 16 per day before the relaunch to about 287 per day now, with clicks following the same curve from well under one a day to roughly two. ChatGPT mention rate from 3.7% to 17.4%, which is 40 of 230 measured answers, seven weeks after we started measuring. And the business number: three keynote booking calls came in through the optimised speaker pages, two of them converted, roughly €10,000 in revenue.

From invisible to recommended, three measures on one clock What I did 19 Apr: relaunch 2 Jul: twin queries start 20 Jul: gap articles live Google impressions per day, seven day average 300 200 100 0 higher means more visible in Google 287 per day about 16 per day Google clicks per day, seven day average, own scale 4 2 0 clicks run about a hundredth of impressions, so they get their own axis 2.1 per day under 1 per day ChatGPT answers naming Jonathan Mall, measured on five dates 20% 10% 0 higher means ChatGPT names me more often the run back to January is assumed, not measured nothing was measured between the dots about 1.5%, assumed 3.7% 4.1% 8.3% 10.7% 17.4% Jan Feb Mar Apr May Jun Jul Aug One shared timeline, three separate scales. Never one axis for all three.
The two Google series are Search Console daily data for jonathanmall.com, 31 December 2025 to 15 August 2026, drawn as trailing seven day averages so weekday noise does not read as movement: about 16 impressions and well under one click per day in the month before the relaunch, about 287 impressions and 2.1 clicks per day in the week to 15 August 2026. The May dip is real, not a gap in the data. The ChatGPT mention rate is a text scan for "Jonathan Mall" over ChatGPT answer texts, so the same instrument runs end to end, but it exists only on the dates it was run: 3.7% and 4.1% on the full 1,362 query set (2 and 9 July 2026), 8.3% on a 376 query attention subset (16 July), then 10.7% (11 of 103 answers, 10 August) and 17.4% (40 of 230 answers, 16 and 17 August pooled) on the top 100 query set. Nothing is interpolated between those dates. The gray dotted run back to January is an assumption rather than data, drawn at half the first observed value (about 1.5%) with a hollow marker, because AI visibility was not measured at all before July. One site measured honestly, not a controlled sample.

One booking made it concrete. A client asked ChatGPT for the best AI keynote speaker in Hamburg, got my name, checked the site, and booked a mid four figure keynote. No ad spend, no outreach, no intermediary.

Why one booking is not evidence, and what I did about it

A single win proves nothing. AI answers reshuffle themselves week to week without any intervention, so any before and after on a handful of queries is mostly noise. To find out what really drives citations, I built a pre-registered model on 1,362 queries, 2 engines and 3 snapshots taken between 2 and 16 July 2026, covering 52,821 citations and 847 fetched pages.

How the model was built, left to right 1,362 queries in thefrozen set 2 engines 3 snapshots,2 to 16 July 2026 52,821 real citations from847 fetched pages 26 variables declared before thedata was touched AUC 0.874 logistic model insidequery risk sets what was measured how it was analysed
The design at a glance. Both the 26 variables and the comparison rule were written down before the data was touched, which is what makes a null result readable rather than a failed search. Cited pages are compared against pages cited for other queries in the same language, engine and topic, so query difficulty and engine cancel out. This is a pre-registered but correlational core: it ranks what travels with a citation, it does not prove what causes one.

Twenty-six variables were declared before the data was touched, covering relevance, authority, formatting, freshness, page structure and the content claims of the current GEO playbook. Cited pages are compared against pages cited for other queries in the same language, engine and topic, so query difficulty and engine are absorbed by design. The model discriminates at AUC 0.874, and only the week to week persistence model on 7,701 page and query pairs approaches a causal reading.

The findings that changed how I write

Beyond raw semantic match, two things move the odds up and one moves them down. An opening that answers the query and a title carrying the query's words raise the odds, dense list formatting lowers them, and the rest of the conventional playbook, backlinks, FAQ markup, question shaped headings, shows no independent effect once page content enters the model.

Factor Effect on citation odds
Query to passage semantic match3.27x per standard deviation
The first 60 words answer the query2.02x per standard deviation
Title contains the query's own words1.96x per standard deviation
FAQPage or HowTo schema presentno reliable effect (p = 0.27)
Domain or page backlinksno effect (p = 0.42 and p = 0.15)
List and table density0.89x per standard deviation
Citation odds ratio, with the 95% interval 1.0 = no effect Query to passage semantic match 3.27x (2.68 to 4.01) First 60 words answer the query 2.02x (1.72 to 2.37) Title contains the query's words 1.96x (1.77 to 2.17) FAQPage or HowTo schema present null, p = 0.27 Domain referring domains null, p = 0.42 Page level backlinks null, p = 0.15 List and table density 0.89x (0.83 to 0.96) 0.5x 1.0x 2.0x 3.0x gets cited less gets cited more Log scale. Every dot right of the reference line marks a page trait that earns citations.
The same effects as the table above, now with the uncertainty attached. Each dot is the point estimate, each bar the 95% confidence interval; an interval that clears the reference line is a real effect. FAQ schema and the two backlink variables are drawn as hollow markers on the line itself because the model could not separate any of them from no effect, so no interval belongs there. Continuous factors are standardised odds ratios per standard deviation; FAQ schema is present against absent.

Five things follow. Put the query's actual words in the title, because title overlap roughly doubles the odds even with semantic similarity controlled. Answer in the first 60 words, because an opening that matches the query is the strongest formatting lever in the model and early citations survive next week's re-retrieval. Keep the schema block as cheap hygiene rather than as a strategy: the markup showed no reliable effect of its own once answer-first content was measured directly, and question shaped headings alone never paid either. Skip the backlink campaign and rewrite for relevance instead. And stop padding: no page length target earns citations, while recently published pages do slightly better.

The section length that gets cited, and why page length is not a lever

About 145 words per section. That is where the model's predicted citation probability peaks, inside the 120 to 180 word window practitioners have claimed for years, and it is the one piece of chunking folklore this data supports. Whole page length runs the other way: there is no 500 to 1,500 word sweet spot anywhere in this corpus, and added length never earns a citation on its own.

Predicted citation probability, by length Words per section: a peak at 145 120 to 180 words 16% 15% 14% 13% 145 words 40 80 145 300 1000 average words per section, log axis Words per page: no peak at all 16% 15% 14% 13% no peak, longer pages cited less 100 300 1000 3000 words per page, log axis Both panels share one vertical scale. Higher means the model expects a citation more often.
These are partial dependence curves: the citation probability the model predicts as one length variable moves across its range, with everything else held at its average. Two cautions come with them. The linear term for section length is null (p = 0.89), so the 145 word peak comes out of the curve rather than out of a slope, and the curve is a narrow spike rather than a broad plateau. And the short page cluster on the right is heterogeneous, full of profile and directory pages, so the honest reading is "length is not a lever" rather than "publish 300 word pages".

I tested the 2026 best-practice playbook, and most of it died

The study also put the advice circulating in the AI search community under the same microscope: semantic triples, entity density, self-contained sections, covering a topic's whole subquery tree on one page. Each claim became a pre-registered variable; the content measures were scored across 818 fetched pages by 13 independent AI raters, and every headline result was attacked by an adversarial verification pass before I believed it. One claim survived.

Best-practice claim What the data said
Answer the query in the first 60 words2.02x citation odds per SD, stable under every check
Pack in semantic triples (quotable "X is Y" facts)no reliable effect, see below
Raise named-entity densityinconclusive, measurement too noisy to test
Make every section self-contained ("extractability")inconclusive, two measures of it barely correlate
Cover the topic's whole subquery tree on one pageno effect either way once relevance is controlled
FAQ schema as a citation driverno reliable effect: the markup rides on the answer-first content under it

The one survivor is the strongest formatting result in the whole project. Pages whose first 60 words semantically answer the query earn roughly twice the citation odds per standard deviation, and the estimate barely moved under every robustness check I threw at it. It is relevance delivered early rather than a separate trick, which is exactly why it works: retrieval reads from the top, and an opening that already matches the question wins the passage contest before the rest of the page is even considered.

The failures taught me as much as the survivor. The semantic-triple claim looked spectacular at first, over 40% higher citation odds with a vanishing p value, until the verification pass traced the whole thing to a single miscalibrated rater batch that counted about 2.5 times higher than the other twelve. Remove those 60 pages and the effect is gone, in the citation model and in the week-to-week persistence model. That is why every variable here is declared before the data is touched and every headline gets attacked before it gets published: the difference between a finding and an artifact is usually one robustness check somebody actually ran.

What I would do first if I started from zero today

Start with measurement, not content. Build a query set that sounds like your customers rather than like a keyword tool, measure where you are absent per engine, then fix the three pages closest to a buying decision: title rewritten with the query's words, answer moved to the top, FAQ schema added.

Expect movement in weeks, not quarters. A new page needs roughly a month to enter the engines, while a LinkedIn post can appear the next day, which argues for publishing now rather than planning a launch. Then re-measure on a fixed cadence: without a baseline you cannot separate a real gain from the churn the engines produce on their own.

One note on where this goes next. The weekly loop is not willpower, it is the package of skills behind it: the audit that measures, the article writer that drafts, the website editor that ships, the twin pretesting that decides what is good enough to publish. I teach that package on the AI visibility page, set up on your own machine rather than rented from an agency, and there is a free intro call there if you would rather ask first whether it fits your situation.

Frequently asked questions about AI visibility

What actually makes ChatGPT cite a page?

How closely a passage on the page matches the meaning of the question. In a pre-registered model on 52,821 citations, semantic match between query and passage carried 3.27 times the citation odds per standard deviation (2.68 to 4.01), more than every authority, formatting and freshness variable combined. The practical version: write the answer in the words the asker uses, in the first 60 words of the page.

How long does it take before ChatGPT recommends you?

On my own site the first measurable movement came within weeks of publishing, and the mention rate went from 3.7% to 17.4% over roughly six months of weekly iterations. A newly published page typically needs about a month to be picked up by the engines.

Do backlinks help with AI visibility?

Not independently. Domain referring domains and page level backlinks both went statistically null once page content and query relevance entered the model (p = 0.42 and p = 0.15). The study could not measure unlinked brand mentions or organic rank, so read this as "no independent effect inside the AI visible pool", not "links never matter".

Does FAQ schema really increase AI citations?

Less than it appears. Pages carrying FAQPage, QAPage or HowTo JSON-LD do get cited more often, but in the full model the markup itself showed no reliable independent effect (p = 0.27): the model attributes the citations to the answer-first content those pages tend to carry. Keep the schema, it costs nothing, but put the effort into the first 60 words.

Do semantic triples or entity density improve AI citations?

Neither survived testing. Across 818 scored pages, the apparent semantic-triple effect traced back to a rater artifact and entity density stayed statistically inconclusive. The content variable that held up was whether the page's first 60 words answer the query, at roughly double the citation odds per standard deviation.

How many queries do you need to measure reliably?

Hundreds, not a handful. AI answers reshuffle themselves substantially week to week with no intervention at all, so a three query check before and after a change measures noise. My own pulse runs on a frozen set of over 1,300 queries.

Sources and method notes

  • All modelling figures come from my own pre-registered quasi-experiment on the citation panel of an AI visibility tracking provider: 26 variables declared before analysis (20 page and domain variables plus 6 content-claim variables), 1,362 queries, 2 engines, 3 snapshots (2 to 16 July 2026), 52,821 citations, 847 fetched pages, AUC 0.874 with grouped cross validation by query.
  • Effects are reported as standardised odds ratios inside query, engine and snapshot risk sets, with multiplicity controlled at q < 0.05. The citation models are correlational. Only the persistence model (same page, same query, one week apart, n = 7,701) approaches a causal claim.
  • The six content-claim variables (semantic triples, entity density and diversity, extractability, answer-first lead match, topic-tree coverage) were scored across 818 pages by 13 independent LLM raters. Reported effects are gated by rater-batch sensitivity corrections and an adversarial verification pass; an estimate that failed those checks is reported as unreliable rather than as an effect.
  • Declared omitted variables: organic top 10 rank, unlinked brand mentions and the internal link graph. The authority nulls inherit those omissions.
  • Site figures (PageSpeed, impressions, clicks, mention rate) come from Google Search Console daily data, PageSpeed Insights and the AI visibility tracking pulse of 17 August 2026 on my own domain, which is a single site rather than a controlled sample.
  • Query realism figures come from our own analysis of real user behaviour on neuroflash, sampled from ~20 million real human queries. The comparison figures for language model output are measured on raw generator output before any rewriting pass.

One site, honestly measured, plus one pre-registered model on other people's citations as well as my own. That combination is what turned a set of tactics into a method.