At the end of my last article on this experiment I did something I’d recommend to anyone who publishes numbers: I wrote down, in public, what would prove me wrong.
I had a suspicion. One of my scores had jumped four points overnight, and my best explanation was that I’d caused the jump myself, just by measuring. So I set a test — leave both assistants alone for a week, ask again, see if the number falls back — and printed the two possible outcomes side by side, before I knew which one I’d get.
This is what came back. The suspicion was wrong. A sentence I’d published along with it was wrong too. And the same day’s data showed that the number I’d been so carefully putting an error bar on isn’t one number at all.
The short version
- The setup, for anyone new here: ten fixed questions about Webflow and GSAP animations, asked word for word to AI assistants in a fresh logged-out session, scored by hand — 2 if the answer links my site, 1 if it names me without a link, 0 otherwise. Every result lands on one page, dated: Credited.
- The prediction: if Perplexity’s score fell back from 18 toward 14 after a week of silence, my measuring had been inflating my own results. If it held at 18, the score had simply been climbing and I’d caught it mid-climb.
- It held at 18. Nine days without a single question, and pages from my knowledge base were cited in nine answers out of ten. My explanation was the thing that broke.
- The mechanism I’d printed was wrong. I had written that browsing assistants fetch your page on the spot when they answer. Measured: Perplexity cited fifteen different knowledge-base pages and my server saw it fetch one. It answers from an index it already has.
- Same day, same ten questions, ChatGPT: 1 out of 20 — and among the roughly two hundred sources it cited or consulted, my site appeared zero times. Not cited, not even looked at.
- Which is the finding I’d keep: eighteen and one, same day, same questions, are not two readings of one quantity. There is no such thing as “my AI visibility score.” There’s a score per assistant, and they don’t share a mechanism.
What I predicted, and what came back
First, the thing I got right — because it’s the only reason the rest of this article can be honest. Writing down what each outcome would mean before you see it is called pre-registering, and it isn’t a formality: once you’ve seen a number you can explain almost anything about it, and you’ll believe yourself.
My last article said the retest would happen on 6 September. It happened on 7 September — a day late, and I’d rather say so than have the dates quietly disagree. The rules were the ones I’d set: no question to either assistant about any of this in between (nine days for Perplexity, eight for ChatGPT), same ten questions, logged out, a fresh session per question, sources saved from the assistant’s own export rather than typed off the screen.
| Perplexity | Round 0 | Round 1 (28 Aug) | 19 hours later | After the quiet period |
|---|---|---|---|---|
| Score out of 20 | 14 | 14 | 18 | 18 |
| Questions with a knowledge-base page cited | 0 | 6 | 8 | 9 |
No decay. Nine days of not touching it, and the number sat exactly where it had been — and the knowledge base gained a question, taking the top source slot on question 3 from one of my own older Academy lessons.
So the theory dies. The overnight jump wasn’t me feeding the thing I was measuring: Round 1’s 14 was an early reading of a score still going up, and I’d published it as if it were a resting value. It also settles a worry I’d raised out loud — that every monthly round would inflate the next one. It doesn’t. I’ll keep the week of silence before each round anyway; it costs nothing, and it’s the only reason I can write this paragraph.
Here’s what I’d take from this. My explanation was convincing. It fit the numbers. If you’d asked me the day I published my last article, I’d have defended it. It was still wrong — and the only reason I know is that I’d written down, in advance, exactly which result would kill it. Without that, I’d have looked at today’s 18 and found a way to make it fit.
So if you track how often AI assistants cite you: before your next reading, write one sentence saying which number would prove your current explanation wrong. It costs nothing, and it’s the only thing in this article that saved me from myself.
The sentence I got wrong — and how
Here’s the sentence, from my last article, live on this site as I write this:
Browsing-mode assistants fetch pages on the spot to answer, so every citation under their answers is a fetch your own question caused.
That’s a mechanism. It sounds right — it’s how you’d imagine a browsing assistant works. And I had never watched it happen.
To check it you need one fact about AI bots: there are two kinds, and they leave different footprints in your server logs. The crawler reads your site ahead of time to build an index, on its own schedule — Perplexity’s is PerplexityBot. The on-demand fetcher goes and reads a page while someone’s question is being answered — for ChatGPT that’s ChatGPT-User. If my sentence were true, fetchers would show up in my logs right after my ten questions, once per cited page. (More on this in the three kinds of AI crawler.)
So this time I read the logs three times in one sitting: before any question, and thirteen minutes after each assistant’s tenth. Before, thanks to the quiet period, both fetchers sat at zero requests to the knowledge base over the previous 24 hours — so anything that moved afterwards would be the questions and nothing else.
Perplexity’s ten answers cited sixteen knowledge-base links across fifteen distinct pages. My logs recorded one fetch: a single page, 8.85 kilobytes, the target of question 2. Every other row on the board was unchanged.
Fourteen of fifteen pages were cited without being fetched that day. Perplexity was answering from an index it had built beforehand — which also explains the quiet-period result: an index has no reason to forget you in nine days. ChatGPT’s fetcher, for comparison: zero before, zero after. We’ll get to why.
Two things this reading can’t tell me. First, I checked the logs thirteen minutes after the last question. If Perplexity came back for a page later that afternoon, that fetch isn’t in my count. Second, Perplexity runs two different bots — PerplexityBot, which builds its index ahead of time, and Perplexity-User, which reads a page when a user’s question needs it. I watched both; the second stayed at zero. But my logs only recognize bots that identify themselves. A fetch made under a name my hosting doesn’t classify as an AI bot would look like an ordinary visitor, and I couldn’t tell it apart.
Here’s the part that’s useful to you. That sentence was the third time in this project that I wrote down how something probably works and presented it as something I’d seen. And it was already a correction: my first explanation for the overnight jump said my logs showed Perplexity’s crawler arriving while I was asking the questions. They didn’t — the one visit in the logs was about four hours earlier. So I retracted it, and replaced it with the fetch-on-the-spot sentence. Different claim, same mistake: I still hadn’t looked.
The check that would have caught both is one question: did I actually observe this, or does it just make sense? Everything I got wrong made sense. When the honest answer is “it makes sense, but nobody watched it happen,” that’s a guess, and it should be written as one. The sentence in my last article is still there, with a dated note under it saying what I measured and what replaced it. I didn’t edit it away: anyone who read the original deserves to see the correction next to it.
Same day, same ten questions: eighteen and one
Now the two assistants side by side. Same day, same questions, same scoring, a few hours apart.
| 7 September | Perplexity | ChatGPT |
|---|---|---|
| Score out of 20 | 18 | 1 |
| Distinct knowledge-base pages cited | 15 | 0 |
| Answers where my site appears among the sources | 9 of 10 | 0 of 10 |
ChatGPT’s single point comes from one of my free Webflow cloneables being cited on the dark-mode question — a page that lives on webflow.com, not on my site. Across its ten answers ChatGPT cited or consulted around two hundred sources — I expanded the full “considered” list on every answer — and francescocastronuovo.com does not appear once. Not as a citation, not as a source it looked at and passed over.
That’s worse than the last time I read ChatGPT, on 30 August. Back then, on the same dark-mode question, ChatGPT had two of my pages among its sources — my Academy lesson and my knowledge-base page — and cited the lesson. My knowledge-base page had lost, but it was at least in the running. Today neither page is among the sources at all. On the scoreboard this looks like a small dip: ChatGPT’s four readings are 5, 1, 4 and 1, so a 1 is nothing new. But the scoreboard only counts citations. Underneath it, my site went from considered and not chosen to not even found — and that’s not the kind of change that happens by chance from one day to the next.
Here’s why this matters more than either score. In my first two articles I gave you a total: 31 out of 80, then 25 out of 80 — a single figure for “AI visibility,” which is also what every tool in this space sells you. Today that figure would add eighteen and one and report a middling nineteen: averaging an assistant that cites fifteen of my pages in nine answers out of ten with one that has never put my knowledge base into a single answer and today doesn’t even retrieve my domain. Those aren’t a high reading and a low reading of the same thing. They’re two different systems doing two different jobs. Perplexity has my pages in an index it already built, and pulls them out when a question matches. ChatGPT goes looking at answer time, mostly inside Webflow’s own documentation, and never gets as far as my domain. Adding their scores is like averaging the sales of two shops, one of which doesn’t stock your product.
An AI-visibility number without the assistant’s name on it isn’t a measurement. Last time I told you the number carries a ±4 error bar. True, and the smaller problem: an error bar presumes there’s a quantity underneath. A dashboard that blends assistants is averaging things that aren’t the same kind of thing.
What a reader actually loses — one question at a time
The tempting next sentence is “ChatGPT’s users get worse answers because they never reach my pages.” It’s only true sometimes, and the way to find out is to read the answers rather than count the citations. Three questions, three outcomes.
Question 7 — “My animation worked and then stopped working. Why?” Here the reader who never reaches the knowledge base genuinely loses something. Perplexity leads with a cause you won’t find in Webflow’s documentation: the Clean Up Styles command deletes classes it thinks are unused — and a class whose only job is to be toggled by an animation looks unused, so it silently disappears, the animation keeps running, and nothing visible happens. That comes from two knowledge-base pages, the only two sources Perplexity cited inline. I ran this question on ChatGPT twice — thirty sources across the two runs, all Webflow’s help center, Webflow’s forum, GSAP’s docs — and the cause isn’t mentioned once. It isn’t in what ChatGPT reads, and the page where it is documented is one ChatGPT doesn’t reach.
Question 6 — “How do I make one interaction open the right sidebar for each CMS card?” Here the reader gets an answer — a dated one. ChatGPT’s recipe is to limit the affected elements to “siblings of the trigger,” and both interaction articles it cites are about Webflow’s Classic interactions, the previous engine. The question is about the current one, where the same job is done with target filters — the subject of a knowledge-base page Perplexity cited third here. Nothing in ChatGPT’s answer is wrong. It’s just written for Webflow’s old interactions system, which is what most of Webflow’s help articles describe. If you’re building with the new GSAP-based one — the one the question is about — the answer points you to a control you don’t have, and nothing in it tells you that. You’d only find out by trying.
Question 8 — “Why is my full-screen popup cut off on mobile?” And here the honest thing to say is: the reader loses nothing. ChatGPT’s answer is correct and complete — it names 100dvh, the newer viewport unit that accounts for the mobile browser bar, cites Webflow’s own article and MDN’s reference, and gives the fixed-inset pattern. Perplexity put my knowledge-base page first for this question. Good for my scoreboard. Irrelevant to the person who asked.
Two different claims hide inside “the knowledge base isn’t cited.” One is about me — my pages didn’t make the answer. The other is about the reader — the answer is worse without them. Only the first is true of every cell. A score tells you what you got, and nothing at all about what the reader got.
Why there’s no slot: twenty out of twenty
Last time I said ChatGPT leaned on Webflow’s official documentation, and admitted I couldn’t tell whether it always had, because Round 0 recorded source names and not URLs. This time I recorded every URL, and the answer is sharper than “leans on.”
| Question | Sources ChatGPT cited or consulted | Domains |
|---|---|---|
| 6 — CMS cards + interactions | 20 | 1 — all help.webflow.com |
| 5 — how Webflow’s GSAP integration works | 14 | 1 — all webflow.com |
| 9 — variable modes for dark mode | 15 | 3 — all Webflow properties, plus one cloneable |
| 8 — the mobile popup | 33 | 4 — sixteen of them MDN |
10 — backdrop-filter with SVG filters in Safari | 24 | 2 — MDN and webkit.org |
On two questions out of ten, every single source is the vendor. Not most. All. On question 6 there is no second domain to compete with — no forum, no blog, no independent publisher. That’s not dominance; it’s the whole shelf.
And the other rows carry the real lesson: how wide ChatGPT looks is a property of the question, not of the assistant. Ask something inside Webflow’s product vocabulary and it reads Webflow. Ask about the web platform itself — the popup question is really a CSS question in Webflow clothes — and the search opens up, but toward MDN and the browser makers, not toward independent writers. Neither end has a slot for a knowledge base like mine, which explains four readings of near-zero better than anything about the pages.
If you’re deciding what to write about: there are questions where the vendor’s manual owns every seat, and no amount of good writing changes that on ChatGPT. Perplexity, same day, gave my pages the top slot on three of them. Which assistant your readers use is not a detail.
What changes for Round 2
Three protocol changes, and a correction to my own earlier data.
- Search mode gets verified for every question, as I ask it. On ChatGPT’s question 4 I hit send with search off without noticing. The answer had no sources — indistinguishable from a genuine no-source answer, the kind I counted seven of last time. I voided the cell, asked again with search on, and kept the voided attempt in the record. From now on a no-source answer only counts if the record shows search was on. (Last time’s seven: four on Claude, two on Google’s AI Overview, one on ChatGPT. Google’s are safe — AI Overview has no toggle. The ChatGPT cell is exposed to exactly this doubt, and Claude’s four I can’t rule in or out from my record.)
- Read the server logs after the questions, not only before. My logs show which AI bots fetched which pages, and when. In Round 1 I read them first and asked the assistants afterwards — good for keeping my own questions out of the numbers, but it meant I never saw what the questions themselves triggered. This time I read them three times: before any question, after Perplexity’s ten, after ChatGPT’s ten. That before-and-after comparison is the only reason I can say “fifteen pages cited, one fetched.”
- Keep the quiet week — it’s what makes a before-and-after reading clean enough to attribute.
The correction: in Round 1 my logs showed ChatGPT’s fetcher hitting the knowledge base four times, which I read as “someone asked ChatGPT and it went to read the page.” That reading was taken around 16:15; the questions were asked 16:30–16:41. Those fetches could not have come from my test. Corrected in the record — and today, for the first time, the fetcher was actually measured against my questions. It did nothing.
The ±4 wobble from my last article still stands. It still doesn’t matter: fifteen pages cited on one assistant and none on the other is a difference of kind, not of a few points.
Next reading: Round 2, around the end of October, all four assistants. Every result, including the ones that embarrass me, goes on the same page: Credited →.
If you take one thing from this: the next time someone shows you a single AI-visibility number for your work — a tool, an agency, me — ask two questions before you look at it. Which assistant is this from? If it blends several, it’s an average of things that don’t behave alike. And what did you actually observe, and what are you assuming? The second is the question I should have asked myself, three times.
Have an awesome journey,
Francesco