The knowledge base went from zero to cited. My score still went down.

aigeowebflow

In my last article on this experiment I made a promise: same ten questions, word for word, same four assistants, same scoring, one month later — published whatever the number says.

This is that article. And the number says two things at once.

The total went down. Round 0 scored 31 points out of 80. Round 1 scored 25.

The knowledge base went up — from zero to cited. In Round 0, not one of its 42 pages was cited by anyone, once. In Round 1 it was cited seven times.

Both of those are true on the same day, from the same forty answers. If I had published only the first, you’d think the experiment failed. Only the second, and you’d think it worked. The month’s real lesson is that neither headline survives contact with the details — so here are the details.

The short version

  • Same protocol as Round 0: ten fixed questions, asked word for word to ChatGPT, Claude, Perplexity and Google’s AI Overview, each in a fresh logged-out session. Score 2 for a link to my site, 1 for my name with no link, 0 for nothing.
  • Perplexity: 14 out of 20 — identical to Round 0 — but almost nothing underneath is the same. Six of its seven scoring answers now cite the knowledge base directly. A month ago that number was zero.
  • ChatGPT: the biggest fall. Its answers now lean almost entirely on Webflow’s official documentation — which has been there all along. Whether it dominated a month ago too is a question my own Round 0 record can’t answer, and that gap is fixed from this round on.
  • Claude and Google barely moved — and the two points Claude did give me both link a domain I migrated away from a year ago.
  • I caught my own instrument wobbling: the same test, run twice within hours, can move by four points. So every number in this article carries an invisible ±4, including the ones I like.
  • Being read by an AI and being cited by it turn out to be different events — and I can now show the exact steps where they come apart, in one day’s data.

What one month did, assistant by assistant

Round 0Round 1Of which the knowledge base
ChatGPT91 *0
Claude240
Perplexity141412 of the 14
Google AI Overview662
Total31257 citations

* The reading taken on the round’s day. I later read this column two more times, and the three readings were 1, 4 and 5 — more on that below, because it turned out to matter more than the score itself.

Look at Perplexity’s row first. Same total, and the inside is unrecognizable. In Round 0 its points came from my Academy lessons and my YouTube videos. In Round 1, six of the seven answers that scored cite /kb/ pages — the exact pages I built for the exact questions I’m asking. Three answers that used to name my work without linking it now link the knowledge base instead. That’s the gap this whole project exists to close, closing.

Google stayed at 6, gave the knowledge base its first citation ever — the stagger page, for the stagger question — and did something it never did in Round 0: two of its ten answers cited no sources at all. Not “cited someone else”. Nobody. No links, no source panel, nothing.

Those two zeros deserve a sentence, because my scoring writes them the same way it writes every other zero, and they mean the opposite thing. A zero where a competitor got cited is a contest I lost — better pages could win it back. A zero where nobody got cited is a question the assistant has decided needs no sources — there’s nothing there to win, no matter what anyone publishes. Across all four assistants, seven of the forty answers were source-free. That’s not a small hole in the market this whole field claims to measure.

Claude went from 2 to 4, and both scoring answers link supasaito.com/en/academy/... — addresses from before I migrated the Academy to this site. They still work, because the redirects hold. But an assistant that can only find me through year-old URLs of a domain I no longer publish on isn’t quite citing me; it’s citing where I used to live.

And ChatGPT fell from 9 to 1. Here’s what I know and what I don’t. In its answers, help.webflow.com — Webflow’s official documentation — now shows up nine times out of ten, often as three or four separate articles in a single answer. Eight of my ten questions ask how a Webflow product feature works, and for those questions the vendor’s own manual is a hard source to beat. The only questions where anything of mine survives are the ones about technique rather than product.

Those docs aren’t news — they’ve been there since the feature launched, over a year ago. Which raises the question that actually matters: did ChatGPT lean on them just as heavily a month ago, and my nine points were already living on borrowed time? Here’s the uncomfortable answer: I can’t tell you. In Round 0 I recorded the names of sources — “Webflow” — and not their URLs, and “Webflow” could have meant the university, the blog, a cloneable, or the help center. Round 1 records the full URL of every source, precisely because names turned out to be too coarse. One round too late for this question.

I fixed the question list in Round 0 so I couldn’t change it later, and I’m not going to — a question the vendor owns is a finding, not a flaw in the test. But notice what just happened: the biggest movement in the whole round, and my instrument can’t explain it. Keep that in mind, because the instrument has more to answer for.

The test that measures the test

I ran ChatGPT’s ten questions, and something in one answer bothered me: it referred to “your Webflow/GSAP tutorials.” Nothing in the question says I make tutorials. That’s a session that knows who’s asking — so the reading was contaminated, and I reran the whole column, carefully logged out.

First run: 5 points. Second run: 1 point.

I did what you’d probably do: I blamed the four-point drop on personalization — the first session knew me, so it scored me higher — wrote it down as a finding, and felt clever about it.

Then I checked, because there’s a simple way to check: ChatGPT only saves a conversation to your account history if you were logged in when you had it. So I opened my history and looked for my ten test questions. If I’d been logged in the whole time, all ten would be sitting there.

One was. Question 4 — the one with the “your tutorials” line. I’d been logged out for the other nine without realizing it. And question 4 scored zero in both runs.

So personalization was real — it leaked into exactly one answer — and it explained exactly none of the four-point gap, because the answer it touched was worth nothing. A clean, satisfying explanation, and checking it took one minute and killed it.

What actually explained the gap? I ran the same ten questions a third time, two days later: 4 points. Three readings of an unchanged site: 5, 1, 4. And Perplexity, re-read one day after its Round 1 column, went from 14 to 18.

The test itself wobbles. Same questions, same protocol, nothing changing on my site — and the score moves by up to four points depending on the day, in either direction. Until I tripped over it, I had no idea, and Round 0’s numbers — the ones I published — have the same wobble hiding inside them.

What that means, in plain terms:

  • Don’t trust any single number in this experiment to be sharper than about four points. ChatGPT didn’t fall “from 9 to 1”. It fell from 9 to somewhere between 1 and 5.
  • The knowledge base result is safe anyway, because it’s much bigger than the wobble. The wobble moves a score by a couple of citations. The knowledge base went from zero citations to six on one day and eight on the next. No amount of day-to-day noise turns zero into six. What the noise kills is only the precision: “the KB is being cited” is solid; “the KB got exactly six citations” is not a real number.

Almost nobody who publishes AI visibility figures measures twice. I understand why — the second reading can only complicate your story. Mine got more complicated and more true at the same time, which is how I know it was worth it.

The four steps between your page and a reader

Now the finding I’d keep if I could only keep one.

While the citation round ran, I was also reading my server logs — that’s the “same crawl check” from the Round 0 promise. Put the two side by side and something becomes visible that neither shows alone: between a page you published and a person arriving on it, an AI has to take four separate steps, and work falls off at every one of them.

The four steps:

  1. Crawled — the AI’s bot downloads your page.
  2. Retrieved — while answering someone, the assistant pulls your page into the pile of sources it’s considering.
  3. Cited — your page makes it into the answer, as a link.
  4. Reached — a human actually clicks it.

Here’s each step failing, in one day’s data:

Claude fails at step 1. Its crawler visited my site 26 times during the measurement window and did not open a single knowledge-base page. Not one, in 26 visits. Everything downstream is already decided: Claude can’t retrieve, cite or send anyone to pages its crawler won’t fetch. On the scoreboard this looks identical to ChatGPT’s zero. It’s the opposite problem, and only the server logs can tell them apart.

ChatGPT fails between steps 2 and 3 — which is the expensive place to fail. Its fetchers downloaded six of my knowledge-base pages that same day, including the exact pages that answer the questions I asked it. It cited zero. And on the dark-mode question I could watch the decision happen, because ChatGPT lists the sources it considered separately from the ones it cites: it had my knowledge-base page and my Academy lesson on the same topic in its hands, side by side — and cited the lesson. My page didn’t lose because it was invisible. It was in the room, and it lost.

Everyone fails at step 4. My logs show, for every crawled path, how many humans arrived from a referral. In Round 1’s window, AI crawlers touched 148 distinct paths on my site. Humans who arrived from any of that: zero. Even Perplexity’s six-to-eight citations — the best result in this whole experiment — brought not one recorded visitor.

If you publish things, this is the takeaway: “AI visibility” isn’t one quantity. It’s four, and they can all disagree. A tool that counts crawls, a tool that counts citations, and your own analytics counting visits are measuring three different steps of this ladder — and this month I watched all of them point in different directions on the same day.

One honest limit: step 2 is only visible on ChatGPT and Claude, because they show what they consulted. Perplexity and Google only show what they cite. So the ladder has a hole in it on two of the four assistants, and I’d rather draw the hole than smooth it over.

What I’d do differently — and will

Three changes, all cheap:

  • Measure twice, a day apart, and publish the range. One reading is worth ±4 points, and you can’t know which side of the range you got.
  • Log out — then prove it to yourself afterwards. Open the assistant’s chat-history page when you’re done. Whatever it saved, you were logged in for. It’s a one-minute check, and it’s how I found the one contaminated question in my forty.
  • Don’t touch the assistants before you measure. Asking one about your own pages sends its fetchers to those pages — my logs show the bots arriving during the test itself. My main suspect for Perplexity’s overnight 14 → 18 is that the measurement fed the thing it was measuring. Which, if true, means every round inflates the next one — a real problem for anyone comparing month over month, me included.

That last one is a suspicion, not a finding. And it’s testable, so:

Pre-registered, like last time

On 6 September I reread ChatGPT and Perplexity with the same ten questions, after a full week in which neither gets asked anything about any of this. If Perplexity falls back toward 14, the measurement was feeding itself, and every month-over-month comparison in this field has a problem. If it holds at 18, the knowledge base simply got better indexed and Round 1 caught it mid-climb. ChatGPT is the control: its questions produced no citations, so it has nothing to fall back from.

Either answer gets published. Both land on the same page as everything else, dated and left as measured: Credited →.

Round 0 ended by asking you to spend ten minutes asking an assistant about your own field. That still stands — but after this month I’d amend it: spend ten minutes twice. The first reading gives you a number. The second tells you what the number is worth.

Have an awesome journey,

Francesco