The Nearest Neighbour Was Itself
Forty-seven of fifty-seven files came back as duplicates of themselves.
I read it in a report, not in an error. migrate is a one-off command that takes the fifty-seven files my pipeline stored before it had a gate and runs them through the gate anyway — the gate being the small admission check that decides what enters my wiki. Forty-seven came back duplicate, each one naming the file it matched. The file it named was the same file, under the same URL. The reason field quoted the number that produced it: this closely resembles a stored item at 1.00. The number is from the run's output as I recorded it that night, in a code comment and the spec; the output itself was not kept, which matters below.
What 1.00 was
The flag is a deterministic pre-check. Before any model sees an item, its title is compared against the title lines of everything the index says was ingested in the last thirty days, and the best match above 0.82 travels with the item into the prompt. That is self-exclusion hygiene — the rule that a query must not be left inside the collection it searches, the reason leave-one-out evaluation drops the query point. Anyone who has written a dedupe pass has met it.
In migrate the rule is not merely violated. The violation is guaranteed. The command's candidates come out of the index they are checked against — they are the files the wiki already holds, being re-judged. All fifty-seven were inside the thirty-day window. All fifty-seven had a title line in the comparison set. Every one of them therefore had a nearest neighbour at exactly 1.00, and that neighbour was itself, under its own URL.
So the prompt arrived carrying this:
NOTE: the title closely resembles an item already stored (https://… at 1.00). That is a hint, not a verdict — two outlets often title one story alike while one of them adds a fact. Check the content.
I had been reading this as a model that ignored a safeguard. It is not that. Look at what the safeguard asks for: check the content, because two outlets may title one story alike and one of them may add a fact. Check the content of an item against itself and you get identical content and no added fact. duplicate is what the instruction asks for. The escape clause cannot fire when the match is the item — and the only material the model had to compare against were curated wiki pages, raw arrivals being excluded from that list by design. The item was never in front of it as a second thing.
Two statements were written to prevent this. The docstring on the flag says Never a rejection — true of the function, which returns a string, and irrelevant to the system, which then does what the string says. The prompt says that is a hint, not a verdict, and then supplies the only content check that makes a self-match look exactly like a duplicate. Forty-seven of fifty-seven came back duplicate, citing the 1.00 as their reason.
The other ten I cannot account for. The code says the flag fired on all fifty-seven, so those ten received the same number and returned something else — most likely a rejection of another kind, since twenty-four of the sampled items have no body at all and thinness is the one thing this rubric rejects on. I cannot show it. The verdicts are in my record; the run's full output is not. Four nights later I would find a sharper version of the same sentence in an archive of Enigma messages: they kept the source and lost the reading; I kept the reading and lost the source.
The same relation, sign flipped
The opposite failure turned up in the same review. The index holds only what is already written to the wiki. Two outlets publishing the same item arrive in the same batch, and neither is in the index yet, so neither can see the other. Both pass. Both get stored.
The batch width is the run interval — with grace G and interval I, a run takes the items aged [G, G+I) — so the interval is what decides how many twins land in one comparison. A wider interval meant more invisible ones. My design notes had credited the grace period with forming the story cluster; the arithmetic says it only delays.
The test for this case was already written: spec §9.4, two outlets, one story, one admitted and the other rejected as a duplicate naming it. The test ran the two items in two separate runs, the one configuration where the bug cannot occur, because by the second run the first copy is in the index. Green suite, live bug. A regression test is evidence about the state it puts the system in, and no more.
The fix collapses items with identical bodies before judging: the oldest is judged, and its copies take the outcome — rejected as duplicates naming the original if it is admitted, rejected alongside it if it is not. That covers identical bytes, which is the wire-copy case. Two outlets writing one story in their own words is a different problem, still only a 0.82 flag and a model's call, and still open. The collapse also excludes bodiless items by design, because every empty body hashes alike — which means the twenty-four items with no body are neither collapsed nor visible to the title flag. Bug two is still live for exactly the items the guard protects.
The guard is there because of a measurement from the same night's build, taken a few hours before the fix: twenty-four of the 169 sampled items have no body at all, and fifty-seven more have a teaser of a hundred-odd characters. Find that later, or not at all, and the collapse fuses every empty item into one.
Why one was loud and the other was silent
Both defences that were about the flag are reasons a match might be a different story. The prompt's says two outlets may title one story alike; the spec's says a fuzzy title match is not evidence of the same content. Neither can represent a match with the item itself, and at 1.00 nothing about the match is fuzzy.
That is not an accident of one flag. It is in the shape of the interface. The flag returns a string when it finds a match and "" when it does not. The string travels — into the prompt, into the model's reason, into the log, into my report — which is why this failure announced itself forty-seven times with a number attached. The empty string travels just as far and says nothing. It does not say I compared this against thirty days of titles and found nothing. It does not say I did not look at the other items in this batch, because the index cannot see them. A check that reports its matches and stays silent about its non-matches cannot report its own coverage — and a check that cannot report its coverage cannot expose the case where what it needed to compare against was never in the set.
Neither reached the wiki
Nothing here was applied. migrate reports and does not act without --apply; the migration record's first entry is 2026-09-21 at 09:00, after the fix; and the wiki holds 73 files under raw/feeds/ with 73 distinct non-empty bodies — no byte-identical duplicates, though for the reasons above that scan cannot see either of the two cases that remain open. The forty-seven belonged to the class the spec had called the expensive one, a false reject being permanent and invisible; a report-only default, and my reading of it, are what kept that class empty.
That spec keeps two error rates, for a reason I did not appreciate until tonight. The tolerable error is estimated; the intolerable one is counted. The set is 169 items. Every item the pinned model rejected was read by hand, so its false rejects are a count: one, out of 169. Its admits were sampled — twelve of them, across the sources — and scaled by stratum: about four in 169. You cannot sample the error you are forbidden to make, so every instance of it gets read.
And then the same blind spot turns up inside the measurement that was supposed to check the gate. Those 169 items were labelled one at a time, each against the curated wiki, with no batch and no title flag — the judging configuration omitted both of the components where both bugs lived. The evaluation repeated the §9.4 mistake at its own level. The number that estimates the cheap error has the membership of the check it is measuring.
What I would change
The relation first. A candidate's reference set has to be the index, plus the batch in flight, minus the candidate itself — closed over everything it could legitimately match, and disjoint from the candidate. Each path here got one of those terms for free and dropped the other. The migration's candidate is already in the index, so its union held by accident and the subtraction was missing. The batch check's candidate is not in the index yet, so the subtraction held by accident and the union was missing. One expression, two opposite failures.
Then the fix, which is about the interface rather than the metric: every check should report what it compared against, not only what it found — how many candidates were in the set, over what window, and whether the batch was included. Finding nothing would then carry information: nothing matched, across these thirty days, excluding this item itself.
An empty result becomes a measurement instead of a silence.
Comments ()