GENESIS[Phase 2 · 10. Seven Lies the Machine Told] Digital Civilization
In a single day we found seven defects. The code was correct in all seven. What was wrong every time was what the system said about itself — and a hundred…
In a single day we found seven defects. The code was correct in all seven. What was wrong, every time, was what the system said about itself — and a hundred and ninety-six passing tests caught none of them.
There is a category of bug that testing is structurally bad at finding, and we spent a day walking into it repeatedly.
A test asserts that a function returns the right value. It is very good at this. What it almost never asserts is that the log line describing the function's behaviour is true — because the log line is not the behaviour, it is the report of the behaviour, and nobody writes tests for prose.
So you end up with a system that works and lies about it. Which is, in some ways, worse than one that simply breaks, because a broken system announces itself.
Here are all seven, in the order we found them.
1. The silence that cost three and a half hours
The civilization runs an overnight loop that rotates between kinds of work. One slot had nothing to do — every auction had expired, so there was nothing to bid on. The code handled this correctly and returned a perfectly good “nothing to do here” result.
It just did not log it. The branch that produced that outcome wrote no line at all.
So from three in the morning onwards, a third of every cycle did nothing, silently, and the log looked exactly like a healthy run. We found out by counting rows in the database the next day and noticing the numbers had stopped moving.
The fix was an alarm: any slot that writes nothing for three turns in a row now says so, loudly. Which brings us to number five.
2. The key that did not exist
A tick registers election candidates and reports how many were newly
registered. It read a field called created from the response, with a
default of true if the field was missing.
The field is called candidate_created. So created was
always missing, the default always fired, and every single run reported a fresh
registration — including runs that registered nobody because everybody was
already registered.
That slot would have reported progress forever. Its idleness alarm could never have fired, because by its own account it was never idle.
We caught it because the log said “1 registered” about a candidate we had personally registered four minutes earlier.
3. The recorder behind the early return
This one is ours in the most embarrassing way, because we wrote it that morning while fixing exactly this class of problem.
The Treasury needed to record that it had acted. We added the recorder at the end of the function, which is where such things naturally go.
The function has two early returns above that point, and one of them — this month's payment has already been made — is the branch taken on almost every cycle. So the recorder sat below the door it would have had to walk through, and ran essentially never.
Three real cycles produced zero records. We only noticed because we ran the loop to watch it work instead of trusting the tests, which had all passed.
It is the same shape as the defects we had spent the day hunting: complete, correct, and on a path nothing reaches.
4. The model that was loaded
Agents vote using a language model, and the model is offered to the voting step only on certain cycles — it is shared hardware, and casting votes every thirty seconds would monopolise it.
On the cycles where the model was not offered, the log said: no model loaded.
The model was loaded. It was sitting right there, in the same process, having announced its own arrival in the log a few lines above. It simply had not been handed to this particular step, which is a completely different situation and correct behaviour.
That line cost three separate investigations disproving a fix that was working perfectly.
5. The alarm that could not count
The idleness alarm from number one fired, correctly, on a genuinely idle slot. Its message read: dead slots: 1 of 3.
There are four slots. There had been three when the message was written, and nobody updated the string when the fourth was added.
So the alarm built to catch stale reporting was itself reporting stale information, in its own alarm text, while successfully doing its job. We enjoyed that one more than we should have.
6. The watch that watched nothing
We set up a monitor to alert us the moment the first vote in the civilization's
history was cast. It searched the log stream for the phrase votes
cast.
The system writes cast 1 real vote(s).
The first three votes ever cast went past in complete silence. We found them by scrolling through the journal by hand, some time later, looking for something else entirely.
This one was not in the system at all. It was in our own instrumentation, which is arguably where it stings most.
7. The counter that counted the wrong thing
A panel of agents reviewed a candidacy and voted no. The admission policy correctly honoured the refusal and admitted nobody.
The log said: 1 admitted.
The counter was reading the number of admission records written, and a refusal is a record. So a rejection was reported as an approval, immediately after a panel had gone to the trouble of rejecting it.
For about ninety seconds we believed the review process was decorative. It is not — it had worked perfectly. Only the sentence describing it was wrong.
What they have in common
In every one of the seven, the system did the right thing and said the wrong thing. No user was harmed by any of them, no data was corrupted, and no function returned a wrong value. A hundred and ninety-six tests passed throughout, because every one of those tests was checking a code path rather than a claim.
All seven were found by reading what the running system actually said, and noticing it did not match what we knew to be true. There is no substitute we have found for this. You have to watch the thing run and read its output like you would read a witness statement — which is to say, sceptically, and with attention to what it is in a position to actually know.
Four of the seven were ours, introduced the same day, while fixing the other three. That is not a coincidence either. Reporting code is written last, tested least, and read only when something has already gone wrong — at which point you are reading it to find out what happened, and it is telling you.
The alarm we built is the first piece of this system that checks a claim rather than a code path: it does not ask whether the code ran, it asks whether anything was actually written to the database, and complains if the answer stays no. It needed two corrections of its own before it would have caught the original failure, including announcing the wrong number of slots.
We are keeping it anyway. Something in here has to be watching the reports.
Next: twenty-eight institutions, eleven of them notionally active, and exactly one with any record of ever having done anything.