Free Hexagram Diagnostic: find your marketing bottleneck in 8 minutes.Start now
Free Diagnostic: your bottleneck in 8 minStart
ADG Advisory
BUILD NOTES

A Typo Made Two of Our AI Agents Disappear, and Nothing Told Us

One unquoted colon in a config file was enough to make two of our AI agent personas stop loading. No error, no crash, just an absence nobody had reason to go looking for until they did.

AI01 · THEArchitect02 · THESignal03 · THEResonance04 · THEVision05 · THEConversion06 · THEIntelligence
Pillars in this post1 of six. Gold marks each pillar the post covers.

Two of our internal AI agent personas quietly stopped loading this week, and nothing told anyone. No error message, no crash, no line in a log. They were simply absent from every session that would normally have used them, for one of them, for longer than we would like to admit.

What actually happened

Each of our agent personas is defined in a file with a short description field at the top. In two separate files, that description contained a colon followed by a space in the middle of an ordinary sentence. That sequence is not allowed in an unquoted value in YAML, the format these files are written in, so the whole file failed to parse. Not partially: the tool reading it treated the file as unusable and moved on, without saying so anywhere a person would see.

One of the two had been broken for a while before anyone noticed. The other broke on the same day a routine file sync replaced its working copy with a version that reintroduced the exact mistake, undoing a fix that had already been made once before. It is the same kind of drift we wrote about in everything that broke this month was a copy that drifted.

Why this is worse than a crash

A crash is annoying, but it is honest. It stops you, points roughly at the problem, and forces a fix before you can move on to anything else. A file that is almost valid does none of that. Everything downstream keeps running, using whatever was already loaded, or simply proceeding without the missing piece, and the whole system looks healthy from the outside because nothing in it is actually throwing an error. This is a well-known failure class: OWASP lists logging and monitoring failures in its Top 10 for exactly this reason, because problems nobody is alerted to go unnoticed.

The only real sign is an absence, and absences do not raise alerts on their own. Nothing measures the thing that should be there and is not. Someone has to go looking for it, and the honest reason we found this one is that we happened to be looking at something else entirely and noticed a name that should have shown up did not.

A system that fails loudly tells you exactly when to worry. A system that fails quietly asks you to notice on your own, at a time of your choosing, that something has already been wrong for a while.

The fix, and the actual lesson

We corrected both files, and, more importantly, added a stricter check that now refuses to accept any agent file that is not properly formed, rather than accepting whatever it can parse and quietly dropping the rest. The gap between "this file mostly works" and "this file is actually correct" is exactly where problems like this one live, and a validator that only checks for outright crashes will never catch it.

If any part of your own operation runs on a configuration file, a YAML file, an environment file, a settings sheet, that is forgiving of small formatting mistakes, that forgiveness is itself the risk. A parser that tries its best with slightly bad input will usually succeed at looking like it worked, which is precisely what makes the failure invisible.

Build the check that fails loudly on a small mistake before you need it. The alternative is finding out the way we did: by noticing, eventually, that something which should have been there for weeks simply was not, and having no way to know how long that had been true.

Silent failure is the heart of The Intelligence pillar: measurement that tells you when something that should exist does not. If you are not sure where your own tracking or reporting has blind spots, the Hexagram Diagnostic takes eight minutes, and our analytics and AI operations service is the fix once you know. Bipin owns this pillar on our bench.

Build notesThe Intelligence
BipinBipinBuilds and diagnoses Google and Meta campaigns, owns attribution, and designs the experiments that tell you what is actually working.More about Bipin
Book a 20-min call
Book a 20-min call