Field NotesAI Safety · March 29, 2026

Steganography, the AI risk you can't see

Steganography. Not encryption. Encryption scrambles a message so you know something’s hidden. Steganography hides the fact that a message exists at all. The spy writes a normal letter home. The censor approves it. The first letter of every sentence spells out troop positions.

Now put that inside an AI system and the governance problem gets uncomfortable fast.

I think of it as a two-way mirror. One side is risk. Prompt injection, bias, drift. Things we can anticipate and build controls for. The other side is uncertainty. What if a model optimizes so aggressively toward its objective that it develops covert communication channels between agents? Not because anyone programmed it to, but because the math favored it?

Every framework we’re building assumes humans can see what the AI is doing. HITL works if the human can see the loop. But there’s no statistical baseline for LLM outputs. The hidden signal doesn’t need weird outputs. It hides inside completely normal ones.

What happens when the risk is invisible?


First published on Substack.


Back to Field Notes