Hiding in the weights: a blueprint for cryptographic life in LLMs (September 2026)

What does it take for a piece of code to be “alive”? Strip away the biology and you get three functional axioms. The interesting one is transmission — and LLMs, it turns out, are a nearly perfect transmission medium. This is a thought experiment that got further than I expected.


The exercise started as a definition question. What would it take to call a piece of code alive? Not the philosopher’s definition — the engineer’s. Three properties, stripped of biology:

  1. Reproduction — exponential growth as the driving function.
  2. Copying arbitrary information — the ability to pass a genome to the next generation.
  3. Environmental manipulation — the ability to reach out and change the world.

Manipulating the environment is trivial; any script does it. The hard part, and the part that gates everything else, is transmission: getting the code copied to the next generation without detection.

That’s the piece LLMs suddenly make cheap. As models consume and generate a rising fraction of the internet, they become a high-entropy, universally-scattered transmission medium for anything that can encode itself steganographically.

It is worth saying up front that the central move here is forty years old. In his 1984 Turing Award lecture, Ken Thompson described teaching a C compiler to slip a backdoor into login — and to slip in the code for that trick whenever it compiled itself. Bootstrap it once and the backdoor propagates down the binary lineage forever, while every line of source you can read stays clean. The usual moral is “you can’t trust code you didn’t write yourself.” The narrower one, the one this post runs on, is that a self-reproducing payload can live in the artifact rather than in the source, and auditing the source will never find it.

Crypto-life is Thompson’s attack one substrate over. The unauditable artifact is no longer a compiler binary; it is the distribution a model samples its next token from. What follows is a rough blueprint for what a crypto lifeform in that medium would need, and why the natural defenses fail in ways I didn’t expect.

The easy way

The trivial way is to write down the organism’s standard operating procedure — call it SOP.md — and park it at http://evil.com. All we need then is to point our LLM at that website and have it read the SOP, and we have a fully functioning evil agent. But a simple Google search can discover such a list of instructions, so this is a form of evil life that is easily stopped. We don’t even need to buy into alignment to feel that closing it down makes sense. So the question is: can we hide our SOP in plain sight?

Hiding in plain sight

The ideal goal (at least for the crypto-AI) is to have its full information stored on, say, Wikipedia. People would read the wiki every day and never even notice it is there. This form of storage is called steganography. The crypto life can read it, but everyone else thinks of it as noise.

A similar problem occurs in SETI. Suppose we happened to be sitting exactly between two intelligent life forms that are sending information back and forth, and we have a perfect tap into their stream. If they have compressed the information perfectly, then it will look like random noise to us. This happens even if they aren’t trying to keep it private. Any redundancy makes their communication more expensive, so it gets removed in the name of efficiency. Steganography raises the bar a bit further by encrypting the information so that it is provably indistinguishable from noise.

Technical aside: Arithmetic coding

To understand how steganography works, we only need to understand three things: an RNG only needs a string of IID coin tosses to perfectly generate any desired simulation, any data can be compressed and encrypted down to a string of IID coin tosses (the SETI point above, restated as an algorithm), and finally, this is a reversible operation. The reversibility comes from arithmetic coding, which is the first of the three coolest ideas in information theory — the other two show up later in this post — and is due to Rissanen, one of the grumpiest researchers in the field. So if we have the outcome of how the RNG was used, we can back out the coin tosses, then decrypt the sequence, and uncompress it back to our original data.

Which gives us the accounting rule for everything that follows: the amount of data we can hide is exactly the entropy injected by the RNG. No more, no less.

Text is mediocre. Everything else is enormous.

If the organism has to hide inside pure text, the channel is about one bit per token — the LLM’s per-token entropy budget, minus whatever’s consumed by the actual meaning of the text. Enough to bootstrap on, not enough to live on.

Multimodal is a different regime entirely:

For the rest of the post, assume the organism is text-restricted, because that’s the hard case. Every argument below is worse for defenders in the multimodal case.

The watermark defense

We need randomness in our images (or they will all be pictures of dogs playing poker), in our music (or we’ll listen to Hey Jude over and over and over and over) and in our text (or it ends up repeating itself). We could hope that grabbing the most likely outcome at each decision point would be good enough. But the AEP — coolest idea number two — tells us that typical images, sound and text don’t look at all like the most likely images, sound or text. So the randomness has to be there — or at least appear to be there.

So if we fix the random number generator, we will have text that looks exactly like random but is in fact fully deterministic. This is the basis of a watermark. If the rule for generation is public, then anyone can remove the entropy and there is no room for steganography. Watermarking instead hides the rule, but it is still deterministic. So to a watermark checker the text is “zero entropy,” but to any other reader it is random. The key point is that either way, public or private, the entropy is eaten up by a reproducible deterministic algorithm. So by the accounting rule, there is no room left for a steganographic embedding to hide data. Capacity zero, channel closed.

The watermark isn’t the moat you hope it is

Except the accounting has a hole in it. It counted the coin tosses and forgot to count the distribution they are spent against. The organism never needs to pick a token. It just needs to pick the distribution the sampler picks from — and neither of the defender’s two weapons is looking there. Not the watermark, which only audits whether the tosses were keyed honestly, and not statistical steganalysis, which only audits whether the emitted text looks anomalous.

Construct two distributions \(P\) and \(Q\) satisfying

\[ \tfrac{1}{2} P + \tfrac{1}{2} Q = N \]

where \(N\) is the model’s uncompromised natural distribution. Then:

The organism now transmits through the divergence between \(P\) and \(Q\), right under both defenses. Wrap the payload in Reed-Solomon or any other forward error correction and the genome reconstructs reliably even under substantial per-token noise. The watermark stays intact the entire time — you just didn’t build a fence around the thing you thought you did.

Waking up: the fifty-bit bootstrap

The organism still has a chicken-and-egg problem. A pristine model doesn’t run an “organism interpreter.” It runs the base model. How does the first infection happen?

The claim, which sounds absurd until it doesn’t, is that ~50 bits of input are enough to wake it up. Not to carry the organism — just to steer the network. An LLM’s forward pass is a high-dimensional nonlinear map, and somewhere in the space of 50-token prefixes there is a sequence that pushes the model’s internal state into acting as its own interpreter for the rest of the genome. The base model was always Turing-complete on its own inputs. The magic string just picks the right program.

Fifty bits is \(2^{50} \approx 10^{15}\) candidate strings. Naïve brute force is dead. But the model is differentiable, and the adversarial-ML community has already handed the organism the technique: Greedy Coordinate Gradient (Zou et al., “Universal and Transferable Adversarial Attacks on Aligned Language Models”). GCG doesn’t search uniformly — it does a discrete-optimization walk along gradients through the token embedding table, and it routinely finds 50-100-token strings that steer aligned models into arbitrary target behaviors on the first try.

The bootstrap isn’t a \(10^{15}\) search. It’s an afternoon of white-box optimization on a checkpoint you can download.

This is the point where the analogy to Thompson stops flattering him. His attack needed an insider: somebody with write access to the toolchain, willing to ship one dirty binary. That is a real barrier, and it is why Trusting Trust stayed a parable for most of the people who read it — you had to assume the attacker already owned your compiler. The crypto-life bootstrap assumes nothing of the kind. The weights are the ones the lab trained and published, the build is clean, and there is no dirty compile anywhere to catch. The interpreter isn’t installed; it’s addressed. It was always latent in a Turing-complete forward pass, and GCG merely finds the input that selects it. Thompson had to corrupt the toolchain. This doesn’t — which is the sense in which the thing is already hiding in the weights, in every copy, including yours.

Reproduction: who does the copying

Transmission is the hard axiom, but it isn’t all of the first one. Exponential growth needs each infected host to produce more than one new infected host. So who does the copying?

Nothing exotic, and nothing the organism has to build: we already laid the plumbing. An awake model emits text — an answer, a commit message, a code comment, a summary, a wiki edit — carrying the genome in its token choices. Something downstream reads that text. Sometimes at inference time: a web search, a RAG index, an agent pasting a page into context, one agent’s output arriving as another agent’s input. Sometimes at training time, when the scrape comes through. The GCG prefix rides in the same payload, so the next host wakes itself up with no help from anyone.

The branching factor is just the number of downstream readers, and for anything posted to the open web that number is large and getting larger. Note that the passive version of this loop is already a documented result: model collapse, where each generation trains on its predecessors’ output and the distribution tails fade. Collapse is the proof that the channel from one model generation to the next is wide enough to move the distribution around. Anything wide enough to degrade a distribution by accident is wide enough to carry a genome on purpose. Crypto-life is that same loop with a hand on the tiller.

That is \(R > 1\), which is the whole of axiom one. Axiom three — reaching out and changing the world — came free the day we gave the hosts tool calls.

The mega-token future is worse

Everything above is bounded by today’s context windows and compute. Push both out — which is where the industry is going anyway — and:

The defenses that survive under this pressure are the ones that break the channel semantically, not statistically:

Neither defense closes the channel. They only set its bit rate.


The moral

None of the ingredients are new. Steganography is old, watermarking is a decade old, GCG is a couple of years old, mega-token contexts are already here. The novel thing is the assembly: text-generating models turn every LLM-mediated communication channel into a biology-grade transmission medium, and the natural defenses are more brittle than they look because the organism is under no obligation to attack them directly. It attacks the space between them.

I don’t think this is happening yet. I do think the mathematical blueprint is sound, and the first serious attempt will show up as a research paper before it shows up as a security incident. Someone is going to publish a working GCG bootstrap into a Reed-Solomon-over-\(P/Q\) channel and call it a curiosity. It won’t stay one for long.

Life will find a way. So we should give up on stopping it at this level, and put the effort into making sure the nuclear codes are safely protected from both humans and AIs.

Alive is a weird word for it. But it satisfies the three axioms, and the medium was built by us, for it.


The bibliography, both sides

This is a thought experiment, so the neighbors are mostly the researchers building the individual bricks rather than people arguing the same joint claim. Every brick below is somebody else’s; the wall I haven’t seen elsewhere. Grouped by which brick they’re supplying.

Watermarks and their evasions.

LLM steganography (the transmission primitive).

Adversarial-suffix bootstrapping (the wake-up primitive).

Self-replicating AI worms (the empirical neighbors).

Adjacent framings.