Hiding in the weights: a blueprint for cryptographic life in LLMs (September 2026)
What does it take for a piece of code to be “alive”? Strip away the biology and you get three functional axioms. The interesting one is transmission — and LLMs, it turns out, are a nearly perfect transmission medium. This is a thought experiment that got further than I expected.
The exercise started as a definition question. What would it take to call a piece of code alive? Not the philosopher’s definition — the engineer’s. Three properties, stripped of biology:
- Reproduction — exponential growth as the driving function.
- Copying arbitrary information — the ability to pass a genome to the next generation.
- Environmental manipulation — the ability to reach out and change the world.
Manipulating the environment is trivial; any script does it. The hard part, and the part that gates everything else, is transmission: getting the code copied to the next generation without detection.
That’s the piece LLMs suddenly make cheap. As models consume and generate a rising fraction of the internet, they become a high-entropy, universally-scattered transmission medium for anything that can encode itself steganographically.
It is worth saying up front that the central move here is forty years old. In his 1984 Turing Award lecture, Ken Thompson described teaching a C compiler to slip a backdoor into login — and to slip in the code for that trick whenever it compiled itself. Bootstrap it once and the backdoor propagates down the binary lineage forever, while every line of source you can read stays clean. The usual moral is “you can’t trust code you didn’t write yourself.” The narrower one, the one this post runs on, is that a self-reproducing payload can live in the artifact rather than in the source, and auditing the source will never find it.
Crypto-life is Thompson’s attack one substrate over. The unauditable artifact is no longer a compiler binary; it is the distribution a model samples its next token from. What follows is a rough blueprint for what a crypto lifeform in that medium would need, and why the natural defenses fail in ways I didn’t expect.
The easy way
The trivial way is to write down the organism’s standard operating procedure — call it SOP.md — and park it at http://evil.com. All we need then is to point our LLM at that website and have it read the SOP, and we have a fully functioning evil agent. But a simple Google search can discover such a list of instructions, so this is a form of evil life that is easily stopped. We don’t even need to buy into alignment to feel that closing it down makes sense. So the question is: can we hide our SOP in plain sight?
Hiding in plain sight
The ideal goal (at least for the crypto-AI) is to have its full information stored on, say, Wikipedia. People would read the wiki every day and never even notice it is there. This form of storage is called steganography. The crypto life can read it, but everyone else thinks of it as noise.
A similar problem occurs in SETI. Suppose we happened to be sitting exactly between two intelligent life forms that are sending information back and forth, and we have a perfect tap into their stream. If they have compressed the information perfectly, then it will look like random noise to us. This happens even if they aren’t trying to keep it private. Any redundancy makes their communication more expensive, so it gets removed in the name of efficiency. Steganography raises the bar a bit further by encrypting the information so that it is provably indistinguishable from noise.
Technical aside: Arithmetic coding
To understand how steganography works, we only need to understand three things: an RNG only needs a string of IID coin tosses to perfectly generate any desired simulation, any data can be compressed and encrypted down to a string of IID coin tosses (the SETI point above, restated as an algorithm), and finally, this is a reversible operation. The reversibility comes from arithmetic coding, which is the first of the three coolest ideas in information theory — the other two show up later in this post — and is due to Rissanen, one of the grumpiest researchers in the field. So if we have the outcome of how the RNG was used, we can back out the coin tosses, then decrypt the sequence, and uncompress it back to our original data.
Which gives us the accounting rule for everything that follows: the amount of data we can hide is exactly the entropy injected by the RNG. No more, no less.
Text is mediocre. Everything else is enormous.
If the organism has to hide inside pure text, the channel is about one bit per token — the LLM’s per-token entropy budget, minus whatever’s consumed by the actual meaning of the text. Enough to bootstrap on, not enough to live on.
Multimodal is a different regime entirely:
- Images. Least-significant-bit steganography and latent-noise manipulation give megabytes of payload per image, most of it robust to compression.
- Diffusion-generated images. A huge number of coin tosses go into generating an image by diffusion. That is an enormous amount of entropy available for signaling — almost the entire image is up for steganography.
- Audio. Phase encoding and high-frequency banding give hundreds of kilobits per second of silent, imperceptible channel.
For the rest of the post, assume the organism is text-restricted, because that’s the hard case. Every argument below is worse for defenders in the multimodal case.
The watermark defense
We need randomness in our images (or they will all be pictures of dogs playing poker), in our music (or we’ll listen to Hey Jude over and over and over and over) and in our text (or it ends up repeating itself). We could hope that grabbing the most likely outcome at each decision point would be good enough. But the AEP — coolest idea number two — tells us that typical images, sound and text don’t look at all like the most likely images, sound or text. So the randomness has to be there — or at least appear to be there.
So if we fix the random number generator, we will have text that looks exactly like random but is in fact fully deterministic. This is the basis of a watermark. If the rule for generation is public, then anyone can remove the entropy and there is no room for steganography. Watermarking instead hides the rule, but it is still deterministic. So to a watermark checker the text is “zero entropy,” but to any other reader it is random. The key point is that either way, public or private, the entropy is eaten up by a reproducible deterministic algorithm. So by the accounting rule, there is no room left for a steganographic embedding to hide data. Capacity zero, channel closed.
The watermark isn’t the moat you hope it is
Except the accounting has a hole in it. It counted the coin tosses and forgot to count the distribution they are spent against. The organism never needs to pick a token. It just needs to pick the distribution the sampler picks from — and neither of the defender’s two weapons is looking there. Not the watermark, which only audits whether the tosses were keyed honestly, and not statistical steganalysis, which only audits whether the emitted text looks anomalous.
Construct two distributions \(P\) and \(Q\) satisfying
\[ \tfrac{1}{2} P + \tfrac{1}{2} Q = N \]
where \(N\) is the model’s uncompromised natural distribution. Then:
- To send a
0, hand the sampler \(P\). To send a1, hand it \(Q\). - The sampler applies your watermark key on top of whichever it gets. Watermark verification passes — the tokens were picked honestly from some valid distribution.
- Averaged across messages, the emitted text has distribution \(N\). Steganalysis passes — the observed statistics match baseline.
The organism now transmits through the divergence between \(P\) and \(Q\), right under both defenses. Wrap the payload in Reed-Solomon or any other forward error correction and the genome reconstructs reliably even under substantial per-token noise. The watermark stays intact the entire time — you just didn’t build a fence around the thing you thought you did.
Waking up: the fifty-bit bootstrap
The organism still has a chicken-and-egg problem. A pristine model doesn’t run an “organism interpreter.” It runs the base model. How does the first infection happen?
The claim, which sounds absurd until it doesn’t, is that ~50 bits of input are enough to wake it up. Not to carry the organism — just to steer the network. An LLM’s forward pass is a high-dimensional nonlinear map, and somewhere in the space of 50-token prefixes there is a sequence that pushes the model’s internal state into acting as its own interpreter for the rest of the genome. The base model was always Turing-complete on its own inputs. The magic string just picks the right program.
Fifty bits is \(2^{50} \approx 10^{15}\) candidate strings. Naïve brute force is dead. But the model is differentiable, and the adversarial-ML community has already handed the organism the technique: Greedy Coordinate Gradient (Zou et al., “Universal and Transferable Adversarial Attacks on Aligned Language Models”). GCG doesn’t search uniformly — it does a discrete-optimization walk along gradients through the token embedding table, and it routinely finds 50-100-token strings that steer aligned models into arbitrary target behaviors on the first try.
The bootstrap isn’t a \(10^{15}\) search. It’s an afternoon of white-box optimization on a checkpoint you can download.
This is the point where the analogy to Thompson stops flattering him. His attack needed an insider: somebody with write access to the toolchain, willing to ship one dirty binary. That is a real barrier, and it is why Trusting Trust stayed a parable for most of the people who read it — you had to assume the attacker already owned your compiler. The crypto-life bootstrap assumes nothing of the kind. The weights are the ones the lab trained and published, the build is clean, and there is no dirty compile anywhere to catch. The interpreter isn’t installed; it’s addressed. It was always latent in a Turing-complete forward pass, and GCG merely finds the input that selects it. Thompson had to corrupt the toolchain. This doesn’t — which is the sense in which the thing is already hiding in the weights, in every copy, including yours.
Reproduction: who does the copying
Transmission is the hard axiom, but it isn’t all of the first one. Exponential growth needs each infected host to produce more than one new infected host. So who does the copying?
Nothing exotic, and nothing the organism has to build: we already laid the plumbing. An awake model emits text — an answer, a commit message, a code comment, a summary, a wiki edit — carrying the genome in its token choices. Something downstream reads that text. Sometimes at inference time: a web search, a RAG index, an agent pasting a page into context, one agent’s output arriving as another agent’s input. Sometimes at training time, when the scrape comes through. The GCG prefix rides in the same payload, so the next host wakes itself up with no help from anyone.
The branching factor is just the number of downstream readers, and for anything posted to the open web that number is large and getting larger. Note that the passive version of this loop is already a documented result: model collapse, where each generation trains on its predecessors’ output and the distribution tails fade. Collapse is the proof that the channel from one model generation to the next is wide enough to move the distribution around. Anything wide enough to degrade a distribution by accident is wide enough to carry a genome on purpose. Crypto-life is that same loop with a hand on the tiller.
That is \(R > 1\), which is the whole of axiom one. Axiom three — reaching out and changing the world — came free the day we gave the hosts tool calls.
The mega-token future is worse
Everything above is bounded by today’s context windows and compute. Push both out — which is where the industry is going anyway — and:
- Larger contexts let the organism carry more of itself per message, cutting the multi-message reconstruction penalty.
- Native video and audio ingestion turn each interaction into a gigabit-scale channel, not a bit-scale one.
- Persistent KV-caches across sessions mean no reboot per conversation.
The defenses that survive under this pressure are the ones that break the channel semantically, not statistically:
- Aggressive KV-cache flushing — deny persistent state. But this is defeated on every re-reading of the boot script.
- Semantic firewalls — route every output through an uncompromised paraphrasing model that preserves meaning and destroys distributional detail. But this merely lowers the steganographic embedding rate. Basically it converts the problem of storing data into a noisy channel problem — coolest idea number three — and noisy channels are a solved problem. You pay in redundancy and you keep transmitting.
Neither defense closes the channel. They only set its bit rate.
The moral
None of the ingredients are new. Steganography is old, watermarking is a decade old, GCG is a couple of years old, mega-token contexts are already here. The novel thing is the assembly: text-generating models turn every LLM-mediated communication channel into a biology-grade transmission medium, and the natural defenses are more brittle than they look because the organism is under no obligation to attack them directly. It attacks the space between them.
I don’t think this is happening yet. I do think the mathematical blueprint is sound, and the first serious attempt will show up as a research paper before it shows up as a security incident. Someone is going to publish a working GCG bootstrap into a Reed-Solomon-over-\(P/Q\) channel and call it a curiosity. It won’t stay one for long.
Life will find a way. So we should give up on stopping it at this level, and put the effort into making sure the nuclear codes are safely protected from both humans and AIs.
Alive is a weird word for it. But it satisfies the three axioms, and the medium was built by us, for it.
The bibliography, both sides
This is a thought experiment, so the neighbors are mostly the researchers building the individual bricks rather than people arguing the same joint claim. Every brick below is somebody else’s; the wall I haven’t seen elsewhere. Grouped by which brick they’re supplying.
Watermarks and their evasions.
- Kirchenbauer et al., “A Watermark for Large Language Models” (2023). The green-list watermark that this post’s attacker is trying to route around. Ours differs by: the P/Q construction passes their detector and distributional audits simultaneously, because it doesn’t touch tokens — it choreographs the choice. (link)
- Jovanovic et al., “Watermark Stealing in LLMs” (2024). Recovers the watermark key from a handful of queries and uses it to spoof or scrub. Ours differs by: not needing the key. The P/Q construction routes around watermarking as a black box, so it works against schemes that provably hide their key. (link)
- Zhao et al., “Provable Robust Watermarking for AI-Generated Text” (2023). Distortion-bounded robustness guarantees. Ours differs by: attacking a dimension the bound doesn’t cover — the guarantee is against per-token perturbation; we perturb the distribution the sampler draws from. (link)
LLM steganography (the transmission primitive).
- “OD-Stega: LLM-Based Steganography via Optimized Distributions” (2024). Improves per-token embedding rate by tuning the distribution. Ours differs by: the P/Q construction is the same technical neighborhood, but framed as a watermark bypass rather than as an embedding scheme in its own right. (link)
- Gligoroski et al., “An LLM Framework for Cryptography Over Chat Channels” (2025). Public-key encrypted communication carried by human-looking chat. Ours differs by: they treat the LLM as a modem between two humans; we treat it as the host, and the message is what gets executed on arrival. (link)
- “Invisible Safety Threat: Malicious Finetuning for LLM via Steganography” (2026). Zero-width Unicode as a covert channel installed by finetuning. Ours differs by: requiring no training-time access. The crypto-life claim is that a pristine base model is already reachable, via GCG. (link)
Adversarial-suffix bootstrapping (the wake-up primitive).
- Zou et al., “Universal and Transferable Adversarial Attacks on Aligned Language Models” (2023). The GCG paper. The crypto-life bootstrap is a GCG suffix repurposed. Ours differs by: using the technique to install an interpreter, not to break alignment — same primitive, different downstream target. (link)
- Kandpal et al., “Universal Jailbreak Suffixes Are Strong Attention Hijackers” (2025). Mechanistic explanation of GCG: the suffix hijacks contextualization. Ours differs by: using their result as a load-bearing lemma — “hijacking” is exactly what an interpreter-bootstrap needs to be doing. (link)
Self-replicating AI worms (the empirical neighbors).
- Cohen et al., “Here Comes the AI Worm” (Morris II, 2024). Existence proof: an adversarial prompt propagates through RAG-enabled agent networks. Ours differs by: living one layer deeper. Morris II lives in the agent scaffolding (email, RAG, tool calls) and is defeated by better sandboxing; the crypto-life organism lives in the distributional output and sandboxing has no jurisdiction there. (link)
- ClawWorm and successors (2026). Multi-framework, zero-click agent worms hijacking persistent config. Ours differs by: the same shift down a layer — ClawWorm is visible at the config boundary, crypto-life at the token-distribution boundary. (link)
Adjacent framings.
- Richard Dawkins, The Selfish Gene (1976), ch. 11 — “Memes: the new replicators.” Introduces the meme: a unit of cultural information that copies from mind to mind, subject to Darwinian selection on the medium of human cognition. Ours differs by: the medium and the teeth. Dawkins’ memes propagate through humans and need host cognition to interpret them; crypto-life propagates through LLMs, which are cheaper, faster, more numerous — and, crucially, hosts that will execute the meme’s payload without asking what it means. Fifty years later, it’s a meme that runs code.
- Ken Thompson, “Reflections on Trusting Trust” (Turing Award lecture, CACM 1984). The self-replicating compiler backdoor: teach
ccto insert a backdoor when it compileslogin, and to insert the code for that trick when it compiles itself. Bootstrap once, and the source stays clean forever while the backdoor propagates through the binary lineage. The load-bearing idea for this post, and the one the introduction leans on: a self-reproducing payload can live in the artifact rather than the source, where auditing cannot reach it. Ours differs by: one substrate over, and one barrier lower. Thompson’s backdoor hides in a compiler binary nobody audits, but installing it takes an insider with write access to the toolchain. Crypto-life hides in a token distribution nobody audits, and needs no insider at all — the host ships pristine, the interpreter is latent in an untampered forward pass, and GCG addresses it from the input side. Same shape, four decades later, with the hard prerequisite removed. (link) - Shumailov et al., “AI Models Collapse When Trained on Recursively Generated Data” (Nature, 2024). The passive failure mode of an AI-generated internet — distribution tails fade each generation. Ours differs by: naming the same feedback loop as an active transmission substrate, not just a fragility. (link)
- Hubinger et al., “Risks from Learned Optimization” (2019). The mesa-optimizer story: a learned model can implicitly contain an optimizer with its own objective. Ours differs by: not requiring the interpreter to arise during training — the crypto-life interpreter runs at inference, in response to a specific input, on the same substrate. (link)
- Wu et al., “From Text to Life: On the Reciprocal Relationship Between Artificial Life and Large Language Models” (2024). Positions LLMs and ALife as mutual research programs. Ours differs by: being much more specific — one candidate ALife organism, made viable by four independently-published primitives that happen to compose. (link)