Two-eyed tools: pair programming when one of you can’t see (May 2026)
On building software for partners with disjoint perceptual stacks.
We’re building a verified coding agent in Lean 4. The agent runs tools, the user watches, the LLM and the user trade ideas. Last week the statistics library we depend on (lean-stats) shipped a small architectural revolution: panels are observers of a shared point-state model, with cross-panel highlighting via WebSocket. Click a point in any cell of a scatterplot matrix and it lights up in every other cell. The pattern is a direct descendant of Luke Tierney’s xlispstat from 1990.
That work made me notice something I’d been missing about how the agent and I actually work together.
The pattern, briefly
The xlispstat insight: every plot is an observer of a shared data model. Selection, exclusion, hiding — these aren’t local to a panel; they’re properties of points themselves, and points have stable identity across views. When you brush a region in one panel, the message select(ids) updates the shared state, and every other observer re-renders with the same selection highlighted.
The cynical reading: this is just MVC.
The honest reading: it’s a species of MVC distinguished by two characteristics. First, many co-equal views with no privileged controller. Second — and this is the load-bearing part — identity is part of the model. Point 7 has an identity. The selection set is Set PointId, a thing all views consult. Linked plots work because they all share the same identity space.
The pattern fits anywhere with stable identities and interactive observation: code editors, vector graphics tools, DAWs, debuggers, parametric CAD, reactive notebooks. It doesn’t fit tight inner loops, batch pipelines, or anywhere identity is unstable. None of this is news to anyone who’s used Smalltalk.
But the question that got my attention wasn’t “where else does this fit.” It was: what happens when one of the observers is an LLM?
The asymmetry
Pair programming as a discipline assumes both partners can see the screen. With an LLM partner, drop that assumption. You and the LLM have radically different perceptual stacks:
- Working memory. Yours: about seven items, but with three and a half billion years of pattern recognition compensating. The LLM’s: hundreds of thousands of tokens, but no visual gestalt.
- Scan rate. You absorb a scatter plot in 200ms. The LLM absorbs 30,000 lines of code in seconds. Each of you would suffocate at the other’s input rate.
- Anchoring. You hold context spatially: this function is near that one in the file tree. The LLM holds it by similarity: this function is like that one across embeddings.
- Fatigue. You get tired. The LLM doesn’t, but its context degrades.
When I plot a scatter plot, the LLM is at a disadvantage. My visual system catches outliers, clusters, trends, and heteroscedasticity in a single glance — pre-attentively, without effort. The 15-number summary we send the LLM is brilliant engineering — counts, quartiles, slopes, residual diagnostics — but a paltry shadow of what I see.
When the LLM browses a 30,000-line codebase, I’m at the disadvantage. It can hold “every place a function is called” without effort. It can notice that two functions have similar structure across the project. It can enumerate every case of a closed inductive. I can do none of these without grinding.
This is what manifests are doing for me. They project some of the LLM’s consistency advantages — names, types, theorems linking functions to claims — into a form my human cognition can audit at a glance. I don’t read 30,000 lines; I read the manifest. The kernel reads the lines on my behalf.
The general principle
Both partners have radically different perceptual apparatuses. Both partners are first-class collaborators. Neither should be left behind.
Every artifact in the shared model needs at least two projections: one optimized for visual cognition, one for symbolic cognition. Where a single projection serves both, it’s a gift.
The 15-number summary works because I can also read it. It’s a degraded version of my visual experience, but I don’t lose anything that wasn’t already in the data. The LLM gains; I don’t lose.
The reverse — a pixel-perfect rendering of a plot — IS lossy for the LLM. We need the LLM’s analog of the 15-number summary, going the other direction. We barely have any of these.
What’s missing on the LLM side, that I see immediately:
- Outliers. Pre-attentive for me. The LLM gets “max distance from regression line = 4.2σ,” which is reasoning about the outlier, not perceiving it.
- Cluster structure. I see three blobs. The LLM gets WSS ratios and reasons from them.
- Trend shape. “It curves up and plateaus” is one human glance. The LLM needs piecewise regression and inflection- point detection to say the same thing.
- The shape of an error message. A wall of red text vs. one tidy red line — I read the urgency before the words. The LLM reads only the words.
- The aesthetic of a project. I tell well-organized from messy in five seconds of
ls -R. The LLM gets the same bytes but no gestalt.
What’s missing on my side, that the LLM gets for free:
- Cross-file consistency. Every call site of every function, always.
- Token-level structural similarity. “These two modules are shaped the same way” without reading both.
- Exhaustive enumeration. Every case of every inductive, always closed.
- Lexical recall over the conversation. What you said 100 turns ago, exactly.
- API surface awareness. What stdlib functions exist, before Googling.
Manifests project the second list into a form I can use. Nothing yet projects the first list into a form the LLM can use, beyond the summary numbers we already send. There’s headroom.
What this looks like in code
Each tool’s output should have at least two render channels:
| Channel | Audience | Optimized for |
|---|---|---|
renderResult |
LLM | Token efficiency, structured fields, named values |
renderUserVisible |
User | Pre-attentive perception, layout, color, shape |
renderShared |
Both | The rare case where one form serves both |
Today our regression tool’s LLM-facing summary is “n=32, R²=0.74, slope=−3.88, p<0.001.” Clean. Useful. But the human gets a graphical view at a URL, and the LLM gets the summary, and neither sees the other’s. The human can read the summary; the LLM can’t open the URL.
The fix isn’t to merge the channels. It’s to make each channel strictly contain the other’s information plus the projection- specific gain. The LLM-facing summary should describe what’s visible in the plot (“residuals show heteroscedasticity, two outliers at indices 14 and 27”). The human-facing view should include the summary numbers prominently on the page. No partner learns anything the other doesn’t get to learn — just in their preferred form.
The discipline is symmetric: whenever one partner is seeing something the other isn’t, that’s a bug in the projection.
Shared projections — the rare wins
Some artifacts have the property that one form serves both partners. Diffs are the canonical example: humans read them at a glance; LLMs parse them perfectly. No projection needed; no information loss. Pure win.
Other candidates worth noticing:
- Tables with sparklines. Numbers carry the precision; the sparkline carries the shape. Both partners read the same row.
- Annotated plots. A scatterplot with text callouts saying “outlier at index 27” or “elbow at k=3.” The LLM reads the callouts; the human sees the plot; the words and the picture point at the same thing.
- AST diffs in dual form. Tree representation for the human; edit-script for the LLM. Same change, two projections.
- Manifests. Doc-comment for the human; type for the kernel; structured text for the LLM. One artifact, three readings, all honest.
The manifest format isn’t just documentation. It’s a successful design of a shared cognitive surface — like diffs and tables with sparklines. They’re rare. When we find one, we should notice and not waste it.
I think there are more of these than we currently exploit. The search for them is its own design discipline.
The risks
Two honest caveats.
The two-channel approach can become makework. Not every tool’s output benefits from two views. A git commit succeeded; the one-line confirmation works for everyone. Adding a graphical commit-confirmation panel would be silly. The pattern fits where the data has intrinsic dimensionality — a regression has structure that benefits from visual projection; a commit hash doesn’t.
The two channels can drift apart. If the human-facing view says “the data looks linear” and the LLM-facing summary says “R² = 0.31,” one of them is lying. Keeping them honest is real maintenance cost. The way lean-stats handles this is instructive: when the user brushes a SPLOM, the WebSocket forwards the same human-readable text the panel would show (“now 12 selected, x range 2.1–4.7”) to the LLM. The protocol becomes the keep-them-honest layer. Both partners hear the same words about the same event.
There’s also a deeper risk. The metaphor “the LLM perceives text the way you perceive plots” is partly true and partly flattering. A 15-number summary IS perception for the LLM in a way that mirrors how a plot is perception for me — sign patterns, correlations, magnitudes. But the LLM doesn’t have the equivalent of a peripheral-vision warning system. When I brush a SPLOM and notice “huh, that cluster came apart,” the LLM doesn’t get the “huh.” It gets whatever message we decide to send. We have to pre-decide what’s worth saying. My visual system pre-decides constantly without effort.
That gap is real. Designing around it is part of the work.
What this means for our agent
We’re partway down this road already, mostly accidentally. Manifests are a shared projection. The recent push to use WorldClaim — environmental assumptions threaded as explicit hypotheses through the theorems that depend on them — turned several “trust me” axioms into named, debuggable, falsifiability- documented claims. That’s the LLM’s consistency advantage projected into a form I can audit.
Going the other direction is what we haven’t pushed on. The agent’s tool surface for plots returns visual artifacts the LLM can’t see. Stats’s WebSocket protocol — emitting human-readable event descriptions — is the right shape; we should match it on our side. Every interactive view should be feeding both partners.
The pattern’s substrate is xlispstat-style: a shared model, multiple observers, atomic typed messages, identity baked in. Both partners are observers. Neither is privileged. The protocol becomes the contract that keeps the projections honest.
Pair programming, but with a partner whose visual cortex is in a different building. You wouldn’t dictate a plot to a blind collaborator and call it pairing. You’d both look — at whatever each of you can see — at the same artifact. We owe the LLM the same courtesy.
This is a thought piece, not a finished design. The next move is concrete: identify the tools that currently render to one partner only, and add the missing channel. Stats’s Plot.Protocol — typed inductive for panel events — is the template. If the discipline survives a few months of accreting tools without the protocol churning into a versioning headache, we’ll know it works at l3m’s scale. If not, we’ll know the abstraction was too ambitious and we’ll back off to ad-hoc two-channel rendering. Either way, the framing — that pair programming with disjoint perceptual stacks is its own design problem — is durable. The implementation question is what substrate makes it expressible. The xlispstat pattern is our current bet.