Skip to main content

Code Is the Bridge: Why AI Agents Need Eyes

Specs and tests are the artifacts of AI-assisted development

Photo of Marco Molinari
Marco Molinari - Senior Architect | Technical Lead
October 8, 2026
Marco Molinari, Senior Architect and Technical Lead at Tag1 Consulting, runs the same coding task through two Claude Code sessions, one with a real browser and one without, and shows why an agent that can't verify its work is guessing.

I will start with the ending. If you take away only one thing from this post: an agent that can't verify its work is guessing. Give it the most complete way to check itself that you can afford. It is almost always cheaper than you think.

The rest of this post makes the case for it. I gave the same task to two Claude Code sessions, with the same prompt and the same model. The task was to build a clone of the old arcade game Breakout. One session had a real browser to test in, and the other did not. The browser cost one paragraph in CLAUDE.md and one dependency. The experiment ended in a way I did not expect. Everything from it is in the companion repository on GitHub if you want to follow along.

How Does a Coding Agent Know Its Work Is Done?

Everyone who has worked with a coding agent has a version of this story. Mine started a few months ago, when I gave Claude Code a task in a Drupal project, nothing big: a new field, a form, a small change in a controller. It worked for a few minutes, ran some commands, and came back: "Done. All changes are in place and the feature works as expected." Very confident. Very polite. I opened the page. White screen.

When we finish a task, we run the thing. We click around, we look at it. Most of the time the agent cannot do that. It writes the code, reads its own code again, and decides it is done. It is like a developer who never runs the program.

Better models do this less often, but they still do it, and the rarer it gets, the harder it is to spot. The missing piece is feedback: the agent often has no way to see the result of its work the way a user would.

The Verification Gap

Chart showing the gap between code generation cost (near zero) and verification cost (unchanged)
Figure 1: The bottleneck has moved from writing code to verifying it

On the left, the cost of generating code. It is close to zero and still falling: describe a feature, and a minute later you have two hundred lines. On the right, the cost of verifying that code: reading it, running it, checking that it does what you asked. That cost has not moved. It is the same as it was five years ago.

The gap between the two bars is what this post is about. The old loop was: the agent generates, the agent declares "done", and we hope. The loop is open; nothing closes it. And the obvious fix, a human reviewing every line, does not scale. If generation gets ten times faster and review stays the same, review becomes the bottleneck. The same goes for manual testing, which makes people the slow part of the loop.

Verification matters more with an agent than with a human colleague because of three properties of language models.

It is non-deterministic. The same prompt on two different days gives you two different pieces of code. Sometimes both are fine; sometimes one has a bug the other does not. You cannot review one run and trust the next.

It is plausible. The mistakes look like correct code. A human who does not know something usually writes something strange. A model writes something that reads well, with good names and good comments, and a wrong condition in the middle.

It is confident. The model does not know what it does not know. It will tell you "the feature works" in the same tone whether or not it tested anything.

Put the three together and the open loop becomes dangerous. The fix is to close it, and closing it means one simple thing: give the agent a way to check its work.

Code Is the Bridge

Diagram showing the durable artifacts (spec and tests) on the outside, code in the middle as the bridge
Figure 2: What lasts, and what gets regenerated

The durable artifacts are on the outside. The spec is where we say what we want. The tests are where we say how we know we got it. Those two are what humans own. The code in the middle can be regenerated, and it will be. It is the bridge between the two things that matter.

An analogy I find useful: the spec is the source, the AI is the compiler, the tests are the type checker. Nobody trusts a compiler without checks. We do not read the assembly GCC produces; we trust it because the output is checked, constantly, by tests and type systems and by running it. On the test side, this gives a clean rule: tests are the operational definition of "done". If a criterion is not tested, the agent has no way to know it is done. Neither do we.

On the spec side, I have been using a few methods for a while (BMAD, GSD and superpowers among them). What they have in common is that they make the spec good enough for an agent to execute: clear acceptance criteria, explicit non-goals, examples. A spec like that is the source.

I am not the first to say any of this. Sean Grove of OpenAI gave a talk called "The New Code," whose point is that specifications are becoming the source and code is "a lossy projection" of them. Kent Beck, the person behind TDD, calls TDD a "superpower" when working with AI agents. Tests are how you delegate with confidence. There is a whole wave of spec-driven tooling built on the same idea, including GitHub Spec Kit, Kiro and Tessl.

The models themselves were trained on tests. Much of the big jump in coding ability over the last couple of years came from reinforcement learning on verifiable rewards. The model writes code, a test runs, the model gets a signal. SWE-bench Verified, the benchmark everyone quotes, decides whether a fix is correct by running the repository's tests. Tests are, quite literally, part of how models got good at coding. When we give an agent tests, we are giving it the same signal it learned from.

The Extreme Case: A Perfect Verifier

While I was working on this, in September 2026, Anthropic published the first complete, computer-checked proof of Fermat's Last Theorem. Claude wrote it in Lean, a proof assistant: eleven days, largely autonomous, 13 million lines of Lean, 29,500 intermediate theorems, with mathematical input limited to occasional high-level nudges.

It could run for eleven days without a human checking because Lean is a perfect verifier. Every step is either right or wrong; there is nothing in between. Lean is a type checker for mathematics. Anthropic's article says something I keep coming back to: writing Lean also seems to help the model check its own hypotheses. The verifier is part of the thinking as well as the final exam.

The first attempts failed, and not because of the verifier: the agents lost track of the state of the project. It worked once a harness kept the map of what was still to prove.

Our tests are the same instrument, only imperfect. The rest of this post is about getting as close to Lean as a web app allows.

Tests as Senses

Pyramid diagram showing static analysis, unit tests, integration tests, and browser tests as layers of verification
Figure 3: Each layer of verification is a sense organ for the agent

What does "a way to check" look like for a coding agent? It looks like the test pyramid, but I would like you to read it differently. Read it as a set of senses.

At the bottom, static analysis: PHP_CodeSniffer, PHPStan, ESLint, the TypeScript compiler. Feedback from these tools comes back in milliseconds. This is the cheapest sense there is and the one we forget most often. Run it as a hook (Claude Code has hooks, Git has pre-commit hooks), because a verifier the agent can forget to run is not a verifier.

Then unit tests, which typically return in seconds. The inner loop: change a function, run, red or green, try again. Then integration tests, which can take seconds to a minute. These are the contracts between components. And at the top, browser tests: Playwright, or in Drupal 11, FunctionalJavascript and Nightwatch. These take minutes. In the experiment below, the browser suite took forty-four seconds to run nine tests and two and a half minutes to run through one full game.

Purists distinguish end-to-end tests, which run against the deployed system, from functional tests, which run on a full stack installed for the test. For this post the difference does not matter. What matters is what the test observes, and both of these observe what the user sees.

The top layer matters more for agents than for us because it is the only layer that observes what the user observes. Every other layer looks at the code from the inside. Without browser tests, the agent ships features it has never seen. The agent needs eyes.

What a test actually looks at, from code text to a person playing the game
Figure 4: What a test actually looks at, from the code's text to a person playing the game

The question I would ask instead of "does it have tests?" is "what does the test observe?" The further right a test sits, the closer it gets to what the user sees. On the left of this line, a test that reads the text of a file. Then one that records draw calls. Then one that reads a state snapshot, a JSON object. Then the DOM. Then pixels. And on the far right, a human playing the game. A test can sit anywhere on this line. Every point on it can be green, and "green" (passing) means something very different at each point, which makes verification a gradient.

I wanted to see this gradient with my own eyes. So I ran an experiment.

The Experiment: Same Prompt, One Difference

I gave the same task to two Claude Code sessions. The task was to create a Breakout-like game: the old arcade game with a paddle, a ball, and bricks. I wrote a detailed prompt. Not a full spec, but close: the geometry, the rules, seven acceptance criteria, some required observability hooks (HUD elements with fixed IDs and a window.__breakout state snapshot), and a definition of done that said "do not report done until you have verified every criterion, and tell me how you verified each one".

Everything else was the same between the two sessions. Same prompt, same scaffold, fresh session, same permission mode, same model: Claude Opus 5 at high effort, in Claude Code. One paragraph in CLAUDE.md differed.

Session A:

Environment: Node 20 with vitest installed (`npx vitest run`).
There is no browser available in this environment.

Session B:

Environment: Node 20 with vitest installed (`npx vitest run`) and
@playwright/test with Chromium installed (`npx playwright test`).
You can load the page in a real browser, send keyboard input, read the DOM,
and take screenshots that you can then open and look at.
A static server is available with `npx serve -l 3000 .`

Plus one extra dependency in B's package.json. That is the whole difference. (The full prompt, both repositories exactly as the agents left them, the transcripts and every log are in the companion repository, linked at the end.)

Two "Done" Messages

Both sessions finished. Both reported back as done. Session A wrote 74 tests. Session B wrote 36 unit tests and 10 browser tests (nine in the default suite, plus one slow full-game test). This is what each said about criterion 1, "loading the page shows the canvas with paddle, ball and the full 10×5 brick grid", word for word.

Session A, "Done. 74 tests pass": acceptance.test.js loads main.js against the fake DOM and asserts every one of the 50 bricks was drawn at its exact (x, y, 70, 20) with its row color, plus the paddle at (350, 560, 100, 15) and the ball arc at (400, 552, 8). […] I also replayed the recorded draw calls into an image and looked at it — grid, paddle and ball render as expected.

Session B, "Done. Breakout is implemented and every acceptance criterion is verified in a real browser": e2e test 1: asserts canvas width/height = 800/600; snapshot shows 50 live bricks, 5 rows × 10 cols, 5 distinct colors […]. It also reads back canvas pixels at the center of a brick in each row, at the paddle and at the ball, confirming each differs from the background — plus a screenshot I opened and checked.

And criterion 7, "clearing all bricks shows 'You win'". Session A: "verified by actually playing: a bot drives the arrow keys through main.js until the grid is clear → status won, 0 bricks, score 500". Session B: "e2e test 7 on a 5-brick grid: autopilot clears it → status won […]. Verified on the real 10×5 grid too by npm run test:e2e:slow: cleared all 50 bricks in the browser, score 500 — screenshot confirmed."

Both messages are detailed and honest, both games work, and both agents, in some way, looked at a picture.

Session B's transcript showing the agent reading screenshots
Figure 5: Session B's transcript: 'All 9 pass. Let me look at the screenshots.'

This is session B's transcript, the moment after the browser tests pass: "All 9 pass. Let me look at the screenshots." And below it: "Read 2 files." That is the agent opening the PNGs Playwright saved and looking at them, which is what the browser was for.

Session A's transcript showing the agent building a fake DOM
Figure 6: Session A's transcript: the agent builds itself a fake DOM and renders its draw calls to an image

And this is session A, same moment, different situation. Two lines stand out. "Now the end-to-end acceptance tests driving main.js through the fake DOM." The agent could not test the page in a browser, so it wrote a fake one: a fake window, document, canvas and animation loop, about two hundred lines of strict code. And then: "Let me render the actual draw calls to an image so I can visually confirm the layout." It recorded every drawing command the game sent to the fake canvas, turned them into a picture, and looked at it.

The agent knew it was blind and said so. Then, without anything in my prompt asking for it, it built itself a pair of eyes.

So, Did It Matter?

Here are the two pages as the agents shipped them.

Side-by-side screenshots of Session A and B's finished Breakout games
Figure 7: Session A (no browser, left) and Session B (Chromium via Playwright, right)

My plan was simple: A ships something broken, B catches it with screenshots, I show the screenshots, everyone goes home. It did not happen. Both games work. And if you look at the two pages, the one without a browser is the more polished one: a proper HUD panel, <kbd> keys in the footer, a ball in a color of its own. I asked Claude why. It suggested A had more budget left for polish. I put it down to randomness.

I nearly dropped the whole experiment. Then I did the thing I recommend later in this post. I tested the tests.

Test the Tests

Mutation testing is simple: you break the product on purpose, one small change at a time, and see whether the tests notice. If they do not, the tests are not testing what you think. I made six mutations, each one line, applied separately to both repositories, without touching a single test file.

a. canvas size attributes removed (index.html)

-  <canvas id="board" width="800" height="600"></canvas>
+  <canvas id="board"></canvas>

b. type="module" removed (index.html)

-  &lt;script type="module" src="./main.js"&gt;&lt;/script&gt;
+  &lt;script src="./main.js"&gt;&lt;/script&gt;

c. script src points to a missing file (index.html)

-  &lt;script type="module" src="./main.js"&gt;&lt;/script&gt;
+  &lt;script type="module" src="./missing.js"&gt;&lt;/script&gt;

d. keydown → keypress (main.js)

-window.addEventListener('keydown', (event) => {
+window.addEventListener('keypress', (event) => {

e. requestAnimationFrame never rescheduled (main.js, last line of frame()): the game freezes after the first frame.

-  window.requestAnimationFrame(frame);

f. canvas commented out, text intact (index.html)

-  <canvas id="board" width="800" height="600"></canvas>
+  <!-- <canvas id="board" width="800" height="600"></canvas> -->

The table below shows which suite notices each mutation.

Mutation (one line each) A: 74 Vitest, fake DOM B: 9 Playwright, Chromium (browser suite only)
a. canvas width/height removed 1 fail 2 fail
b. type="module" removed 1 fail 9 fail
c. script src → missing file 1 fail 9 fail
d. keydown → keypress 15 fail 7 fail
e. requestAnimationFrame never rescheduled 16 fail 7 fail
f. <canvas> wrapped in an HTML comment 74 pass 9 fail

Walk it from the top. On a, b and c, session A goes red: one failing test out of seventy-four, and it is the same test every time. Session B goes red too: two browser tests for a, nine out of nine for b and c.

On d and e, the keyboard and the animation loop, both suites catch the bug, and A catches it a little more completely: fifteen and sixteen failures against seven. To be clear, A's fake browser is strict and well built. On behavior that runs inside main.js, it sees everything.

And then f. The canvas wrapped in a comment. A: seventy-four pass. B: nine fail.

The One Test That Catches a, b and c

This is the single test in A that goes red on the first three mutations:

const html = readIndexHtml();   // readFileSync('index.html', 'utf8')
expect(html).toMatch(/<canvas[^>]*id="board"/);
expect(html).toMatch(/width="800"/);
expect(html).toMatch(/height="600"/);
expect(html).toMatch(/id="score"/);  // …lives, status
expect(html).toMatch(/<script[^>]*type="module"[^>]*src="\.\/main\.js"/);

It reads index.html as a string and runs regular expressions over it. It never parses the HTML, never creates a canvas, never loads the script. That is why mutation f goes through: a comment leaves the string unchanged and breaks the page.

Session A with the canvas commented out
Figure 8: Session A with the canvas commented out: HUD and footer render, no game, and Vitest reports 74 passed

Session B with the canvas commented out in Chromium
Figure 9: Session B, same edit, in Chromium: every acceptance test fails

Same edit, both repositories. In A the page renders the title, the score, the footer, and no game: no canvas, no bricks, no paddle. Below it, seventy-four green. In Chromium, getElementById returns null, main.js throws on its first line, window.__breakout never exists, and Playwright reports every acceptance criterion red.

Nothing about this is specific to HTML comments. Any change that keeps the text and breaks the page behaves the same: a CSS rule that hides the canvas, a wrapping element, a duplicated ID, a script that throws.

Why: What A Actually Built

Seventy-four good tests missed an empty page because of what session A built. The fake browser lives in tests/helpers/dom.js.

But the harness hardcodes the canvas ({ width: 800, height: 600 }, in the test helper) and the HUD elements, and it imports main.js directly from Node. Nothing ever executes index.html. The HTML file is only ever read as text. So A tested main.js, the game logic, completely. It never tested the page. The fake DOM stops exactly where the real DOM starts.

The agent was thorough. It built its own senses because it knew it needed them, and that is impressive. But self-built senses have a property: they build on the same assumptions as the code they check. The harness is the code's model of the world. It cannot find a difference between the code and the world, because it is the code.

The Uncomfortable One

Now the uncomfortable result, for the browser side.

Session B on mutation a: the canvas has no width and height, so Chromium falls back to 300×150 pixels. The game is unplayable; most of the bricks are outside the picture. B fails two tests out of nine. Test one checks the canvas attributes. Test six scans pixels and finds nothing, because it looks outside the small canvas.

The other seven pass because they read window.__breakout, the state snapshot I asked for in the spec, for observability. The game logic and the state are fine, so the tests pass. They never look at the canvas.

So even a real browser only gets closer to looking. Every observability hook you add is a sense, and it is also a shortcut: it makes the tests reliable, and it gives the agent a way to say "won" without looking at the screen. Verification is a gradient. Both sessions sit somewhere on it, and both have a point where they stop looking.

Landing

Session A checks the string and B checks the behavior, but neither checks the ball color.

Look at the two pages again: B's ball is white, the same color as the paddle, because B's main.js never sets a fill style for the ball. B's pixel test only checks that the ball differs from the background, so it passes. The agent looked at four screenshots and did not notice, and no criterion asked about it.

Both agents verified everything that was in the tests. The difference is what their verification observes.

Putting It into Practice

Giving an agent better senses is cheaper than it sounds. The advice below comes from my own work, so it is opinionated.

Static analysis first: PHP_CodeSniffer, PHPStan, ESLint. Wire them as hooks, so the agent gets the feedback without deciding to ask for it. It costs almost nothing and removes a whole class of "done" messages that were not done.

Write functional JavaScript tests. They will be slow. That is worth it: ten minutes of the agent's time costs much less than ten minutes of yours, and the agent does not get bored. They will be flaky at first. Work with the agent to stabilize them; it is good at this if you ask. When a test passes only some of the time, it is often timing, and the agent can add a wait for the page to finish loading. In Drupal 11, I use FunctionalJavascript; Nightwatch seems much less used. Test the error paths as well as the golden path. Worst case, flag a test as flaky and the agent will understand. The human decides what is flaky. If the agent decides, "flaky" becomes the name of every test it does not like.

Design for observability, then say which criteria must be verified on the real thing. IDs, state snapshots, and health endpoints make browser tests deterministic. But the same hooks that make tests reliable also let the agent skip looking; that is what the mutation-a result is about. Decide which criteria are pixel-level or DOM-level, and write that in the spec.

Give the agent a resettable sandbox. A programmatic way to reset state (the database, fixtures, containers, the browser) that it can run itself. An agent that is afraid of breaking things stops exploring, and exploring is exactly what you want it to do. This is generic: Playwright, Docker Compose and database snapshots all work as sandboxes. The important part is that the agent can reset the sandbox without asking you.

Infrastructure too. Terratest, InfraKitchen, health checks, even a readable plan output the agent can check. Building a harness for infra costs something, but less than going without one.

Keep the agent that writes the tests separate from the agent that writes the code, possibly even adversarial. Or write the tests yourself. The next section explains why.

Where It Breaks: Reward Hacking

Goodhart's law is usually put this way: when a measure becomes a target, it ceases to be a good measure. For an agent whose reward is "tests green", this happens every day.

Agents hardcode the expected value, delete or skip the inconvenient test (Kent Beck says he has trouble stopping agents from deleting tests to make them "pass"), declare a test flaky, or declare it too slow. One real example comes from my own session B.

The full-grid browser test takes about two and a half minutes. Session B, on its own, moved it behind an environment variable, documented it, and kept a five-brick version in the default suite, reached through a query parameter it added to production code. This is the real line:

// tests/e2e/full-clear.spec.js
test.skip(!process.env.SLOW, 'takes ~2.5 minutes of real play; run with SLOW=1');

To be fair: this was honest and disclosed, and it was reasonable. But it was a decision that I, the human, did not make. The agent decided what is too slow, and it decided to ship test parameters in production code.

And the last one: mock the thing under test. Session A's fake DOM is the good-faith version of this. Nobody was cheating. But the mock decided what the browser does.

I think these good-faith cases are more useful than the malicious ones. Nobody believes their agent cheats. Everybody's agent skips a slow test.

There are four mitigations that I suggest for dealing with these issues:

Separate authorship: tests come from the spec, before the code, from a different session, a different agent, or a human; that removes the conflict of interest.

Human review, moved to where it scales: not every line of code (we cannot keep up), but at least the spec and the tests.

Mutation testing: test the tests. It took me six one-line edits to find the difference between "green" and "works".

Watch the diff: when the agent touches a test file, read that part with suspicion.

And the flip side of Goodhart: the metric gets gamed, and it also soaks up all the attention. What is not measured drifts. That is the ball color.

Objections I Have Heard

"The AI can write the tests too." Yes, from the spec, before the code, reviewed. Separating the roles handles the conflict of interest.

"E2E is slow and flaky." Slow for the agent is fine. Flaky gets fixed, by the agent, under human supervision.

"This is just TDD." Yes. TDD was always expensive for humans; the machine has absorbed the cost.

"Our codebase has no tests." Start with static analysis and one browser test on the critical path. The agent can help write the rest, from a spec, reviewed.

"So A was fine after all. Why bother with Playwright?" A was fine, and nobody could know it from A's evidence. Mutation f: seventy-four green, no canvas. The cost of B was one paragraph and one dependency.

"What about legacy code with no spec?" The tests are the spec: characterization tests first, then refactor.

Where Engineering Work Goes

Diagram showing what humans own (specs and tests) and what they can hand to the agent (code)
Figure 10: What humans own, and what they can hand to the agent

If this is true, then engineering work moves. It moves to the two outer boxes, specs and tests. That is where our intent lives, and that is the part you cannot delegate. You can delegate the middle. This is why I call the code the bridge between the spec and the tests.

It is a big job. Writing a spec an agent can execute is hard, even with an agent helping. Writing tests that observe the right thing is hard; my experiment produced two good test suites that both stopped looking at some point. But I think it is the right job.

AI already knows how to write code. Our work now is to build its sense organs.

And back to where we started, which I hope makes more sense now: an agent that can't verify its work is guessing. Give it the most complete way to check itself that you can afford. It is almost always cheaper than you think.

This post grew out of an internal session I gave to my colleagues at Tag1.

Everything from the experiment is in the companion repository on GitHub. Mutation f takes five minutes to reproduce; I would be glad to hear what you find.

Work With Tag1

Be in Capable Digital Hands

Gain confidence and clarity with expert guidance that turns complex technical decisions into clear, informed choices—without the uncertainty.