The System Only One Person Understands
If every critical change starts with the same name, your delivery system is carrying a quiet executive risk.
7 min read
01.10.2026, By Stephan Schwab
Most teams say they want to move faster. Then they make a large change, wait days to learn whether it works, and call the resulting anxiety a quality problem. Kent Beck offered a less glamorous bargain: make the next change small enough to understand and check it while the decision is still fresh. Test-driven development is the best-known part of that bargain, but the point was never a trophy cabinet of tests. It was the ability to keep changing a system without losing control of it.
Kent Beck helped turn test-driven development into a practical working discipline. Its familiar rhythm is simple: write a test for the next bit of behavior, see it fail, write enough code to make it pass, then improve the structure without changing the behavior. Martin Fowler’s account of TDD traces its development through Beck’s work on Extreme Programming in the late 1990s.
The order matters. A failing test proves that the test can detect the missing behavior. A passing test after a small change provides evidence about that change. Refactoring under a passing suite lets the design improve without quietly rewriting the agreement. None of those steps is magic. Together they make uncertainty visible while it is still cheap.
Consider an internal billing workflow. A request says, “Never send an invoice without a customer reference.” The easy move is to add a field, update the form, and discover three weeks later that the batch import still creates reference-free invoices. A better first question is executable: what happens when the batch import receives a record without that reference? Make that case fail. Add the rule at the boundary both paths use. Then test the form and the import together.
The point is not to collect more green checkmarks. It is to find out whether the business rule lives in the system, rather than in a meeting note.
Beck’s career is often reduced to a list of branded ideas: JUnit, TDD, refactoring, Extreme Programming. In his own accounting, he stresses that the work was collaborative and explains how the ideas supported one another. JUnit made tests cheap to write in the same language as the production code. Tests made small design changes safer. Refactoring kept a growing system changeable. Extreme Programming connected that technical discipline to people working together and to users who could respond to working software.
The earlier work was collaborative too. In their 1989 paper, Beck and Ward Cunningham described CRC cards: cheap index cards used to reason about the responsibilities and collaborators of objects. The cards mattered because people could move them, challenge an incomplete design, and change their minds without a formal handoff. Cunningham later carried the same appetite for shared understanding into the wiki and the technical debt metaphor. His story belongs beside Beck’s, not in a footnote to it.
Beck also signed the Manifesto for Agile Software Development. That fact is less useful than the working practices behind it. A team can hold every ceremony on the calendar and still defer its first honest technical feedback until release week. The calendar will be immaculate. The release may not be.
“Small steps” can sound like a plea to slow down. Beck makes the opposite argument. In his comparison of TDD and Kanban, limiting work in progress makes trouble visible sooner. One failing test focuses attention on one needed behavior. A passing test says the code now meets that particular demand. The existing suite checks that the new behavior did not break what already worked.
That rhythm changes how a team handles a difficult requirement. Instead of designing every rule for a future billing platform on a whiteboard, it can take one real invoice scenario, express the expected outcome, implement it, and ask operations where the scenario is wrong. The next example may expose an exception. The design changes while it is still small.
This is where companies that insist they are “only automating a workflow” tend to get into trouble. A spreadsheet becomes an integration. The integration becomes a customer promise. Suddenly the organization owns software behavior whether or not anyone put “software product” on the budget line. Fast feedback is now an operational need, not a developer preference.
Test-driven development can turn into theater too. Teams can write tests around implementation details, mock away the risky integration, and congratulate themselves on a green dashboard. A suite that confirms the wrong invoice rule with perfect reliability is still wrong.
Beck himself has described choosing to fix and ship one defect without an automated test when writing that test would have consumed hours of investigation. In his account, the choice depended on the kind of product and the feedback he needed at that moment. This is more useful than a commandment that every conceivable case must be tested the same way.
Use the smallest check that answers the next real question. Sometimes that is a unit test. Sometimes it is an integration test, a conversation with the person who reconciles invoices, or a careful production experiment. The discipline is to keep the question and the evidence close together. Recovery still matters when a test cannot anticipate reality.
An AI coding agent can produce a plausible invoice flow in minutes. That reduces typing time. It does not tell the team whether the customer reference belongs on every invoice, whether legacy imports are exempt, or who has authority to decide. Faster implementation simply lets an unresolved decision spread through more code.
Beck’s approach gives people a way to keep judgment in the loop. State a concrete behavior. See the current system fail it. Let a developer or agent make a small change. Check the result. Refactor while the evidence is still available. Then ask the next question. The same discipline that helped teams avoid a large speculative design can stop AI-generated code from turning one guessed requirement into five integrated guesses.
That does not mean a prompt containing “use TDD” is a safety system. Tests beat instructions for AI coding agents when they express actual expectations and run against the change. Someone still has to select the expectations, notice missing cases, and keep the system understandable.
A CTO cannot demand confident releases while rewarding only visible feature output. If developers must defend every test as a delay, they will learn to postpone the checks. If a team cannot revise its design without a separate political process, it will learn to protect a bad design. The resulting “speed” is a loan paid back during the next incident.
Give the team a short route from business question to executable example to working change. Keep the people who know the exceptions available while the software is being built. Run the tests on every change, and make room to improve the structure when the tests expose friction. Ask what the team learned this week, not how many tickets it moved.
Beck’s lesson was never that every change must be tiny forever. It was that a team earns the ability to make consequential changes by building trust through frequent evidence. That trust is made one small, honest step at a time.
Tell me what is happening. I listen, ask a few practical questions, and reflect back what I see: where the risk may sit, what may be blocking delivery, and what looks worth checking next. No pitch, no obligation. Confidential and direct.
Talk it through. Practical reflection, no pitch.
Start a ConversationA senior developer for your team
Our Embedded Delivery Partner writes production code with your team, improves the pipeline, and accelerates delivery. 60-70% coding, 30-40% coaching. A temporary teammate who ships from day one.