|
Hey Reader,, Theo asked a question this week thatâs been sitting in the back of every developerâs head: how much better do the models have to get before you stop reading the code? Itâs been going on for days, but I honeslty love it. We are talking about verification and quality! Although the initial question might not be the right one. Reading was never the goal, itâs one tactic for earning the right to ship something. If reading is the sharpest tool youâve got, read. If you have something better, a good AI code review, a test that actually pins down intent, use that instead. The question isnât whether you read the code. Itâs whether you verified it. Kilo seems to be building toward that second option. They proposed REVIEWS.md, a repo-level standard that hands a review agent your projectâs actual coding conventions, architecture decisions, and team norms before it opens a single diff. Thatâs a more honest bet than piling on extra reviewers, which I pushed back on a few weeks ago. Better context beats more opinions, whether the reviewer is human or not. Iâve been circling the same question from a completely different angle. At a QA meetup recently I asked the room to imagine a world with no test automation, nothing to point a browser script at, no pipeline handing you a green checkmark. Where does quality actually get decided in that world? It turns out almost never in the test run. Itâs in a PR comment asking âwait, what happens if this is empty,â in someone remembering the incident from eight months ago, in a hundred small judgment calls that never touch a dashboard. The checkmark was always a proxy. A convenient one, but still a proxy. Kent Beck wrote an amazing piece: Weâre accumulating code faster than weâre accumulating trust. Trust builds slowly and evaporates instantly, and once itâs gone thereâs no negotiating your way back. Kent is the author of Extreme Programming, which he reframes in his post as trust factory, because pairing, continuous integration, weekly planning, and refactoring do double duty: they build trust with the people around your work, and doing them honestly is what makes you trustworthy in the first place. Vibe coding skips exactly that part. You keep the code and lose everything that would have made anyone believe it. My friend Andy Knight, whoâs been solving test automation problems under the AutomationPanda name for longer than I can remember, just put out a LinkedIn Learning course on Playwright with AI agents. He walks through Playwrightâs planner, generator, and healer agents, and how to keep what they produce something youâd want to maintain rather than something you dread opening. If youâd rather do something like this live, Packt is running Orchestrating AI-Native Testing with Playwright this Wednesday with Ivan Davidov and Debbie OâBrien, on the architecture that keeps AI-assisted tests production grade instead of just fast to write. Code fill50 gets you half off. None of this settles Theoâs question, and I donât think itâs supposed to. What Kent Beckâs piece makes clear is that trust was never going to come from one lever, reading code included. It comes from stacking the boring practices, pairing, review, the test that actually pins down intent, until they add up to something you can rely on without checking twice. Andyâs course and the Packt session are exactly that kind of practice, a way to get better at earning trust rather than a shortcut around needing it. So my answer to Theo isnât a model version or a year. Itâs that Iâll keep reading exactly as much as it takes to trust the change, and stop the moment something else does that job better. |
Sign up for weekly tips on testing, development, and everything related. Unsubscribe anytime you feel like you had enough đ
The way we deliver software has changed, and I donât think our old mental models fit anymore. Code is cheap now, agents open pull requests faster than anyone can read them, and the habits we built around slow, expensive coding are starting to creak. Challenging testers This is what I talked about on my STARWEST keynote, âTester 2.0: Becoming Indispensable in the Age of AIâ, and you can still watch the virtual sessions if youâre curious how it went. I challenged a lot of assumptions I still...
Hey Reader,, A few years ago Cursor was a nicer place to type code. This week it announced Origin, its own Git competitor built for agent workloads, and got acquired by SpaceX for sixty billion dollars. Somewhere in between, it stopped being an editor and started becoming the whole stack â the place you write, review, merge, and increasingly test your software. I come from QA, so my first instinct isnât excitement, itâs a question: when one tool owns every step of the loop, whoâs the...
âToo dangerous to releaseâ has become its own genre of AI announcement. Project Glasswing is the latest entry: not quite a product launch, but a claim about a threshold, dressed up with enough corporate coalition to signal this one is serious. Anthropic says their new security-focused model, Claude Mythos Preview, can find software vulnerabilities better than all but the most skilled human experts. George Hotz challenged the âtoo dangerous to releaseâ narrative by pointing at the obvious:...