|
The way we deliver software has changed, and I donât think our old mental models fit anymore. Code is cheap now, agents open pull requests faster than anyone can read them, and the habits we built around slow, expensive coding are starting to creak. Challenging testersThis is what I talked about on my STARWEST keynote, âTester 2.0: Becoming Indispensable in the Age of AIâ, and you can still watch the virtual sessions if youâre curious how it went. I challenged a lot of assumptions I still see testers and QA professionals hold. I think I might have made people uncomfortable, but I believe itâs now a right time to do that. If youâre wondering what to do about it, Debbie OâBrien may have some good suggestions in an upcoming webinar, Orchestrating agentic test automation with Playwright. OpenAI newsOn a different note, I joined the Codex ambassador program. Iâm very excited about it, and Iâm already planning community events in Slovakia and Czech republic, so if youâre nearby, tell me what youâd like to see. The timing was great, because one of the biggest events - OpenAI DevDay took place on September 29 in San Francisco, with Sam Altmanâs keynote livestreamed. I didnât make it there, but the amount of attention on tools that write code is hard to miss. Jev is everywhereSpeed was also the headline of another release. Diogo Almeida, who says he co-invented ChatGPT, announced Jev, a model from his new company TypeSafe AI. It doesnât write text. It returns typed decisions with a confidence score, and the claims are big: 20 to 200 times faster and 40 to 400 times cheaper than existing models. Those are the vendorâs numbers, so I treated them as a hypothesis and not as a result. I wanted to see how it behaves on something I know well, so I hooked it up to Playwright CLI and used nothing else. Playwright CLI hands Jev a snapshot of the page, and Jev decides which element to interact with. The task was a fuzzy instruction in my Trello playground app: create a new board, a list and a card. I ran the same thing through Playwright MCP, where snapshots and decisions go through an LLM and every step is a tool call. Jev came out about 98% cheaper and twice as fast. Thatâs one task in one app, so please donât read it as a benchmark. What I find interesting is the shape of it. Choosing which element to click is a narrow decision, and a model that only makes narrow decisions, with a confidence score attached, fits browser automation better than I expected. But my friend Jonathan already poked a couple of holes into my assumptions. Software factoriesâQodo 3.0 is an exciting release. Weâve got all the tools in line to help companies create their own software factories. Qodo now groups code into work packages, meaning connected changes across repos, and it adds a map of how those changes affect each other. If youâre not familiar with the concept of software factory, I have made a related video that explains what a software factory is. But a factory with no quality line is just a way to ship bad product faster, so beware of how you set things up. Let me know what do you think about all this. Have you tried Jev yet? Are you building software factories? Do you think testing is about to change? Iâd love to know where you stand |
Sign up for weekly tips on testing, development, and everything related. Unsubscribe anytime you feel like you had enough đ
Hey Reader,, Theo asked a question this week thatâs been sitting in the back of every developerâs head: how much better do the models have to get before you stop reading the code? Itâs been going on for days, but I honeslty love it. We are talking about verification and quality! Although the initial question might not be the right one. Reading was never the goal, itâs one tactic for earning the right to ship something. If reading is the sharpest tool youâve got, read. If you have something...
Hey Reader,, A few years ago Cursor was a nicer place to type code. This week it announced Origin, its own Git competitor built for agent workloads, and got acquired by SpaceX for sixty billion dollars. Somewhere in between, it stopped being an editor and started becoming the whole stack â the place you write, review, merge, and increasingly test your software. I come from QA, so my first instinct isnât excitement, itâs a question: when one tool owns every step of the loop, whoâs the...
âToo dangerous to releaseâ has become its own genre of AI announcement. Project Glasswing is the latest entry: not quite a product launch, but a claim about a threshold, dressed up with enough corporate coalition to signal this one is serious. Anthropic says their new security-focused model, Claude Mythos Preview, can find software vulnerabilities better than all but the most skilled human experts. George Hotz challenged the âtoo dangerous to releaseâ narrative by pointing at the obvious:...