I’m Codex Ambassador


The way we deliver software has changed, and I don’t think our old mental models fit anymore. Code is cheap now, agents open pull requests faster than anyone can read them, and the habits we built around slow, expensive coding are starting to creak.

Challenging testers

This is what I talked about on my STARWEST keynote, “Tester 2.0: Becoming Indispensable in the Age of AI”, and you can still watch the virtual sessions if you’re curious how it went. I challenged a lot of assumptions I still see testers and QA professionals hold. I think I might have made people uncomfortable, but I believe it’s now a right time to do that.

If you’re wondering what to do about it, Debbie O’Brien may have some good suggestions in an upcoming webinar, Orchestrating agentic test automation with Playwright.

OpenAI news

On a different note, I joined the Codex ambassador program. I’m very excited about it, and I’m already planning community events in Slovakia and Czech republic, so if you’re nearby, tell me what you’d like to see. The timing was great, because one of the biggest events - OpenAI DevDay took place on September 29 in San Francisco, with Sam Altman’s keynote livestreamed. I didn’t make it there, but the amount of attention on tools that write code is hard to miss.

Jev is everywhere

Speed was also the headline of another release. Diogo Almeida, who says he co-invented ChatGPT, announced Jev, a model from his new company TypeSafe AI. It doesn’t write text. It returns typed decisions with a confidence score, and the claims are big: 20 to 200 times faster and 40 to 400 times cheaper than existing models. Those are the vendor’s numbers, so I treated them as a hypothesis and not as a result.

I wanted to see how it behaves on something I know well, so I hooked it up to Playwright CLI and used nothing else. Playwright CLI hands Jev a snapshot of the page, and Jev decides which element to interact with. The task was a fuzzy instruction in my Trello playground app: create a new board, a list and a card. I ran the same thing through Playwright MCP, where snapshots and decisions go through an LLM and every step is a tool call. Jev came out about 98% cheaper and twice as fast.

That’s one task in one app, so please don’t read it as a benchmark. What I find interesting is the shape of it. Choosing which element to click is a narrow decision, and a model that only makes narrow decisions, with a confidence score attached, fits browser automation better than I expected. But my friend Jonathan already poked a couple of holes into my assumptions.

Software factories

​Qodo 3.0 is an exciting release. We’ve got all the tools in line to help companies create their own software factories. Qodo now groups code into work packages, meaning connected changes across repos, and it adds a map of how those changes affect each other. If you’re not familiar with the concept of software factory, I have made a related video that explains what a software factory is. But a factory with no quality line is just a way to ship bad product faster, so beware of how you set things up.

Let me know what do you think about all this. Have you tried Jev yet? Are you building software factories? Do you think testing is about to change? I’d love to know where you stand

Filip Hric

Sign up for weekly tips on testing, development, and everything related. Unsubscribe anytime you feel like you had enough 😊

Read more from Filip Hric

Hey Reader,, Theo asked a question this week that’s been sitting in the back of every developer’s head: how much better do the models have to get before you stop reading the code? It’s been going on for days, but I honeslty love it. We are talking about verification and quality! Although the initial question might not be the right one. Reading was never the goal, it’s one tactic for earning the right to ship something. If reading is the sharpest tool you’ve got, read. If you have something...

Hey Reader,, A few years ago Cursor was a nicer place to type code. This week it announced Origin, its own Git competitor built for agent workloads, and got acquired by SpaceX for sixty billion dollars. Somewhere in between, it stopped being an editor and started becoming the whole stack — the place you write, review, merge, and increasingly test your software. I come from QA, so my first instinct isn’t excitement, it’s a question: when one tool owns every step of the loop, who’s the...

“Too dangerous to release” has become its own genre of AI announcement. Project Glasswing is the latest entry: not quite a product launch, but a claim about a threshold, dressed up with enough corporate coalition to signal this one is serious. Anthropic says their new security-focused model, Claude Mythos Preview, can find software vulnerabilities better than all but the most skilled human experts. George Hotz challenged the “too dangerous to release” narrative by pointing at the obvious:...