Natural Language Test Automation: Karim Jouini on TestGuild

What if your automated tests weren’t code anymore? In episode 578 of the TestGuild Automation Podcast, Joe Colantonio sits down with Karim Jouini, our CEO and co-founder, to separate hype from reality in AI test agents. Karim makes the case for an unusual approach: natural language test automation, where every step written in plain English is interpreted at run time, with no Selenium or Playwright script generated behind the scenes. He explains why that removes most of the maintenance work, what humans stay in charge of, and puts numbers on two customer projects, including a European bank that cut a three-year testing effort to four months. By the end of this page, you’ll know what an AI test agent actually does, and what it doesn’t.

Episode card

Podcast
TestGuild Automation Podcast
 - episode
A578
Host
Joe Colantonio
Guest(s)
Karim Jouini, CEO and co-founder of Thunders
Published
February 24, 2026
Duration
42 min
 ·
English
Listed
YouTube
 ·
TestGuild Automation Podcast

Watch the episode: what an AI test agent really does

This episode is for testers, automation engineers, QA leads and DevOps leaders who see a new “AI” testing tool every week. If you only have ten minutes: the opening demo (Joe runs a test on his own site from nothing but a URL), the “there is nothing underneath” explanation at 19:07, and the bank case study at 34:59.

Chapters

  • 00:00 – Intro and demo: a test generated from a URL
  • 02:20 – From Microsoft to Expensya: why testing
  • 05:19 – The four dimensions of the QA problem
  • 08:37 – Manual vs automated testing: “a manual tester on steroids”
  • 11:41 – Test agent or QA assistant? The jobs question
  • 13:44 – Module 1: from spec to test plan
  • 14:53 – Module 2: “Selenium in plain English”
  • 16:59 – Inside CI/CD: failure analysis, Jira ticket, AI fix
  • 19:07 – No code generated under the hood
  • 21:18 – Git, JSON and the two types of users
  • 22:18 – Guardrails and the learning loop
  • 24:23 – Writing a good test: the 10-humans rule
  • 26:28 – Test strategy stays human
  • 28:36 – Web, API, Citrix, native mobile: what you can test
  • 29:39 – QA managers: everyone becomes a manager
  • 31:50 – Case 1: a SaaS vendor with no QA team
  • 34:59 – Case 2: a European bank, from 3 years to 4 months
  • 38:13 – Domain testing modules built by partners
  • 40:20 – The pitch in three numbers
  • 41:25 – Karim’s advice: define success before AI

Scroll inside the list to see all chapters.

The episode in 2 minutes

Karim Jouini starts from what he lived through at Expensya, his previous company (250 people, 60 countries): software quality slowed growth more than anything else, with close to a million lines of Playwright code to maintain. His answer with Thunders is natural language test automation. Tests are written in everyday English, and an AI runs them by interpreting each step on screen, without generating any intermediate code.

Three ideas shape the conversation:

  • Natural language survives change. If the “Next” button becomes an arrow, the test doesn’t break: the AI understands the intent.
  • Humans own the strategy. They decide what to test, with what data, and confirm what is or isn’t a defect. The AI executes, analyses and learns.
  • ROI is measured in speed and coverage: ship twice as fast and cover ten times more scenarios with the same resources, according to Karim. A European bank used this approach to bring a core banking migration down from three years of testing to four months.

Why traditional test automation holds growth back

Before talking about AI, Karim describes the problem he faced as a CEO. Expensya, “the Expensify of Europe” as Joe puts it, shipped updates every day, in 60 countries and 17 languages. With regulations changing country by country, that meant two to four regulatory changes a week to absorb without regressions. He breaks the problem into four dimensions.

  1. Cost: Around a million lines of Playwright code, written and maintained by developers and QA engineers. Time that didn’t go into the product.
  2. Campaign length: With corporate cards, bank connections and travel agencies to integrate, a major release needed four to five weeks of test campaigns, which capped the number of major releases per year.
  3. Maintenance: After a rebrand, getting the automated tests working again took six months. The selectors had changed; the product’s behaviour hadn’t.
  4. Collaboration: The company was agile, but testing stayed waterfall. And no one could confidently answer an enterprise customer asking: “What did you test, in our setup?”

That diagnosis is what convinced Karim and his co-founder Jihed Othmani, after selling Expensya, that software quality was the problem worth rebuilding from scratch.

“A manual tester on steroids”: codeless, natural language test automation

Karim sees two families of testing. Manual testing delivers high quality, because a human notices what no script checks: a UX issue, a slowdown, an inconsistency. But it doesn’t scale. Automation scales, but it’s expensive to build, needs technical profiles, and creates a permanent hand-off between the people who know how the product should work and the people who can code the tests.

At Thunders, we aim for the best of both: the flexibility of a great manual tester, someone you can walk through a feature or hand a document to, with the speed of automation. The platform has two modules.

From spec to test plan

You start from a user story, a Jira ticket, a functional spec or a simple prompt. The AI proposes test sections, then scenarios, cases and finally steps, as detailed as a Selenium script. The document is collaborative: several people, including a customer, can edit it, or ask the AI to.

“Selenium in plain English”

The second module runs those steps. “Go to this page, click here, enter this, check the page matches the Figma design.” AI personas follow the script and act like a user. Karim’s example: a test says to click “Next”, but the button has become an arrow. A Selenium script fails because it can’t find the element. The AI understands the arrow means “next”, carries on, and leaves a comment so you can sharpen the instruction. Tests can also be created by recording: the recorded journey is turned into a plain-English story rather than click coordinates.

No code under the hood: what changes for maintenance

This is where Joe pushes: if the test is in English, what runs underneath, Selenium or Playwright? “There is nothing underneath,” Karim answers. Many tools generate code, then try to repair it when it breaks (the familiar self-healing of selectors). Thunders interprets each step at run time by looking at what’s in the browser. Caching and optimisation keep it fast and avoid burning tokens, but no code is produced.

Two practical consequences:

  • Less maintenance: a UI change that doesn’t change the feature doesn’t break the test. It’s what self-healing tests promise, taken all the way.
  • Tests stay a versionable asset: automation engineers used to Git can sync every test to their repository through the API, as a JSON file containing the English. PMs, POs and manual testers care more about owning and debugging their tests.

Inside the CI/CD pipeline

When a test that passed yesterday fails, the AI describes the failure (“this form should be read-only, one field is editable”), compares it with the previous run and traces the cause, say a test user whose role has changed. If a human confirms the defect, a Jira ticket is created with repro steps. Many customers then connect Jira to a coding assistant (Cursor, Devin) and rerun the test to validate the fix. An end-to-end loop.

Guardrails

By default, a written test runs without sign-off. Human control applies to what counts as a defect, with a learning loop. The example: checking TestGuild’s partner pages for typos. If a partner has an invented brand name, the AI will flag it as a typo; you tell it that’s expected, and it won’t flag it again.

How to write a good natural language test

Do you need good English to get good results? No, says Karim: typos are handled. The real enemy is ambiguity. His rule: if you give your instruction to ten different people and they don’t all do the same thing, it’s a bad instruction.

“Click the blue element” will run, but it’s weak: how many shades of blue, and is it a button, a link or a piece of text? A prompt-scoring helper, announced for the next major release, flags that kind of step and explains why it’s fragile.

Joe draws the conclusion for domain experts: the person who knows the application’s rules naturally writes the best tests. The AI does the execution, the expert does the thinking. Karim’s way of putting it: Thunders is “your army of interns”, testing in parallel without limits, but the test strategy (what to test, when, with what data) stays human. On the technology side, anything a human can open in a browser can be tested: web apps, APIs, apps delivered through Citrix, and native mobile through a device farm.

Two case studies: a bank and a SaaS vendor

Joe admits it: his audience is sceptical. Karim answers with two cases with numbers, warning upfront that one of them makes him “slightly uncomfortable”.

A European bank: three years of testing cut to four months

A large European bank needed to upgrade its core banking system, heavily customised and surrounded by in-house apps. It had planned three years of testing with a team of ten QA engineers. With Thunders, the project took four months, for three reasons:

  1. Domain experts wrote tests by recording their own journeys, closing a knowledge gap ten QA engineers couldn’t close alone.
  2. Existing user sessions were turned into scenarios by an in-house tool.
  3. The old Selenium, Cypress and Playwright code was converted into natural language (see how to import your existing automated tests).

Karim’s line: “A mortgage is a mortgage regardless of the version of your core banking.” The form may go from one page to a four-step flow and the components may change: functionally, you’re still applying for a mortgage. He’s candid about the limit: it still took four months of work, “we are 80, 90% in that direction”.

A SaaS vendor that handed testing to its product managers

The second case is a software vendor with around 50 million in revenue, growing fast, that had already decided to let its manual and automation testing teams go. Its request to us: stop quality from collapsing. Nine product managers now write the specs and build every test case in Thunders. Sceptical at first, they found they spent less time testing themselves than they used to spend explaining to another team what to test, then analysing its results. Karim owns the discomfort: the ROI is real, but people lost their jobs in the process.

Joe sees a direction Karim confirms: reusable domain testing modules. A consultancy specialising in car leasing software is building, on Thunders, what it means to “test a leasing system”, to reuse it from one car brand to the next.

QA managers: from individual contributor to AI manager

The jobs question comes up twice. Karim doesn’t think natural language test automation destroys roles in the short term, but he acknowledges that a team with strong functional leaders can use it instead of hiring QA engineers.

His conviction goes beyond testing: AI turns every individual contributor into a manager. He describes his own day: picking in the morning the ten tasks to hand to his agents, writing his prompts for two hours, then reviewing, correcting, sending back. For a QA manager, the question becomes: is my team ready to become a team of QA leads, who delegate execution while still owning quality? Delegating to Thunders doesn’t remove ownership of the result, just as a developer stays accountable for code written by AI.

6 key takeaways

  • A natural language test interpreted at run time doesn’t break when the UI changes but the feature doesn’t.
  • “No code under the hood” is an architectural difference, not a detail: generating then repairing code is not the same as interpreting at every run.
  • Ambiguity is the enemy of AI testing, not the quality of your English: a step that can be understood ten ways is a bad step.
  • Test strategy stays human: what to test, when, with what data, and what counts as a defect.
  • Domain experts become test authors, which is how a European bank went from three years to four months.
  • Define your success metric before launching an AI project: according to Karim, that’s the first reason AI pilots fail.

Key moments from the episode

  1. “Thunders is a manual tester on steroids.” (10:41)
  2. “The power of natural language is its immunity to maintenance issues.” (15:56)
  3. “There is nothing underneath. […] This is really interpreted as we go. […] Our code is English.” (19:07)
  4. “If the instruction you have written, you give it to 10 different humans and they don’t all do the same thing, then that’s a bad instruction.” (24:23)
  5. “A mortgage is a mortgage regardless of the version of your core banking.” (36:01)
  6. “Don’t get into AI before you know why you’re doing it, and then get into AI quickly, because otherwise testers that use AI will replace you.” (41:25)

About the TestGuild Automation Podcast

Hosted by Joe Colantonio, the TestGuild Automation Podcast is one of the reference podcasts of the English-speaking software testing community, with more than 570 episodes on automation, tools, performance testing and, increasingly, AI. TestGuild also runs Automation Guild, a yearly online conference, where Joe got to know Karim and Thunders better in 2026. Its editorial line is to separate hype from real-world practice, for a famously demanding audience of automation engineers. Listen to the episode on testguild.com.

About Karim Jouini

Karim Jouini is our CEO and co-founder, alongside Jihed Othmani, our CTO. An engineer with a master’s degree in artificial intelligence, he spent seven years at Microsoft, on the Visual Studio and then cloud teams, where he learned that testing can be “harder than building the product”. He then founded Expensya, an expense management platform used in 60 countries, which was sold for more than 100 million. That experience as a CEO facing a million lines of Playwright tests is what drives his vision of test automation without code. All articles by Karim Jouini.

Frequently asked questions

What is natural language test automation?

It’s an approach where tests are written in everyday language (“go to the cart page, apply the promo code, check the total”) and run by an AI that interprets each step on screen. At Thunders, no Selenium or Playwright script is generated: the text itself is the test, which makes it readable by product and business teams.

How is it different from self-healing tests?

Most self-healing tools generate code, then try to fix broken selectors after the fact. Natural language interpretation doesn’t depend on selectors: the AI looks for the element that matches the intent (a “Next” button that became an arrow, for example) and flags the difference so you can sharpen the instruction.

Is it the same as codeless test automation?

It goes one step further. Many codeless tools still generate code behind a visual recorder or low-code editor, and that code is what breaks. With natural language test automation, there is no hidden script: the plain-English steps are interpreted at every run.

Can an AI test agent replace a QA team?

Not entirely. According to Karim Jouini, AI takes over execution, failure analysis and ticket writing, but strategy stays human: deciding what to test, with what data, and what counts as a defect. The role shifts towards a QA lead who delegates execution to AI.

How do you write a good natural language test?

Avoid ambiguity. An instruction is good if ten different people would all carry it out the same way. “Click the blue element” is weak, because several elements may match; “click the ‘Place order’ button” is precise. Typos, on the other hand, are handled.

See your first natural language test in 40 seconds

Enter your app’s URL: our agents explore it, write a first test case in plain English and run it in front of you. No account, no code.

External sources

Fortune, on MIT NANDA’s The GenAI Divide report (Aug 2025): 95% of generative AI pilots at companies are failing, quoted by Karim at 41:25

Playwright docs - Locators: what a selector is, and why a coded test depends on page structure