TLDR
- Claude reads the ticket and PR first: It checks the description, acceptance criteria, comments, and code changes before testing.
- It checks that the change is actually deployed: Claude makes sure the right code is in the right environment before running the test.
- Thunders tests it like a real user: The test runs in a real browser, with real clicks, pages, and assertions.
- No test? Claude creates one: It can turn the bug’s reproduction steps into a Thunders test.
- The test stays connected to the ticket: So you can run it again on future releases instead of testing the same thing from scratch.
- The verdict comes with evidence: Claude helps tell the difference between a real product failure, a test issue, and temporary environment noise.
Every engineering team has a column where tickets go to wait. The code is merged, the pull request is green, and the ticket sits there until someone has time to prove it actually works.
We decided to stop waiting. Our team now hands that whole column to Claude and Thunders. One Claude agent takes each ticket, all of them in parallel. Each agent reads the ticket, checks the code, and tests the change through Thunders in a real browser, the way a user would.
This is what Claude software testing looks like when it is connected to the application rather than limited to reading code or generating test cases. Claude can understand what changed, find the right checks, create a test when one is missing, and use Thunders to verify the change in the running product. Thunders provides the end-to-end testing layer that exercises the application in a real browser.
A few minutes later, every ticket has a verdict, with the evidence behind it.
This article explains how the workflow works, what it catches, and what we learned from testing 13 tickets at once with Claude and Thunders.
Why does "merged" not mean "working"?
A green pull request tells you the unit tests pass. It does not tell you the fix works in the product.
Between the two sits a long list of things that go wrong in real life. The fix is merged into the wrong branch. The deploy failed and nobody noticed. The code works, but only behind a feature flag that is off. The bug is fixed on the happy path and still happens on the page the customer actually uses.
This is where Claude QA testing becomes more useful than simply asking Claude to review code. The software still needs to be exercised in the environment where a customer will use it. Thunders is built for end-to-end testing, so the verification happens against the running product rather than only against the code.
Manual verification catches these problems, but it is slow and it does not scale. Someone has to read the ticket, find the pull request, check the deploy, open the app, reproduce the original bug and confirm it is gone.
Multiply that by a full sprint of tickets and you lose days. And the check happens once. Next release, nobody runs it again.
What does Claude software testing with Thunders look like?
For each ticket, one Claude agent follows the same steps:
- Read the ticket: Review the description, acceptance criteria, comments and linked pull requests.
- Check the code is really there: Confirm that the change is merged into the branch the environment runs, CI passed and the deploy succeeded.
- Pick the tests: Find the existing Thunders tests that exercise the affected feature.
- Write a test when none exists: Turn the bug's reproduction steps into a new Thunders test with setup, actions, assertions and cleanup.
- Run the test as a real user: Use Thunders to execute the test in a real browser against the target environment.
- Read the evidence: Review step results, screenshots and application telemetry to separate real failures from issues such as a network blip.
- Return a verdict: Mark the ticket ready, not ready or in need of a human decision, with the reason and supporting evidence.
The agents work in parallel, so a full column of tickets can take about as long as a single one.
This is the part that changes the Claude test automation story. Instead of using Claude only to generate a script or suggest test cases, the agent can participate in a loop that starts with a ticket and ends with evidence from the running application. For teams that want this same flow in their delivery process, Thunders can also run tests through CI/CD integrations.
What can Claude test automation actually catch?
Bugs that only exist in a real browser: Unit tests check that a function returns the right value. Users do not call functions. They click buttons inside iframes, scroll grids that only render visible rows, and click on things a popup is quietly covering. A mocked page never sees any of that. A Thunders test does, because it reads the page, acts on it and checks what it sees, the way a person does.
Fixes that are only half there: A configuration change that was never applied. A deploy that did not run. A feature that works only when a flag is on. Claude checks the code, the deploy and the running app together, so "done in the pull request" and "done in the product" finally mean the same thing.
Scope that quietly changed: When a fix drops part of what the ticket asked for, the verdict says so and names who needs to agree. The gap is visible before release, not after a customer finds it.
Each of these would otherwise reach production and come back as a support ticket.
What happens to the tests after release?
This is where the time saving compounds.
Every test Claude writes is linked to its ticket. The next time anyone ships a change near that code, the test runs again. So does the one after that. This turns a one-off check into reusable test management that stays connected to the change it was created to protect.
A bug is not verified once and forgotten. It becomes a permanent check that the fix, and every improvement built on top of it, keeps working.
Over time, your regression suite grows out of the real bugs your team fixed, not out of a test plan someone wrote once and never updated. And because Thunders provides self-healing tests when the interface changes, that suite does not turn into a maintenance job.
Who saves time, and how?
Developers: Stop being interrupted to demo their own fix. They merge, and the verdict arrives with the evidence: which test ran, what it clicked and what the telemetry showed. When something fails, the report says whether it is a product bug, a test problem or a flaky environment. Thunders also provides test analysis and insights to help teams understand failures and act on the evidence.
QA engineers: Stop re-clicking the same flows every release. They can spend more time on what needs judgment: edge cases, exploratory testing and reviewing the tests Claude drafts.
Product managers: Get a clear answer to "can this ship?" for every ticket, in plain language. When the answer is "needs a human", the report says exactly which decision is waiting and who owns it.
Customers: Get fewer regressions. The fix they asked for stays fixed next month because a test now guards it on every release.
Does Claude replace people in QA testing?
No. Some questions are not for a machine to answer: is a reduced scope acceptable, can a release go out with a known gap, or should a feature ship now or wait?
Claude finds these questions and puts them in front of the right people. It does not answer them on anyone's behalf.
The verdicts are also honest about their own limits. When a test only proves that nothing else broke, the report says so instead of calling the fix verified.
That distinction matters in Claude QA testing. Automation is useful only when people can understand what the evidence actually proves.
How does this turn into a happier business?
The chain is simple:
- Every change is tested as a real user before release: The running product is checked instead of relying only on code-level tests.
- Every fix gets a test that runs on future releases: The regression check stays connected to the original change.
- Fewer bugs and regressions reach customers: Problems are found before they become production issues.
- Customers get a more reliable product: Fewer regressions mean fewer support issues and less disruption after releases.
- The team spends more time building: Developers and QA spend less time repeating manual verification.
Faster, safer releases are not only an engineering win. They are part of keeping customers confident in each release.
Where Thunders fits in Claude software testing
For a deeper look at the differences between the two approaches, see our Thunders vs Claude comparison.
Thunders is the part of the workflow that behaves like a real user. It runs end-to-end tests in real browsers, lets anyone describe a test in plain language, and keeps tests working when the interface changes.
Claude does the reading, the checking and the writing around it. Thunders provides the execution and evidence from the running application.
Together, they turn your testing column from a waiting room into a step that takes minutes.
You can start the same way we did: pick the tickets waiting in your testing column and let Claude and Thunders test them as your users would.
Create a free Thunders account or book a demo to see it on your own app.





