Build a QA Strategy for Tests That Can Write and Run Themselves
September 29, 2026
Executive Brief
Summary
Agentic automation can help your QA team plan tests, explore an application, and investigate failures with less manual coordination. The opportunity is to expand what you can examine while reducing some of the effort required to create and maintain useful tests. Your strategy needs to establish what agents should investigate, how expected behavior is defined, and which changes require review. With those decisions in place, you can introduce agentic capabilities into your existing testing process and evaluate whether they work better.
Questions Answered in This Article
- What is agentic automation in software testing?
- Agentic automation uses AI agents to carry out testing tasks through a sequence of actions. An agent may explore an application, generate executable tests, or investigate a failure.
- How does agentic testing fit with existing automation?
- An agent can create and maintain tests that run through an established automation framework. Your team checks for known behavior while using agents to investigate additional scenarios.
- Can agents repair tests when an application changes?
- Some systems can propose and evaluate repairs. You need to confirm that a repair preserves the behavior the test was intended to verify.
- What responsibilities remain with your QA team?
- Your team defines important risks, validates expected behavior, and decides whether the results greenlight a release.
- How should you measure its value?
- Look at what the agent helps uncover and the effort required to review and maintain its work. Test volume offers context, but doesn’t define the quality of the coverage.
When the System Can Change the Test
If a test fails after a change to your customer portal. You could have an agent investigate, adjust the test, and run it again. That might represent a useful maintenance improvement, or it could mean the test has stopped checking something important.
The difference depends on what changed. Perhaps a button was renamed and the agent updated how the test locates it. Perhaps a required confirmation disappeared and the agent removed the step that expected it. Both changes might produce a passing result, but they would have very different implications for your confidence in the application.
This is the question at the center of an agentic QA strategy. When a system can participate in writing, running, and revising tests, you need to understand how it determines what success means. You also have an opportunity to use that capability productively. An agent can take on some of the repetitive investigation and preparation that consume your team’s attention, giving you more room to examine behavior that has received too little coverage.
Start by making those responsibilities explicit. Decide what the agent can change independently, what evidence it should retain, and where your team needs to review a decision. You can then evaluate the benefits against a testing process you understand.
Understand What You Are Giving the Agent to Do
Agentic testing capabilities vary. Some tools explore a running application and propose scenarios; others generate executable tests or investigate failures. What they can accomplish depends on the instructions, product context, and tools available to them. Playwright’s test-agent documentation offers a concrete example through separate agents for planning, generation, and repair. That arrangement shows how agentic capabilities can work with an established test framework. An agent can help produce a test that your team inspects and then runs repeatedly through familiar automation.
When you assess a tool, ask to see the work between the request and the result. You should be able to examine the proposed plan and understand the checks it produces. If a repair occurs, you need to see what changed and why. A demonstration that ends with a passing suite gives you less information than a workflow whose decisions you can follow.
Begin With One Behavior You Understand Well
A focused assignment makes it easier to judge the agent’s contribution. Choose a workflow that matters to your product and that your team understands well enough to evaluate. Asking an agent to test an entire application at once can produce a large amount of activity before you have established whether its expectations are sound.
Consider a customer portal where someone can update their contact information. At first glance, the task appears straightforward. Enter a new value, save it, and confirm that the change succeeded. But your expectations extend beyond that visible sequence. The correct account should receive the update, invalid information should receive a useful explanation, and another customer should be unable to change the record.
Those expectations give you a basis for reviewing a proposed test plan. If the agent produces several variations of a successful update but overlooks account access, you can identify the missing behavior before generating more tests. You can also explain which failures carry the greatest consequence, helping the agent concentrate its effort.
Keep the first assignment manageable enough to inspect closely. You are learning how effectively the tool translates your requirements into useful testing work. That knowledge will help you decide where to use it next.
Give the Agent a Basis for Knowing What Should Happen
An application can demonstrate its current behavior without establishing that the behavior is correct. If an agent observes an existing defect and treats it as the expected result, the generated test may preserve the mistake. Your requirements need to provide an independent reference.
For the contact-information workflow, explain which fields a customer may change and which require additional review. Supply an approved example if the rules are difficult to express. A short, current description of that distinction can be more useful than a large collection of documents containing conflicting instructions.
Defect reports and support requests can add context about where to investigate. They may reveal a particular combination of actions that has caused trouble or a point where people repeatedly become confused. Use relevant information within your organization’s data-handling requirements, and remember that a rare behavior can still carry substantial risk.
Someone also needs to resolve disagreements between the requirements and the running product. An agent may expose that disagreement, but your team has to establish which behavior is intended. Leaving the ambiguity unresolved makes it harder to trust any test built around it.
Check Whether the Test Could Catch the Problem
A generated test can perform many actions while verifying very little. It may navigate through the portal, enter a new phone number, and confirm that a success notification appears. If the application displays that notification without saving the change, the test could pass while the feature fails.
Ask what would have to go wrong for the test to fail. That question helps you examine the assertion, which is the part of the test that checks the result. In this example, you would want evidence that the appropriate account retained the updated information. The notification might be worth checking too, but it answers a narrower question.
You also need to examine the starting conditions. A test using information left behind by a previous run may appear reliable until the execution order changes. Clear test data and a known account state help make the result repeatable and the failures easier to interpret.
During your initial trial, review a small set of generated tests in enough detail to understand these choices. You may discover that the agent produces useful scenarios but needs better guidance on assertions. That is actionable feedback you can use to improve the workflow before expanding it.
Keep the Testing Methods That Already Provide Useful Evidence
An agent can help write a conventional automated test. Once reviewed, that test can run repeatedly without asking the agent to plan the same scenario again. This is a useful connection between newer capabilities and the repeatable checks your team already depends on.
Exploration serves another purpose. You might ask an agent to investigate variations in a workflow and report behavior that deserves attention. Your team can then reproduce the findings and decide which discoveries belong in the regression suite. Human exploratory testing remains valuable for evaluating whether an experience is understandable and noticing problems outside the agent’s assigned scope.
AI-powered features also contain behavior with precise expectations. A generated answer may vary, while the rules governing access to private account information remain clear. You can use repeatable checks for those rules and additional evaluation methods for the quality of the answer itself. In The 2026 Software Testing Playbook for Manual Automated and AI Testing, I explain how these approaches contribute different evidence. Agentic capabilities fit into that combination according to the work they can perform dependably.
A Repair Should Preserve the Reason the Test Exists
Test maintenance is an appealing place to apply agents because small changes can create repetitive work. If an interface control has moved or been renamed, an agent may help identify it and propose an appropriate update. The benefit is real when the revised test continues to check the original requirement.
Return to the missing confirmation in our opening example. Before changing the test, your team needs to establish whether the confirmation was intentionally removed. If the product requirement still calls for it, the failing test has supplied useful evidence. Repairing the test by accepting its absence would hide the issue.
Ask for an explanation that connects the failure to the proposed change. A reviewer should be able to see why the previous check failed and why the revised version still tests the intended behavior. Changes that remove assertions or skip steps deserve particular attention because they may reduce coverage.
Intermittent failures need similar care. A longer wait or another retry can make a run finish successfully while leaving a genuine application problem unresolved. Keep the original failure evidence available so your team can investigate the cause rather than judging the repair only by its final status.
You can make this review practical by distinguishing routine maintenance from changes to the test’s meaning. Updating how an equivalent control is located may be a limited adjustment. Changing what counts as a correct result requires someone to confirm the underlying expectation.
Give Exploration a Place to Work Safely
An agent exploring your application needs a defined environment and appropriate access. Use test accounts and data suited to the task, and enforce the boundaries through the environment’s controls. Written instructions alone should not be the only thing preventing a test from affecting a real customer.
Be specific about the assignment. You might ask the agent to investigate how a form handles incomplete entries or what happens when someone returns to an earlier step. Set reasonable limits on execution time and repeated attempts so the exploration can stop and report when it becomes stuck.
The resulting report should give your team enough information to investigate.