
Manual Test Agent: from test case to Playwright
Year:
2026
Service:
AI-assisted Engineering
Industry:
Software
Team:
Internal build, senior engineers
An internal tool built on the open-source browser-use library executes plain-language manual test cases with a browser agent. The design around it adds independent verdict checks, Playwright scripts for CI, human review gates and evidence for every run.
Introduction
Many software teams still keep a large part of their quality knowledge in manual test cases: a title, preconditions, test data, numbered steps in plain language and an expected result, stored in a test management tool or a spreadsheet. Running them by hand before every release is slow. Rewriting all of them as scripts is expensive, and many cases run too rarely to repay it.
The Manual Test Agent began as an internal tool built on the open-source browser-use library, which lets a language model drive a real Chromium browser. The tool takes a manual case as it was written and executes it. What makes it a business-grade solution is not that an agent can click through a case. It is an honest comparison of two ways to automate, and a design that uses each one where it is strong.

Challenge
Manual regression does not scale, and both obvious fixes stall.
Classic automation stalls on cost. Every case has to be rewritten as code by an engineer and then maintained each time the interface changes. Many cases run a few times a year, never repay that effort and stay manual.
Execution by an agent alone stalls on trust. An agent that reads a case and clicks through the application looks impressive in a demo, but a test verdict has to be believed. The library's own documentation warns that the success flag is only the agent's self-assessment. In an academic study of 113 manual test cases, the best research agent gave the correct verdict in about 60% of cases, using 2024-generation models. A false pass is worse than a false fail: an agent that finds a workaround can hide a real defect.
There is a security problem too. A browser agent reads untrusted page content at every step, so prompt injection and credential leaks have to be contained by design, not left to luck.
Solution
We first compared the two options on the same criteria.
Option A - agent at run time (browser-use). The plain-language case is the test. Cheap to author and tolerant of UI change. But every step is a model call: slower, paid per run and not deterministic.
Option B - scripted Playwright. Fast, deterministic, built for CI, with traces that reproduce a failure. But expensive to write and maintain, and it checks only what was scripted.
The design combines them: the agent explores and authors, scripts run in CI, people approve.
Ingest and triage. Cases are normalised into steps, data and explicit expected results. Frequent and high-risk cases go to the script path; the long tail stays with the agent.
Agent run. browser-use drives an isolated Chromium through the Chrome DevTools Protocol (it has not used Playwright since August 2025), on a test environment, with a domain allow-list, test accounts, secrets as placeholders, step and cost limits, telemetry off and full recording.
Independent verdict. Deterministic checks first, then a judge model; unclear results go to a tester.
Promotion. Stable cases become Playwright tests that an engineer reviews in a pull request. CI runs them with no model calls.
Supervised healing. When a script breaks, an agent classifies the failure and proposes a patch; a person approves. A suspected defect ends as a bug report.
Evidence. Every run records the case version, executor, model version, artefacts and who confirmed the verdict.
Playwright now ships its own test agents (planner, generator and healer, since version 1.56), which work at authoring and repair time and leave ordinary test files; the design can use them for promotion and healing.
Stack: Python, browser-use (pinned), Chromium over CDP, Playwright Test, CI.

Result
What exists today is an internal tool built on browser-use and a design we offer as the Manual Test Agent accelerator. For each case the pipeline is designed to give a team a verdict per step (pass, fail or unclear) with a plain-language reason, screenshots and a recording, an action log and run metadata. Promoted cases arrive as Playwright code in the team's repository. Cases too vague to automate are listed, which is useful feedback in itself.
We do not publish pass rates or savings. No reliable public figure exists for the share of a manual suite an agent executes correctly, and vendor benchmarks do not transfer to a specific application. Each deployment therefore measures four things on its own suite: agreement between agent and human verdicts on a sample, cost and duration per case, the share of cases promoted to scripts and the share of healing patches accepted.
One limit up front: the agent path is Chromium only, so cross-browser regression belongs to Playwright.

