Keter Notes: browser-use or Playwright? Two ways to automate manual tests

Published date:

Share directly to:

Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI
Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI
Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI
Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI
Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI

Published date:

Share directly to:

Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI
Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI
Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI
Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI
Keter Notes: browser-use or Playwright? Two ways to automate manual tests - Keter AI

Any team with a large manual test suite faces the same question: should an AI agent run the tests, or should they be scripted in Playwright? Our Manual Test Agent began as an internal tool built on browser-use, and designing it meant comparing both. Our answer: complements, not substitutes.

Where both tools stand

  • browser-use is an MIT-licensed Python library that lets an LLM drive a real Chromium browser. Current stable version: 0.13.10, released on 4 September 2026. It is still pre-1.0, so APIs can break.

  • Playwright is Microsoft's Apache-2.0 testing framework for Chromium, Firefox and WebKit. Current version: 1.63.0, also released on 4 September 2026.

A common error: browser-use is not built on Playwright. In August 2025 the project dropped it and now speaks the Chrome DevTools Protocol (CDP) directly through its own client.

Playwright Test Agents (planner, generator and healer, introduced in version 1.56 in October 2025) use an LLM when tests are written or repaired and leave ordinary test files behind, so CI runs need no model calls.

The numbers, and who measured them

  • Vendor-run: Browser Use reported 89.1% on WebVoyager in 2024, on a task set it had trimmed and re-reviewed itself, and 97% on Online-Mind2Web in March 2026 for its paid cloud agent, scored by its own judge.

  • Independent: the Online-Mind2Web paper (300 tasks, human evaluation, April 2025) measured the open-source agent at 30.0%. It predates current models.

  • Manual test cases: in an academic study the better agent gave the correct verdict on about 60% of 113 cases, with 2024-generation models.

  • Speed: the vendor reports about 3 seconds per step and 68 seconds per task. Stagehand, another agent framework, documents 2 to 3 seconds for a cached replay without a model, against 20 to 30 seconds for a first agent run: roughly an order of magnitude.

  • Cost: on the vendor's benchmark, 100 hard tasks cost about 10 USD on the basic plan and close to 100 USD with a top model. Scripts cost only CI compute.

  • Script path: one practitioner's Playwright agent pipeline generated 47 tests; 62% passed at first and 96% after healing, for about 327k tokens.

We found no public head-to-head measurement on the same suite.

Side by side

  • Authoring: for the agent, the plain-language case is the test. In Playwright each case becomes a script someone must review.

  • UI change: the agent re-reads the page on every run, but may quietly work around a real regression. A script breaks, and every repair is a visible diff.

  • Verdict and evidence: scripts repeat the same steps and leave traces. The agent leaves a narrative with screenshots, and its success flag is only its own assessment.

  • Coverage: agents suit exploratory passes and long-tail cases. Scripts suit regression and cross-browser runs, which a CDP-only agent cannot provide.

Security cautions

An agent reads untrusted page content on every step, so prompt injection is a live risk. A critical allow-list bypass (CVE-2025-47241, since fixed) was published in May 2025, and an academic analysis demonstrated credential exfiltration. Telemetry is on by default. Run agents sandboxed, on test environments only, with a domain allow-list, test accounts, placeholder secrets and telemetry off.

The design we recommend

  • The agent runs each imported manual case first, sandboxed and fully recorded.

  • Verdicts are checked independently: deterministic checks where possible, a tester for unclear results.

  • Stable, frequent or high-risk cases become Playwright scripts, reviewed by an engineer in a pull request.

  • CI runs scripts only, with no model calls. When one fails, an agent proposes a repair and a person approves it.

  • The agent stays for exploratory passes and long-tail suites. Every run leaves evidence.

Vendors are converging on this split: Browser Use's cloud offers rerunnable scripts, and Playwright ships agents.

Sources

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.