Help and version
NO_COLOR is set; FORCE_COLOR keeps them. e2e without a command prints
the help on stderr and exits 2. A misspelled command exits 2 with a suggestion
(e2e rnu answers Did you mean run?).
e2e init
Scaffolds a project and adds@e2edev/e2e to devDependencies when it is not
already declared. Existing config and test files are reported and skipped.
Existing dependency versions and other package fields are preserved. With a
directory argument the project is scaffolded there, the directory is created
when missing, and the closing next: line starts with cd into it.
For a new config, choose Web (Playwright, the default), Mobile
(iOS/Android) (agent-device), or None. Web adds @e2edev/playwright;
Mobile adds @e2edev/agent-device.
Next the wizard asks which model gateway agent steps use: Vercel AI
Gateway (the default), OpenRouter, OpenAI-compatible endpoint, or
None. The OpenAI-compatible choice prompts for a base URL, which must be
HTTPS or loopback. A gateway adds ai@^7.0.0, the gateway’s own provider
package when it has one, and an agents block that constructs the model.
Then the wizard offers the agent skill (see e2e guide) for .agents/skills/, the directory
Codex, Cursor, Copilot, Gemini CLI, OpenCode, Zed, and most other agents
read, and .claude/skills/ for Claude Code, both preselected. After
confirming the file changes, choose once whether to install the selected
dependencies.
Device setup defaults to iOS on macOS and Android elsewhere, using one worker
and opening Settings. Change the platform in e2e.config.ts when needed.
Replace the engine’s app value with your app’s bundle ID or package name,
or use appPath to install a build. iOS needs Xcode and a simulator; Android
needs the Android SDK and an emulator. See Testing iOS and Android.
packageManager field or lockfile, then the
invoking package manager, falling back to npm. Declining installation still
saves the dependencies to package.json, and the closing next: line starts
with the install command, followed by the run command for the engine:
APP_URL=http://localhost:3000 npx --no-install e2e run for Playwright.
Device setup runs without APP_URL. When the project has no
tsconfig.json, init prints a one-line suggestion to add one.
Engine and AI choices apply only when creating a config. Re-running init
keeps an existing .ts or .mts config and adds no optional packages for it.
The skill prompt appears only while no known directory holds a copy. Once one
does, later runs rewrite the copies whose files differ from the installed
version, so init after an upgrade refreshes the skill without asking, and
never adds a directory the project did not choose. Files the skill does not
ship are left alone.
Cancellation leaves files untouched and exits 0. An invalid manifest exits 2
before any write, quoting the parser’s position or the field with the wrong
shape. A failed installation keeps the scaffold, prints the retry command, and
exits 2. Without a terminal on both stdin and stdout (a CI step, a pipe) the
prompts cannot be answered, so init without --yes exits 2 at once and
names the flag instead of waiting for input that never comes.
e2e guide
init installs: how to add e2e to a project, write
tests, use agent steps, run the CLI, and read a failing run. guide alone
prints the overview, which lists the topics; guide <topic> prints one of
setup, writing-tests, agent, running, explore, debugging, or mcp. An unknown
topic exits 2 and names the valid ones. The text is the same as the installed
skill’s, so a coding agent working in a project without the skill files can
read it from the CLI, and e2e --help points there.
The same skill installs from the repository with
npx skills add tester-army/e2e. Coding agents covers how
an agent uses it.
e2e mcp
open_session opens a session
on one target (any config the agent names), call runs any tool of that
session (observe, tap, type, press, select, scroll, navigate,
type_secret, locate, screenshot, and the project’s own tools), tools
describes them, and close_session ends it. The agent looks at the real
screen and checks a locator before writing a test. e2e init registers the
server for Claude Code and Cursor. The tools, the session rules, and the
flags are on the MCP server page.
e2e run
Positional arguments
- a file path (
tests/signup.e2e.ts) selects that file; - a directory (
tests/agent,.) selects every test file the config globs discover beneath it; - a name that exists nowhere selects every discovered file whose path ends with
it at a segment boundary:
signup.e2e.tsandagent/signup.e2e.tsboth selecttests/agent/signup.e2e.ts, whilegent/signup.e2e.tsselects nothing. Without a/or a.it is matched against the file name with every extension dropped, sosignupselects it too; - a glob (
'tests/**/*.smoke.e2e.ts', quoted so the shell leaves it to the runner) uses the same grammar as the configtestsglobs:*,?, and a complete**segment. A malformed glob isINVALID_GLOB.
tests/Agent finds tests/agent on
Windows and macOS but not on Linux), while names and globs are case-sensitive
everywhere. A path outside the project root, including one on another drive,
is COLLECTION_ERROR.
A positional that starts with - and names no existing entry is a usage error
(exit 2), not a file. It gets there when a package manager forwards the --
separator: pnpm test:e2e -- --headed reaches e2e as run -- --headed, and
the CLI would otherwise take --headed for a file and run headless. The
message names the direct form for the detected package manager with everything
from that flag on, so -- --tag smoke gets pnpm exec e2e run --tag smoke. A
file that really starts with a dash (-smoke.e2e.ts) still selects as before.
When nothing is left to run, the NO_TESTS message says why, from the most
upstream cause: the globs matched no file (naming look-alike files such as
tests/login.test.ts beneath the globbed directories, which is usually the
file meant), a positional matched no file (with the nearest discovered path or
file name when there is one, so tests/agnet.e2e.ts gets did you mean tests/agent.e2e.ts?), a matched file registered no tests (its test import
is from somewhere else), or every collected test was filtered by tag or
platform, skipped, or is a setup test.
Flags
--retries and --workers obey the bounds of the config keys they replace
(retries 0-10, workers 1-1024). The parser accepts any nonnegative
integer and rejects anything else with exit 2; the config resolver then
applies the bounds, so an out-of-range flag fails with INVALID_CONFIG
before anything starts, the same as the config value would.
--debug appends two tables to stderr after the reporter’s output. Phase
timings are summed across the runner and every worker and sorted by total
time; the agent step table lists every agent step in execution order with its
per-phase split, model calls, tokens, and the gateway’s cost when it reported
one. A run without agent steps prints the timings alone.
cached is the share of the step’s input tokens the provider served from its
prompt cache, with the count in parentheses; - when the provider reports no
cache split. The list reporter’s usage line carries the same share for the
run (12.4k tokens · 38% cached · $0.01), and the report records it per step
(model.cacheReadTokens, model.cacheWriteTokens) and per run
(usage.modelCachedTokens).
Output
Every run atomically writes
.e2e/report.json under the artifact parent
regardless of the selected reporters. The three ids name
reporters on the same contract a config can supply its
own on: junit writes junit.xml beside the report from that same document
once the summary has printed, and a junit.xml that could not be written is a
line on stderr, never a run error. A reporter can never change run status; the
rows it returns print under the list summary.
With --ai-trace, the run also writes .e2e/ai-trace.json next to the
report: one run per agent step, named after the test and the step, with one
entry per model round trip (an act turn, a waitFor poll, a repair round)
carrying the exact prompt, tool definitions, response,
usage, and provider metadata of every round trip. It is the model’s view of
the run: observations are already redacted before they reach the model.
Inline bytes and base64 in SDK image and file message parts are replaced
with decoded byte counts. URLs and text remain readable. Encoded strings
outside those parts, including arbitrary tool result JSON, are preserved. Open it
with npx unbox-ai .e2e/ai-trace.json; see Debugging a run.
Signals
Signals escalate. The firstSIGINT or SIGTERM (Ctrl-C) interrupts: the
running test ends at once with status interrupted, even a body that is not
calling the harness at that moment; its afterEach hooks and session close
each run within the cleanup budget, every worker disposes its engine and
exits, and the runner writes a partial report with status: interrupted and
exits 130. A second signal forces: every worker disposes its engine right
away instead of finishing its test, and is killed once the cleanup budget is
spent; the report is still written. A third signal exits on the spot, the last
resort for a teardown that is itself stuck; before exiting it kills the process
groups of every app command and service the run started, so no dev server or
database container survives to fail the next run with APP_ALREADY_RUNNING. A
process reused through reuseExisting was not started by the run and is left
alone. An interrupt that lands during setup names the step it cut short
(interrupted while starting service "postgres": tearing down) instead of
claiming to stop a running test. The final summary names the interrupt too: a
run cut before its plan arrived ends with Test Files none started (interrupted) rather than no test files, and a run cut after it keeps its
counters against the planned total. This ladder
belongs to the CLI; the runner itself never handles process signals.
A worker whose runner disappears (killed, crashed) never runs on by itself: it
disposes its engine and exits within the cleanup budget, so a device or a
browser is not driven by a process nobody is listening to.
e2e explore
report_finding tool, and ends with an assessment. The goal is one quoted
sentence; without one it is Explore the app and find bugs. The model is the
selected agent’s, as for run. Exploring without a test
describes what a run does and how to read it.
The run has one result, under the virtual file
explore, titled by the goal.
The trace cache is off and retries are zero whatever the config says.
.e2e/report.json gains run.explore: the goal, the budgets, why the run
ended (finished, step-limit, time, stuck, or aborted), the
assessment, the steps, and the findings.
Exit codes: 0 when steps ran and no finding of kind issue was reported,
1 when one was, or when no step ran and nothing was found (the run is
blocked, never a zero-coverage pass), 2 and 3 as for run, 130 on
Ctrl-C.
e2e list
list collects and selects tests exactly as run does, prints
one line per test-target pair, and exits. Nothing starts: no app process, no
engine, no worker, and nothing is written under .e2e/. Use it to check what
a set of files, tags, and targets selects before paying for the run.
The positional arguments and the selection flags are those of run:
--config, --target, --tag, --tag-mode, and --pass-with-no-tests,
with the same meaning. --reporter accepts list (the default) or json;
junit exits 2. The same NO_TESTS, UNKNOWN_TARGET, and config errors
apply, with the same exit codes.
Pairs that a tag, target, or platform filter removed are not listed; they are
not part of the run either.
e2e cache
Reads and empties the trace cache the runs write under.e2e/cache/. A cache
entry is a file named after the digest of its key, so a directory listing says
nothing about what is in it; these commands are the reader. All three resolve
the store the same way a run does, through the discovered config or
--config <path>, honoring cache.dir, and all three exit 0 on an empty or
absent store, because a cache nobody has filled yet is not an error.
- in those columns; it still replays.
A file the store cannot read (truncated, over the 1 MiB entry ceiling, or
from a newer schema) is not an entry: ls and stats count it separately on
stderr, a run treats it as a miss, and clear removes it. Files the runner
never wrote are left alone and named on stderr, so a cache.dir pointing at a
directory that holds anything else cannot lose it.
A project that configures
cache.store replaced the file store with its own,
and these commands read files only, so they exit 2 and say so rather than
reporting an empty cache.e2e telemetry
Shows whether anonymous usage telemetry is on, and switches it. The CLI sends one event per command and one per run, built from counts, versions, and the names your config declares; Telemetry lists every property.E2E_TELEMETRY_DISABLED=1 and DO_NOT_TRACK=1 turn telemetry off for one
shell or one CI job without touching the file. E2E_TELEMETRY_DEBUG=1 prints
every event to stderr as [telemetry] {...} and sends nothing. The first
command on a machine prints a two-line notice on stderr, once; in CI nothing is
printed and no file is written.
Exit codes
Precedence for a mixed run is
130 > 4 > 3 > 2 > 1 > 0. The report keeps every
individual result; precedence affects only the process exit code.
Argument parsing failures (an unknown flag, a non-integer --retries, an
invalid --tag-mode, an unknown --reporter) exit 2 with a one-line error on
stderr followed by (add --help for usage).
Diagnostics
Every message that rejects a name offers the nearest valid one when a typo is plausible: an unknown config, target, agent, cache, limits, or artifacts key, an unknown--target ID, an unknown reporter, an unmatched positional, and a
fixture the target does not have. Keys from other runners’ configs (testDir,
baseURL, webServer, use, projects, or a url on a target) say where
that fact lives here instead. A page, browser, context, request, or
driver fixture is explained in terms of app, screen, and web.
A failing import in the config or a test file is explained past the loader’s
message: a package that package.json declares but node_modules lacks ends
with the install command for the project’s package manager, an undeclared one
with the add command, a wrong subpath with the subpaths the package exports,
and a removed export such as defineConfig with its replacement. A test
failure carries its cause chain (fetch failed: connect ECONNREFUSED 127.0.0.1:3000), and navigation to an address where nothing listens is
APP_UNREACHABLE with the URL and the three ways to fix it, not an opaque
engine failure. A model credential the provider rejects is reported as such,
pointing at the variable the provider package reads (AI_GATEWAY_API_KEY,
OPENROUTER_API_KEY, …) or the key passed at construction, and color codes
in provider messages are stripped before they reach the terminal or the
report.
