Compare two recorded sessions step-by-step. Pairs each step in session A to the same step-id in session B, diffs the captured screenshot and accessibility snapshot, reports the first divergence and an aggregate similarity score.
When to use
Visual regression after a UI change (record before, record after, diff).
Verifying a browser-replay run matches the parent session within tolerance.
Comparing two A/B variants of the same form flow.
Steps
Locate both RVF containers:
bash npx -y ruvector@0.2.25 rvf status <session-id-a>.rvf npx -y ruvector@0.2.25 rvf status <session-id-b>.rvf
Load both trajectories from trajectory.ndjson. Build a step-id → (screenshot_path, snapshot_path) map for each.
Pair steps by step-id. Steps that exist on only one side are flagged as unmatched and contribute to the divergence score.
Pixel diff (--mode pixel|both): compare the two PNGs at each step. Report mse, psnr, and the bounding box of the largest diff cluster. Threshold default 0.02 (2% of pixels).
DOM diff (--mode dom|both): compare the accessibility snapshots node-by-node. Report added / removed / changed nodes with their accessible names.
Aggregate similarity: weighted average across matched steps, weighted by step duration. Verdict goes into a new findings.md under a fresh RVF container so the diff itself is replayable.
Persist the diff verdict in browser-sessions under both source ids' tags so future searches surface "ran a diff against session X".
Caveats
Pixel diff is sensitive to font hinting, antialiasing, and scrollbar position. Keep viewport pinned across both sessions.
DOM diff over Playwright's accessibility tree is more stable than HTML diff. Prefer it.
This skill does not handle dynamic content (clocks, ads); add ignore regions to the field map or pre-process snapshots before diffing.
The browser_screenshot_diff MCP tool is not planned (ADR-0001 §7); the skill operates against locally-saved RVF artifacts and uses browser_eval only for live verification.
Want to see if ruvnet/browser-screenshot-diff would activate on your phrasing? Try it before you install — no LLM call, no signup.
Related skills
Other measured skills in the registry, with their headline benchmark lift.
Sign, verify, and track fix-marker regressions over time using a deterministic Ed25519 witness manifest. Works in any project — clone the toolkit, run init, register fixes, regen on each release.