sk Skill · pixiebrix
agent-browser-shield-diagnose
Diagnose why a specific benchmark task underperformed on guarded vs baseline — whether that's a judge pass/fail regression or a cost regression (extra tokens, steps, or duration). Use when the user asks "why did task X fail on guarded", "why is the guarded run worse than baseline on Y", "why did guarded cost more on Z", "what went wrong with run_<id>", or wants to investigate a flaky or expensive (scenario, task)…
Open on skills.sh ↗read 2026-09-19
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 1
- Connections
- 0
bashTypeScript
- Host repository
- pixiebrix/agent-browser-shield
- Host stars
- 35
- Host language
- TypeScript