BigHugger
sk Skill · pixiebrix

agent-browser-shield-diagnose

Diagnose why a specific benchmark task underperformed on guarded vs baseline — whether that's a judge pass/fail regression or a cost regression (extra tokens, steps, or duration). Use when the user asks "why did task X fail on guarded", "why is the guarded run worse than baseline on Y", "why did guarded cost more on Z", "what went wrong with run_<id>", or wants to investigate a flaky or expensive (scenario, task)…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
bashTypeScript
Host repository
pixiebrix/agent-browser-shield
Host stars
35
Host language
TypeScript