BigHugger
sk Skill · JasonColapietro

suede-ai-eval

Suede Labs AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates. Use when a change ships LLM, RAG, agent, classifier, prompt, or generated-media behavior, or when asked to write evals for an AI feature, design test cases for a model surface, audit existing eval coverage, or judge…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
bashJavaScript
Host repository
JasonColapietro/suede-creator-skills
Host stars
137
Host language
JavaScript