sk Skill · hparreao
score-robustness
Measure accuracy and consistency of an LLM across prompt, sampling, and ordering perturbations using SCORE-style evaluation. Use when assessing reliability beyond a single model response.
Open on skills.sh ↗read 2026-09-19
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 1
- Connections
- 0
Python
- Host repository
- hparreao/Awesome-AI-Evaluation-Guide
- Host stars
- 19
- Host language
- Python