BigHugger
sk Skill · hparreao

score-robustness

Measure accuracy and consistency of an LLM across prompt, sampling, and ordering perturbations using SCORE-style evaluation. Use when assessing reliability beyond a single model response.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
Python
Host repository
hparreao/Awesome-AI-Evaluation-Guide
Host stars
19
Host language
Python