BigHugger
sk Skill · ContextJet-ai

build-eval-dataset

Use this to build a good evaluation dataset for an LLM app, the part everyone underestimates. Trigger on "make an eval set", "what should I test my LLM on", "I don't have test data for my prompt", "build a golden dataset", or before setting up evals. A great eval set beats a great metric; garbage-in means your evals lie to you.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
Python
Host repository
ContextJet-ai/awesome-llm-observability
Licence
CC0-1.0
Host stars
34
Host language
Python