BigHugger
sk Skill · Aperivue

mllm-eval

Design or audit a model-agnostic evaluation harness for an LLM or multimodal LLM on a clinical task (radiology report generation, visual question answering, clinical text extraction/classification) — the adjudicated reference standard, clinical-efficacy metrics (RadGraph-F1 / CheXbert-F1 beyond BLEU/ROUGE), faithfulness and hallucination, pretraining-contamination of public benchmarks, prompt-sensitivity and…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
6
bashPython
Host repository
Aperivue/medsci-skills
Allowed tools
Read, Write, Edit, Bash, Grep, Glob
Host stars
305
Host language
Python