BigHugger
sk Skill · imlrz

evaluate-inmind

Evaluate long-term memory systems on the InMind benchmark using the fixed LME-s background, canonical middle injection, shared answer prompt, and binary judges. Use when an agent needs to integrate a memory method, run InMind tasks, construct per-task timelines, validate result JSONL, build judge requests, compare metrics, or prepare a leaderboard submission.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
bashPython
Host repository
imlrz/InMind
Host stars
17
Host language
Python