BigHugger
GH Repository · truera

trulens

Evaluation and Tracking for LLM Experiments and AI Agents

stars
3,551
30-day movement
+72/day
Related entries
60
Connections
2
python/poetrypythonnodemachine-learningneural-networksllmsexplainable-mlllmopsllm-evaluationevalsmakeai-monitoringai-agentsPythonagent-evaluationllm-evalagent-appai-observabilityagentops

TruLens is a Python toolchain (Poetry-managed, with Make and Node components) for evaluating and tracking LLM experiments and AI agents. Its topics span llmops, ai-observability, and agent evaluation, indicating it monitors and scores LLM and agent behavior.

Reach for it when you need evaluation and tracking around LLM experiments or agent runs rather than ad-hoc logging.

Use it to

  • Evaluate LLM experiment outputs
  • Track AI agent behavior over runs
  • Add observability to LLM applications
  • Measure agent performance with evals
  • Adopt llmops monitoring practices

For Teams building and evaluating LLM apps and agents

Role
agent-app
Language
Python
Licence
MIT
Forks
340
Open issues
20
Last push
2026-09-14
Latest release
0.0.5 · 2021-02-03
topicsllm-evaluationagent-evaluationai-observabilityllmopsevalsmonitoring