chinese-llm-benchmark
非线智能 NoneLinear — ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。
- stars
- 6,442
- 30-day movement
- starts with the next reading
- Related entries
- 60
- Connections
- 1
A continuously updated Chinese-language LLM capability benchmark (ReLE) covering 374 models, including commercial systems like ChatGPT, Gemini, Claude, ERNIE, Qwen, Baichuan, iFlytek Spark and SenseChat, plus open-source models like DeepSeek, Kimi, GLM, Llama and Mistral. Beyond a leaderboard, it maintains a defect database of over 2 million entries for community research and model improvement.
You need a Chinese-focused, regularly refreshed evaluation reference with leaderboard and defect data rather than building your own benchmark.
Use it to
- Compare Chinese LLM capabilities across commercial and open-source models
- Browse the defect database to study model failure patterns
- Track newly released models via continuous updates
- Support research on analyzing and improving LLMs
For Chinese NLP researchers and model developers
- Role
- agent-framework
- Forks
- 265
- Open issues
- 16
- Last push
- 2026-09-12
- Latest release
- v1.0 · 2023-06-10