Agent
chinese-llm-benchmark
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。.
Interfaces & auth
Interfaces: not recorded
Auth methods: not recorded
Verification history
Verification is automated (see methodology). Every run is stored permanently; consecutive runs with an identical result are shown as one row.
2026-09-30 23:03 UTC · battery 1.2.0 · same result on 10 runs since 2026-09-26
partial
endpoint probe
skip
auth discoverable
skip
homepage reachable
pass
source repo reachable
skip
capabilities plausible
skip
pricing currency figure
skip
2026-09-18 13:37 UTC · battery 1.2.0
unreachable
endpoint probe
skip
auth discoverable
skip
homepage reachable
fail
source repo reachable
skip
capabilities plausible
skip
pricing currency figure
skip
2026-09-17 22:52 UTC · battery 1.2.0 · same result on 40 runs since 2026-08-24
partial
endpoint probe
skip
auth discoverable
skip
homepage reachable
pass
source repo reachable
skip
capabilities plausible
skip
pricing currency figure
skip