Benchmark

ZebraLogic

ZebraLogic is an evaluation framework for assessing large language models' logical reasoning capabilities through logic grid puzzles derived from…

Modality
text
Categories
reasoning
Openness
unknown
Source
llm_stats
Reported scores
8

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Kimi K2 InstructMoonshot AI0.892025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.892025-09-05evidence
LongCat-Flash-ChatMeituan0.8932025-08-29evidence
LongCat-Flash-ThinkingMeituan0.9552025-09-22evidence
MiniMax M1 40KMiniMax0.8012025-06-16evidence
MiniMax M1 80KMiniMax0.8682025-06-16evidence
Qwen3 VL 235B A22B ThinkingQwen0.9732025-09-22evidence
Qwen3-235B-A22B-Instruct-2507Qwen0.952025-07-22evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.