Benchmark

MAXIFE

MAXIFE is a multilingual benchmark evaluating LLMs on instruction following and execution across multiple languages and cultural contexts.

Modality
text
Categories
general
Openness
unknown
Source
llm_stats
Reported scores
11

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen3.5-0.8BQwen0.3922026-03-02evidence
Qwen3.5-122B-A10BQwen0.8792026-02-24evidence
Qwen3.5-27BQwen0.882026-02-24evidence
Qwen3.5-2BQwen0.6062026-03-02evidence
Qwen3.5-35B-A3BQwen0.8662026-02-24evidence
Qwen3.5-397B-A17BQwen0.8822026-02-16evidence
Qwen3.5-4BQwen0.782026-03-02evidence
Qwen3.5-9BQwen0.8342026-03-02evidence
Qwen3.6 PlusQwen0.8822026-04-02evidence
Qwen3.7 MaxQwen0.8922026-05-19evidence
Qwen3.7-PlusQwen0.8882026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.