Benchmark
ClawEval-MM
ClawEval-MM is the multimodal variant of ClawEval, evaluating agentic problem solving with visual inputs.
- Modality
- multimodal
- Categories
- multimodal, agents, vision
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 4
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Qwen3.7-Plus | Qwen | 0.557 | 2026-05-31 | evidence |
| Qwen3.8-27B | Qwen | 0.569 | 2026-08-14 | evidence |
| Seed 2.1 Pro | ByteDance | 0.51 | 2026-06-24 | evidence |
| Seed 2.1 Turbo | ByteDance | 0.46 | 2026-06-24 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.