Benchmark

Job Bench

Job Bench evaluates AI agents on realistic professional tasks that require multi-step planning, research, and production of work artifacts.

Modality
text
Categories
productivity, reasoning, agents
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Kimi K3Moonshot AI0.5292026-07-16evidence
Muse Spark 1.1Meta0.5472026-07-09evidence
Qwen3.8 MaxQwen0.5342026-08-02evidence
Qwen3.8-27BQwen0.3342026-08-14evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.