Everyday model
muse-spark-1.3-contributor-free48.10‡paid proxyHighest general intelligence score in the current free roster.
Start with the best free model for a common task, then inspect category detail against one paid reference model. Scores use the public Artificial Analysis dataset.
Reading the scores: ‡ and “paid proxy” mean the free roster model has no free-variant AA record, so the paid variant is shown as a clearly labelled proxy. It is not a free-tier measurement.
muse-spark-1.3-contributor-free48.10‡paid proxyHighest general intelligence score in the current free roster.
muse-spark-1.3-contributor-free93.5%‡paid proxyHighest GPQA score in the current free roster.
muse-spark-1.3-contributor-free75.80‡paid proxyHighest coding score in the current free roster.
qwen/qwen3.8-27b:free33.70‡paid proxyHighest intelligence score among free models accepting image input.
AA Intelligence Index — general knowledge, reasoning, agents, and scientific performance. 4 roster models without this metric are omitted.
| Model | Score |
|---|---|
muse-spark-1.3-contributor-free | Free 48.10‡★paid proxy |
qwen/qwen3.8-27b:free | Free 33.70‡paid proxy |
upstage/solar-pro4:free | Free 28.20‡paid proxy |
| Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic | Paid reference 57.60 |
AA Coding Index — code generation, completion, and review. 5 roster models without this metric are omitted.
| Model | Score |
|---|---|
muse-spark-1.3-contributor-free | Free 75.80‡★paid proxy |
qwen/qwen3.8-27b:free | Free 68.10‡paid proxy |
mimo-v2.5-free | Free 56.80‡paid proxy |
| Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic | Paid reference 81.60 |
Terminal-Bench Hard — agentic terminal usage, tool use, and multi-step workflows. 14 roster models without this metric are omitted.
| Model | Score |
|---|---|
mimo-v2.5-free | Free 41.7%‡★paid proxy |
nvidia/nemotron-3-ultra-550b-a55b:free | Free 36.4%‡paid proxy |
nemotron-3-ultra-free | Free 36.4%‡paid proxy |
| GPT-5.6 Sol (max)OpenAI | Paid reference 65.9% |
GPQA Diamond — graduate-level scientific reasoning benchmark. 5 roster models without this metric are omitted.
| Model | Score |
|---|---|
muse-spark-1.3-contributor-free | Free 93.5%‡★paid proxy |
qwen/qwen3.8-27b:free | Free 90.5%‡paid proxy |
thinkingmachines/inkling-small:free | Free 89.5%‡paid proxy |
| GPT-6 Astra (xhigh)OpenAI | Paid reference 96.3% |
Free models accepting image input, ranked by the AA Intelligence Index because the public AA dataset has no dedicated MMMU field.
| Model | Score |
|---|---|
qwen/qwen3.8-27b:free | Free 33.70‡★paid proxy |
thinkingmachines/inkling-small:free | Free 27.80‡paid proxy |
thinkingmachines/inkling:free | Free 25.00‡paid proxy |
| Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic | Paid reference 57.60 |