Decision guide · free roster vs paid frontier

Benchmarks

Start with the best free model for a common task, then inspect category detail against one paid reference model. Scores use the public Artificial Analysis dataset.

Data source: Live Artificial Analysis dataLast generated: 2026-09-25 14:02

Reading the scores: ‡ and “paid proxy” mean the free roster model has no free-variant AA record, so the paid variant is shown as a clearly labelled proxy. It is not a free-tier measurement.

Everyday model

muse-spark-1.3-contributor-free48.10‡paid proxy

Highest general intelligence score in the current free roster.

Research

muse-spark-1.3-contributor-free93.5%‡paid proxy

Highest GPQA score in the current free roster.

Coding

muse-spark-1.3-contributor-free75.80‡paid proxy

Highest coding score in the current free roster.

Vision

qwen/qwen3.8-27b:free33.70‡paid proxy

Highest intelligence score among free models accepting image input.

Overall intelligence

AA Intelligence Index — general knowledge, reasoning, agents, and scientific performance. 4 roster models without this metric are omitted.

Overall intelligence free models and paid reference
ModelScore
muse-spark-1.3-contributor-freeFree 48.10‡★paid proxy
qwen/qwen3.8-27b:freeFree 33.70‡paid proxy
upstage/solar-pro4:freeFree 28.20‡paid proxy
Paid reference 57.60

Coding

AA Coding Index — code generation, completion, and review. 5 roster models without this metric are omitted.

Coding free models and paid reference
ModelScore
muse-spark-1.3-contributor-freeFree 75.80‡★paid proxy
qwen/qwen3.8-27b:freeFree 68.10‡paid proxy
mimo-v2.5-freeFree 56.80‡paid proxy
Paid reference 81.60

Agentic coding / terminal

Terminal-Bench Hard — agentic terminal usage, tool use, and multi-step workflows. 14 roster models without this metric are omitted.

Agentic coding / terminal free models and paid reference
ModelScore
mimo-v2.5-freeFree 41.7%‡★paid proxy
nvidia/nemotron-3-ultra-550b-a55b:freeFree 36.4%‡paid proxy
nemotron-3-ultra-freeFree 36.4%‡paid proxy
Paid reference 65.9%

Scientific reasoning

GPQA Diamond — graduate-level scientific reasoning benchmark. 5 roster models without this metric are omitted.

Scientific reasoning free models and paid reference
ModelScore
muse-spark-1.3-contributor-freeFree 93.5%‡★paid proxy
qwen/qwen3.8-27b:freeFree 90.5%‡paid proxy
thinkingmachines/inkling-small:freeFree 89.5%‡paid proxy
Paid reference 96.3%

Vision / multimodal

Free models accepting image input, ranked by the AA Intelligence Index because the public AA dataset has no dedicated MMMU field.

Vision-capable free models and paid reference
ModelScore
qwen/qwen3.8-27b:freeFree 33.70‡★paid proxy
thinkingmachines/inkling-small:freeFree 27.80‡paid proxy
thinkingmachines/inkling:freeFree 25.00‡paid proxy
Paid reference 57.60