Compare / head-to-head

MAI-Thinking-1vsSeed1.8

MAI-Thinking-1 leads 2 of 3 shared benchmarks. Both offer a 256K-token context window.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
2 – 1
MAI-Thinking-1 leads
Cheaper per token
—
list price, input + output
Larger context
Tie
both 256K tokens
Providers
2 providers
Microsoft · ByteDance Seed
MAI-Thinking-1
Microsoft · released 2026-08-12
text
Context
256K
Max out
64K
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
15 · 7 core
Seed1.8
ByteDance Seed · released 2025-12-18
textvisionvideo
Context
256K
Max out
64K
Input /1M
$0.25 ↗
Output /1M
$2 ↗
Cached /1M
$0.05
Scores
16 · 6 core
Quality

Benchmark matrix

BenchmarkMAI-Thinking-1Seed1.8ΔEdge
Reported by both · 3
SWE-bench Verified73.5% ↗72.9% ↗+0.6 ptMAI-Thinking-1
MMLU-Pro85% ↗84.9% ↗+0.1 ptMAI-Thinking-1
MultiChallenge53% ↗66.7% ↗-13.7 ptSeed1.8
Only MAI-Thinking-1 reports · 12
GPQA Diamond84.2% ↗not reported——
AIME 202597% ↗not reported——
AIME 202694.5% ↗not reported——
AdvancedIF85% ↗not reported——
BFCL v372% ↗not reported——
GraphWalks (<=128K)90% ↗not reported——
HMMT February 202684.9% ↗not reported——
IFBench69% ↗not reported——
LiveCodeBench v687.7% ↗not reported——
SimpleQA Verified31% ↗not reported——
SWE-bench Pro52.8% ↗not reported——
Terminal-Bench 2.046% ↗not reported——
Only Seed1.8 reports · 13
AIME-25not reported94.3% ↗——
ARC-AGI-1not reported67.9% ↗——
BFCL-v4not reported57.2% ↗——
BrowseComp-ennot reported67.6% ↗——
GPQA-Diamondnot reported83.8% ↗——
HLE (text-only, agentic search)not reported40.9% ↗——
LiveCodeBench(v6)not reported79.5% ↗——
MM-BrowseCompnot reported46.3% ↗——
Multi-SWE-Benchnot reported42% ↗——
OSWorld-Verified (XLANG)not reported61.9% ↗——
Terminal Bench 2.0not reported45.2% ↗——
WideSearchnot reported63.8% ↗——
τ2-Benchnot reported72% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is MAI-Thinking-1 minus Seed1.8 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecMAI-Thinking-1Seed1.8Edge
Context window256K tokens256K tokensTie
Max output64K tokens64K tokensTie
Input price / 1M—$0.25 ↗—
Output price / 1M—$2 ↗—
Cached input / 1M—$0.05 ↗—
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
—$2.25—
Modalitiestexttext · vision · videoSeed1.8
Released2026-08-122025-12-18—
Cited benchmark scores1516—
Reliability

Provider status

All providers →
More matchups

Seed1.8 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run MAI-Thinking-1 and Seed1.8 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.