ModelsCompareBest forBenchmarksStatusPricingAPI
Compare / head-to-head

ERNIE 5.0vsLongCat Flash Chat

ERNIE 5.0 leads 7 of 7 shared benchmarks. Both offer a 128K-token context window.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
7 – 0
ERNIE 5.0 leads
Cheaper per token
list price, input + output
Larger context
Tie
both 128K tokens
Provider uptime (30d)
— · —
Baidu · Meituan LongCat
ERNIE 5.0
Baidu · released 2026-02-06
textvision
Context
128K
Max out
66K
Input /1M
Output /1M
Cached /1M
Scores
21 · 5 core
LongCat Flash Chat
Meituan LongCat · released 2025-09-01
text
Context
128K
Max out
Input /1M
Output /1M
Cached /1M
Scores
31 · 9 core
Quality

Benchmark matrix

BenchmarkERNIE 5.0LongCat Flash ChatΔEdge
Reported by both · 7
GPQA Diamond86.36% 73.23% +13.1 ptERNIE 5.0
MMLU-Pro83.8% 82.68% +1.1 ptERNIE 5.0
AIME 202589.06% 61.25% +27.8 ptERNIE 5.0
HumanEval+94.48% 88.41% +6.1 ptERNIE 5.0
IFEval93.35% 89.65% +3.7 ptERNIE 5.0
MBPP+82.54% 79.63% +2.9 ptERNIE 5.0
ZebraLogic96.5% 89.3% +7.2 ptERNIE 5.0
Only ERNIE 5.0 reports · 14
ACEBench Chinese89.6% not reported
ACEBench English87.7% not reported
BBEH66.63% not reported
BFCL v466.47% not reported
BrowseComp-ZH64.71% not reported
ChineseSimpleQA86.03% not reported
HMMT 202579.58% not reported
Humanity's Last Exam25.81% not reported
LiveCodeBench v676.21% not reported
Multi-IF85.56% not reported
MultiChallenge65.98% not reported
SimpleQA74.01% not reported
SpreadsheetBench40.08% not reported
Tau2-Bench78.79% not reported
Only LongCat Flash Chat reports · 24
SWE-bench Verifiednot reported60.4%
MMLUnot reported89.71%
LMArena Elonot reported1422.9
AceBenchnot reported76.1%
AIME 2024not reported70.42%
ArenaHard-V2not reported86.5%
BeyondAIMEnot reported43%
C-Evalnot reported90.44%
CMMLUnot reported84.34%
COLLIEnot reported57.1%
DROPnot reported79.06%
GraphWalks-128knot reported51.05%
LiveCodeBenchnot reported48.02%
MATH500not reported96.4%
Meeseeks-zhnot reported43.03%
Safety: Criminalnot reported91.24%
Safety: Harmfulnot reported83.98%
Safety: Misinformationnot reported81.72%
Safety: Privacynot reported93.98%
Tau2-Bench (airline)not reported58%
Tau2-Bench (retail)not reported71.27%
Tau2-Bench (telecom)not reported73.68%
TerminalBenchnot reported39.51%
VitaBenchnot reported24.3%
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is ERNIE 5.0 minus LongCat Flash Chat in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecERNIE 5.0LongCat Flash ChatEdge
Context window128K tokens128K tokensTie
Max output66K tokens
Input price / 1M
Output price / 1M
Cached input / 1M
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
Modalitiestext · visiontextERNIE 5.0
Released2026-02-062025-09-01
Cited benchmark scores2131
Reliability

Provider status

All providers
Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run ERNIE 5.0 and LongCat Flash Chat on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.