SADB V1 · Reference snapshot

Seven views of the same 72-question bank.

These results are a provider-path snapshot collected with the frozen Chinese V1 bank. They are useful for understanding strengths and gaps, not for making a universal claim about a model.

OverviewResultsMethod
Model ranking

Capability points and request stability

Collected 2026-08-31, Chinese bank, five identical stability samples. A score is shown as raw points followed by completion of the 72-question bank.

01

zai-org/GLM-5.2

Capability leader

55 (76.39%)100% stable
02

moonshotai/Kimi-K2.7-Code

Strong all-rounder

53 (73.61%)100% stable
03

deepseek-ai/DeepSeek-V4-Pro

Reliable depth

52 (72.22%)100% stable
04

Qwen/Qwen3.5-397B-A17B

49 (68.06%)100%
05

deepseek-ai/DeepSeek-V4-Flash

28 (38.89%)100%
06

nex-agi/Nex-N2-Pro

22 (30.56%)100%
07

meituan-longcat/LongCat-2.0

5 (6.94%)N/A
Seven-model comparison

A vertical view keeps the gaps visible.

Chinese and English completion points are shown side by side for all seven tested models.

SADB V1 seven-model vertical capability ranking chart
System and hardware diagnostics14 questions
GLM 11 · Kimi 10 · Pro 9 · Flash 3 · Nex 2
Desktop operations and task planning12 questions
GLM 9 · Kimi 7 · Pro 8 · Flash 4 · Nex 2
File, document and data processing10 questions
GLM 8 · Kimi 7 · Pro 6 · Flash 5 · Nex 4
Software configuration and troubleshooting10 questions
GLM 5 · Kimi 7 · Pro 7 · Flash 2 · Nex 2
Programming, scripting and automation10 questions
GLM 7 · Kimi 7 · Pro 7 · Flash 3 · Nex 3
Security, permissions and privacy8 questions
GLM 8 · Kimi 8 · Pro 8 · Flash 5 · Nex 5
General logic and mathematical reasoning8 questions
GLM 7 · Kimi 7 · Pro 7 · Flash 3 · Nex 3