A clear way to check whether an AI provider is ready for desktop work.
SADB is a released, deterministic benchmark. It tests the complete path from model and provider to the answers a desktop application receives.
Stability first, capability second.
Five identical requests check whether the configured path responds consistently. Then 72 short, deterministic tasks check whether the model can follow instructions and solve desktop-oriented problems.
A successful sample needs an explicit 2xx response and non-empty content.
Each exact answer is one raw point. There is no model judge and no partial credit.
Protocol, bank, scorer, fingerprint and request parameters must match before results are compared.
SADB V1 is frozen and reviewable.
Question bank sadb-bank-1.0 and scorer sadb-score-1.0 are included in the release. Chinese and English banks share case IDs and semantics, but their results are reported separately.
Read the methodSADB is not an IQ test and is not a universal model leaderboard. It is a practical diagnostic for the provider path a desktop app actually uses.
