A small, fixed protocol that stays easy to inspect.
The V1 design favors reproducibility over spectacle: short prompts, exact answers, explicit transport status and a frozen identity tuple.
72 questions across seven desktop domains.
Every case has a stable ID, a prompt, a canonical answer and a domain. Difficulty labels help describe coverage but do not change the raw one-point scoring rule.
Two suites, one result record.
Send the same request five times. A sample passes only with an explicit 2xx response and non-empty content. Report success rate, P50/P95 latency and structured errors.
Run all 72 independent cases. Normalize the returned text and compare it with the canonical answer. One exact match equals one point.
Chinese and English banks are independent runs. Never merge, average or silently substitute their scores. The Super Optimizer reference uses Chinese only.
API keys are read from the environment, never serialized into results, and raw model responses are not persisted by the standard runner.
sadb/1.0
Direct comparisons require the same protocol, question bank, scorer, SHA-256 fingerprint and request parameters. If any part changes, publish a new identity instead of relabeling an old result.
7133690d67ef91609aaf1810a1d8ebc659cee1c016d381eaf388e015d29b5598
Question bank SHA-256
