Benchmarks / decision evidence
No winner before a run.
This room exposes the experiment design before it exposes a score. There are currently no completed runs, model rankings, or recommendation claims.
Reproducible model comparisons organized around real building decisions.
What the workbench refuses to fake.
- 01Decision first
Every comparison begins with a real use case and binding constraint.
- 02Exact identities
Provider model IDs, settings, region, and reference dates must be recorded.
- 03Failures count
Errors, refusals, and retries remain visible in the denominator.
- 04Scoped verdict
A use-case result cannot establish one universal best model.
Inspect the contract, not a synthetic leaderboard.
Choose a use-case lens, inspect the proposed manifest, and open the empty results contract. Case and repeat counts below are proposed experiment designs, not completed calls.
Benchmark before belief.
Design a reproducible decision, register the run contract, and keep unmeasured cells visibly empty.
Choose a model for a product copilot that turns mixed discovery evidence into traceable decisions.
Synthesize interviews, support evidence, product constraints, and metrics into a decision brief without inventing certainty.
- Binding constraint
- Evidence fidelity before eloquence
- Proposed design
- 24 cases × 3 repeats × candidate count
- Quality bar
- At least 90% deterministic evidence-link checks and a blind rubric score of 4/5 or better.
- Failure tolerance
- No fabricated source, unsupported priority, or silent omission of a blocking constraint.
No benchmark result exists for PM tools.
That is the honest starting state. Exact candidates, immutable inputs, current official pricing, and execution approval must be registered before a single model call is made.
- Exact model IDs + reference datesrequired
- Dataset and prompt versions + hashesrequired
- Official pricing source + UTC check timerequired
- Provider, data, billing, and artifact approvalclosed
Representative, difficult, and known-failure cases.
- messy interview synthesis
- conflicting evidence
- PRD critique
- decision memo
- roadmap constraint
- plausible but uncited claims
- majority-vote prioritisation
- flattened stakeholder disagreement
Use redacted research fixtures unless every approved provider can accept the same private evidence class.
Interactive draft returned within the declared product wait budget.