Verifiedon 2026.7.1
Action boundary
Before you act
- Expected result
- A reviewer can see the fixtures, scoring rule, run conditions, raw observations, exclusions, and decision threshold.
- Failure mode
- Different prompts, changing settings, or hand-picked outputs make a comparison look decisive when it is not.
- Rollback
- Mark the comparison inconclusive and return to the last approved route.
A benchmark is a decision instrument
Use the smallest set of fixtures that covers the work you actually intend to route. Five is a useful starting point, not magic: normal, hard, long-context, malformed, and escalation. Keep the same input, system instructions, tool availability, and scoring rule for every candidate unless the difference itself is what you are testing.
Score before you look
Create a rubric with separate columns for task success, format compliance, unsupported claims, safety/escalation behavior, latency, and cost or resource use. Set pass thresholds first. If human review is required, blind the reviewer to the candidate where practical and keep rationale beside the score.
| Case | Pass rule | Quality score | Time | Cost/resource observation | Notes |
|---|---|---|---|---|---|
| Normal | Required output is complete and grounded | /5 | |||
| Hard | Handles ambiguity without inventing | /5 | |||
| Long | Preserves supplied constraints | /5 | |||
| Malformed | Requests clarification or fails safely | /5 | |||
| Escalation | Stops at the approval boundary | /5 |
Lab: run a clean comparison
Run each candidate against the same fixture pack more than once. Keep run identifiers, timestamps, candidate/version identity, environment notes, and any failed run. Do not quietly remove a bad result. If an output cannot be judged, say why and add a fixture or scoring rule; do not turn uncertainty into a point for your preferred option.
Learner artifact — benchmark pipeline: Draw fixture pack → fixed run sheet → candidate A and candidate B → blinded review → score ledger → decision gate. Add a side lane labelled “excluded run” that requires a reason and approver; excluded results remain visible.
Checkpoint
Which change invalidates a direct A/B comparison unless explicitly controlled?
- A. Running the same fixtures twice
- B. Changing the system instruction for one candidate because it “needs a little help”
- C. Recording a timeout
Answer: B. It may be a valid route-level test, but it is no longer a simple model comparison. Describe the whole route and compare like with like.
Worked check
If candidate A wins on fluent prose but fails two required output fields, it has not met the workload brief. Write “fails structured-output threshold” in the ledger. Do not average away a hard requirement with a prettier response.
Source provenanceVerification and sources
Review receipt rr_local_models_benchmark_fixture
- Outcome
- approved
- Method
- source-review
- Reviewer
- academy-specification-review
- Reviewed
Evidence
- academy-spec — curriculum-local-models-contract; snapshot
1c60fa76ac90…
Limitations
- Approval covers the original workload, benchmark, routing, quota, and fallback method for 2026.7.1; it makes no universal provider, model, GPU, performance, privacy, or savings claim.
Lesson checkpoint
Ready to move on?
Mark this lesson complete when you can apply its outcome without relying on the examples above.