Lesson 3 of 6 · 0%Build a benchmark that can reject your favourite modelNext
Course map

Local Models, Performance and Cost Control

0 of 6 complete0 of 6

Lesson 3.3 · 60 minutes

Build a benchmark that can reject your favourite model

Create a small, repeatable benchmark with fixed fixtures, explicit scoring, controlled runs, and an audit trail strong enough to overturn a preference.

Skip course map
Current lessonBuild a benchmark that can reject your favourite model0% complete · 0/6 lessons

Local Models, Performance and Cost Control

0% complete · Current: Build a benchmark that can reject your favourite model

Verifiedon 2026.7.1

Action boundary

Before you act

Expected result
A reviewer can see the fixtures, scoring rule, run conditions, raw observations, exclusions, and decision threshold.
Failure mode
Different prompts, changing settings, or hand-picked outputs make a comparison look decisive when it is not.
Rollback
Mark the comparison inconclusive and return to the last approved route.

A benchmark is a decision instrument

Use the smallest set of fixtures that covers the work you actually intend to route. Five is a useful starting point, not magic: normal, hard, long-context, malformed, and escalation. Keep the same input, system instructions, tool availability, and scoring rule for every candidate unless the difference itself is what you are testing.

Score before you look

Create a rubric with separate columns for task success, format compliance, unsupported claims, safety/escalation behavior, latency, and cost or resource use. Set pass thresholds first. If human review is required, blind the reviewer to the candidate where practical and keep rationale beside the score.

Case Pass rule Quality score Time Cost/resource observation Notes
Normal Required output is complete and grounded /5
Hard Handles ambiguity without inventing /5
Long Preserves supplied constraints /5
Malformed Requests clarification or fails safely /5
Escalation Stops at the approval boundary /5

Lab: run a clean comparison

Run each candidate against the same fixture pack more than once. Keep run identifiers, timestamps, candidate/version identity, environment notes, and any failed run. Do not quietly remove a bad result. If an output cannot be judged, say why and add a fixture or scoring rule; do not turn uncertainty into a point for your preferred option.

Learner artifact — benchmark pipeline: Draw fixture pack → fixed run sheet → candidate A and candidate B → blinded review → score ledger → decision gate. Add a side lane labelled “excluded run” that requires a reason and approver; excluded results remain visible.

Checkpoint

Which change invalidates a direct A/B comparison unless explicitly controlled?

  • A. Running the same fixtures twice
  • B. Changing the system instruction for one candidate because it “needs a little help”
  • C. Recording a timeout

Answer: B. It may be a valid route-level test, but it is no longer a simple model comparison. Describe the whole route and compare like with like.

Worked check

If candidate A wins on fluent prose but fails two required output fields, it has not met the workload brief. Write “fails structured-output threshold” in the ledger. Do not average away a hard requirement with a prettier response.

Source provenanceVerification and sources

Review receipt rr_local_models_benchmark_fixture

Outcome
approved
Method
source-review
Reviewer
academy-specification-review
Reviewed

Evidence

  • academy-spec — curriculum-local-models-contract; snapshot 1c60fa76ac90…

Limitations

  • Approval covers the original workload, benchmark, routing, quota, and fallback method for 2026.7.1; it makes no universal provider, model, GPU, performance, privacy, or savings claim.

Open the public evidence snapshot

Lesson checkpoint

Ready to move on?

Mark this lesson complete when you can apply its outcome without relying on the examples above.