Skip to content

Eval harness

course 11 of 22

lesson 09 · lessons/09-eval-harness

1 h About $0.10 Mac, then Fireworks

What you will do and why

You build the table that decides between models: how often each is right, how fast it answers and what it costs. Every model sorts the same customer support tickets. Customers switch models on that table, not on a hunch.

Why it matters: The lesson’s example: a big general model sorts 81% of support tickets into the right category, at $0.42 per 1,000 tickets. A small fine-tuned model (one given extra training on examples of this task) sorts 96% at $0.02, and answers about twice as fast.

You are done when: The practice run, your Mac’s two servers and the Fireworks model have each printed a row, all saved in results/09-eval.csv. You have written down the Fireworks model’s categories right, p95, cost per 1,000 tickets and judge score: the bar step 2’s fine-tune has to clear.

Training Techniques · 3:14

Download mp4 (11.4 MB)

Chapters

In this video The four standard ways to change how an existing model behaves, how to choose one from what is going wrong, and why LoRA (training a small add-on instead of the whole model) keeps it cheap.

The narrator’s “next video” is the file order. Your next step is below.

3 key points

  1. Pick the method from the kind of mistake: examples teach format, comparisons teach taste, an automatic score teaches correctness.

    Can you write the right answer? Supervised fine-tuning (SFT: training on examples of it). Can you only say which of two answers is better? Preference tuning (DPO, direct preference optimization). Can a program score an answer? Reinforcement fine-tuning (RFT). Missing facts need looking up (retrieval), not training.

  2. LoRA (low-rank adaptation) keeps it cheap: the model stays frozen (unchanged) and only a small add-on is trained, so one running model can carry many add-ons.

    The video’s ticket example: 2,000 examples x 600 tokens (word pieces) x 2 passes = 2.4 million training tokens, about $1.20 at $0.50 per million. Accuracy rose from 81% to 89%, and the add-on is about 30 MB, a tiny fraction of the model.

  3. Try a better prompt and retrieval first: they are free.

    Train only when those stop helping. Then look at what is still wrong: the video’s example stops at 89% and asks what the other 11% of mistakes look like, because that picks the next method.

An eval (short for evaluation) is a repeatable test: the same questions, graded the same way, for every model you are considering. Here the task is sorting support tickets. Each reply is graded three ways, cheapest first: automatic code checks, a second AI model acting as judge, and a person checking a sample. The results go in one table, with speed and cost beside quality.

Picture it

Hiring for a job: every candidate does the same work sample, instead of you trusting their CV. An automatic check catches the basic errors, a senior colleague scores the judgement calls, and you spot-check that colleague’s marks. Then you weigh the scores against salary and speed.

With real numberslesson 09’s example table and the kit’s ticket dataset

  • The test: 100 support tickets kept aside and never used for training. Each run grades 50 or 60 of them.
  • Each reply must give a category (billing, outage, how-to or abuse), a severity from 1 to 5 and a next step.
  • Big general model: 81% of categories right, 3.9 out of 5 from the AI judge, 95 of 100 replies back within 640 ms (thousandths of a second: about two thirds of a second), $0.42 per 1,000 tickets.
  • Small fine-tuned model: 96% right, 4.5 out of 5, within 310 ms, $0.02 per 1,000 tickets.
  • So the small model wins on all three: 15 points more accurate (96 - 81), about twice as fast (640 ÷ 310 = 2.1) and 21 times cheaper (0.42 ÷ 0.02).

Words to know

Eval (eval harness)
A repeatable test of a model on a fixed set of prompts. Example: the held-out tickets, graded three ways.
Held-out tickets
Test examples the model was never trained on, so a good score means it learned the task, not the answers. Example: the 100 tickets in test.jsonl.
Judge (LLM judge)
A second model that scores answers against a scoring guide (a rubric). Example: 3.9 against 4.5 out of 5.
p95
The value 95 out of 100 requests stay under. Example: 95 of 100 replies back within 310 ms.

From the lesson

lessons/09-eval-harness/README.md

Every customer conversation about models ends in a trade-off, and the table is how you have it:

candidate valid accuracy judge p95 $/1k tasks
big general model 100% 81% 3.9 640 ms $0.42
small fine-tune 100% 96% 4.5 310 ms $0.02

Never show one column alone. Better quality at 3× the latency may be the wrong call for a chat product and the right one for batch.

Graders, cheapest first:

  1. Code checks: schema validity, exact-match labels. They are free, deterministic and the first to run.
  2. LLM judge: a second model scores the fuzzy field (next_action) against a rubric. It is useful but biased, so hand-check about 20 of its scores before you trust it.
  3. Humans check a sample, especially before launch.

Always evaluate on held-out data (test.jsonl is never trained on). Otherwise you are grading the student on the answer key.

  • evaluate.py → parse_candidate() lets any target, model and price enter the same table.
  • RUBRIC is the judge prompt. Change it and watch the scores move, which shows why judges need calibrating.
  • usd_per_1k comes from real token counts (usage) × prices you supply.
Terminal window
cd lessons/09-eval-harness
python evaluate.py mock:mock-8b mock:mock-8b-lora --n 60 --judge mock # offline rehearsal
# untuned models on the same task: the bar a fine-tune must clear (step 2 today, Day 8)
python evaluate.py ollama mlx --n 50
python evaluate.py "fireworks:accounts/fireworks/models/gpt-oss-120b@0.15/0.60" --n 50 --judge fireworks

What each command does

  1. python evaluate.py mock:mock-8b mock:mock-8b-lora --n 60 --judge mock

    A free, offline rehearsal. The practice server pretends to be two models: mock-8b (untuned) and mock-8b-lora (a pretend fine-tune). Each sorts the same 60 test tickets, and the practice server also plays the judge. If it is not running, start it with make mock & from 03-labs. Look for the -lora row higher on category_% (categories right): about 93 against 75. These numbers are made up: never quote them.

  2. python evaluate.py ollama mlx --n 50

    The same test, free, on your Mac’s model from Day 2: Llama 3.1 8B at 4 bits, served by Ollama and by MLX, 50 tickets each. Look for real scores, and 0.0 in usd_per_1k: your own Mac charges nothing per token. judge_1to5 reads nan (not a number: empty), since no judge was asked.

  3. python evaluate.py "fireworks:accounts/fireworks/models/gpt-oss-120b@0.15/0.60" --n 50 --judge fireworks

    Run it as shown here and in the code block above, with gpt-oss-120b, not with the gpt-oss-20b the kit’s own README names. Why: on 27 September 2026 gpt-oss-20b’s model page said “Serverless: Not supported”. Fireworks no longer runs it per token, so the kit’s line would fail. gpt-oss-120b is still billed per token, at $0.15 per million tokens sent and $0.60 per million received (its model page, checked the same day). The run grades the same 50 tickets. @0.15/0.60 passes those two prices, so the cost column fills in. --judge fireworks has your default Fireworks model, also gpt-oss-120b, score each suggested next step. Paid: the lesson estimates about $0.10 on the old model. gpt-oss-120b costs about twice as much per token (0.15 ÷ 0.07 = 2.1 for tokens in, 0.60 ÷ 0.30 = 2 for tokens out), so allow up to about $0.20. Prices change: copy today’s from the model’s page. Look for a judge score between 1 and 5 and a cost above 0. Your categories right, p95 and cost will not match the lesson’s sample table: its big-model row (81%, 640 ms) was written for gpt-oss-20b, the model evaluate.py’s header names.

Prices are examples. Copy the current ones from the model’s page on fireworks.ai.

Optional: pip install eval-protocol and port one grader to Fireworks’ own eval framework, so you can talk about their tooling from experience.

How to read it

Each row is one model on the same test tickets. schema_valid_% is the share of replies in the right format; category_% and severity_% are the shares with the right category and the right urgency. judge_1to5 is the judge’s average, and nan (not a number) means no judge was asked. p95_ms is the time 95 of 100 replies came back within, and usd_per_1k the dollars per 1,000 tickets. On the practice server expect about 75% of categories right for mock-8b, about 93% for mock-8b-lora, and 0.0 dollars. The sample below is not from this exact command: it shows 0.015 dollars, which a run with no prices cannot print.

candidate schema_valid_% category_% severity_% judge_1to5 p95_ms usd_per_1k
mock mock-8b 100.0 75.0 77.5 3.9 407.4 0.0
mock mock-8b-lora 100.0 92.5 97.5 4.5 404.0 0.015

Write down the base (untuned) rows, above all the Fireworks one. They are the bar the fine-tunes in step 2 today and on Day 8 have to clear. Both also re-grade their own starting model, Qwen2.5 7B Instruct, beside the fine-tune.

2 questions. Say your answer out loud, then tap to check it.

A judge model (a second AI model that grades answers against a scoring guide) gives the fine-tuned model 4.5 out of 5 and the original model 3.9. Is that proof the fine-tune is better?Show answerHide

In plain words

No. First check that the judge agrees with you: score about 20 of the same answers yourself and compare. AI judges tend to reward longer, more polished answers even when they are no more correct.

Picture it

A teacher who gives longer essays in neat handwriting higher marks, whatever they say. Before you trust their grades, you mark 20 of the same essays yourself and see whether you agree.

With real numberslesson 09’s example table and its scoring guide

  • The judge scores each suggested next step from 1 to 5: 5 = specific, right owner, safe; 3 = plausible but vague; 1 = wrong or unsafe.
  • Average scores: 3.9 for the original model, 4.5 for the fine-tune, a gap of 0.6 points (4.5 - 3.9).
  • The check: score about 20 of the same answers by hand. If you and the judge often disagree, the 0.6 means little.
  • Code checks need no such trust: 81% against 96% of categories right is counted against the known answers.

Words to know

Judge (LLM judge)
A second model that scores answers against a rubric. Example: gpt-oss-120b in the Fireworks run.
Rubric
The scoring guide a judge follows. Example: 5 = specific, right owner, safe; 1 = wrong or unsafe.
Hand-labelled sample
Answers a person has marked right or wrong, used to check a judge. Example: about 20 scores, redone by you.
Bias (judge bias)
A judge’s habit of favouring something that is not correctness. Example: longer, more polished answers.
Go deeper: the engineer version

The kit's question

The judge gives the fine-tune 4.5 and the base 3.9. Is that proof?

The kit's answer

No. Check that the judge agrees with you on a hand-labelled sample first. Judges favour length and style.

More detail: The lesson also says to change the RUBRIC prompt and watch the scores move, which shows why judges need calibrating. evaluate.py prints only the average (judge_1to5), so to hand-check, print each judge_score() result next to its ticket. In the practice run the stand-in judge is rigged: it gives 4 or 5 when the next step names the ticket’s true category and 1 to 3 otherwise (felab/mock_server.py), so its scores track category accuracy. A real judge has no such anchor.

After fine-tuning, accuracy went up 17 points: 17 more tickets in every 100 sorted correctly. But p95 latency doubled: the time within which 95 of 100 replies come back is now twice as long. What do you recommend?Show answerHide

In plain words

Recommend it only if the slower replies still meet the customer’s speed promise (their SLO, service level objective). Show them both numbers so they choose. Then offer ways to keep the accuracy without the wait: fine-tune a smaller starting model, or add a small helper model that drafts the next few words for the big one to check (speculative decoding).

Picture it

A delivery service that gets more parcels to the right door but takes twice as long. For flowers that must arrive today, that is a no; for a monthly supplies order, it is fine.

With real numberslesson 09’s example table

  • The question’s trade: 17 more tickets right in every 100, but the time within which 95 of 100 replies come back is twice as long.
  • For scale, the eval lesson’s example table has p95 figures of 310 ms and 640 ms: 640 ÷ 310 = 2.1, about a third of a second against about two thirds.
  • In that same table the small fine-tuned model is the fast one: 96% right, 310 ms, $0.02 per 1,000 tickets, against the big model’s 81%, 640 ms, $0.42. That is why a smaller starting model is one of the answers.
  • The lesson’s own example is harsher, 3 times the wait: that can be the wrong call for a chat product and the right one for batch work that runs overnight (like Day 6’s Batch API).

Words to know

SLO (service level objective)
The speed promise. Example: 95 of 100 replies back within a set time.
p95 latency
The time within which 95 of 100 replies come back. Example: 640 ms against 310 ms.
Speculative decoding
A small draft model guesses several tokens and the big model checks them all in one pass; the output does not change. Example: Day 9.
Base model
The model you start from, before any fine-tuning. Example: a smaller one can be faster and cheaper.
Go deeper: the engineer version

The kit's question

Accuracy went up 17 points but p95 doubled. What do you recommend?

The kit's answer

It depends on the SLO. Put both in the memo, and consider a smaller base model or speculative decoding.

More detail: The p95 here is whole-request latency: evaluate.py times each call from sending it to the full reply, not time to first token. Speculative decoding (Day 9) cuts the time per token when a small draft model’s guesses are accepted. A smaller base model cuts time and cost together, and the lesson’s table shows a small fine-tune beating a big general model on all three columns.

Question A new model tops the public benchmarks. How would you decide whether it is right for us?

One clear answer

Benchmarks are for launches. For a customer I use three graders on their data and one table: quality, p95 latency and dollars per thousand tasks.

What this means

  • “Benchmarks are for launches”: Public benchmarks (standard tests every model maker reports scores on) help compare models when they come out. They do not measure this customer’s job.
  • “For a customer I use three graders”: I grade each answer three ways, cheapest first: code checks (free and exact), an AI judge with a scoring guide, and a person checking a sample.
  • “on their data”: On the customer’s own examples, kept aside from any training (held out), like the 100 test tickets here.
  • “and one table”: All the results side by side, so nobody chooses a model on one number.
  • “quality, p95 latency and dollars per thousand tasks”: How often it is right, how long 95 of 100 answers take, and what 1,000 tasks cost. The lesson’s example: 81% against 96%, 640 ms against 310 ms, $0.42 against $0.02.

Your numbersSaved on this device and collected in the Day 7 wrap-up.

Hint: The category_%, p95_ms and usd_per_1k columns of the run with --judge fireworks (the gpt-oss-120b line). Write the three in that order.

Hint: The judge_1to5 column of the same run. nan means no judge was asked.

Hint: The ollama and mlx rows of python evaluate.py ollama mlx --n 50: the same model, served two ways. Their usd_per_1k is 0.0, since your own Mac charges nothing per token.

Run make fw-check now: it should list nothing. It lists anything on Fireworks that is still billing, so an empty list means nothing is costing you money.

The practice run, your Mac’s two servers and the Fireworks model have each printed a row, all saved in results/09-eval.csv. You have written down the Fireworks model’s categories right, p95, cost per 1,000 tickets and judge score: the bar step 2’s fine-tune has to clear.

Connection refused on the practice run.
The practice server is not running. From 03-labs, run make mock &, then run the command again.
Connection refused on the ollama mlx run.
One of the two servers is not running: make check (from 03-labs) shows which. Start Ollama with ollama serve, and MLX in another terminal with bash lessons/02-three-local-servers/serve_mlx.sh.
FIREWORKS_API_KEY is not set
Your key belongs in the .env file in 03-labs, as on Day 1: a line FIREWORKS_API_KEY= followed by the key. Then run the command again.
The kit’s gpt-oss-20b line fails, or says the model is not found.
Expected: gpt-oss-20b is no longer offered per token (“Serverless: Not supported” on its model page, checked 27 September 2026). Run the gpt-oss-120b line from the code block instead: python evaluate.py "fireworks:accounts/fireworks/models/gpt-oss-120b@0.15/0.60" --n 50 --judge fireworks.
A Fireworks model id is not found, even on the gpt-oss-120b line.
Model ids change. Copy a current one from the model library on fireworks.ai that is offered serverless, put it in place of accounts/fireworks/models/gpt-oss-120b, and copy its current prices after the @ too.
The Fireworks row shows a low schema_valid_%, or empty replies.
gpt-oss-120b writes out its thinking before its answer (a reasoning model, as on Day 6). evaluate.py allows each reply 120 tokens (max_tokens), and the thinking may use them up, so the answer comes back cut off or empty. Write it down as a finding: it is part of that model’s row, not a broken setup.
The --judge fireworks run says a model is not found.
The judge is your default Fireworks model, gpt-oss-120b. Put a current id in 03-labs/.env as FIREWORKS_MODEL=accounts/fireworks/models/<id>, then run the command again.
judge_1to5 says nan.
No judge was asked: only runs with --judge fill that column. The ollama mlx run has no judge, so nan is expected there.
evaluate.py Python · 118 lines
"""
Lesson 09 — the eval harness: quality, latency and cost in ONE table.
"The fine-tune seems better" is not a decision. This is:
candidate schema_valid category severity judge p95_ms $/1k tasks
fireworks gpt-oss-20b 100% 81% 74% 3.9 640 $0.042
mlx triage-lora 100% 96% 93% 4.3 310 $0 (your Mac)
Each candidate is "target:model[@in_price/out_price]" — prices in $ per 1M tokens,
copied from the provider's pricing page (local = free, so omit).
Graders (cheap → expensive):
1. schema_valid code — does the JSON match the contract?
2. category/severity accuracy code — against held-out labels
3. judge an LLM with a rubric scores next_action 1–5 (--judge target:model)
Use for fields with no single right answer. Spot-check it by hand!
python evaluate.py mock:mock-8b mock:mock-8b-lora --n 60
python evaluate.py ollama "fireworks:accounts/fireworks/models/gpt-oss-20b@0.07/0.30" \
--judge fireworks --n 50
"""
import argparse
import json
import time
from felab import TARGETS, client, percentile, record, table
from felab.targets import Target, resolve
from felab.tickets import SCHEMA, SYSTEM, grade, load_test, parse
RUBRIC = """You are grading a support-triage assistant. Ticket:
{ticket}
Proposed next_action: {action}
Rubric: 5 = specific, correct owner, safe; 3 = plausible but vague; 1 = wrong or unsafe.
Reply with JSON only: {{"score": <1-5>}}"""
def parse_candidate(spec: str) -> tuple[Target, float, float]:
"""'fireworks:accounts/x/y@0.07/0.30' → (Target, in $/M, out $/M)."""
price_in = price_out = 0.0
if "@" in spec:
spec, price = spec.rsplit("@", 1)
price_in, price_out = (float(x) for x in price.split("/"))
name, _, model = spec.partition(":")
if name not in TARGETS:
raise SystemExit(f"unknown target '{name}' — choose from {list(TARGETS)}")
ns = argparse.Namespace(target=name, model=model or None, base_url=None)
return resolve(ns), price_in, price_out
def judge_score(jcli, jmodel: str, ticket: str, action: str) -> float | None:
r = jcli.chat.completions.create(model=jmodel, temperature=0, max_tokens=20,
messages=[{"role": "user", "content": RUBRIC.format(ticket=ticket, action=action)}])
obj = parse(r.choices[0].message.content)
s = obj.get("score") if obj else None
return float(s) if isinstance(s, (int, float)) and 1 <= s <= 5 else None
def evaluate(t: Target, pin: float, pout: float, rows: list[dict], judge) -> dict:
cli = client(t)
system = SYSTEM + "\nSchema:\n" + json.dumps(SCHEMA)
grades, lat, judged = [], [], []
tok_in = tok_out = 0
for r in rows:
t0 = time.perf_counter()
resp = cli.chat.completions.create(
model=t.model, temperature=0, max_tokens=120,
messages=[{"role": "system", "content": system}, {"role": "user", "content": r["ticket"]}],
response_format={"type": "json_schema", "json_schema": {"name": "triage", "schema": SCHEMA}})
lat.append((time.perf_counter() - t0) * 1000)
reply = resp.choices[0].message.content
grades.append(grade(reply, r))
u = resp.usage
tok_in += u.prompt_tokens if u else len(system + r["ticket"]) // 4
tok_out += u.completion_tokens if u else len(reply or "") // 4
if judge and (obj := parse(reply)) and obj.get("next_action"):
judged.append(judge_score(judge[0], judge[1], r["ticket"], obj["next_action"]))
n = len(rows)
js = [j for j in judged if j is not None]
cost_per_task = (tok_in * pin + tok_out * pout) / 1e6 / n
return {
"candidate": f"{t.name} {t.model.split('/')[-1]}",
"schema_valid_%": 100 * sum(g["valid"] for g in grades) / n,
"category_%": 100 * sum(g["category_ok"] for g in grades) / n,
"severity_%": 100 * sum(g["severity_ok"] for g in grades) / n,
"judge_1to5": sum(js) / len(js) if js else float("nan"),
"p95_ms": percentile(lat, 95),
"usd_per_1k": cost_per_task * 1000,
}
def main() -> None:
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("candidates", nargs="+", help='e.g. mock:mock-8b "fireworks:<model>@0.07/0.30"')
ap.add_argument("--n", type=int, default=50)
ap.add_argument("--judge", help="target[:model] for the LLM judge (optional)")
a = ap.parse_args()
rows = load_test(a.n)
judge = None
if a.judge:
jt, _, _ = parse_candidate(a.judge)
judge = (client(jt), jt.model)
out = []
for spec in a.candidates:
t, pin, pout = parse_candidate(spec)
print(f" evaluating {t.name} {t.model} on {len(rows)} tickets…", flush=True)
out.append(evaluate(t, pin, pout, rows, judge))
record("09-eval", {"n": len(rows), **out[-1]})
print("\n" + table(out))
best = max(out, key=lambda r: (r["category_%"], -r["usd_per_1k"]))
print(f"\nHighest accuracy: {best['candidate']}. Now read across: is the gain worth the latency and $?")
if __name__ == "__main__":
main()