LoRA on Fireworks
course 12 of 22
lesson 10 · lessons/10-lora-fireworks
2 h About $5: cents to train, about $8 an hour to serve Fireworks
What you will do and why
You fine-tune a real model on Fireworks from start to finish: upload examples, train, put it online, test it, shut it down. It proves what customers most need to hear: training is cheap, and keeping the model running for users is the real cost.
Why it matters: The lesson’s sample: training costs about 15 cents, while 38 minutes on a GPU (the chip that runs the model) reserved for it cost $5.07. The fine-tune sorted about 97% of tickets into the right category, against 80% before.
You are done when: 5_teardown.sh has printed deleted after with nothing listed below it, and make fw-check lists no deployments. You have written down the fine-tune’s gain over the base and the total cost.
Start this first
Run it from the 03-labs folder with the Python environment on (source .venv/bin/activate). First check that make fw-check lists no deployments. 1_train.sh uploads the 800 training tickets and starts training, billed per token: about $0.15. Then run bash 2_wait.sh. It checks once a minute and usually finishes in 10 to 30 minutes. Watch the video while it waits.
cd lessons/10-lora-fireworks && bash 1_train.shQLoRA On Your Own Box · 2:49
Download mp4 (10.4 MB)Watch while the training job runs.
Chapters
In this video How to fine-tune on a computer you already own, with the model shrunk to 4 bits and left unchanged, and how that compares with today’s paid cloud run.
3 key points
QLoRA (LoRA on a model shrunk to 4 bits) leaves the model unchanged and trains only a small add-on, so an 8-billion-number model trains on a laptop.
The video’s Mac example: about 5 GB of model left unchanged (frozen), a 30 MB add-on, about 11 GB of memory at the peak, 30 to 50 minutes, $0. Day 8 runs a similar one on your Mac: Qwen2.5 7B on the same 800 tickets.
The data and a test set kept aside decide the result, not the hardware.
200 messy, unchecked examples: 97% right on the examples it trained on, but 74% on new ones, so it memorised. 2,000 clean examples: 93% and 89%. Set the test examples aside before training, never after.
Compare doing it yourself with a managed service on time, privacy and cost.
Your own machine: free, and the data stays home. Managed, as in today’s run: the video says a dollar or two of training (today’s smaller dataset costs cents) and a working model in minutes. But serving a LoRA on a hosted platform usually needs a GPU reserved for it.
In plain words
Section titled “In plain words”Fine-tuning teaches an existing model a new habit from examples. LoRA (low-rank adaptation) freezes the model and trains a small add-on (an adapter) instead, so training is quick and cheap. The expensive part comes after: on Fireworks a fine-tuned model needs a GPU reserved for it, billed for as long as it runs, used or not.
Picture it
Training is like printing a cookbook: you pay once. Serving the model (keeping it running so it can answer) is like renting a stall to sell the book: rent is due every hour the stall is open, buyers or not. So you share the stall with other sellers, make sure enough buyers come, and close up the moment you are done.
With real numberslesson 10’s scripts and sample output, and the lab book’s prices
- Training data: 800 example tickets with the right answers, read twice (2 epochs): about 0.3 million tokens (word pieces).
- Training price: $0.50 per million tokens for models up to 16 billion numbers. 0.3 x $0.50 = about $0.15.
- Serving: a dedicated GPU at about $8 an hour, charged by the second from the moment it starts, used or not.
- The lesson’s sample kept it up 38 minutes: 38 ÷ 60 x $8 = $5.07, over 30 times the training cost.
- Left on for a month (about 730 hours): 730 x $8 = $5,840. That is why you shut the GPU down (delete the deployment) in the same sitting.
- The result: 80% of ticket categories right before, about 97% after, a 17-point gain.
Words to know
- LoRA
- A cheap way to customise a model: train a small add-on instead of changing the whole model. Example: the Training Techniques video’s 30 MB add-on.
- Adapter
- The small add-on LoRA trains, saved as its own file and applied on top of the frozen model. Example: about 30 MB in the Training Techniques video.
- Dedicated deployment
- A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not. Example: about $8 an hour.
- Epoch
- One full pass over the training examples. Example: 2 epochs over 800 tickets.
From the lesson
lessons/10-lora-fireworks/README.md
What and why
Section titled “What and why”LoRA leaves the big model frozen and trains a small “clip-on” adapter, often under 1% of the weights, on your examples. It is cheap to train and small to store, and it teaches format and behaviour well (our triage JSON), while teaching new facts poorly (use RAG for those).
The lesson that matters for a Solutions Architect: training is the cheap part.
train 800 examples × 2 epochs ≈ cents (per training token)serve dedicated deployment ≈ $ per GPU-hour ← the real costshare multi-LoRA: many adapters on one deployment → the cost-sharing answerThe five steps
Section titled “The five steps”| script | does | billing |
|---|---|---|
1_train.sh |
upload data/tickets/train.jsonl, start the SFT job (rank 8, 2 epochs) |
per training token |
2_wait.sh |
poll until COMPLETED | – |
3_deploy.sh |
create a dedicated deployment and wait for READY | clock starts |
4_eval.sh |
step 1’s eval harness: base vs fine-tune, same 60 held-out tickets | per request |
5_teardown.sh |
delete, then prove deployment list is empty |
clock stops |
The ids are kept in .state between steps, so you can close the terminal after 1_train.sh.
cd lessons/10-lora-fireworksbash ../00-setup/fireworks_guardrails.sh # nothing running? goodbash 1_train.shbash 2_wait.shbash 3_deploy.sh # clock startsbash 4_eval.shbash 5_teardown.sh # clock stops: do not skipWhat each command does
bash ../00-setup/fireworks_guardrails.shShows who you are signed in as and anything running on your Fireworks account. Look for an empty deployments list: anything listed bills by the hour. Delete a leftover with
firectl deployment delete <DEPLOYMENT_ID>.bash 1_train.shUploads the 800 training tickets and starts training a small add-on for Qwen2.5 7B Instruct (rank 8, the add-on’s size setting; 2 passes over the data). Billed per training token: about 0.3 million tokens, so about $0.15. Look for
started triage-sft-followed by a date stamp. If you ran it under Start this first, skip it.bash 2_wait.shChecks the job once a minute until it finishes, usually in 10 to 30 minutes. Nothing bills by the hour yet, and the ids are saved, so you can close the terminal and come back. Look for
trained:followed by your model’s name.bash 3_deploy.shPuts your fine-tune on a dedicated deployment (a GPU reserved for you). The billing clock starts now, about $8 an hour, used or not. Look for the time it was created, then a dot every 20 seconds until
ready. Run the next two commands straight away.bash 4_eval.shStep 1’s test on 60 held-out tickets: the base model (Qwen2.5 7B Instruct before your training), paid per token, against your fine-tune on its GPU. Look for the fine-tune’s
category_%well above the base’s, then a last line with the minutes so far and their cost.bash 5_teardown.shDeletes the deployment, which shuts down the GPU and stops the clock, then lists what is left. Your trained model stays stored in your account, so you can put it back online later. Look for
deleted afterwith the minutes and dollars, and nothing listed below it. Never skip this step.
The default BASE_MODEL is a 7B instruct model. Check it appears in Fireworks’ list of
tunable models and override it if needed: BASE_MODEL=accounts/fireworks/models/<id> bash 1_train.sh.
What you should see
Section titled “What you should see”How to read it
Two rows graded on the same 60 held-out tickets: the base model first, then your fine-tune, labelled with its deployment’s id. In the lesson’s sample, categories right rise from 80.0 to 97.0 (17 points), urgency (severity) from 70.0 to 95.0, and p95 drops a little, from 420 to 380 ms. The sample is illustrative: each of the 60 tickets is worth 1.67 points (100 ÷ 60), so your figures land on values like 96.7 or 98.3. judge_1to5 reads nan (no judge was asked) and usd_per_1k reads 0.0 (no prices were given; the fine-tune bills by the hour). The last line is that bill: 38 minutes at $8 an hour = $5.07.
candidate schema_valid_% category_% severity_% p95_msfireworks qwen2p5-7b-instruct 100.0 80.0 70.0 420fireworks triage-lora-…#… 100.0 97.0 95.0 380▸ deleted after 38 min (≈ $5.07 serving + a few cents of training)Check yourself
Section titled “Check yourself”3 questions. Say your answer out loud, then tap to check it.
The customer wants 12 fine-tuned versions of a model, one for each business unit. Served separately, each one needs its own reserved GPU (the chip that runs the model), billed by the hour. How do you keep serving costs sensible?Show answerHide
In plain words
Share one GPU. Keep one copy of the base model running and load the right small add-on (adapter) for each request. This is multi-LoRA: one bill instead of 12.
Picture it
One kitchen with one set of equipment and 12 spice packets, one per restaurant brand, instead of 12 separate kitchens. Each order says which packet to use.
With real numberslesson 10’s deployment rate and the lab book’s prices
- One dedicated GPU: about $8 an hour, charged whether or not anyone uses it.
- Running all month (about 730 hours): 730 x $8 = $5,840.
- 12 separate deployments: 12 x $5,840 = $70,080 a month.
- Multi-LoRA, 1 deployment holding 12 adapters: $5,840, one twelfth, as long as one GPU keeps up with the traffic.
- Each adapter is small: about 30 MB in the Training Techniques video. One deployment holds up to 100 by default, the lab book says.
Words to know
- Multi-LoRA
- Serving many LoRA add-ons on one running copy of the base model, picking the add-on for each request. Example: 12 business units, 1 GPU.
- Adapter
- The small add-on LoRA trains, applied on top of the frozen model. Example: about 30 MB.
- Dedicated deployment
- A GPU reserved for you on Fireworks, billed by the second at an hourly rate, used or not. Example: about $8 an hour.
- Base model
- The model you start from, before any fine-tuning. Example: Qwen2.5 7B Instruct in this lesson.
Go deeper: the engineer version
The kit's question
The customer wants 12 fine-tunes, one per business unit. How do you keep serving costs sane?
The kit's answer
Multi-LoRA: one base deployment with 12 adapters, loaded per request.
More detail: Adapters are loaded per request: each request names its adapter, and the server applies it on top of the shared base weights, batching requests for different adapters together (the field guide: hundreds of fine-tunes on one base at the base model’s price). The lab book adds the Fireworks details: a default quota of 100 adapters per deployment, and multi-LoRA needs a BF16 (16-bit) deployment shape with --enable-addons; FP8 and FP4 (8-bit and 4-bit) shapes cannot host adapters.
A fine-tune gains 17 points of accuracy: 17 more tickets in every 100 sorted correctly. When would you still advise against it?Show answerHide
In plain words
When a cheaper fix gets most of the way there: a better prompt, or a few worked examples inside it. Or when there is too little traffic to pay for a GPU that bills every hour.
Picture it
Hiring a full-time specialist to handle one call a day. First try better instructions for the staff you have; if that closes most of the gap, or the calls are rare, the salary is not worth it.
With real numberslesson 10’s sample, lesson 09’s example table and the lab book’s prices
- The gain: 80% to about 97% of categories right, 17 points (lesson 10’s sample).
- The fine-tune needs its own GPU: about $5,840 a month if it stays on (730 hours x $8).
- Paying per token instead: the big general model in the eval lesson’s table costs $0.42 per 1,000 tickets.
- Break-even: $0.42 buys 1,000 tickets, so $5,840 buys 5,840 ÷ 0.42 = about 13,900 lots of 1,000: 13.9 million tickets a month, about 5 every second, day and night. Below that, paying per token is cheaper.
- Say no if traffic is below that volume, or if a better prompt closes most of the 17 points. Or share the GPU with other fine-tunes (multi-LoRA), so each pays less.
Words to know
- Few-shot examples
- A few worked examples placed in the prompt to show the model what you want, with no training. Example: a cheaper fix to try first.
- Serverless (pay-per-token)
- Fireworks’ shared, always-on models, billed per token you send and receive. Example: $0.42 per 1,000 tickets.
- Break-even
- The volume at which two ways of paying cost the same. Example: about 13.9 million tickets a month.
- GPU-hour
- One GPU for one hour, the unit dedicated deployments are billed in. Example: about $8.
Go deeper: the engineer version
The kit's question
The fine-tune gains 17 points of accuracy. When would you still say no?
The kit's answer
When a prompt change or few-shot examples close most of the gap, or when traffic is too low to justify a dedicated deployment.
More detail: Try the free levers first: the Training Techniques video puts better prompts and retrieval before any training. Then compare the monthly dedicated cost with the pay-per-token bill at the real volume. The lab book’s cost calculator draws that line against one H100 at $8 an hour, 24x7 ($5,840 a month). Multi-LoRA changes the sum by splitting that GPU across adapters. The sum above pairs a big general model’s price with a small fine-tune’s GPU; for a real answer, use the customer’s own candidates.
During training the loss (how wrong the model is on its training examples) keeps falling, but its accuracy on the held-out test tickets stays flat. What does that mean?Show answerHide
In plain words
The model is memorising its training examples instead of learning the task, so it does no better on new tickets. Train less (fewer passes), give it more varied examples, or make the add-on smaller (a lower rank).
Picture it
A student who memorises last year’s exam answers: their practice scores keep rising, but on a new exam they do no better. More varied practice, fewer repeats and a smaller notebook to cram into all help.
With real numbersthe QLoRA On Your Own Box video and lesson 10’s training settings
- The video’s example, 200 messy, unchecked examples: 97% right on the examples it trained on, 74% on held-out ones. A 23-point gap: memorised.
- 2,000 cleaner examples: 93% and 89%. A 4-point gap: it learned the task.
- Today’s settings are already small: 2 passes (epochs) over 800 examples, adapter rank 8. The script suggests 16 to 32 only for harder tasks.
- The fixes: fewer passes, more varied examples, or a lower rank.
Words to know
- Loss (training loss)
- A score of how wrong the model is on its training examples; training pushes it down. Example: it can fall while test accuracy stays flat.
- Overfitting (memorising)
- Training until the model remembers its examples instead of learning the task. Example: 97% on training examples, 74% on new ones.
- Epoch
- One full pass over the training examples. Example: 2 in
1_train.sh. - Rank (LoRA rank)
- The size setting of a LoRA add-on; a higher rank trains more numbers. Example: 8 here, 16 to 32 for harder tasks.
Go deeper: the engineer version
The kit's question
The loss is falling but eval accuracy is flat. What does that mean?
The kit's answer
The model is memorising. Use fewer epochs, more varied data, or a lower rank.
More detail: Falling training loss with flat held-out accuracy is overfitting. 1_train.sh says it in one line: more epochs means memorising, so watch the eval, not the loss. On Day 8 the MLX run prints validation loss (loss on examples kept out of training) every 50 steps: stop when it flattens. A lower rank means fewer trainable numbers, so less room to memorise.
Explain what you learned
Section titled “Explain what you learned”Question Fine-tuning sounds expensive. What will it really cost us?
One clear answer
Training was about a dollar. The thing to plan is serving: a LoRA needs dedicated capacity, so we either share it with multi-LoRA or justify it with volume.
What this means
- “Training was about a dollar”: Teaching the model cost little. This lesson’s 800 examples are about 0.3 million training tokens, about $0.15; the Training Techniques video’s bigger job, 2.4 million tokens, is $1.20.
- “The thing to plan is serving”: The real cost is running the fine-tuned model for users afterwards.
- “a LoRA needs dedicated capacity”: On Fireworks a LoRA fine-tune cannot be served pay-per-token. It needs a GPU reserved for it, about $8 an hour whether or not anyone uses it: the lesson’s 38-minute test cost $5.07.
- “so we either share it with multi-LoRA”: Put many fine-tunes on that one GPU so they split its cost: 12 business units, 1 GPU.
- “or justify it with volume”: Send enough traffic that the GPU is cheaper than paying per token: about 13.9 million tasks a month, at the eval lesson’s $0.42 per 1,000.
Your numbersSaved on this device and collected in the Day 7 wrap-up.
Hint: The category_% column of 4_eval.sh: the first row is the base model, the second your fine-tune.
Hint: The p95_ms column of the same table.
Hint: The line 5_teardown.sh prints once the deployment is deleted, deleted after ... min with its dollar figure. Add the training: about $0.15 (0.3 million tokens at $0.50 per million).
Run make fw-check now: it should list nothing. It lists anything on Fireworks that is still billing, so an empty list means nothing is costing you money.
Done when
Section titled “Done when”5_teardown.sh has printed deleted after with nothing listed below it, and make fw-check lists no deployments. You have written down the fine-tune’s gain over the base and the total cost.
Stuck?
Section titled “Stuck?”1_train.shfails because the base model cannot be fine-tuned, or is not found.- Pick a model from Fireworks’ list of tunable models and run
BASE_MODEL=accounts/fireworks/models/<id> bash 1_train.sh. Pass the sameBASE_MODEL=...tobash 4_eval.shlater: the scripts do not remember it, and would grade the default base instead. 2_wait.shstops with FAILED or CANCELLED.- It prints the job’s details: read the error there. Nothing bills by the hour yet. Once it is fixed, run
bash 1_train.shagain: it starts a new job with new names. 3_deploy.shkeeps printing dots and never reachesready.- The deployment already exists, so it may be billing. Press Ctrl-C and check the deployment’s state in the Fireworks console. If it shows ready, carry on with
bash 4_eval.sh. If it is not coming up, runbash 5_teardown.sh, thenmake fw-checkfrom03-labs. 4_eval.shstops with an error, such as a model not found.- Your GPU is still billing: if you cannot fix this in a few minutes, run
bash 5_teardown.shfirst and come back. If the error names your fine-tune, copy the model string from the deployment’s API tab in the Fireworks console and runpython ../09-eval-harness/evaluate.py "fireworks:<that string>" --n 60. If it names the base model, its id may have changed: copy a current one from the model library and runBASE_MODEL=accounts/fireworks/models/<id> bash 4_eval.sh. - You closed the terminal while the deployment was running.
- Go back to
lessons/10-lora-fireworksand runbash 5_teardown.sh. The ids are saved in the.statefile, and the clock runs until you do. 5_teardown.shstops with an error, or you are unsure whether anything is still billing.- From
03-labs, runmake fw-check. Delete anything it lists withfirectl deployment delete <DEPLOYMENT_ID>, then runmake fw-checkagain until the list is empty.
Go deeper: the lab book's Fireworks lab for this step All fixes
Code in this step
Section titled “Code in this step”1_train.sh Bash · 25 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 10 · step 1 — upload the dataset and start a LoRA SFT job.## Cost: training is billed per training token. 800 short examples × 2 epochs is# ~0.3M tokens → cents. Training is NOT the expensive part of this lesson (step 3 is).## --lora-rank 8 adapter size. Small task, small rank. (16–32 for harder tasks)# --epochs 2 passes over the data. More = memorise; watch eval, not loss# --learning-rate leave default unless loss misbehaves# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"; source _state.shSTAMP=$(date +%m%d%H%M)[[ -f ../../data/tickets/train.jsonl ]] || python ../../data/make_tickets.py
save DATASET "triage-train-$STAMP"save JOB_ID "triage-sft-$STAMP"save OUT_MODEL "triage-lora-$STAMP"
firectl dataset create "$DATASET" ../../data/tickets/train.jsonlfirectl sftj create --job-id "$JOB_ID" --base-model "$BASE_MODEL" \ --dataset "$DATASET" --output-model "$OUT_MODEL" --epochs 2 --lora-rank 8
echo "▸ started $JOB_ID → next: bash 2_wait.sh"2_wait.sh Bash · 11 lines
#!/usr/bin/env bash# Lesson 10 · step 2 — wait for training. Typically 10–30 min for this size.set -euo pipefailcd "$(dirname "$0")"; source _state.shwhile true; do state=$(firectl sftj get "$JOB_ID" | grep -iE "^\s*state" | head -1 || true) echo " $(date +%H:%M:%S) $state" [[ "$state" =~ COMPLETED ]] && { echo "▸ trained: accounts/$ACCOUNT/models/$OUT_MODEL → bash 3_deploy.sh"; exit 0; } [[ "$state" =~ FAILED|CANCELLED ]] && { firectl sftj get "$JOB_ID"; exit 1; } sleep 60done3_deploy.sh Bash · 23 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 10 · step 3 — serve the LoRA. ⚠ THE BILLING CLOCK STARTS HERE ⚠## A fine-tuned LoRA is served on a DEDICATED deployment (GPU-hours), not per token.# That is the real cost of a fine-tune and the conversation to have with a customer:# one deployment can host MANY LoRAs (multi-LoRA), which is how that cost is shared.## Budget for this step: ~1 hour × $DEPLOY_RATE. Run 4_eval.sh then 5_teardown.sh promptly.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"; source _state.shMODEL="accounts/$ACCOUNT/models/$OUT_MODEL"
out=$(firectl deployment create "$MODEL" --deployment-shape default)echo "$out"save DEPLOYMENT_ID "$(echo "$out" | grep -oE 'deployments/[A-Za-z0-9-]+' | head -1 | cut -d/ -f2)"save DEPLOY_STARTED "$(date +%s)"echo "▸ deployment $DEPLOYMENT_ID created at $(date +%H:%M) — clock running at ~\$$DEPLOY_RATE/h"
echo "▸ waiting for READY"until firectl deployment get "$DEPLOYMENT_ID" | grep -qiE "state.*READY"; do sleep 20; echo -n "."; doneecho " ready → bash 4_eval.sh"4_eval.sh Bash · 14 lines
#!/usr/bin/env bash# Lesson 10 · step 4 — the delta: base vs fine-tune, same harness as lesson 09.# The deployed model is addressed as <model>#<deployment>. If your console shows a# different model string on the deployment's API tab, use that instead.set -euo pipefailcd "$(dirname "$0")"; source _state.shTUNED="accounts/$ACCOUNT/models/$OUT_MODEL#accounts/$ACCOUNT/deployments/$DEPLOYMENT_ID"
python ../09-eval-harness/evaluate.py \ "fireworks:$BASE_MODEL" \ "fireworks:$TUNED" --n 60
mins=$(( ( $(date +%s) - DEPLOY_STARTED ) / 60 ))echo "▸ deployment has been up ${mins} min ≈ \$$(python -c "print(round($mins/60*$DEPLOY_RATE,2))") → bash 5_teardown.sh NOW"5_teardown.sh Bash · 11 lines
#!/usr/bin/env bash# Lesson 10 · step 5 — stop the clock. Then PROVE it stopped.set -euo pipefailcd "$(dirname "$0")"; source _state.shfirectl deployment delete "$DEPLOYMENT_ID"mins=$(( ( $(date +%s) - DEPLOY_STARTED ) / 60 ))echo "▸ deleted after ${mins} min (≈ \$$(python -c "print(round($mins/60*$DEPLOY_RATE,2))") serving + a few cents of training)"echo "▸ remaining deployments (should be empty):"firectl deployment list# The trained model and dataset stay in your account (cheap storage) so you can redeploy later.rm -f .state_state.sh Bash · 8 lines
# Shared by the numbered scripts: remembers ids between steps in .state (git-ignored).# shellcheck disable=SC2034STATE="$(dirname "$0")/.state"[[ -f "$STATE" ]] && source "$STATE"save() { echo "$1=\"$2\"" >> "$STATE"; eval "$1=\"$2\""; }ACCOUNT=${FIREWORKS_ACCOUNT_ID:-$(firectl whoami 2>/dev/null | grep -oE 'accounts/[a-z0-9-]+' | head -1 | cut -d/ -f2)}BASE_MODEL=${BASE_MODEL:-accounts/fireworks/models/qwen2p5-7b-instruct} # must be tunable — check the docs listDEPLOY_RATE=${DEPLOY_RATE:-8} # $/hour for the deployment shape you get — check the pricing page