Skip to content

One product criticism

course 18 of 22

30 min Free Anywhere

What you will do and why

Write one specific criticism of Fireworks, backed by a number, from something you hit on Days 1 to 9. The lab book lists it among the four things to bring: field engineers are hired partly to have opinions.

Why it matters: On Day 7 your LoRA fine-tune trained for cents, but serving it needed a GPU reserved by the hour: the lesson’s sample eval ran 38 minutes, about $5.07. That gap, with its numbers, is a criticism.

You are done when: All five boxes are filled in. The criticism names one thing you hit, has at least one number you measured and one source with the date you read it, and ends in a proposal. You have said the short version out loud once.

A good product criticism names one thing you hit yourself, backs it with a number, says who it hurts and offers a fix. It shows you used the product for real and will carry customers’ problems back to the people who build it.

Picture it

A restaurant critic who writes “the food was bad” gets ignored. One who writes “the fish arrived long after everything else at most tables, and a second grill would fix it” gets a call from the owner.

With real numbersDay 7’s LoRA lesson, the lab book’s prices and the Platform Map video’s traffic

  • Training the LoRA: 800 examples read twice is about 0.3 million training tokens, by the script’s own estimate. At $0.50 per million (the price for models up to 16 billion learned numbers), that is about $0.15.
  • Serving it: Fireworks runs a trained LoRA only on a dedicated deployment, $8 an hour for an H100, busy or idle.
  • The lesson’s sample run kept the GPU for 38 minutes to test the add-on, then deleted it: 38 ÷ 60 x $8 = $5.07, about 34 times the training cost ($5.07 ÷ $0.15).
  • Kept ready all month: $8 x 730 hours = $5,840. For scale, the Platform Map video’s 20 million tokens a day, at the base model’s per-token price ($0.20 per million): 20 x $0.20 = $4.00 a day, about $122 a month. The always-ready GPU costs about 48 times more ($5,840 ÷ $122).
  • Fireworks bills a dedicated deployment from the moment it starts, used or not. By default it gives the GPU back only after 1 hour with no requests (the lab book), a setting you can shorten. So each forgotten session costs about $8 more: delete it in the same sitting.

Words to know

LoRA
A small add-on trained on your examples while the big model stays unchanged. Example: Day 7’s ticket-sorting add-on, trained for about $0.15.
Dedicated deployment
A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not. Example: $8 an hour for an H100.
Serverless
Fireworks’ shared, always-on models, billed per token you send and receive. Example: about $122 a month for 20 million tokens a day at $0.20 per million.
Multi-LoRA
One deployment serving many LoRA add-ons on the same base model, so they share its cost. Example: up to 100 add-ons by default on Fireworks.

Five short boxes turn “something annoyed me” into useful product feedback. Fill them in under Explain what you learned below as you work through this page. They are saved on this device and collected in the Day 10 wrap-up.

The strongest criticism comes from your own run, not from a forum post. These are the ones the course’s own steps run into. Pick one you actually met, or one of your own.

  • A fine-tune can only be served on a GPU rented by the hour. On Day 7 your LoRA (a small add-on trained on your examples) cost cents to train. Fireworks then serves a trained LoRA only on a dedicated deployment: a GPU reserved for you at about $8 an hour, busy or idle. The QLoRA On Your Own Box video from Day 7 says the same: serving a LoRA on a hosted platform usually needs dedicated capacity.
  • The docs disagree about it. The lab book found the pricing page saying fine-tuned models serve “at the same price as base models”, while the LoRA docs say trained LoRAs deploy only to on-demand (dedicated) deployments.
  • Two of the course’s experiments needed a GPU billed by the hour. Day 7’s LoRA and Day 9’s speculative decoding test (30 to 45 minutes of one deployment) both run on a dedicated deployment, because LoRA serving and a custom draft model are only available there.
  • Sharing one GPU among many fine-tunes rules out the smaller number formats. A number format is how many bits store each of the model’s numbers; Day 4 showed fewer bits run faster. Multi-LoRA (many add-ons on one deployment) needs a 16-bit setup, BF16. The 8-bit and 4-bit setups, FP8 and FP4, cannot host add-ons. So the cheapest way to share a GPU and the fastest way to run one do not combine.
  • An idle deployment keeps billing for an hour. A dedicated (on-demand) deployment bills from the moment it starts, used or not. By default it scales to zero (gives its GPU back) only after 1 hour with no requests, a setting the lab book says to shorten. So a forgotten eval costs about another $8.
  • Model names change under you. The lab book read Fireworks’ prices on 23 September 2026, and several models were due to retire two days later. The kit’s own fix for “model id is not found” is “ids change; copy a current one”. A customer’s code that names a model breaks when it retires.
  • The request limit is described two ways: the published limits page says a flat 6,000 requests a minute, while a recent changelog entry (Fireworks’ dated list of product changes) describes adaptive limits, ones that change instead of staying at a flat number.

Without a number it is an opinion. Write down three things.

  • Your own figures. How long the deployment ran and what it cost: the last lines of Day 7’s 5_teardown.sh print both. Or the minutes you lost, and how many tries it took.
  • The source. Which page says what, and the date you read it. Fireworks’ prices and features change weekly, so the date matters.
  • The comparison. The same job done another way, in the same unit: dollars a month, minutes or percent.

Name one kind of customer and put a monthly number on it. A team with one fine-tune and light traffic pays up to $5,840 a month to keep one H100 ready around the clock ($8 x 730 hours). Letting it scale to zero saves money, but then the first request after a quiet hour waits for a GPU to start (a cold start). For scale, take the Platform Map video’s 20 million tokens a day. At the base model’s per-token price, $0.20 per million (the lab book’s price for models of 4 to 16 billion learned numbers, which includes Day 7’s 7B base model), that is $4.00 a day, about $122 a month. The always-ready GPU costs about 48 times that ($5,840 ÷ $122).

Offer something the product team could act on, and say how you would know it worked. For the LoRA example:

  • per-token (serverless) serving for LoRAs on the most-used base models. Multi-LoRA already shares one base deployment among up to 100 add-ons inside one account; the proposal extends that across customers;
  • until then, make the pricing page and the LoRA docs say the same thing, and show the hourly cost when firectl deployment create starts a GPU;
  • a shorter default idle time before a deployment scales to zero (gives its GPU back).

You would know it worked when more trained fine-tunes get deployed and fewer deployments sit idle.

Five beats, in about 20 seconds, the pace of the capstone drill:

  1. What works, honestly, with a number.
  2. What you hit, in one specific sentence.
  3. The evidence and the impact: one number you measured, one for the customer.
  4. Your proposal.
  5. A question back, such as the lab book’s: how do field engineers get fixes into the platform?

Five rules:

  • One criticism, not a list. Depth beats breadth.
  • Criticise the product, never the people who built it.
  • Say how you know and when: “in the docs I read last week”.
  • Check it again before sharing it. If it has been fixed, update your feedback.
  • Leave out “always”, “never” and “broken”.

With the lesson’s sample figures from Day 7. Put your own in their place.

  • What works: fine-tuning was easy: 800 examples trained for about $0.15, by the script’s estimate (about 0.3 million training tokens x $0.50 per million). Check the real charge in your Fireworks account.
  • What you hit: serving the LoRA needed a dedicated deployment, billed by the hour.
  • Evidence: the eval deployment ran 38 minutes: 38 ÷ 60 x $8 = $5.07, about 34 times the training. The LoRA docs say dedicated only; the pricing page says fine-tuned models serve at base-model prices. The lab book read both on 23 September 2026: read them again.
  • Who it hurts: a customer with one fine-tune and light traffic, such as the Platform Map video’s 20 million tokens a day. At the base model’s per-token price that is about $122 a month; keeping the GPU ready around the clock costs up to $5,840, about 48 times more.
  • Proposal: per-token serving for LoRAs on the most-used base models. First, make the two pages agree and show the hourly cost when a deployment starts.
  • Short version, out loud:

“Fine-tuning on Fireworks was easy: my LoRA trained for about 15 cents. Serving it cost $5 for a 38-minute eval, because trained LoRAs run only on a dedicated GPU, $5,840 a month if kept ready. I’d propose per-token serving for LoRAs on popular base models. How do field engineers get that onto the roadmap?”

Your numbersSaved on this device and collected in the Day 10 wrap-up.

Hint: One thing, specific: what you were doing, on which day, and what happened. Not “the docs are confusing”.

Hint: Your own figures first (the last lines of Day 7’s 5_teardown.sh print minutes and dollars), then the page that says so and the date you read it.

Hint: Name one kind of customer and put a monthly number on it.

Hint: A fix the product team could act on, and how you would know it worked.

Hint: About 20 seconds, the capstone drill’s pace: what works, what you hit, one number, your proposal, a question back.

All five boxes are filled in. The criticism names one thing you hit, has at least one number you measured and one source with the date you read it, and ends in a proposal. You have said the short version out loud once.

I skipped Day 7, so I have no fine-tune of my own.
Use something you did hit: Day 9’s speculative decoding test also needed a dedicated deployment (30 to 45 minutes of one), or a model name that no longer worked. If you ran into nothing, you can use the docs disagreement in section 1 above, but say you found it by reading, not by running.
Everything worked, so I have nothing to criticise.
Look for what cost you time or money rather than what broke: a step that needed a GPU billed by the hour, a price you had to look up in two places, a limit you only found in a changelog (the provider’s list of recent changes). A small, specific friction beats a big, vague one.
My criticism may be out of date.
The lab book read Fireworks’ pages on 23 September 2026 and warns that prices move weekly. Read the pricing page and the LoRA deployment docs again before sharing your recommendation. If it has been fixed, say so and ask what drove the change: that shows you keep up.
It sounds like a complaint.
Start with what works, with a number, and end with your proposal and a question. Cut “always”, “never” and “broken”; keep your one measured number.