Learn / Prompting that works / Cost and latency
Cost and latency
Where your token spend actually goes, streaming vs waiting for the full response, and the cheap wins before you reach for a smaller model.
Where the tokens go
Every request has input tokens (system prompt + context + conversation history + the user’s message) and output tokens (what the model generates). Providers typically price output tokens higher than input tokens, because generation is the more expensive operation per token.
Two practical consequences:
- A long, static system prompt gets sent - and billed - on every single request, not once. At meaningful volume, trimming a verbose system prompt is a direct, recurring saving.
- Uncapped output length is an open-ended cost. Set an explicit
max_tokens(or equivalent) based on what the task actually needs, not the provider’s default ceiling.
Streaming: latency, not cost
Streaming sends tokens to the client as the model generates them, instead of waiting for the
full response and sending it all at once. This is exactly what this site’s own chat API does
(src/pages/api/chat.ts, Server-Sent Events). It does not reduce total generation time or
total cost - the model still has to generate every token either way. What it changes is
perceived latency: the reader sees the first words in a few hundred milliseconds instead of
waiting several seconds for the whole answer to land at once, which matters a lot for a
conversational UI even though the underlying work is identical.
Cheap wins, in order
Before reaching for a smaller model (which changes output quality, not just cost), try:
- Trim redundant context. Retrieved chunks with heavy overlap (RAG track, lesson 2), conversation history that’s grown longer than the task needs, boilerplate instructions repeated when they could be stated once - all of it is billed input tokens.
- Cap output length explicitly. An unconstrained response for a task that only needs a short answer wastes output tokens (the more expensive kind) on padding.
- Cache what doesn’t change per request. Some providers offer prompt caching for a static system prompt or context block, billing the cached portion at a reduced rate on repeat use - worth checking if your provider supports it before assuming a full-price re-send is unavoidable.
Only then: a smaller or cheaper model
Switching models is a real lever - and the most consequential one, because it can change answer quality, not just price. Don’t swap based on the price sheet alone: run your existing test suite (lesson 4) against the candidate model on your actual task. A model that’s 3x cheaper but fails your golden set 15% more often isn’t obviously a win once you account for the cost of wrong answers reaching users - re-runs, support tickets, lost trust - which rarely shows up on the same invoice as the token bill but is real cost all the same.
Key takeaways
- You pay for both input and output tokens, and output tokens are typically priced higher - a verbose system prompt on every request adds up fast at scale.
- Streaming doesn't reduce total cost or total generation time, but it dramatically improves perceived latency - the reader sees the first words almost immediately instead of waiting for the whole response.
- Cheap wins before reaching for a smaller/cheaper model: trim redundant context, cap max output tokens explicitly, and cache anything that doesn't change per request.
- A smaller/cheaper model is a real lever, but test it against your actual task with your actual eval set (lesson 4) - 'cheaper' that also means 'wrong more often' isn't actually cheaper once you count the cost of bad answers.
Quick check
3 questions - see how much stuck.