Engineer tuning LLM costs

Cost Control Methods for Scaling Your LLM Features

Cost OptimizationLarge Language Models
Published

When building out a new LLM-powered feature, I’m amenable to de-prioritizing cost (within reason). If I’m still trying to validate an idea, I want to know whether the feature solves the problem before optimizing for scale.

That question changes once the feature works and usage grows. The API bill becomes a recurring expense, and reducing it can improve the bottom line. However, a cheaper setup only helps if the feature still meets the requirements that made it useful. Cut corners too much, and you have a feature that costs next to nothing, but also produces poor work.

Note

This blog post will be relying heavily on OpenAI for concrete examples, but the practices apply to other LLM providers as well!

You may not need the biggest/newest model

It’s tempting to switch to the newest/biggest model when public benchmarks make it look better. Those benchmarks can help you find candidates, but they’re situational. Their results may not transfer to the requests your application actually handles.

TypeScript
1// Source: https://developers.openai.com/api/docs/quickstart
2
3import OpenAI from "openai";
4const client = new OpenAI();
5
6const response = await client.responses.create({
7 model: "gpt-6-astra", // Play around with this
8 input: "Write a one-sentence bedtime story about a unicorn.",
9});
10
11console.log(response.output_text);

At work, we recently considered switching to a newer model that was cheaper and faster. However, when we ran it against our evaluations, its accuracy was worse, so we didn’t deploy it. The candidate looked appealing on price and latency, but those weren’t the only things we needed it to improve.

The takeaway here is to run your current model and each candidate on the same evaluation set. Compare quality, latency, and cost, and decide what counts as acceptable before looking at the results. For high-stakes workloads (e.g. health, legal), a replacement needs to be at least as accurate as the model we already use. A less critical feature might have room for a small accuracy tradeoff. The right bar depends on what your product is expected to do.

Use the lowest reasoning effort that passes

Some models let you configure reasoning effort. When they do, try different settings on the same evaluations you use to compare models.

TypeScript
1// Source: https://developers.openai.com/api/docs/quickstart
2
3import OpenAI from "openai";
4const client = new OpenAI();
5
6const response = await client.responses.create({
7 model: "gpt-6-astra",
8 reasoning: {
9 effort: "low", // Play around with this
10 },
11 input: "Write a one-sentence bedtime story about a unicorn.",
12});
13
14console.log(response.output_text);

One thing to keep in mind is that more effort doesn’t guarantee a better answer. In my tests, answer quality was mixed: increasing effort sometimes helped, but the gains could diminish, and quality could get worse at higher settings. Cost and latency generally rose as effort increased, though the relationship wasn’t linear.

That makes this another product-specific tradeoff. Find the lowest effort setting that meets your quality bar and keeps latency acceptable. The available settings and their behavior vary by model, so test the configuration you plan to run.

Prefix your prompts with static instructions

Many applications send repeated instructions, tool definitions, or reference material with each request. When requests share the same prefix, prompt caching can let a provider reuse part of that input instead of processing it from scratch.

A practical starting point is to put stable instructions and shared context first, then put the user’s request and other changing content later. The provider and model determine the minimum prefix length and other eligibility rules, so stable ordering alone doesn’t guarantee a cache hit.

OpenAI’s pricing gives one concrete example: for short-context GPT-6 Luna requests, cached input is currently listed at $0.01 per million tokens, compared with $0.10 per million uncached input tokens. That is a 90% lower rate for the cached input tokens; it does not make the whole request 90% cheaper. Uncached input and output are still billed, and actual savings depend on how much of the prompt is reused. Check the current pricing before using a specific rate.

You can also leverage the OpenAI dashboard to monitor cache activity. The token counts show what was written to and read from cache, but it can be hard to map them back to the exact prompts or prompt sections. OpenAI also recently documented per-request cache diagnostics for the Responses API on GPT-5.6 and later supported models. They can help explain why a request reused fewer tokens than expected.

Putting static content before dynamic content is a practice we’ve followed from the beginning. We haven’t isolated its effect in a before-and-after test, though, so I treat it as a useful way to structure prompts—not as a savings result we’ve measured.

Measuring your spend

Total monthly spend is an easy business metric to compute. When request volume grows quickly, though, total spend alone can hide efficiency gains. The bill may rise while the cost of serving each successful request falls.

In that situation, compare cost per successful request alongside monthly spend. Define success using the quality your product requires, and keep latency within its acceptable range. That gives you a way to see whether a change made each useful interaction less expensive, even when more people are using the feature. OpenAI's Responses API returns a breakdown of the token count used for input and output, so you can combine this with the pricing data of the model you used to construct an estimate of the query cost.

TypeScript
1import OpenAI from "openai";
2
3const client = new OpenAI();
4
5// Define pricing per 1M tokens
6const COST_PER_1M_INPUT = 10;
7const COST_PER_1M_OUTPUT = 50;
8
9const response = await client.responses.create({
10 model: "gpt-6-astra",
11 input: "Write a one-sentence bedtime story about a unicorn.",
12});
13
14// Extract token counts from usage object
15const inputTokens = response.usage?.input_tokens ?? 0;
16const outputTokens = response.usage?.output_tokens ?? 0;
17
18// Calculate exact call cost
19const costUSD =
20 (inputTokens / 1_000_000) * COST_PER_1M_INPUT +
21 (outputTokens / 1_000_000) * COST_PER_1M_OUTPUT;
22
23console.log(`Input tokens: ${inputTokens}`);
24console.log(`Output tokens: ${outputTokens}`);
25console.log(`Cost for this call: $${costUSD.toFixed(6)}`);
26

The goal is to lower the cost of providing a successful answer while preserving the quality and latency your product needs.

This is not an exhaustive list!

There are many more cost control methods you can try to employ, such as:

  • Context trimming - Send the minimum amount of context needed for the LLM to generate a good output. Shorter context -> less tokens sent -> less cost. One way to ensure this is to ensure your system is able to retrieve the right information. To measure retrieval, you can check out my other blog post here.
  • Model routing - Send simpler requests to smaller models and more complex ones to larger models. You can have a small model perform the decision (e.g. Jev), but note that this is another component in your system that needs its own set of evaluations.



Last updated

Share


Like