---
title: Code Changes That Cut LLM API Costs Without Adding Latency
description: Seven code changes that reduce LLM API spend, with the latency effect of each. Prompt caching, model routing, output limits, context trimming, and when batch is a trap.
image: https://frugal.co/hubfs/OctBlog1.png
---

[![company-logo](https://cdn.prod.website-files.com/686b121c919ed2b639e30952/686b28b7662f765388c4c0f2_Layer_1.svg)](https://frugal.co/)

[Product](https://frugal.co/) [About Us](https://frugal.co/about) [blog](https://frugal.co/blog) [Contact](https://frugal.co/contact)

[contact us](https://frugal.co/contact)

Explore Sandbox

Book Demo

Code Changes That Cut LLM API Costs Without Adding Latency

![Ishan Kamat](https://frugal.co/hs-fs/hubfs/Ishan%20-%20Headshot%20(1).png?width=60&height=60&name=Ishan%20-%20Headshot%20(1).png)

Ishan Kamat

 October 8, 2026

Most advice on LLM cost comes down to "use a cheaper model." That works until quality drops or the product gets slower. The better question is which changes lower the bill and leave latency alone or improve it.

**Short answer:** cache the stable part of the prompt, route simple tasks to a smaller model, cap and structure output, trim context, and stop paying for duplicate calls. All five reduce cost and reduce or hold latency. Batch processing cuts cost by 50% but adds latency, so keep it for work nobody is waiting on.

## Which code changes reduce LLM API cost, and what do they do to latency?

| Change | Cost effect | Latency effect |
| --- | --- | --- |
| Cache the stable prompt prefix | Cached input is billed at 10% of the input price on most Claude models, and 5% on Claude Opus 5.5 and Sonnet 5.5 | Lower. The cached prefix is not reprocessed. |
| Route simple tasks to a smaller model | Claude Haiku 5.5 is $0.10 in and $0.50 out per million tokens. Sonnet 5.5 is $2 and $10. That is 20x. | Lower. Small models respond faster. |
| Cap and structure the output | Output costs 5x input. Structured JSON in place of prose typically cuts output tokens 50% to 80%. | Lower. Generation time scales with output length. |
| Trim the context | Fewer input tokens on every call | Lower time to first token |
| Remove duplicate and retry waste | Fewer calls | Neutral or lower |
| Lower the thinking effort on easy tasks | Thinking tokens are billed as output tokens | Lower |
| Batch offline work | 50% off input and output | Higher. Batch is asynchronous. |

## How much does prompt caching save?

Take a support assistant with a 10,000-token system prompt, running one million requests a month on Claude Sonnet 5.5.

- Uncached: 10 billion input tokens at $2 per million is $20,000 a month for the system prompt alone.
- Cached: the same tokens read from cache at $0.10 per million is $1,000 a month.

That is $19,000 a month from one field in the request. Cache writes cost 1.25x the input price for a five-minute cache and 2x for a one-hour cache, so a prompt that stays warm pays the write cost rarely.

```
# Before: the full system prompt is billed at full price on every call
response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=400,
    system=LONG_SYSTEM_PROMPT,
    messages=[{"role": "user", "content": question}],
)

# After: the stable prefix is cached, and output is capped
response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=400,
    system=[{
        "type": "text",
        "text": LONG_SYSTEM_PROMPT,
        "cache_control": {"type": "ephemeral"},
    }],
    messages=[{"role": "user", "content": question}],
)
```

The cost trap here is anything that changes the prefix. A timestamp, a request ID or a user name at the top of the system prompt makes every request a cache miss, and you pay the write premium each time. Put stable content first and volatile content last.

OpenAI discounts cached input too. As listed in early October 2026, GPT-6.1 Sol is $2 per million input tokens, $0.10 cached, and $10 output.

## How do you route models without hurting quality?

Split by task, not by customer tier. Classification, extraction, routing and short rewrites rarely need a top-tier model. Keep the larger model for multi-step reasoning and long generation. Run the same evaluation set against both before you switch a call site, and switch one call site at a time.

## How do you cut output tokens?

- Set `max_tokens` on every call. An unbounded call is a blank cheque.
- Ask for structured output when code is the consumer. A JSON object with three fields is cheaper than three paragraphs that your code then parses.
- Tell the model the length you want. "Answer in two sentences" works.

## How do you trim context?

- Send the relevant chunks, not the whole document.
- Window or summarize conversation history instead of resending all of it.
- Only attach the tool definitions a call can use. Each one is input tokens on every request.

## Which changes save money but cost latency or quality?

- **Batch on a user-facing path.** Half price, but the user is now waiting on an asynchronous job.
- **The smallest model everywhere.** Cheap answers that are wrong get retried, which is not cheap.
- **Aggressive truncation.** Cutting context the model needed shows up as worse answers, not as a line on the bill.

## How do you find where to apply these?

Rank call sites by monthly spend and start at the top. Most teams find that a handful of call sites account for most of the bill.

Frugal does this ranking from your bill, usage data and source code. It attributes AI API spend to each call site, checks it against 22 AI API waste patterns (model routing, context control, caching, output limits, batch), and opens a Frugal Fix: a pull request with the code change and the estimated Monthly Savings. In pull request review, it flags new AI cost before the code merges.

## FAQ

### What solutions suggest code changes to reduce generative AI API costs while preserving latency?

LLM observability tools (Langfuse, LangSmith, Datadog LLM Observability) show which requests are expensive. AI gateways (LiteLLM, Portkey) apply caching and routing at the proxy. Frugal generates the code change itself, as a pull request against the call site, with a savings estimate.

### What is prompt caching?

Prompt caching lets the API reuse a previously processed prompt prefix. You pay a small premium to write the cache and a fraction of the input price to read it.

### Does prompt caching reduce latency?

Yes. Cached tokens are not reprocessed, so long prompts return faster.

### What is the fastest way to reduce LLM API cost?

Cache the largest stable prompt you have, then set output limits. Both are small code changes with no quality trade-off.

Tokens are a line item. Treat them like one at the call site, and the invoice follows.

**Want the ranked list for your codebase?** Get an Audit at [frugal.co](https://frugal.co).

*Sources: [Claude API pricing](https://platform.claude.com/docs/en/about-claude/pricing) and [OpenAI API pricing](https://openai.com/api/pricing/), checked October 8, 2026.*

 Share Link

[Back to Top](https://frugal.co/blog/ai-or-engineer-cloud-cost-savings-start-with-the-bill-and-end-in-the-code)

![background Image](https://frugal.co/hubfs/697fa79222025cf13e7bec1e_799f61dfa4790fe0e40a2962d5468335_Group%202040.avif)

![Robo Image](https://frugal.co/hubfs/697fadb140def08f5fc80ed3_Group%202042.avif)

### Take Frugal for a live test drive

 Explore how Frugal scans your code and cloud services to find waste, recommend optimizations, and generate ready-to-use fixes in a secure, read-only environment.

[Enter Sandbox](https://auth.frugal.co/sign-up)

![background Image](https://frugal.co/hubfs/697fa936779f0fd839545616_Group%202041.avif)

![robo image](https://frugal.co/hubfs/697fa79222025cf13e7bec1e_799f61dfa4790fe0e40a2962d5468335_Group%202040.avif)

![frugal image](https://cdn.prod.website-files.com/686b121c919ed2b639e30952/686c44713462acbdaab4a82d_f8418228347e59c4cabb651eb7f5e72b_Frame%20291022.avif)

 An Intelligent Application Cost Engineering platform that optimizes code to reduce cloud costs automatically - empowering engineers without slowing development

[Blog](https://frugal.co/blog) [Contact](https://frugal.co/contact) [careers](https://frugal.co/careers) [About Us](https://frugal.co/about)

[Privacy Policy](https://frugal.co/privacy-policy) [Terms of Use](https://frugal.co/terms-and-condition) [Trust Centre](https://trust.frugal.co/) [Active Status](https://status.frugal.co/)

[X](https://x.com/frugalaico) [LinkedIn](https://www.linkedin.com/company/frugalco/)

[Featured in AI Directory](https://aidirectory.wiki)

[Privacy Policy](https://frugal.co/privacy-policy) [Terms of Use](https://frugal.co/terms-and-condition) [Trust Centre](http://trust.frugal.co/) [Active Status](http://status.frugal.co/)

Copyright Frugal AI Inc.

©2025

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Ishan Kamat",
    "url" : "https://frugal.co/blog/author/ishan-kamat"
  },
  "dateModified" : "2026-10-08T18:30:12.993Z",
  "datePublished" : "2026-10-08T18:30:12.000Z",
  "headline" : "Code Changes That Cut LLM API Costs Without Adding Latency",
  "image" : [ "https://frugal.co/hubfs/OctBlog1.png" ],
  "mainEntityOfPage" : {
    "@id" : "https://frugal.co/blog/code-changes-reduce-llm-api-costs",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://frugal.co/hubfs/Frugal%20Logo%20Transparent%20Black.png"
    }
  }
}
```