← Blog AI Cost Management

Claude Code Using Too Many Tokens? 7 Ways to Cut Your Usage

By FavTray · · Updated

Written by the team behind FavTray, a Mac developer dashboard for the menu bar. We check every app we cover against its own site, docs or a hands-on test, and each post says which.

Short answer: Claude Code burns tokens because it resends your whole conversation, every file it has read and every command output with each message, so long sessions get expensive. Type /clear between unrelated tasks (or /compact mid-task), block node_modules and build output with permissions.deny rules, keep a short CLAUDE.md, switch to Haiku with /model for simple work, and name the exact file and function in your prompts.

TL;DR

  • Every message resends the whole session, so context size, not your prompt, drives the cost.
  • Agents cost more than chat because each tool call is another request carrying that context.
  • /clear between tasks, permissions.deny for bulky files, Haiku for simple work, specific prompts.

You opened your Anthropic billing page, or hit your plan’s usage limit, sooner than expected. Much of that token usage is avoidable. This guide covers seven specific techniques that cut Claude Code’s token use without changing how productively you use the tool.

Why does Claude Code burn through tokens so fast?

Claude Code consumes tokens aggressively because every interaction carries full conversation context, system prompts, tool call results, and file contents — often resending 100,000+ tokens of prior context with each new message. A single coding session that spans 15-20 messages can accumulate over 500,000 input tokens. Claude Code’s prompt caching bills most of that resent context at a fraction of the normal input price, but it still adds up, and on a Pro or Max plan it all counts against your usage limits.

Here’s where the tokens actually go in a typical Claude Code interaction:

Token CategoryTypical SizeResent Each Turn?Cumulative Cost Impact
System prompt1,500-2,500 tokensYesLow per turn, adds up
Conversation history5,000-200,000 tokensYesMajor cost driver
File read results2,000-50,000 per fileYes (in history)Major cost driver
Tool call context500-5,000 per callYes (in history)Moderate
Your new message50-500 tokensOnceNegligible
Claude’s response500-5,000 tokensIn future turnsModerate

The compounding effect is the critical insight. Every file Claude reads, every command it runs, every tool call it makes — all of that becomes part of the conversation history. By message 10 of a session, you might be sending 150,000 tokens of context for a 100-token question. At Sonnet 5.5’s $2-per-million input price, that’s $0.30 per message before caching discounts.

Why does an agent cost more than chat?

A chat turn is one request: your question, a little history, one answer. An agent task is a loop of reading a file, editing it, running the tests, reading the output and trying again, and every step is a separate request that resends the whole working context.

Here is an illustration, not a measurement. Say a session carries 150,000 tokens of context on Sonnet 5.5. Sent uncached at $2 per million input tokens, that is about $0.30 per request. Read from the prompt cache, which costs 10% of the input price ($0.20 per million), the same context is about $0.03. Output is billed on top at $10 per million, so a 2,000-token reply adds $0.02. A chat turn with a 3,000-token prompt and a short answer costs about a cent. Caching keeps agents affordable, and a cache miss after a break (the cache lasts five minutes by default on an API key) pays the full input price again, plus a cache-write premium.

Parallel agents multiply this. Anthropic says agent teams use about 7x more tokens than standard sessions when teammates run in plan mode, because each teammate keeps its own context window. For scale, Anthropic’s figures from enterprise deployments average about $13 per developer per active day and $150-250 per month, with 90% of users staying under $30 per active day.

Sources: Claude Code cost docs and Claude models overview, both checked 30 September 2026.

Understanding this cost structure is the foundation for all seven optimization techniques below.

Tip 1: Use /clear between unrelated tasks

The /clear command resets your conversation context to zero, eliminating accumulated history that inflates every subsequent message. Use it whenever you switch tasks, finish a debugging session, or notice your session has exceeded 10-15 turns. It pays off most for developers who work on several tasks in one session.

Most developers treat a Claude Code session like a continuous conversation, asking about authentication, then switching to database queries, then moving to frontend styling — all in one session. By message 20, every new question carries the full weight of all previous topics.

The rule of thumb: if your next question has nothing to do with your previous conversation, type /clear first. You lose the conversational continuity, but you save thousands of tokens per message for the rest of that task.

A related command, /compact, compresses conversation history into a summary rather than deleting it entirely. This is useful when you’re mid-task but the accumulated context has grown unwieldy. It preserves the key decisions and context while dramatically reducing token count.

Tip 2: Block large directories with permissions.deny rules

Claude Code has no .claudeignore file, despite what many guides say. The supported way to stop it reading files that add bulk without value (node_modules, build outputs, generated code, test fixtures, binary assets) is a permissions.deny list of Read(...) rules in .claude/settings.json at your project root. This reduces the tokens consumed by file-read tool calls, which are one of the largest cost contributors in codebase-aware sessions.

Example .claude/settings.json for a typical web project:

{
  "permissions": {
    "deny": [
      "Read(./node_modules/**)",
      "Read(./dist/**)",
      "Read(./build/**)",
      "Read(./.next/**)",
      "Read(./coverage/**)",
      "Read(./package-lock.json)",
      "Read(./**/*.min.js)",
      "Read(./**/*.map)"
    ]
  }
}

Without these rules, Claude Code may read package-lock.json (often 50,000+ lines) or crawl into node_modules when searching for patterns. One accidental lockfile read puts thousands of lines of JSON into your session, and every later message resends them.

For monorepos or large projects, the impact is even more dramatic, because there is simply more generated code for Claude to wander into.

Tip 3: Write a focused CLAUDE.md file for your project

A well-written CLAUDE.md file gives Claude Code the architectural context it needs upfront, reducing the number of file reads and exploratory tool calls it makes to understand your codebase. Instead of Claude reading file after file to figure out your project structure, it reads your CLAUDE.md and knows immediately where things are and how they fit together.

An effective CLAUDE.md includes:

  • Project structure: Key directories and what they contain
  • Tech stack: Frameworks, libraries, and versions
  • Coding conventions: Naming patterns, file organization rules
  • Common tasks: How to run tests, build, deploy
  • Architecture decisions: Why things are structured the way they are

The token savings come from reduced exploration. Without a CLAUDE.md, Claude Code might read several files to understand your project before it can answer a question about where to add a new feature. With a good CLAUDE.md, it can often go straight to the right file.

This technique pairs perfectly with the deny rules from Tip 2: together, they ensure Claude reads the right files and skips the wrong ones.

Tip 4: Switch models mid-session with /model

Not every task requires Sonnet-level reasoning. Type /model to switch to Haiku mid-session (or start with claude --model haiku) for simple tasks like generating boilerplate, writing documentation, formatting code, or answering quick factual questions. Haiku 4.5 costs $1/$5 per million input/output tokens compared to Sonnet 5.5’s $2/$10, half the per-token price (list prices checked 30 September 2026).

Tasks well-suited for Haiku:

  • Generating repetitive code patterns (CRUD endpoints, test boilerplate)
  • Writing or updating documentation strings
  • Simple refactors (renaming, extracting functions)
  • Answering questions about syntax or API usage

Tasks that need Sonnet or Opus:

  • Complex debugging across multiple files
  • Architecture design decisions
  • Large-scale refactoring with interdependencies
  • Security review and vulnerability analysis

The cost difference adds up. At half the per-token price, whatever share of your tokens moves to Haiku costs half as much: move 40% and the day’s total falls by about a fifth.

You can also set model preferences in your CLAUDE.md to remind yourself which model to use for which types of tasks.

Tip 5: Write compact, specific instructions

Vague prompts cause Claude Code to do more work — reading more files, running more exploratory commands, and generating longer responses. A specific, constrained prompt produces a targeted response with fewer tool calls and less output. The difference between “fix the authentication” and “fix the JWT expiry check in src/auth/middleware.ts line 45” can be several exploratory file reads and searches.

Token-efficient prompting patterns:

  • Name the file: “In src/api/routes.ts” instead of “in the routes file”
  • Specify the function: “Update the validateToken function” instead of “update the validation logic”
  • Constrain the scope: “Only change the return type” instead of “refactor this”
  • Set output limits: “Show me just the changed function” instead of “show me the updated file”

Each unnecessary tool call (file read, directory listing, command execution) adds 1,000-10,000 tokens to your session. A well-scoped prompt that requires zero exploration saves those tokens entirely.

Tip 6: Avoid reading large files when you know the specific section

When you ask Claude Code to work on a specific part of a large file, tell it exactly which lines or section you mean. Otherwise, it reads the entire file into context. For a 2,000-line file, that’s approximately 8,000-10,000 tokens per read — and it stays in your conversation history for every subsequent message.

Instead of: “Look at the UserService class in services.ts”

Try: “Look at the createUser method around line 150-180 in src/services/user-service.ts”

For very large files (1,000+ lines), consider whether Claude Code needs the file contents at all. Often, you can paste just the relevant 20-30 lines directly into your message, avoiding a tool call entirely. This is especially useful for quick questions about specific logic — paste the function, ask your question, get the answer, no file read required.

If you find yourself frequently working with large files, that’s also a signal to refactor. Smaller, focused files are better for both human readability and AI token efficiency.

Tip 7: Monitor token usage in real time and set session budgets

You can’t optimize what you can’t measure. FavTray tracks your Claude Code spending from your macOS menu bar by reading local log files, showing today’s, 7-day and 30-day costs, and /usage in Claude Code shows what the current session has cost. When you see a session crossing $5, that’s your cue to apply the techniques above — clear context, switch models, or scope your next prompt more tightly.

Effective session budgets follow a simple framework:

  1. Set a daily target: Monthly budget divided by 22 working days
  2. Set a session warning: One-third of your daily target per session
  3. Review sessions that exceed the warning: Were the extra tokens productive or wasteful?

If you’re on a subscription and still hit your limits every week, work out whether Claude Max is worth it for your hours before you upgrade.

FavTray reads the ~/.claude/ log files that Claude Code already creates on your Mac, calculates costs based on current token pricing, and shows the running totals one click from your menu bar. There’s no account, and no FavTray server ever sees your usage.

Putting it all together: a token-efficient Claude Code workflow

Combining all seven techniques creates a workflow that’s both productive and cost-conscious:

  1. Start each task fresh with /clear
  2. Let CLAUDE.md provide context so Claude doesn’t need to explore
  3. Use permissions.deny rules to keep junk out of file reads
  4. Scope your prompts tightly with specific file and function references
  5. Use Haiku for simple tasks, Sonnet for complex ones
  6. Paste small code snippets instead of triggering file reads for quick questions
  7. Watch your costs in FavTray and adjust when a session runs hot

The same ideas apply to any model API; see how to reduce AI API costs for the provider-side levers such as caching and batch discounts. The key insight is that most token waste comes from accumulated context and unnecessary exploration — problems that are easy to fix once you understand where the tokens go.

Frequently Asked Questions

Why does Claude Code use so many tokens?

Claude Code uses more tokens than expected because every interaction includes system prompts (~2,000 tokens), full conversation history, file contents read during the session, and tool call results. A 10-turn conversation can accumulate 200,000+ input tokens as prior context is resent with each message.

Does clearing conversation history reduce Claude Code costs?

Yes, using /clear or /compact resets or compresses the conversation context, dramatically reducing input tokens for subsequent messages. A conversation with 150K tokens of accumulated context resends all of it with every message: about $0.30 per message at Sonnet 5.5's list input price, less when it's served from the prompt cache, and it still counts against your plan's usage limits. After clearing, the next message starts fresh at a fraction of that cost.

What is a .claudeignore file and how does it save tokens?

Claude Code doesn't read a .claudeignore file. To keep it out of node_modules, build output, lockfiles and test fixtures, add permissions.deny rules such as Read(./node_modules/**) to .claude/settings.json in your project. Claude Code then can't load those files into context.

How much can I save by switching models mid-session in Claude Code?

Haiku 4.5 costs half as much per token as Sonnet 5.5 ($1/$5 versus $2/$10 per million input/output tokens, checked 30 September 2026), so a boilerplate task that costs $0.30 on Sonnet costs about $0.15 on Haiku. How much a day saves depends on your mix: if simple tasks are 40% of your tokens, moving them to Haiku cuts the day's total by about a fifth.

Why do AI coding agents cost more than chat?

A chat turn is one request with a short prompt. An agent works in a loop of reading files, editing, running commands and reading the output, and every step is a separate request that resends the whole conversation so far, so one task can send hundreds of thousands of tokens. Prompt caching bills most of that resent context at 10% of the input price, which keeps it affordable, and Anthropic says agent teams use about 7x more tokens than a standard session when teammates run in plan mode.

Are AI agents worth the extra cost?

It depends on the task. Agents earn their cost on work they can read, edit and verify on their own, such as a multi-file refactor or a failing test you can't trace. For a question about one function, pasting it into chat is cheaper and often just as quick. Anthropic puts the average across enterprise deployments at about $13 per developer per active day, so weigh that against the time the agent saves you, and scope tasks tightly so it doesn't spend tokens exploring.

Something new: My Dock Buddies