01
Why tokens matter
Every message Claude reads and writes uses tokens. That includes your whole conversation history, files, tool results, and instructions, not just your latest message.
Saving tokens means:
- Hitting usage limits less often
- Faster answers
- Lower API bills
- Better focus, because Claude reads less noise
Long, messy chats are the biggest hidden cost.
02
How Claude usage limits work
All Claude plans use a rolling five-hour usage window, and paid plans also have weekly limits. Paid plans can buy extra usage credits at API rates once they hit a limit.
Bigger models, higher effort, long chats, large files, and tools all use more of your limit.
03
Pick the right model
- Haiku
- Fastest and cheapest. Simple tasks, summaries, extraction.
- Sonnet 5.5
- Most everyday work and coding.
- Opus 5.5
- Hard reasoning and complex projects.
- Fable 5.1
- Long-running agents and the hardest problems. Uses limits fastest.
Start with a smaller model and move up only when quality needs it.
04
Lower the effort
Claude’s effort levels range from low to max. Higher effort means more thinking tokens.
- Low or medium for routine tasks
- High for planning and debugging
- Max only for the hardest problems
In Claude Code, change it with /effort.
05
Start fresh chats often
- Start a new chat for each new topic
- Ask for a short summary before you switch chats
- Paste only the summary into the new chat
- Do not keep one chat going for days
Every new message in a long chat re-reads everything before it.
06
Edit instead of correcting
If an answer misses, edit your original message and resend it instead of adding a correction. That replaces the wrong turn instead of stacking another one.
07
Write tighter prompts
- Give the context Claude needs, once
- Ask several related questions in one message
- Ask for the length and format you want
- Ask for bullet points or tables instead of long prose
- Say “no preamble” when you only need the answer
08
Share files carefully
- Share only the pages or sections that matter
- Convert huge PDFs to text when possible
- Remove repeated boilerplate
- Use Projects for files you reuse, instead of re-uploading
09
Turn off what you do not need
Every connector, tool, and plugin adds instructions Claude must read. Turn off connectors and web search for tasks that do not need them.
10
Keep CLAUDE.md lean
Claude Code reads CLAUDE.md at the start of every session. Keep it short:
- Only rules Claude actually needs
- No long histories or pasted documentation
- Link to files instead of copying them in
- Remove rules that no longer apply
11
Use Claude Code commands
- /clear
- Start fresh when you switch tasks
- /compact
- Summarize a long session to free up context
- /model and /effort
- Use cheaper settings for simple work
- /mcp
- Check which servers are connected, and disconnect unused ones
Point Claude to specific files instead of asking it to explore the whole codebase.
12
Use subagents wisely
Subagents can read and search in their own context, keeping your main session lean. Run them on a cheaper model, such as Haiku, for simple research and file reading.
13
Save on the API
- Use prompt caching for repeated instructions and documents
- Use the Batch API for non-urgent jobs, at 50% off
- Choose Haiku 5.5 for high-volume tasks
- Keep Haiku 5.5 prompts under 100K tokens to avoid higher rates
- Set maximum output lengths
- Use token counting before sending large requests
14
Habits that stretch your limits
- Plan big work in one message instead of many small ones
- Do heavy tasks early in your usage window
- Save good prompts as skills instead of retyping them
- Ask for a plan before long tasks
- Stop a task early if it is going the wrong way
15
Master token-saving prompt
Answer concisely, with no preamble or repetition.
Use bullet points. Keep the answer under [length].
If you need more information, ask one short question instead of guessing.
Task: [task].
The golden rule
Do not keep feeding Claude more context. Give it less, but give it exactly what it needs.
- Model
- Effort
- Fresh chats
- Lean files
- Commands
- Caching
Start a fresh chat for your next task, with one clear message. Notice how much longer your limits last.