Chelle Code Michelle Hallworth

chelle code / guide

Working with Claude Code: what it costs, and what I got wrong

Five principles and a running log of lessons, including advice that turned out mostly wrong and a setup that filled up the AI's memory on the first message.

For: Intermediate Claude Code users whose sessions feel slow, expensive, or forgetful. Updated September 28, 2026. Free to use and adapt.

Five principles

  1. Context is the scarce resource. Every token in the window costs money or usage, latency, and attention. Most quality problems and most cost problems are context problems first.
  2. Cost follows tokens, not requests. A small question in a session already carrying a long history and big file reads is an expensive question. Know what is riding along.
  3. Capability is layered: model, then harness, then context. Pick the model for how hard the reasoning is, the harness (chat, CLI agent, subagents, scheduled job) for the shape of the work, and curate the context for both.
  4. Durable knowledge goes in files, not chat history. Sessions are disposable; repos are not. If a session taught you something, write it down before closing it. (The starter kit is where.)
  5. The model is a collaborator with no memory and a lot of confidence. Constrain it with written rules and verify anything load-bearing.

Lesson: audit tools generate recommendations, not conclusions

I ran a token-optimization plugin over my whole setup. The measurement was useful. Three of its four recommendations were wrong, each in a different way:

  • It contradicted something I already knew. It said disabling my connectors would save hundreds of tokens per server. The CLI already defers connector schemas and loads them on demand, so the real saving was a couple hundred tokens total, and the setting was global, so it would have switched off integrations I use every day. I applied it and reverted it the same session.
  • The mechanism it proposed does not exist. It suggested moving my model choice out of settings and into CLAUDE.md. CLAUDE.md cannot set the main session’s model; only settings and /model can.
  • Its numbers were off by an order of magnitude. One setting was billed as recovering about 2,000 tokens; the real figure was a few hundred.

The realistic total across everything that held up was a few hundred tokens per session. Config-level optimization is not where the savings are. Session length is. Clearing between topics beat every config change combined.

Lesson: connectors plus a heavy skill can fill the window before you type

A new conversation in the desktop app compacted on its first prompt. Account-level connectors (email, calendar, drive, notes) were attaching their tool schemas to every conversation, and I had invoked a skill that injects a very large reference document by design. Either alone was fine. Together they blew past the threshold. The fix: turn connectors on per project, not per account, and spend heavy skills in otherwise lean sessions.

Habits that stuck

  • One task, one session. Start fresh rather than steering a long session into a new topic.
  • If a session compacts repeatedly, the working set is too big for the approach. Restructure (subagents, a written handoff, narrower scope) instead of pushing through.
  • Send broad searches to subagents so the main session keeps the conclusion, not the file dump.
  • For big files, extract the slice you need (a page range, a grep) instead of reading the whole thing.
  • Automated “always do X” behavior belongs in hooks, not in hoping the model remembers.

← All free things

Subscribe to Chelle Code

New writing and new free tools. Free, whenever there's something to send.

Draft: signup is not wired to Buttondown yet.