Token Efficiency Checklist¶
Qin Yu, 21 May 2026, updated 10 Aug 2026
Token efficiency means spending tokens on useful context and reasoning while removing avoidable overhead.
Keep Active Context Small¶
- Lead with the outcome — state the desired result, scope, constraints, stop condition, validation steps, and expected summary. Short prompts are useful; vague prompts are expensive.
- Treat context as working memory, not storage — attach only relevant files and URLs. Start a new session for unrelated work.
- Keep IDE context focused — Copilot can use open files for inline suggestions, so close irrelevant tabs.
- Reference specific files — reference durable information in files or docs instead of pasting large blocks of text. Avoid screenshots unless the visual information matters.
- Keep code modular — split oversized files and functions along meaningful boundaries.
- Work in phases — use Research → Plan → Implement → Verify so each phase receives only the context it needs.
- Isolate noisy work — use sub-agents for exploration, parallel tasks, and specialised reviews. Return concise findings.
Keep Instructions and Tools Lean¶
- Use the right configuration layer — keep mandatory repo rules in versioned instruction files, reusable workflows in skills, specialised behaviour in scoped agents, deterministic checks in hooks or CI, and temporary details in the session prompt.
- Keep persistent instructions short and current — record only durable conventions and recurring fixes. Review and prune them periodically. Use memory as a convenience, not as the source of mandatory policy.
- Load specialised context only when needed — use skills and file references for progressive disclosure instead of placing every workflow and reference in the default context.
- Remove unused tools — every unnecessary MCP server or tool adds overhead, and it is charged before the agent does anything: enabling tools at all adds a system prompt of a few hundred tokens, plus every tool name, description, and schema, on every request. Prefer a direct API, native CLI, or script for narrow, stable, high-frequency operations.
Control Tool Output¶
- Think in code, not prose — analyse large files or datasets with a script when possible. Scripts are deterministic, reusable, and produce compact output.
- Request bounded outputs — ask for the exact artifact you need, such as one function or a concise diff summary.
- Prefer structured output — use summaries, tables, and focused JSON instead of raw logs. Compress verbose shell output with tools such as rtk-ai/rtk.
- Collapse related tool calls — batch independent reads and actions where supported. Plugins such as jsturtevant/copilot-codeact-plugin can help.
Match the Workflow to the Task¶
- Route by task shape — use cheaper models for extraction, formatting, search-heavy work, and simple transforms. Reserve stronger models for ambiguity, architecture, difficult debugging, and high-risk review.
- Judge cost by completed work — account for retries, long loops, repeated tool output, reasoning tokens, cache costs, and image tokens rather than headline model prices alone.
- Tune effort for the task — where available, start at the model's documented default and step down for quick, low-risk work, reaching up only when extra reasoning is likely to reduce retries or review churn. Effort settings do not transfer across model versions: re-run the sweep instead of carrying them over.
- Parallelise independent work — run separate reviews, comparisons, or checks concurrently when their results do not depend on each other.
- Add critique loops selectively — use Generate → Critique → Revise → Validate when quality matters more than latency.
Validate and Improve¶
- Catch errors early with deterministic guardrails — use tests, linters, type checkers, security scanners, hooks, and CI so mistakes do not compound across long workflows.
- Analyse usage regularly — use Codex
/status; Copilot CLI/context,/usage,/session, and experimental/chronicle tips; or Claude Code/context,/usage, and/cost. - Benchmark representative tasks — re-baseline when changing prompts, model versions, tools, effort settings, or routing rules. Different models respond differently to context ordering, instruction style, and thinking controls.
- Treat misses as incidents — when a workflow fails, update the relevant instruction, skill, hook, eval, or routing rule instead of repeatedly retrying the same approach.
API Workflow Optimisations¶
When building or scripting agent workflows via API:
- Put stable prompt prefixes first — place invariant instructions before variable task data to improve prompt-cache reuse.
- Preserve cache continuity — when the API supports mid-conversation system messages, use them within the documented placement rules for changing permissions, token budgets, or environment context instead of rebuilding the whole prompt.
- Hold reasoning controls constant within a session — settings that shape the rendered prompt, such as effort level or a client-managed task budget, invalidate the cached prefix when they change between requests. Vary them across workloads, not inside one cached conversation.
- Budget the whole loop, not just the request — where the API offers an advisory per-task token budget, size it from measured spend on representative tasks and pair it with a hard
max_tokensceiling. A budget set too low reads as an impossible task and causes early stops. - Keep tool contracts compact — design tools to return concise, structured results rather than verbose text.
References
- OpenAI prompt caching
- OpenAI Codex CLI slash commands
- GitHub Copilot prompt engineering
- GitHub Copilot content exclusion
- GitHub Copilot CLI command reference
- GitHub Copilot CLI best practices
- GitHub Copilot Chronicle
- VS Code context management
- Claude Code commands
- Anthropic: Writing tools for agents
- Claude effort
- Claude task budgets
- Introducing Claude Fable 5 and Claude Mythos 5
- OpenAI: GPT-5.6 — Frontier intelligence that scales with your ambition