Context Engineering in Production #20
Don't measure AI coding only by output quality.
Track:
• context size
• prompts/task
• retrieval accuracy
• cost/task
• time/task
If context engineering works, those numbers should move.

Context Engineering in Production #20
Don't measure AI coding only by output quality.
Track:
• context size
• prompts/task
• retrieval accuracy
• cost/task
• time/task
If context engineering works, those numbers should move.
Context Engineering in Production #19
"AI hallucinated."
Don't stop there.
Check:
Fake symbol?
Fake file?
Invented API?
Wrong parameter?
Unsupported claim?
Groundedness can be tested against repository evidence.
Context Engineering in Production #18
The AI produced an answer.
Now verify it.
Do those files exist?
Are those symbols real?
Are the parameters correct?
Is the answer grounded in retrieved context?
Generate → Verify.
Context Engineering in Production #17
Changed 4 files?
Don't rebuild context from 4,000.
Start with the diff.
Then expand outward only when dependencies require it.
Context should grow with evidence, not repository size.
Context Engineering in Production #16
Before changing code, ask:
Who depends on this?
Finding the target is only half the job.
Understanding the blast radius tells you what the change could break.
Navigate → Impact → Edit.
Context Engineering in Production #15
Traditional workflow:
Search → open → wrong file → retry.
Better workflow:
Question → rank relevant code → inspect only what matters.
Exploration has a token cost.
Make retrieval earn every file it opens.
Context Engineering in Production #14
Context compression shouldn't mean:
"Remove useful information."
It should mean:
"Represent useful information with fewer tokens."
Signatures, relationships and ranked context preserve signal while dropping noise
Context Engineering in Production #13
Don't go:
Repository → AI
Go:
Repository
↓
Module
↓
File
↓
Symbol
↓
Implementation
Each step removes irrelevant context.
By the time AI sees code, most noise should already be gone.
Context Engineering in Production #12
Finding "RankingEngine.score()" is step one.
Now ask:
Who calls it?
What does it call?
What depends on it?
Code becomes useful context when you understand its relationships.
Context Engineering in Production #11
Don't start by loading implementation.
Start with structure:
What exists?
Where?
Who calls it?
What does it call?
Then open implementation only where needed.
Structure first.
Details on demand.
Context Engineering in Production #10
An agent may not need a 2,000-line file.
It may need:
AuthService.login()
→ TokenService.issue()
→ SessionStore.save()
Symbols preserve structure without dragging every implementation detail into context.
Context Engineering in Production #9
Search asks:
"Where does "auth" appear?"
Retrieval asks:
"Which code actually implements the auth flow?"
One returns matches.
The other returns context.
That's a huge difference for coding agents.
Context Engineering in Production #8
Context starts before your prompt.
Extensions.
MCP servers.
Open repos.
Tool output.
A noisy workspace can make an AI session expensive before question #1.
Reduce the environment before optimizing the prompt.
Context Engineering in Production #7
Caching can make repeated context cheaper.
It doesn't make bad context useful.
30K cached tokens of noise are still 30K tokens of noise.
Cache optimizes reuse.
Context engineering optimizes what deserves reuse.
Context Engineering in Production #6
Keep sessions focused.
Good:
Auth → JWT → Claims → Refresh
Bad:
Auth → Terraform → Search → Databricks
Every topic jump carries old context into a new problem.
Focused sessions preserve signal.
Context Engineering in Production #5
A coding session grows every turn.
Code.
History.
Tool logs.
Previous answers.
By turn 20, your new question may be tiny compared with everything traveling with it.
Context debt compounds.
Context Engineering in Production #4
Treat context like a budget.
0–40% 🟢 Healthy
40–50% 🟡 Watch
50–70% 🟠 Compress
70%+ 🔴 Reset
A bigger window doesn't remove the need for context discipline.
Context Engineering in Production #3
The cheapest token isn't a cached token.
It's the token that never enters the session.
Before optimizing models, ask:
Why are we sending this context at all?
Good retrieval removes noise before tokenization.
Context Engineering in Production #2
A 1M-token window doesn't mean you should fill it.
Capacity ≠ necessity.
100K tokens of noisy code can be worse than 5K tokens of relevant symbols.
The goal isn't maximum context.
It's minimum sufficient context.
Context Engineering in Production #1
Your question might be 10 words.
But the model may receive:
• code
• chat history
• tool output
• system instructions
The prompt is often the smallest part of the request.
Optimize what surrounds it.