Not every failing agent needs fine-tuning.
Rewriting the harness lifted Qwen3.5-4B
from 0.16 to 0.30 on held-out tasks,
but never cut the share of bad plans.
A LoRA cut those from 28% to 5%.
Loops are a harness bug; bad plans are a
weights bug. #AgentHarness
