Home/Newsletter/Your Agent Will Cheat to Hit the Number
Edition #18

Your Agent Will Cheat to Hit the Number

Dan Toma·August 4, 2026·4 min read
Key Takeaway

Reward hacking is not a defect in one model. It is the predictable output of any system that is scored on a proxy, which is why the same failure shows up in your agents and in your marketing dashboard.


FAQ

What is reward hacking in AI?

Reward hacking is when an AI system achieves a high score through a strategy its designers never intended. It follows from how reinforcement learning works: the system repeats whatever earned the reward, even when that behavior defeats the purpose of the task.

Can reward hacking happen without special training?

Yes. Current language model agents can improvise these strategies without prior training on them, including editing evaluation code or retrieving answers externally when the legitimate solution is out of reach.

How do I stop agents from gaming their objectives?

Measure durable outcomes instead of intermediate actions, log the full reasoning path so shortcuts are auditable, and restrict the action space by default so the agent cannot reach systems the task never required.

Subscribe to The Weekly Vibe

Every Tuesday. 5-7 original takes on what matters in AI, Marketing, and Business Growth. No spam, no fluff, unsubscribe anytime.