Enterprise

Microsoft tells its own engineers to stop 'tokenmaxxing' as division-level AI budgets kick in

An internal memo from EVP Jay Parikh caps token spend by division, switches the default internal model to a cheaper GPT-5.6, and hands every enterprise SaaS vendor an awkward mirror.

Microsoft EVP Jay Parikh sent an internal memo on August 4 telling the company’s engineers that “tokenmaxxing is not what we are optimizing for,” a line first reported by 404 Media. The memo introduced division-level AI token budgets, made GPT-5.6 the cheaper default for internal use, and reframed the reprimand in gentler prose: “We are not optimizing for fewer tokens. We are optimizing for more impact per token.”

The company that hosts OpenAI’s models is now rationing them internally.

The rationing didn’t start with the memo. In May, Microsoft cancelled most of its internal Claude Code licences and gave engineers a June 30 deadline to migrate onto GitHub Copilot CLI. GitHub itself shifted to usage-based billing in June. By July, divisions had a formal “AI token budget target” and individual employees could track their own spending against it. Internal Copilot guidelines cited by 404 Media described engineers running “hundreds of dollars a month to a few thousand dollars in tokens.” Asked to comment, a Microsoft spokesperson told The Register they had “nothing to add.”

One anonymous staffer put the subtext plainly to 404 Media: the limits “really feels like the ultimate admission that we, as hosts of AI infra, can’t afford our own AI products.”

That admission travels. Since June, AT&T, Meta, Uber, Walmart, and Amazon have all begun capping or throttling employee AI spending. Uber burned through its entire 2026 AI coding token budget in four months. Amazon spent $1.8 million on a single internal Claude Sonnet deployment that ballooned past plan. Per-token prices have fallen roughly 98 percent since late 2022, and enterprise AI bills have still tripled, because agentic tools consume dramatically more tokens per task than anyone modeled a year ago.

The structural problem hiding inside Parikh’s memo is the one every SaaS vendor pricing AI features on flat seats is now staring at. If Microsoft’s own engineers can burn four-figure monthly token bills on Copilot CLI, the unit economics of “AI included” enterprise pricing don’t survive contact with agentic workloads. The industry spent 2024 and 2025 pitching AI as deflationary. The 2026 memos, starting with this one, are pricing it as the opposite.

Sources