How to reduce AI token costs across your business
Ramp reports that AI token spending for businesses surged 572% from June 2025 to June 2026.
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.


MUNGKHOOD STUDIO // Shutterstock
How to reduce AI token costs across your business
AI token spend for business grew 572% year over year from June 2025 to June 2026, according to proprietary data from Ramp, which processes AI vendor payments on behalf of thousands of businesses. Finance teams are catching up to what engineering already knows: AI isn’t a line item you set once and forget. It’s a recurring, fast-moving cost that behaves more like infrastructure than SaaS, and it needs the same oversight.
As Ramp explains below, controlling AI token spend requires action at two layers: in the engineering stack and in the finance stack.
Why AI token costs are harder to manage than other software expenses
Most software costs are predictable. SaaS is seat-based, cloud storage is capacity-based, and both follow patterns your finance team can budget against. As Ramp data shows in figures below, AI token costs don’t behave that way.
Token spend varies by model, prompt length, the number of calls your applications make, and which teams are running which workloads. Month-to-month swings are common: The typical median business’s AI spend swings by about 58% month to month. Sixty-one percent of businesses average swings of 40% or more.
The median business uses two distinct AI vendors. The average is 2.8, spread across Anthropic, OpenAI, Google, and others. Without a unified view, you’re reading multiple invoices without the context to know what changed or who drove it.
That’s the core management problem. The spend is significant: 8.4% of companies now exceed $10,000 per month, the level where AI becomes a formal budget line. But the visibility to govern it sits in systems that finance teams don’t control.
The two layers of AI cost control
Reducing AI token costs requires action at two distinct levels:

Ramp
Most cost reduction guides focus entirely on the infrastructure layer because that’s where tokens are generated. But the structures that make cost reduction stick across teams live in the finance layer.
What drives unexpected AI token cost increases
AI token cost surprises follow predictable patterns: model drift to premium tiers, long-context inflation, multi-model sprawl, and missing team-level visibility. These drivers of unexpected AI token cost increases compound because no single team owns visibility across the stack—the invoice arrives only after the spend has already happened.
Model drift upward. Teams start with a budget model for internal tooling and upgrade to a premium model for production use without a corresponding increase in budget. The gap between token share and cost share is where budget overruns hide.
Long-context inflation. Applications that pass large documents, full conversation histories, or rich system prompts in every request consume tokens at a rate that’s hard to anticipate from initial testing.
Multi-model sprawl. Without a clear policy on which models teams can use for which cases, they add models incrementally, and spend accumulates across contracts you didn’t know you had.
How finance teams can track AI API costs by team and model
Tracking AI costs at the team and model level means closing the gap between what your AI providers report and what your finance system understands.
Your providers bill by token consumption. Your finance system needs to see cost by business unit, project, cost center, and model. That level of detail lets you hold teams accountable and spot where optimization is warranted.
A basic tracking framework covers four dimensions:

Ramp
The cost of goods sold (COGS) vs. operating expenses (OpEx) distinction matters more than it sounds. AI tokens that directly serve customers—a product feature, a support chatbot, or a code generation tool in your product—belong in COGS. Tokens consumed by internal tools belong in OpEx.
Mixing them distorts gross margin and makes it harder to model AI unit economics as usage scales.
How to set AI budgets that hold up month to month
If you set a single AI budget at the company level, it’s likely to fail. AI costs aren’t generated at the company level—they’re generated by specific teams running specific workloads.
A more durable approach: Budget by team and model category, not just in aggregate.
Start with a baseline. Pull the last three months of AI spend by provider and map it to the teams that generated it. Your own actuals are the right starting point. Industry medians are useful benchmarks, but your internal pattern is what matters for forecasting.
Set per-team spend targets. Give each team a monthly AI budget based on their current spend and planned workloads. This doesn’t have to be a hard cap initially. A soft target with visibility is enough to change behavior.
Classify new workloads before they scale. Before a new AI feature or application goes to production, require a cost model that includes estimated tokens per request, expected call volume, and projected monthly spend at three growth scenarios. This is standard practice for cloud infrastructure. AI tokens warrant the same rigor.
Review monthly, not quarterly. AI spend compounds fast. Month-to-month swings of 40% or more mean quarterly reviews catch problems too late to prevent budget overruns.
What the data shows about AI cost efficiency
The gap between high- and low-cost AI teams isn’t about which tactics they know. Ramp’s data across thousands of businesses makes the cost difference concrete.
Prompt caching. If your applications send the same system prompt or document context repeatedly, caching stores that input and charges a fraction of the normal rate on subsequent calls. According to Anthropic, cache reads cost 90% less than standard input tokens, a substantial reduction for high-volume, repetitive workloads.
Model selection and routing: Not every request needs your most expensive model. Routing simpler tasks—classification, summarization, short-form drafts—to smaller, cheaper models while reserving frontier models for complex reasoning cuts spend without touching output quality.
Batch processing. Many AI providers offer reduced pricing for asynchronous batch requests—tasks where a response in seconds isn’t required. Internal reporting, document processing, and bulk analysis are common candidates. Anthropic’s Message Batches API reduces costs by 50% for async workloads.
According to Ramp Token Spend Data, as of April 2026: About 74.5% of businesses connected to Anthropic’s API have enabled prompt caching, compared with 50.8% for OpenAI. The share that haven’t are still paying full input token prices.
At the finance layer
These tactics create the structures that make infrastructure optimizations sustainable and catch spending problems before they compound.
Spend limits by team. A monthly limit for each team’s AI budget creates a forcing function: Teams must choose which workloads to prioritize and which to optimize. Start with soft limits (alerts, not hard stops) and move to enforcement as your tracking matures.
Chargeback to cost centers. When AI spend is allocated back to the teams that generated it, those teams have a direct incentive to reduce waste. This is the same mechanism that made cloud cost management effective in engineering organizations, and it works for AI spend.
Anomaly alerting. A 40% month-over-month spike is hard to catch when reviewing invoices manually. Automated flagging of spend that deviates materially from a team’s baseline gives finance the early warning to intervene before an overrun becomes a budget crisis.
This story was produced by Ramp and reviewed and distributed by Stacker.
![]()
