Token Costs Drive AI Gateway Deals in Fintech
- Cormac Repman

- 11 minutes ago
- 2 min read
We booked three AI gateway deals today, and every prospect led with the same pain: token costs are out of control.
Phillip Vetter at Exclaimer walks 24 billion emails annually through their system. He didn't ask about consolidating vendors or reducing tool sprawl. He asked how we could show him what tokens were costing him and give him control over the spend. A 20-minute call about cost visibility turned into a booked demo.
Claude Binns, moving into infrastructure at a logistics company, spent half his call time comparing token pricing across Tetrate's Envoy, Open Router, and Databricks. He wasn't shopping for a new platform. He was shopping for the one that let him see where every token dollar went. He booked a meeting to dig into our routing and cost reporting.
Ian Lester at LMS didn't mention consolidation either. Infrastructure teams don't care about fewer vendors. They care about predictability. Token costs look like a variable expense that scales with traffic in ways they can't forecast. One $50K month becomes $200K the next month when traffic spikes. That unpredictability sells deals faster than efficiency metrics ever will.
Here's what we got wrong in our early pitches: we led with "simplify your AI stack" and "route to the cheapest model." Engineering teams already know which models are cheap. They know Claude is more expensive than Open Router. They know local inference costs nothing but runs slower. Stating the obvious doesn't book calls.
What actually moved Phillip and Claude and Ian was specificity: showing them the exact cost per token by model, by day, by endpoint. Showing them the ability to set spend guardrails. Showing them that they could AB test routing strategies without guessing at the financial impact. That's not consolidation. That's cost intelligence.
The secondary themes were real but tertiary. Phillip mentioned data privacy alongside token costs. Claude mentioned comparing inference quality across models. But both came back to the same question: can you show us what we're actually spending and give us control?
We're selling into organizations where the CFO knows an AI bill arrived but has no visibility into what drove it. They're running inference at scale with no choke point, no spending cap, and no way to measure ROI per dollar spent. Token cost reduction isn't a secondary feature. It's the entire use case.
The pitch changes tomorrow. We stop leading with platform consolidation and start leading with: "Here's what you're probably spending. Here's how you make it visible. Here's how you control it without breaking inference quality."
Phillip books Wednesday. Claude books Wednesday. Three calls, three token cost pain signals, three closes. We're onto something real.

Comments