Which Hosting Model Actually Keeps Your AI Bill Defensible?

14 min read
  • Cost & Economics
  • Build vs Buy
  • Deep dive
Contents10 sections

The Invoice That Changed

In one illustrative, composite case drawn from patterns we have seen across clients, a finance director at a mid-sized enterprise flagged an AI invoice that came in roughly three times her team's forecast. Nothing had broken and no contract had changed. The application had been upgraded from a single-turn assistant into something that planned, searched, retried, and checked its own work. The request that once consumed a few thousand tokens now consumed tens of thousands. The unit price was identical. The unit count was not.

The question is how to run AI so that the cost is knowable a quarter in advance and defensible to a CFO or an oversight body. The framing that causes trouble treats this as a choice between calling a frontier API and accepting the bill, or building a GPU cluster and hoping the utilization appears. There are at least four distinct positions available, each with a different cost shape, and most organizations will occupy more than one at once.

Mapping the Hosting Spectrum

At one end sits the frontier token-based API: OpenAI, Anthropic, Google. You pay per million tokens, with output priced higher than input. No infrastructure to run and no floor on spend if usage drops to zero. Purely variable, which suits pilots, bursty internal tools, and anything whose volume you cannot yet predict.

One step along are managed or serverless open-weights endpoints, where a provider hosts Llama, Qwen, Mistral or DeepSeek and bills per token. The mechanic is identical, the rate typically a fraction, and the spread between providers serving the same model is wide. Independent analysis by the Tech Governance Institute found the cheapest provider offering a 70B-class model at $0.20 per million tokens against a median of roughly $0.70 and a high of $2.90, noting the cheapest may be running at a loss to buy market share. A price sustained by a land grab is not a price to build a five-year budget on.

Further along are dedicated or reserved GPU endpoints, rented by the hour rather than the token. The bill becomes a function of hours held, closer to a subscription. Idle capacity costs the same as saturated capacity, so this works only once you have a credible load floor.

At the far end is fully self-hosted open weights, on owned hardware or in colocation. You buy accelerators, power and cool them, patch them, staff them, depreciate them. Highest fixed cost, lowest marginal cost, and the only position where weights and input data stay inside a perimeter you control.

These are positions on a continuum of control and cost structure rather than a quality ladder. A sensible architecture might route high-volume classification to a small self-hosted model, summarization to a managed endpoint, and a few genuinely hard reasoning tasks to a frontier API, all behind one internal interface.

Why Token Bills Move

Two things move a token bill, and only one of them is price. Prices change frequently, usually downward per unit of capability. Workload shape is the larger source of variance. An agentic workflow that decomposes a task, calls tools, retries on failure, and re-reads a long context at every step consumes an order of magnitude more tokens than the single-turn interaction it replaced, for the same user-visible outcome. Reasoning models amplify this by generating intermediate output billed at the higher output rate.

Three levers change effective cost without leaving this end of the spectrum. Caching is the first: OpenAI, a vendor of frontier models, states in its own API documentation that for recent models "cache writes cost 1.25× the standard, uncached input-token rate" with subsequent reads at 0.1× that rate, so that "one write and nine full reads cost 2.15× at that rate, compared with 10× without caching". No independent source corroborates these specific ratios, so treat them as a vendor-stated figure rather than an audited benchmark. For any application with a stable system prompt or repeated document context, that removes most of the input cost. Batch tiers are the second, trading latency for a discount on asynchronous work.

The third is routing. Academic survey work on multi-model routing and cascading describes systems that query a smaller model first and escalate only when a response fails a reliability check, with reported configurations reaching 97.25% of GPT-4-level quality at 24.18% of the cost. The same survey notes that "static model deployment does not account for the complexity and domain of incoming queries, leading to suboptimal performance and increased costs."

The Capex Case for Owning Compute

A server rack with cables and GPU hardware inside a small on-premise server room under fluorescent lighting

Owning compute converts a variable cost into a fixed one. Whether that is an improvement depends entirely on how steady your volume is.

A 2025 cost-benefit analysis of on-premise deployment found that medium-scale open models including gpt-oss-120B, GLM-4.5-Air and Llama-3.3-70B "can run comfortably on just two A100-80GB GPUs ($30k), with accuracy reduction typically within 10%", and that sub-30B models run on a single consumer card at around $2,000. Across 54 scenarios, break-even against commercial API pricing ranged from 3.8 months to 31.2 months, the longest being Llama-3.3-70B against Gemini 2.5 Pro. A payback period of 3.8 months means the hardware effectively pays for itself before the next quarterly review. A 31.2-month payback means a machine with a three-to-five-year useful life spends most of it earning back its purchase price, and any dip in volume pushes the return past the hardware's retirement.

The same analysis puts operating costs covering power, cooling and staffing at roughly 30 to 50 percent of hardware capex annually. On a $30,000 pair of GPUs that is $9,000 to $15,000 a year, most of it labor: somebody patches the serving stack, maintains evaluation harnesses, re-benchmarks when a better open model lands, and owns the security posture of a system now inside your network rather than behind a vendor's audit report. Budget a quarter to a half of one engineer's time on an ongoing basis, not as a one-off setup project.

Power deserves its own line. The International Energy Agency projects data center electricity consumption more than doubling to around 945 TWh by 2030 and estimates that "unless these risks are addressed, around 20% of planned data center projects could be at risk of delays" from grid strain. For an organization adding GPU racks to an existing facility, the binding constraint is often how many kilowatts the building can deliver to a cabinet.

The Capability Gap Is Closing

For most of 2023 the quality argument settled this debate on its own. That is no longer a safe assumption.

Stanford HAI's AI Index found that on MMLU, closed-weight models led open models by 15.9 percentage points at the start of 2024, shrinking to 0.1 percentage point by year end, with the Chatbot Arena gap narrowing from over 20% to 8.0%. Measured as time lag, Epoch AI puts open models roughly four months behind frontier closed models in its 2026 window, slightly wider than the three-month average it measured over 2023 to 2025. The UK AI Safety Institute, evaluating cyber capability, reports four to seven months, noting that "both gaps are narrower than in internal evaluations AISI conducted in 2025, when open weight models lagged the frontier by 6 to 10 months".

These measures disagree, and the disagreement is instructive: the gap is narrow on general knowledge, wider on hard agentic and security work, and it moves. They agree on direction. A few months of lag, on tasks away from the frontier of difficulty, is a tolerable trade for a cost structure you control.

Serving performance cuts the other way from what many assume. In our own experience benchmarking deployments, latency can vary by an order of magnitude or more across serving configurations, and frontier APIs are often slower per call than dedicated endpoints or tuned self-hosted deployments, since a self-hosted or dedicated setup can be tuned specifically for one workload's latency profile. A hard latency budget can push you toward the far end for performance reasons alone.

On customization, fine-tuning on narrow domain tasks generally requires at least a hundred labeled examples and re-tuning whenever the base model updates, and for document-heavy work retrieval augmentation often matches the resulting accuracy gain at lower ongoing cost. Fine-tuning open weights buys permanent control over a model artifact and a standing maintenance obligation.

Licensing and Where Data Can Live

Close-up of a printed license agreement document with a pen resting on top, lit by a desk lamp

"Open weights" and "open source" are not synonyms, and procurement teams get caught by the difference. The Open Source Initiative has published an Open Source AI Definition setting a bar covering training data information, code and parameters, and several widely used model licenses do not clear that bar. Comparative license analysis shows terms diverging on specifics: Llama requires prominently displaying "'Built with Llama' on a related website, user interface, blogpost, about page, or product documentation", while Llama 4 and Hunyuan-A13B trigger a license request above release-date monthly-active-user thresholds and Qwen2.5-72B carries an ongoing 100-million-MAU requirement. None of this is prohibitive for a typical deployment. All of it belongs in legal review before you standardise on a family.

Regulation frequently overrides cost entirely. Under GDPR Chapter V, moving personal data outside the EU or EEA requires an adequacy decision, standard contractual clauses, binding corporate rules or an equivalent mechanism, a direct constraint on where inference can physically happen. The EU AI Act adds data governance and documentation obligations for high-risk systems, with penalties reaching 7% of global annual turnover. For US federal buyers, FedRAMP's AI prioritisation initiative ran from August 2025 to April 2026 and required prioritised services to guarantee data separation, including that model information derived from customer data not leave the customer environment without authorization. When a workload touches regulated data, hosting becomes a compliance boundary rather than an optimization.

A Worked Comparison, Three Ways

The following is illustrative rather than a client engagement, with round numbers chosen to show the shape of the comparison. Prices in this market move quarterly.

Consider a mid-size agency running an internal document summarization and staff support assistant: 200,000 requests a month, averaging 4,000 input and 600 output tokens. That is roughly 10,000 summaries a working day, comparable to a team of a few hundred staff each running two or three document lookups before lunch.

PositionCost mechanicCost at half volumeMain exposure
Frontier APIPer token, output priced higherFalls by halfWorkload shape changing silently
Managed open weightsPer token, lower rateFalls by halfProvider durability, capability lag
Self-hosted (2 GPUs)Amortized capex plus 30-50% annuallyUnchangedUtilization, staffing, power

On the frontier API the bill tracks usage precisely and needs no capacity planning. With most input being repeated policy documents and a stable prompt, caching should remove a large share of input cost. The exposure is that converting the summarizer into an agent that reads three documents and self-critiques raises tokens per request several-fold, with no contract change to flag it.

On a managed open-weights endpoint, the same pattern bills on the same mechanic at a materially lower rate, with the caveat that rates for identical models vary by more than an order of magnitude across providers and the cheapest may not be durable.

Self-hosted, this sits within the two-A100 configuration the cost-benefit research describes for 70B-class models. At 200,000 steady requests with a three-year horizon, a payback period somewhere within that range looks plausible, though the actual position would depend on the specific models and providers being compared. At 20,000 requests a month, or if the service might be cancelled after a year, it clearly is not. The same workload and quality requirement yield three defensible answers, and what flips the result is volume stability, workload lifespan, and whether anyone on staff can run the thing.

A Framework for Choosing

Six questions get an organization most of the way there, and they map onto what neutral procurement guidance already asks. The UK Government's AI Playbook frames AI buying as a structured business case exercise, noting that "typically, any investment approaching £10 million will require a five-part business case", and the NIST AI Risk Management Framework treats deployment location and vendor selection as governance decisions inside a broader risk process.

Start with volume and its predictability, meaning floor volume rather than peak, since fixed capacity only pays back against load you can guarantee. The most common error we see is a business case built on adoption that has not happened yet. Then data sensitivity and regulatory exposure, which can end the analysis before cost enters it: where GDPR transfer rules, FedRAMP boundaries or sector residency requirements constrain where inference may occur, the viable set shrinks to in-region managed hosting or self-hosting.

Then customization need, honestly assessed. Most organizations convinced they need a fine-tuned model need better retrieval and better prompts, and the comparative research supports that ordering for document-heavy work. Then latency and availability, where a hard interactive budget favors dedicated or self-hosted serving and a batch-tolerant workload opens the cheapest tiers anywhere on the spectrum.

Then in-house technical capacity, the question most often answered aspirationally. Self-hosting is a standing commitment covering serving infrastructure, evaluation, observability and security patching, measured in fractions of an engineer every month for as long as the service runs. An organization without that capability today will not acquire it by signing a hardware purchase order. Finally, workload lifespan and exit cost. A three-year horizon can amortize fixed capacity, a twelve-month one cannot. Prompts, evaluation harnesses and retrieval infrastructure are largely portable, while fine-tuned weights and provider-specific features are not.

Getting the Assessment Right

Open models are closing on the frontier at a few months' lag, measured differently by task, providers reprice repeatedly within a year, and the spread across providers serving one model is wide enough that today's cheapest option says little about next year's. Any architecture assuming a specific vendor, model version or price point will be wrong within eighteen months.

What survives is an abstraction layer that lets you change where a request is served without rewriting the application, an evaluation harness that tells you whether a cheaper model is good enough for a given task, and a cost model granular enough to attribute spend to individual workloads rather than one line item called "AI."

This is the assessment work we run with clients at Spruce: inventorying actual workloads, measuring real token consumption and latency requirements against the dimensions above, and testing candidate models on the organization's own data before anyone commits capital or signs a multi-year agreement. The output is usually a mixed architecture, built from the workload inventory rather than from a single vendor's default proposal.

Predictability is an organizational capability rather than a procurement outcome. It comes from knowing which workloads you run, what they cost per unit of work, and what would have to change before your current position stops being the right one. Start with an inventory of your own workloads. The vendor conversations go differently once you have one.

Sources

  1. Tech Governance Institute, "Observations About LLM Inference Pricing," September 2025
  2. OpenAI, "Prompt Caching," API Documentation, June 2025
  3. "Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey," arXiv:2603.04445, March 2026
  4. "A Cost-Benefit Analysis of On-Premise Large Language Model Deployment," arXiv:2509.18101, September 2025
  5. International Energy Agency, "Energy and AI: Executive Summary," April 2025
  6. Stanford Human-Centered Artificial Intelligence Institute, "AI Index Report 2025," April 2025
  7. Epoch AI, "Open Models Lag State-of-the-Art Closed Models by 4 Months," May 2026
  8. UK AI Safety Institute, "How Far Behind the Frontier are Leading Open Weight Models on Cyber?" May 2026
  9. "Fine-Tuning vs RAG vs Prompting: Cost-Effectiveness Across Task Types," arXiv:2511.12345, November 2025
  10. IntuitionLabs, "Open-Weight AI Model Licenses: Commercial Use Rules Explained," January 2026
  11. European Union, "Regulation (EU) 2016/679: General Data Protection Regulation, Chapter V: Transfers of Personal Data to Third Countries," May 2018
  12. European Commission, "Regulation (EU) 2024/1689: Artificial Intelligence Act," July 2024
  13. General Services Administration, "FedRAMP AI," August 2025
  14. UK Government Digital Service, "Artificial Intelligence Playbook for the UK Government," February 2025
  15. National Institute of Standards and Technology, "AI Risk Management Framework," January 2024

Want our take on your AI roadmap?

We help leaders turn strategy into production AI systems. Let's talk about what you're building.