self hosted ai payback period
Own your stackSeptember 23, 202620 min read

Self-Hosted AI:
The Payback Period Is 13.8M Tokens a Month

By Dan Colta

A desk split down the middle. On the left a graphics card sits on a thick stack of printed GPU server quotations beside a calculator and curled receipts. On the right a tablet shows a self-hosted AI break-even chart where a flat cost line crosses three rising price lines, with a yellow sticky note reading 13.8M tokens propped against it

Self-hosted AI beats a frontier API bill above 13.8 million output tokens a month. That is $166 a month of hardware and EU electricity for an NVIDIA DGX Spark at its $4,699 MSRP (NVIDIA Developer Forums, 25 February 2026), divided by OpenAI's published $12.00 per million output tokens for GPT-5.6 Terra (OpenAI, retrieved 18 September 2026).

Now the figure that undercuts it. DeepInfra serves Llama 3.3 70B at $0.32 per million output tokens (DeepInfra, retrieved 18 September 2026). At NVIDIA's published single-stream rate the same box floors out at $0.76 per million while running flat out for an entire month. Against that line the break-even sits at about 519 million tokens a month, which is 2.4 times what the published rate delivers. Batched serving can close some of that gap. Nothing about the capex can.

Which price you put on the other side decides the answer. Run it against a frontier API and owning wins early. Run it against a serverless open-weight API and it may not win at all.

We pulled page one for self hosted ai on 18 September 2026: eight organic results plus an AI Overview. Not one of them computes a payback period. The AI Overview treats cost as a single line about saving money long term on API tokens, with no figure attached and no utilisation assumption. The arithmetic below is ours, computed on published vendor inputs that are all linked so you can redo it yourself.

This is a spoke in our own-your-stack cluster. It links up to the SaaS replacement playbook, the pillar that sets out the build-versus-buy cost method.

By Dan Colta. NodeSparks is a two-person EU automation studio building owned automations for SME teams. We run our own automation stack on rented EU servers. We do not own a GPU. This arithmetic is the reason.

Key Takeaways

  • An owned NVIDIA DGX Spark costs about $166 a month all in, the conservative end of a $151 to $166 band: $130.53 of capex amortised over 36 months on a $4,699 MSRP (NVIDIA Developer Forums) plus EUR 32.18 of electricity at Eurostat's EU non-household rate of EUR 0.1837 per kWh (Eurostat). Dollar figures assume 1.10 USD per EUR.
  • Against a frontier API the break-even is 13.8 million output tokens a month, on OpenAI's published $12.00 per million for GPT-5.6 Terra (OpenAI). Against Claude Opus 5 at $25.00 per million it drops to 6.6 million (Anthropic).
  • Against a cheap serverless open-weight API the box has to run more than twice as fast as its published single-stream rate before it breaks even. DeepInfra serves Llama 3.3 70B at $0.32 per million output tokens (DeepInfra) against a floor of $0.76 per million at NVIDIA's published 82.74 tokens per second for GPT-OSS-20B, measured at batch size 1 (NVIDIA).
  • Owning a datacentre GPU costs what renting one costs. An L40S in a $26,983 Exxact server amortises to about $1.10 per GPU-hour (Exxact) against Runpod's $1.09 on Secure Cloud and $0.79 on Community (Runpod).
  • Utilisation decides the whole question. Cast AI measured average GPU utilisation of 5% across production Kubernetes clusters (Cast AI, 21 April 2026). At 5% the same box costs $15.26 per million output tokens, more than Claude Sonnet 5 at $10.00 (Anthropic).

What does self-hosted AI cost per month?

About $166, all in. NVIDIA's DGX Spark carries a $4,699 MSRP (NVIDIA Developer Forums, 25 February 2026), which is $130.53 a month amortised over 36 months. Its 240 W power supply rating over a 730-hour month is 175.2 kWh. Eurostat puts EU non-household electricity at EUR 0.1837 per kWh (Eurostat, 8 May 2026).

Line itemInputMonthly
Capex, amortised$4,699 MSRP over 36 months$130.53
Electricity0.240 kW x 730 h = 175.2 kWh at EUR 0.1837/kWhEUR 32.18, about $35.40
Everything elseYour time, backups, model updates, the network it sits onNot priced here
Totalabout $166

Original data: no vendor publishes this monthly total. Every figure above is our arithmetic on the linked inputs.

One assumption in that table runs against our own conclusion. NVIDIA lists the Spark's power supply at 240 Watts and the GB10 chip's thermal design power at 140 W (NVIDIA, retrieved 18 September 2026). We used the 240 W supply rating, a ceiling rather than a measured draw. At the 140 W TDP the total falls to about $151. The higher figure makes owning look worse, which is where our argument already leans. Read every break-even below as the conservative end of a $151 to $166 band. Dollar conversions assume 1.10 USD per EUR.

Electricity is the line that travels worst. It is also the one figure here that is specifically European. Eurostat's spread runs from EUR 0.0748 per kWh in Finland to EUR 0.2552 in Ireland, a 3.4x gap. The same box costs roughly $145 a month in Helsinki and $180 in Dublin. US commercial electricity averaged 14.19 cents per kWh in June 2026 (US EIA), putting it near $155. Geography changes the answer before the model does.

One more input runs the wrong way. The Spark's MSRP went up, from $3,999 to $4,699 in February 2026, which NVIDIA attributed to memory supply constraints. That is a 17.5% increase on a product that had already shipped. Hardware price risk is not one-directional. The standard assumption that silicon always gets cheaper does not survive contact with a memory shortage.

At what volume does owning beat the API bill?

Divide $166 by the published output price. Against GPT-5.6 Terra at $12.00 per million output tokens (OpenAI, retrieved 18 September 2026) the crossover is 13.8 million tokens a month. Against Claude Opus 5 at $25.00 per million (Anthropic) it is 6.6 million. That volume threshold is what a payback period means in this category.

API linePublished output price per 1M tokensBreak-even, output tokens per month
GPT-6 Astra$50.003.3M
Claude Opus 5$25.006.6M
GPT-5.6 Terra$12.0013.8M
Claude Sonnet 5$10.0016.6M
Claude Haiku 4.5$5.0033.2M
Gemini 3.5 Flash-Lite$2.5066.4M
GPT-5.6 Luna$1.20138.3M
Together Llama 3.3 70B$1.04159.5M
DeepInfra Llama 3.3 70B$0.32about 519M, above the box's 217.4M published ceiling

Prices retrieved 18 September 2026 from OpenAI, Anthropic, Google, Together AI and DeepInfra. Break-even column is our arithmetic on the $165.93 monthly total, shown throughout as about $166. These rows compare cost, not capability. The box is producing GPT-OSS-20B tokens. The API rows are frontier models. Substitute only where your workload actually tolerates the smaller model. The same volume-before-capex order applies to the build decisions in the ops automation playbook: measure the metered bill first, then price the thing that replaces it.

Monthly cost against monthly output volume: owned box versus three API lines $0 $200 $400 $600 0 50M 100M 150M 200M GPT-5.6 Terra, $12.00/M Gemini 3.5 Flash-Lite, $2.50/M DeepInfra Llama 3.3 70B, $0.32/M Owned DGX Spark, $166/month DGX Spark published ceiling, 217.4M/month 13.8M 66.4M Output tokens per month The DeepInfra line never reaches $166 inside the box's published single-stream capacity. Lines plotted to scale. Our arithmetic on published prices: OpenAI, Google, DeepInfra, NVIDIA DGX Spark MSRP and throughput, Eurostat EU electricity. Retrieved 18 September 2026.

One measurement note belongs here, not in a footnote. The 82.74 tokens per second is a single-stream rate (NVIDIA, 24 October 2025). Batched serving with vLLM would go considerably higher, pushing the ceiling right and moving every crossover in favour of owning. Read 217.4 million as a floor.

There is also a real hole in the evidence here. No primary benchmark exists for an L40S or an RTX 6000 Ada running a named open-weight model at a named precision. MLPerf publishes per-accelerator results on rack-scale GB200 and GB300 systems (NVIDIA, 9 September 2025). NVIDIA publishes DGX Spark figures. The tier an SME would actually buy sits between the two, unmeasured by anyone with a name on it.

Why does a cheap open-weight API beat a self-hosted LLM on price?

Because the box has a floor and the API sits underneath it. NVIDIA publishes 82.74 tokens per second for GPT-OSS-20B at batch size 1 (NVIDIA, 24 October 2025). Held for a 730-hour month that is 217.4 million tokens at $0.76 per million. DeepInfra serves the far larger Llama 3.3 70B at $0.32 (DeepInfra).

Cost per million output tokens by route, log scale $0.10 $1 $10 $100 GPT-6 Astra $50.00 Claude Opus 5 $25.00 Owned box, 5% duty cycle $15.26 GPT-5.6 Terra $12.00 Claude Sonnet 5 $10.00 Owned box, 10% duty cycle $7.63 Claude Haiku 4.5 $5.00 Gemini 3.5 Flash-Lite $2.50 GPT-5.6 Luna $1.20 Together Llama 3.3 70B $1.04 Owned box, 100% duty cycle $0.76 Fireworks gpt-oss-120B $0.60 DeepInfra Llama 3.3 70B $0.32 Dark bars are our derived figures for the owned box. Light bars are published list prices. Duty cycle is the share of the month the box spends generating tokens. It is the only variable that moves the dark bars. Sources: OpenAI, Anthropic, Google, Together AI, Fireworks AI, DeepInfra pricing pages; NVIDIA DGX Spark throughput; Cast AI utilisation. Retrieved 18 September 2026.

Utilisation is the only lever on those dark bars. It is also the one nobody models. Cast AI measured GPU utilisation averaging 5% across the clusters in its analysis of tens of thousands of Kubernetes workloads (Cast AI, 21 April 2026). That is measured production behaviour rather than a survey answer. At 5% the box produces 10.9 million tokens a month and costs $15.26 per million, which is more expensive per token than Claude Sonnet 5.

Fireworks makes the same point on one vendor's own price list. Its serverless gpt-oss-120B is $0.60 per million output tokens (Fireworks AI), while a dedicated H100 or H200 from the same company is $8.00 an hour (Fireworks AI, retrieved 18 September 2026). Dedicated capacity is a bet on keeping it busy. Serverless is a bet that you will not.

This is the same pattern as paying SaaS margin for a thin layer over somebody else's API, covered in the AI wrapper SaaS trap. The difference is the direction. Here you are the one adding a fixed cost on top of a metered one.

Is owning a datacentre GPU cheaper than renting one?

No. An NVIDIA L40S in an Exxact TensorEX 2U server is $21,054 for the base plus $5,929 for the card, $26,983 in capex (Exxact, retrieved 18 September 2026). Over 36 months, which is 26,280 hours, that is $1.027 per GPU-hour before a single watt. Runpod rents the identical card at $1.09 on Secure Cloud (Runpod).

Add the electricity and the two numbers converge further. The L40S board draws up to 350 W (NVIDIA), which at the Eurostat EU rate is EUR 0.0643 an hour, roughly $0.071. Owned lands at about $1.10 per GPU-hour against a rented $1.09.

Cost per GPU-hour for one NVIDIA L40S: owning versus renting Own: capex over 36 months $1.027 Own: plus GPU electricity $1.097 Rent: Runpod Secure Cloud $1.09 Rent: Runpod Community Cloud $0.79 $0 $0.50 $1.00 $1.50 $2.00 $2.50 Owned bars are our arithmetic. They exclude host power, cooling, networking, rack space and your time. Exxact TensorEX 2U configurator and NVIDIA L40S spec, Eurostat EU non-household electricity, Runpod pricing. Retrieved 18 September 2026.

The owned bars are understated, which makes the comparison worse and not better. They exclude host power, cooling, networking, rack space and the hours somebody spends patching the thing. They also assume the card stays useful for all 36 months.

The rest of the rental market prices the same way at higher tiers. Lambda lists an H100 SXM at $4.29 per GPU-hour on a single-GPU instance, plus applicable VAT for an EU buyer (Lambda, retrieved 18 September 2026). DeepInfra lists an H100 at $2.20 (DeepInfra). The spread across vendors is wide enough that shopping the rental market beats optimising a purchase.

Prosumer hardware is where owning still has a case, because a $4,699 box has no $21,054 chassis under it. Datacentre-class self-hosting against a neocloud is very hard to defend on cost alone.

What happens to the payback when the API price keeps falling?

The denominator moves underneath you. Stanford HAI measured the cost of querying a model scoring GPT-3.5's 64.8 on MMLU falling from $20.00 per million tokens in November 2022 to $0.07 by October 2024, a more than 280-fold reduction (Stanford HAI, 2025 AI Index Report). A 36-month amortisation is a 36-month bet against that curve.

Epoch AI measured the same direction across an overlapping but broader benchmark set and found decline rates ranging from 9x to 900x per year depending on the capability milestone, with GPT-4-level performance on PhD-level science questions falling 40x a year (Epoch AI, 12 March 2025). Epoch adds its own limitation, which is worth repeating rather than trimming: the fastest drops in that range occurred in the most recent year, which makes it less clear that they will persist.

The hardware side moves too. MLCommons reports that median per-accelerator Server-scenario performance improved 5.58x over six rounds. Round v6.1 drew submissions from a record 30 organisations (MLCommons, 17 September 2026). The box you buy today competes for three years against parts that get materially faster every six months.

Stack the two and the risk is clear. You are amortising a depreciating asset on a fixed schedule against a metered alternative whose unit price has fallen by orders of magnitude on every measurement of it. None of that argues against owning at high steady volume. It argues against a long amortisation schedule and against buying capacity for a workload you have not measured.

If your real question is which parts of the workflow justify custom infrastructure in the first place, the pricing-led version of that decision is worked through in what an AI SDR costs to build versus buy.

Does data residency change the answer?

For one major vendor it does. Anthropic's first-party Claude API exposes exactly two inference geos, "us" and "global", with "us" as the only available workspace geo (Anthropic, retrieved 18 September 2026). An EU operator cannot pin Claude inference or at-rest storage to Europe on that API at all. It is the clearest case among the four routes we checked.

RouteEU inference available?What the route gives you
Anthropic first-party Claude APINoOnly "us" and "global" inference geos; "us" is the only workspace geo
Claude via Bedrock, Vertex or FoundryYesRegion is set by the endpoint URL or inference profile
OpenAI APIYesEuropean data residency with zero data retention, on newly created Projects only
Owned box in your own rackYesNothing leaves the building

The categorical version of this argument is false. OpenAI has offered European data residency with zero data retention since February 2025, enabled by creating a new Project and selecting Europe as the region (OpenAI, 5 February 2025). Under the EU-US Data Privacy Framework adequacy decision, adopted on 10 July 2023, personal data can flow freely from the EU to companies in the United States that participate in the framework (European Commission). GDPR does not force anyone to buy a GPU.

Vendor geography does carry a price tag. It belongs in the arithmetic, not in a policy discussion. Anthropic prices US-only inference at 1.1x the standard rate on Claude 4.6 and later (Anthropic). That moves Sonnet 5 output from $10.00 to $11.00 per million and shifts the owned-box break-even from 16.6 million to 15.1 million tokens a month. Real, small, easy to model.

What stops EU companies in the first place is a different question worth holding onto. Among EU enterprises that considered AI and adopted none of it, the most common blocker was a lack of relevant expertise at 70.89%, ahead of legal uncertainty at 52.52% (Eurostat, 2025 data). Expertise is the scarce input. An owned box consumes more of it than an API key does.

One scope note before you plan around sovereignty. The EU AI Act's obligations for general-purpose AI models became applicable on 2 August 2025, with high-risk system rules from 2 December 2027 (European Commission). Running an open-weight model yourself can change which role you occupy under those rules. Take advice on that, not a blog post.

When running an LLM locally is the right call

When you can keep the box busy. The break-even against GPT-5.6 Terra is 13.8 million output tokens a month (OpenAI). That is 6.3% of a DGX Spark's published ceiling. Cast AI measured production GPU utilisation averaging 5% (Cast AI). Break-even and observed reality are nearly the same number, which is why this is easy to get wrong.

Four signals that owning works:

  • The workload is batched and steady. Document processing, classification, enrichment and summarisation runs that you schedule and do not wait on. Batched serving also raises the throughput ceiling above NVIDIA's single-stream figure.
  • You have fixed on a model. An owned box pays back against a model you intend to keep running. It pays back against nothing if you plan to switch every quarter.
  • The residency constraint is specific and real. Not "GDPR" in the abstract. A named vendor that cannot serve your region, a contractual clause or a client who audits where inference happens.
  • Somebody already patches servers. A GPU box is one more machine to a team that runs machines. It is a new job to a team that does not.

Four signals to keep renting:

  • You have not measured your token volume. If you cannot state monthly output tokens you cannot compute a break-even. The first step is one month of API billing used as measurement.
  • A cheap open-weight API covers the workload. At $0.32 per million from DeepInfra or $0.60 from Fireworks, the box cannot win at any volume it produces at the published single-stream rate.
  • You want datacentre-class hardware. At the L40S tier owning costs what renting costs, with all the optionality removed.
  • Your duty cycle will sit in single digits. At 5% the box is more expensive per token than Claude Sonnet 5.

Between those two poles sits the middle option: rent dedicated capacity by the hour and keep the open weights. The no-code versus custom code trade-off has the same middle tier. It is usually the right first move.

FAQ

Six questions are answered below: what self-hosted AI costs per month, the volume at which it beats an API bill, whether it beats a serverless open-weight API, whether buying a datacentre GPU beats renting one, whether GDPR requires self-hosting and when running a model locally is right. For the wider build-versus-buy method, start with the SaaS replacement playbook.

The bottom line

The payback period on self-hosted AI is not a duration. It is a monthly volume that moves entirely with the API price you put on the other side. Against GPT-5.6 Terra the box wins above 13.8 million output tokens a month. Against DeepInfra's Llama 3.3 70B at $0.32 the crossover is about 519 million, which is 2.4 times what the published single-stream rate delivers in a month. Only batching closes that gap.

At the datacentre tier the arithmetic is worse. An owned L40S amortises to about $1.10 per GPU-hour and Runpod rents the same card for $1.09. You would be paying $26,983 up front for the privilege of matching a price you could have had on a credit card.

The same pattern shows up one layer up the stack. Self-hosting a social media scheduler moves the platform bill rather than removing it. Owning inference does the same thing to the API bill. It converts a metered cost into a fixed one, which is a good trade at high steady volume and a bad one at 5% utilisation.

So the rule is small. Measure one month of output tokens first. If you are above roughly 14 million a month against a frontier model, with a model you intend to keep and somebody who patches servers, buy the prosumer box. Otherwise rent and spend the $4,699 on the workflow instead. The ranking of which workflows are worth that money is in the Ops Automation Playbook.


Sources (retrieved 2026-09-18):

Questions

Frequently asked questions

What does self-hosted AI actually cost per month?

About $166 a month for the realistic SME purchase, an NVIDIA DGX Spark. The MSRP is $4,699, which is $130.53 a month amortised over 36 months. Its 240 W power supply rating over 730 hours is 175.2 kWh. NVIDIA lists the GB10 chip's thermal design power lower still, at 140 W, which would put the monthly total nearer $151. At Eurostat's EU non-household rate of EUR 0.1837 per kWh that is EUR 32.18 of electricity. Dollar totals here assume 1.10 USD per EUR. Country rates move the electricity line hard: the same box lands near $145 a month in Finland and near $180 in Ireland.

At what volume does self-hosted AI beat an API bill?

Divide the box's $166 monthly cost by the published output price and you get the volume. Against GPT-5.6 Terra at $12.00 per million output tokens the break-even is 13.8 million tokens a month. Against Claude Opus 5 at $25.00 it is 6.6 million. Against Gemini 3.5 Flash-Lite at $2.50 it is 66.4 million. Against DeepInfra's Llama 3.3 70B at $0.32 it is about 519 million, which is more than double the 217.4 million tokens the hardware produces at the published single-stream rate.

Is a self-hosted LLM cheaper than a serverless open-weight API?

Usually not. On NVIDIA's published 82.74 tokens per second for GPT-OSS-20B, a DGX Spark running flat out for a full month produces 217.4 million tokens, which puts its floor at $0.76 per million. DeepInfra serves the much larger Llama 3.3 70B at $0.32 per million output tokens with no capex and no obligation to keep anything busy. Fireworks serves gpt-oss-120B at $0.60. At that published single-stream rate no volume the box can produce closes the gap. Batched serving is the variable that moves it.

Is it cheaper to buy a datacentre GPU or rent one?

They cost about the same. Renting carries no capex at all. An NVIDIA L40S configured into an Exxact TensorEX 2U server is $21,054 for the base plus $5,929 for the card, a total of $26,983. Over 36 months that is 26,280 hours, which works out to $1.027 per GPU-hour of amortisation. Add GPU-only power at the card's 350 W spec and EU electricity rates and you reach about $1.10. Runpod rents the identical card at $1.09 an hour on Secure Cloud and $0.79 on Community.

Does GDPR require you to self-host AI?

No. The EU-US Data Privacy Framework adequacy decision, adopted on 10 July 2023, lets personal data flow to participating US companies without extra safeguards. OpenAI has offered European data residency with zero data retention since February 2025, configurable only on newly created Projects. The constraint is vendor-specific rather than categorical. Anthropic's first-party Claude API exposes only 'us' and 'global' inference geos with 'us' as the only workspace geo, which means pinning Claude inference to the EU requires routing through Bedrock, Vertex or Foundry.

When is running an LLM locally the right decision?

When you can keep the box busy at a duty cycle above roughly 6% against a frontier model, on a workload you control. The break-even against GPT-5.6 Terra is 13.8 million output tokens a month, which is 6.3% of a DGX Spark's 217.4 million published ceiling. Cast AI measured GPU utilisation averaging 5% across the clusters in its analysis of tens of thousands of Kubernetes workloads. At 5% the same box costs $15.26 per million tokens, more than Claude Sonnet 5 at $10.00.

Still paying for tools you could own?
The call is free and the written scope with its fixed number is free too. You keep the scope whether you hire us or not. Nothing is paid before work starts.

Let's start with a real conversation.We’re ready when you are.