Nutanix and Ooma run AI inference in-house to end per-token bills

Nutanix and Ooma each told investors on 26 August 2026 that their own AI now runs in-house on open-weight models, ending the per-token bill.

Nutanix, Inc. (NTNX) and Ooma, Inc. (OOMA) each told investors on 26 August 2026 that their own AI work now runs in-house, on machines they own, using open-weight models, and that they no longer pay an outside model vendor per unit of usage. Both gave cost as the only reason [1][2].


The AI bill used to scale with usage; now it is a one-time machine purchase

Software companies have paid for AI one call at a time. The step where a large model actually does the work is called inference: a question goes in, an answer comes back. When the model belongs to an outside vendor, that step is billed per "token," meaning per unit of text sent in and returned, so the bill grows with use. Nutanix sells software that lets companies build and manage server clusters in their own data centers. Ooma sells cloud phone and communications service to small and mid-sized businesses. The two businesses do not overlap, but both call models repeatedly inside their products and their internal processes.

What changed is the price of renting compute against buying it. Only the strongest frontier models used to be good enough, and those can only be rented per token. Open-weight models, whose parameters are published so anyone can download and deploy them, are now good enough for most of the work, so a company can buy the servers once and run the model itself. Nutanix's CEO said the company started on frontier and cloud models like most of its peers and moved to its own clusters running open-weight models only after both cost and usage rose sharply [1].

The arithmetic holds for similar companies, which is why it can spread. As long as AI usage grows faster than the per-unit price falls, the rented bill keeps getting larger, while purchased capacity is paid for once and spread over several years. It also does not depend on size: Nutanix's quarterly revenue is roughly nine times Ooma's, and both arrived at the same answer.


Both have already switched, and the largest company on this architecture still saw margin fall

Nutanix said the work running on its own clusters with open-weight models no longer carries a per-token charge and becomes a one-time investment it can use to the fullest for many years [1]. Ooma's CEO said the same day that most of the company's AI runs on its own machines, custom-tailored for what it needs, and that this is the only way to get the cost structure as low as it wants; every outside option it examined was more expensive than running the work internally [2]. Ooma's subscription and services gross margin was 72% in the quarter against 71% a year earlier [2].

The larger company on this architecture reads the other way. Zoom (ZM) described the same federated setup on its call one day earlier, moving high volumes onto its own small language model and citing that in support of its long-term 80% gross margin confidence. Even so, non-GAAP gross margin fell to 79.1% from 79.8% a year earlier, which the company attributed to AI use spiking with some of its new products [3]. Taken together the three support the conclusion that the switch is done and the reason is consistent, but not the claim that owning inference lifts gross margin.


The spending is moving from a model vendor's invoice onto the buyer's own balance sheet

The control point is moving downstream. What used to decide where this money went was whose model was strongest; increasingly it is who can run a good-enough model cheaply inside a company's own data center. For the software companies themselves, AI cost moves out of usage-linked operating expense and into capital expenditure and depreciation, which hands management a lever it controls, though it does not stop the cost from growing. Demand also gains a class of buyer: companies assembling small inference clusters in their own facilities, with the servers, the storage and the channel that deploys them all on that line. What gets subtracted is frontier-model interface revenue sold to mid-sized software vendors, but those labs are private, so that leg cannot be verified from disclosure.

The boundary to keep is that none of this is in anyone's reported numbers yet. Across the eight quarters through April 2026, Ooma's quarterly capital expenditure stayed between $1.2 million and $1.7 million with no inflection, and Nutanix's ran between $5.9 million and $34.6 million against roughly $650 million to $720 million of quarterly revenue [4]. Ooma's two closest peers, RingCentral (RNG) and 8x8 (EGHT), said nothing at all about AI cost architecture on their most recent calls [5][6]. Two things can be checked later: whether either company's quarterly capital expenditure leaves those ranges, and whether Ooma's subscription gross margin holds at 72% once the AI feature pack due in its third fiscal quarter is in volume use [2].


Companies exposed to this change

  • Penguin Solutions (PENG): It sells the integrated servers and memory systems companies need to build their own AI compute, which is exactly what the buyer class described above purchases. It disclosed a Tier 1 financial services customer adding memory AI servers for an on-premise "AI factory" whose initial use is inference for code generation on open-weight models, the same category of internal use Nutanix described. Its non-hyperscale AI infrastructure net sales grew 81% year over year and reached 58% of Advanced Computing net sales against 33% a year earlier, though that growth also comes from sovereign AI and neocloud customers and cannot be read directly as software vendors self-hosting inference [7].
  • Insight Enterprises (NSIT): It is a large IT distributor and integrator, and server purchasing and deployment for a company building an inference cluster in its own data center usually passes through it. On its 6 August 2026 call the CEO was asked whether enterprises are moving AI workloads back on-premise for security, latency and cost, and answered that the server business is very strong and that he agreed with that hypothesis. On the same call the CFO immediately restated it conditionally as "If workloads do start significantly repatriating," so this is only an unconfirmed signal from the channel [8].

Sources

[1] Drillr · Nutanix, Inc. (NTNX) · 2026-08-26 · FY2026 Q4 earnings call

"And for all the stuff that runs on our clusters with open-weight models, we no longer have to pay on a per-token basis. That's a one-time investment that we make and and then we can use it to the maximum extent over many years."

[2] Drillr · Ooma, Inc. (OOMA) · 2026-08-26 · FY2027 Q2 earnings call

[3] Drillr · Zoom (ZM) · 2026-08-25 · FY2027 Q2 earnings call

[4] Drillr · Ooma (OOMA), Nutanix (NTNX) · through 2026-04-30 · quarterly financial data, capital expenditure series (Nutanix data through 2026-01)

[5] Drillr · RingCentral (RNG) · 2026-07-23 · Q2 2026 earnings call

[6] Drillr · 8x8 (EGHT) · 2026-08-04 · FY2026 Q1 earnings call

[7] Drillr · Penguin Solutions (PENG) · 2026-07-07 · FY2026 Q3 earnings call

[8] Drillr · Insight Enterprises (NSIT) · 2026-08-06 · Q2 2026 earnings call


This is only meant to help you spot industry changes and companies that may have been overlooked - it is not a stock recommendation.

Want deeper analysis?

Ask drillr anything about EGHT, NSIT, NTNX, OOMA, PENG, RNG, ZM — powered by SEC filings, earnings calls, and real-time data.

Try drillr.ai for free

Drillr can make mistakes. Information only — not investment advice. Learn more