GPU-as-a-service: Should enterprise IT rent or own AI compute?

GPU-as-a-service lets enterprises skip high CapEx but vendor lock-in is a concern

Cloud native security concept image showing cloud symbol pictured on top of a digitized cube representing a circuit board and GPU.
(Image credit: Getty Images)

GPUs have become the hardware stars of the AI era, allowing massive parallel processing of data that shortens lead times from weeks or months on a CPU to days. For organizations that want to take advantage of this technology for their own AI inference, however, there’s a major hurdle: cost.

According to an article by Server Parts, a company specializing in selling refurbished IT hardware, 32-128 GPUs are needed for fine-tuning and 4-32 for inference. Purchasing this hardware can come at a significant cost, however: NVIDIA H100 GPUs, for example, cost US $250,000 to $400,000 per unit. Additional hardware such as storage and networking cards add to the hefty procurement bill.

Vasily Mazin, CRO and co-founder of Mind Simulation Lab, a Maths PhD building AGI architecture and LLM alternatives, told ITPro: “Modern enterprise data centers are built to handle 10-20 kW per rack. The latest generation of AI hardware (like Nvidia’s GB300 architecture) demands up to 150 kW per rack, requiring direct-to-chip liquid cooling and massive power substations”.

According to Mazin, existing enterprise server rooms aren’t built to support high-end GPUs for AI training and operations. “Enterprises literally can’t plug these machines into their existing server rooms without melting the infrastructure.” Even if IT leaders choose to allocate high CapEx budgets, electricity and cooling costs don’t justify RoI. Building and running a data center becomes a separate business branch for such enterprises.

Latest Videos FromIT Pro

AI-native companies, including Anthropic and OpenAI, don’t own GPUs but rent them through multi-billion-dollar partnerships with hyperscalers or compute providers. Gigawatts are locked in office towers for a decade, or perhaps two, with continuous demand and use. The case is different for enterprise AI adopters because they don’t use GPUs once the job is done.

That’s where GPU as a service (GPUaaS), a cloud computing model, can be a reliable choice for AI-adopting companies. GPUaaS is a rental service for enterprises to access GPUs via the internet. The rental benefit eliminates the need to allocate capital and maintain physical infrastructure.

Enterprise subscribers can gain access to GPUs on demand in exchange for a fee. On-demand GPUaaS allows enterprise users to pay only for used resources. Enterprises can pay as low as US $2 or $10 per hour for the same NVIDIA H100 GPU that costs tens-of-thousands of dollars in procurement.

Kevin O’Connor, founder of an AI security consultancy, TKOResearch, and former technical director at the NSA, tells ITPro: “There’s been nearly a double in price for the same flagship tier card [GPU], partly due to supply chain but also demand with the explosion of AI”.

Tech products, including GPUs, tend to exhibit high launch prices. As a general rule, hardware prices tend to decrease with time–due to newer launches, but that’s not the case for GPUs. O’Connor shared that the decade-long shortage of consumer GPUs has increased current pricing for 2026 beyond the launch price.

Supporting GPUaaS, O’Connor said: “There have been some really small gaps between certain recent card [GPU] generation releases or even revisions on cards that have made buying less appealing.

Pricing models to consider

GPUaaS isn’t just about CapEx avoidance, it’s a different consumption model. O’Connor asks: “Why manage the infrastructure required to run, manage, and make use of [the GPU] when you can essentially use it with the same cost modeling as a SaaS product?”

That’s why 'GPU as hardware' transitioned into a cloud computing ‘as a service’ model. “The cost of the [GPU] is only fractional compared to what it costs in total hardware, operations, and maintenance - not to mention energy prices”. The IEA reports that servers, whether equipped with GPUs or CPUs, account for 60% of electricity consumption in data centers, followed by cooling.

Instead of allocating a single physical GPU to a customer, the service provider runs software for many of them to share the same hardware. Each gets a virtual slice of the GPU. Hyperscalers, data center operators, hardware providers, and a few emerging startups sell GPU-as-a-service in three pricing models: on-demand, reserved, and spot-pricing.

  • On-demand GPUaaS: Enterprises can access GPUs on demand, paying only for used resources, whether for seconds or hours. There is no commitment, only flexibility. Service providers tend to quote the highest price in the demand-based pricing model. Experimental, one-time, or less frequent projects run on on-demand GPUaaS. Retail, media, and financial services tend to choose on-demand GPU-as-a-service.
  • Reserved GPUaaS: Enterprises can rent GPUs at a discounted rate in exchange for a time-bound commitment. The reserved model offers low cost but requires repeated expenditures. Steady, 24/7 (regular), predictable, and sustained machine learning workloads can execute inference for a pre-defined period.

    Healthcare and clinical trials often opt for a reserved pricing contract. The downside of this model is the lack of flexibility – enterprises still have to pay for the projects continuously even if they stop and restart stop and restart.
  • Spot-priced GPUaaS: Enterprises can access a provider's spare/unused center capacity at a steep discount, which might be as low as 70%. If the demand spikes, the service provider can interrupt or revoke GPU access with or without prior notice. Spot pricing can support fault-tolerant, interruptible, low-priority, and batch workloads.
  • Dedicated/bare metal GPUaaS: Enterprises can physically rent an entire GPU infrastructure from the service provider. As the GPU virtualization layer is absent, bare metal is the priciest but safest GPUaaS model.

    Dedicated pricing models can support compliance-heavy, performance-sensitive, and critical workloads for industries such as defence, aerospace, banking, pharmaceutical, and government. No other tenants on a virtual machine also means greater security.

Untold story of GPUaaS

The future of GPUaaS is a function of enterprise choice, AI workload goals, confidentiality, and budget.

From a buyer’s perspective, IT leaders should be able to choose GPUaaS as a managed 360-degree service. In addition to GPU access, enterprises should be on the lookout for vendor lock-in with SLAs governing high-performance storage, fast networking, security updates, and orchestration. A simple GPU login provided for GPUaaS pricing is a GPU reseller in disguise.

Mazin said: “Right now, companies are locking themselves into multi-year GPUaaS contracts at peak market prices just to support brute-force AI.“ The warning comes after longer lead times. “But silicon depreciates rapidly and, more importantly, software paradigms shift,” adds Mazin.

Much of the GPUaaS market is occupied by “neoclouds” – specialized GPU compute providers. While cloud providers and hyperscalers package GPUs with dozens of other unrelated services in a bundle, neoclouds are purely built to serve customers looking for different generations in GPU-as-a-service.

The optimal solution is to classify the AI workflow and prioritize. On-premises GPUs make sense for regular workloads with data sovereignty and compliance-heavy requirements. Enterprises can choose the best of both worlds, a hybrid strategy – buy GPU hardware and opt for GPUaaS from trusted providers when the demand spikes.

Venus Kohli
Freelance writer

Venus is a freelance technology writer specializing in IT, quantum physics, electronics, and among other technical fields. She holds a degree in Electronics and Telecommunications Engineering from Mumbai University, India.

With years of experience in writing for global media brands and IT companies, she enjoys translating complex content into engaging stories. When she’s not writing about the latest IT trends, Venus can be found tracking enterprise trends or the newest processor in town.