GPUs Are Scarce and Sitting Idle at the Same Time

Most large enterprises now own a GPU cluster that cost several million rand and spends most of its working life doing nothing. Not broken, not being upgraded, just idle, waiting for data that can’t reach it fast enough. A widely cited industry estimate published this year put average enterprise GPU utilisation at around 5%, with the shortfall traced to data pipelines and scheduling rather than to the GPUs themselves. Cloudera and VAST Data‘s newly announced partnership, a joint effort to unify data infrastructure for AI workloads across on-premises and cloud environments, is worth reading as a response to that number rather than as the announcement it’s dressed up as.

The deal itself is less exciting than “AI factory,” the term both companies keep reaching for. Cloudera is bringing its next-generation lakehouse services, the containerised tools that handle data engineering, governance and analytics, and combining them with VAST’s AI Operating System, which unifies storage, database and global namespace functions into a single layer built on NVIDIA’s AI Data Platform reference design. The pitch is consistency: an enterprise should be able to run the same AI pipeline whether its data sits in a private data centre or across several public clouds, without rebuilding the plumbing every time. Cloudera calls the underlying problem GPU starvation, a phrase blunt enough to survive contact with a marketing department mostly intact.

It’s also, as it turns out, the phrase of the year. Data Center Knowledge reported similar findings from a Cast AI study of enterprise Kubernetes clusters, where one infrastructure executive summarised the cause bluntly: this isn’t a hardware problem, it’s a systems problem. Dell has been running a near-identical argument through its own Forbes-sponsored coverage, HPE has certified its own storage platform against the same problem, and WEKA has built a company around it. When four or five vendors who compete with each other converge on the same diagnosis within months of one another, that’s usually a sign the diagnosis is correct rather than a coincidence of marketing calendars.

What changed to make this the year everyone noticed isn’t really the GPUs. It’s the shape of the workload. Training a model is a bounded job: it runs for a defined stretch, consumes a known amount of compute, and finishes. Most enterprise data architecture, built for the batch-analytics era of quarterly reports and overnight jobs, was designed around exactly that kind of boundedness. Inference doesn’t behave that way. It’s continuous, unpredictable in volume, and increasingly agentic, meaning a single user request can trigger a cascade of retrieval, reasoning and follow-up queries that never quite stops. A data platform built to answer a report request once a day has no real chance of feeding that. The GPU isn’t underpowered. It’s waiting on a system that was never asked to run at this pace before.

None of that means the announcement deserves an uncritical read. Cloudera and VAST describe a combined 60 exabytes of customer-managed data across their two installed bases, framed as evidence of scale. It’s worth noticing that this is two separate customer footprints being added together for effect, not a single pool of data that’s now somehow unified. And the phrase both companies use to describe the full stack, from NVIDIA silicon through to Cloudera’s application layer, is “silicon-to-application,” a phrase doing considerably more work than the deployment reality it’s standing in for. Large enterprises rarely experience infrastructure partnerships as seamless. They experience them as two vendors’ roadmaps slowly learning to tolerate each other, with the customer absorbing the friction in between.

This story plays out differently here than it does in Cloudera and VAST’s home market. ITWeb reported on the local version of this gap earlier this year, with HPE South Africa’s own MD noting that many local organisations are collecting data without the architecture in place to use it, even as data centre investment picks up. South African firms chasing AI capability need to fix the data layer before the GPU layer, not after. That’s not simply a cost-saving argument here. Local banks, insurers and telecoms operate under POPIA and sector-specific regulation that makes data residency a genuine constraint rather than a preference, which is exactly what Cloudera and VAST are gesturing at when they talk about private, sovereign AI that behaves consistently whether it’s running in Johannesburg or a public cloud region overseas. For a South African enterprise trying to keep sensitive data onshore while still wanting cloud-scale AI economics, that consistency is closer to the actual requirement than anything novel about the technology underneath it. It’s one of the rare AI infrastructure stories this year where the hybrid deployment pitch isn’t decoration.

What the partnership really marks is a shift in what enterprise AI conversations are about. Two years ago, they were about acquisition, about how many GPUs a company could get its hands on and how quickly. Cloudera and VAST are betting the next phase is about admission, about enterprises finally being willing to say the compute was never the hard part. That’s a less flattering story for the industry to tell about itself, because it means the last two years of AI infrastructure spending built a lot of expensive capacity before anyone worked out how to use it properly. It also means the companies that solve the boring problem, the data plumbing nobody wanted to budget for, end up mattering more than the ones that sold the GPUs in the first place.

Zeen Social Icons