A multi-year deal with the startup Together AI will put Nvidia Blackwell systems on IBM Cloud, and the wager is that enterprises now care more about the cost of running AI than the prestige of the model doing the running.
IBM has decided that the money in artificial intelligence is no longer only in building the cleverest model, but in running it cheaply, and it has put $240m behind that conviction.
The company has signed a multi-year agreement with Together AI, a San Francisco startup, to stand up a large-scale AI inference cluster on IBM Cloud, and the pitch is squarely aimed at enterprises trying to trim their AI bills.
The cluster will run on Nvidia HGX B300 systems, built on the Blackwell architecture that Nvidia markets as tuned specifically for inference, knitted together with the company’s Spectrum-X Ethernet networking, which is the same recipe that a wave of specialists have been chasing as they bet that AI’s real profits lie in cheap inference rather than in ever-larger training runs.
Together AI is an interesting partner to pick. Valued at $8.3bn as of July, its platform lets companies train and run workloads on open-source models, including DeepSeek, MiniMax and Kimi, and it sells itself as a cheaper, more flexible alternative to the closed, proprietary systems that dominate the headlines.
It reportedly serves around 400 trillion tokens a month, a figure that gives some sense of how much raw inference now sloshes through these pipes, and of why a legacy vendor such as IBM would want a share of that traffic running on its own cloud.
Inference, the business of actually answering queries once a model is trained, has quietly become one of the largest drivers of demand for computing capacity, which is why the layer has attracted a scramble of money and talent.
Nebius recently paid $643m to absorb a 20-person team working on inference optimisation, a price that only makes sense if you believe that shaving cost per token is where the margins now live.
The other half of the story is the slow enterprise drift towards open source. Businesses want to cut what they spend on AI, open models are gaining genuine traction inside large organisations.
And there is a security dimension too, since cybersecurity worries about closed models from Anthropic, OpenAI and Meta are cited as one reason some firms would rather run something they can inspect and host themselves.
None of that is a purely technical preference, because for a bank or a hospital the ability to keep sensitive data on infrastructure it controls is often the whole point.
For IBM, hosting cheap, open-source inference is a way to pick a fight it might actually win.
The company was never going to out-muscle Amazon, Microsoft or Google on sheer cloud scale, so positioning IBM Cloud as the place to run open models affordably lets it compete on economics rather than size, and it dovetails neatly with a wider European appetite for infrastructure that is not locked to a single American giant.
That appetite is showing up elsewhere on the map. The push to build inference capacity that enterprises and governments can trust and control has spawned outfits such as TensorX, which raised €8m to build sovereign AI inference for Europe on Nvidia Blackwell, a reminder that the same silicon underpinning IBM’s bet is being wired into a broader argument about who owns the plumbing.
There is a whiff of gold rush about all this, and IBM knows it. The open-versus-closed contest in enterprise AI has often been framed as a debate about capability.
Yet the more telling battle is now about price, and a $240m cluster full of Blackwell chips is IBM’s way of saying it would rather sell the picks and shovels than the promise.
Whether cheap, open inference proves as sticky and as profitable as its backers hope is the wager the whole sector is quietly making, and a $240m cluster is one of its larger stakes yet.
Get the TNW newsletter
Get the most important tech news in your inbox each week.