Capital Signal
IBM and Together AI Sign $240M Inference Deal
IBM confirmed a $240 million multi-year agreement with Together AI to build an open-model inference cluster on IBM Cloud using NVIDIA HGX B300 systems.
IBM confirmed a $240 million multi-year agreement with Together AI to build an AI inference cluster on IBM Cloud using NVIDIA HGX B300 systems and Spectrum-X Ethernet for open-source models. Expected to become available in the first quarter of 2027, the agreement turns future open-model serving demand into a specific infrastructure build plan.
A long-term agreement locks in an infrastructure plan
IBM says the agreement is worth $240 million over multiple years and covers a large-scale AI inference cluster on IBM Cloud. Together AI plans to use it for open-source model inference. Compared with short-term consumption purchasing, a multi-year agreement binds expected demand to a specific cloud platform, network architecture and GPU system. It is therefore a concrete shift from resource consumption to infrastructure commitment. The cluster is expected in the first quarter of 2027 rather than already operating.
Inference capacity requires compute and networking
IBM disclosed NVIDIA HGX B300 systems and Spectrum-X Ethernet as part of the planned configuration. The former supplies accelerated compute, while the latter connects the cluster. Commercial open-model inference must manage model scale, concurrent requests, latency and service cost together. The agreement is therefore more than a GPU purchase: it commits a combined compute-and-network capacity stack for sustained serving. Together AI is securing planned capacity dedicated to its inference business.
IBM Cloud seeks an open-model serving role
The partnership gives IBM Cloud a named workload in open-model inference and connects Together AI’s model-serving business to IBM’s enterprise cloud infrastructure. Editorially, the contract value alone does not define the market, but it signals more explicit long-term spending on open-model inference. The strongest countercase is that IBM has not disclosed cluster size, utilization, initial customers or unit economics, so durable demand remains to be tested after launch.
What to watch next
The next evidence to watch is whether IBM and Together AI disclose GPU counts, deployment milestones and service regions; whether Together AI identifies models, customers or inference volumes supported by the cluster; and whether launch is followed by additional procurement or similar long-term cloud agreements. Those observable developments would strengthen or weaken the case that open-model inference is moving toward dedicated contracted capacity.