IBM and Together AI’s $240 Million Deal: What It Means for Enterprise AI Inference

7498afde f6e4 424e aafc a94373772cc0

IBM and Together AI’s $240 Million Deal: What It Means for Enterprise AI Inference

IBM and Together AI have announced a multiyear, $240 million agreement to expand open-source AI inference on IBM Cloud. The planned infrastructure combines NVIDIA HGX B300 systems, Spectrum-X Ethernet networking, and Together AI’s inference platform. IBM expects the new cluster to be available in the first quarter of 2027.

For businesses exploring AI, the useful question is simple: how will this help them run models reliably and affordably once a proof of concept becomes a real product?

What is AI inference?

Training teaches a model patterns. Inference is what happens when an application uses a trained model to answer a question, summarize a document, analyze an image, or take another action.

Every customer support response, document extraction, and AI agent step may require one or more inference calls. As usage grows, response speed, capacity, and cost per request become practical concerns. That is the problem this partnership aims to address.

We divided the work into six phases. Each phase had defined deliverables, and each had to be complete before the next could build on it.

What have IBM and Together AI announced?

IBM plans to deploy a large cluster of NVIDIA HGX B300 systems on IBM Cloud. Together AI intends to use that capacity to serve open-source model inference for its customers. NVIDIA Spectrum-X networking is part of the planned infrastructure.

This is an infrastructure and service expansion, rather than the release of a new AI model. The expected availability is Q1 2027, so organizations should not assume the specific new cluster is available today. Together AI already provides AI services; the announcement concerns additional capacity hosted on IBM Cloud.

 

Why do open models matter to businesses?

Open models can give development teams more choice over the model they use for each task. A business may test several models against its requirements for accuracy, response time, cost, and deployment controls. In some cases, it can customize or fine-tune a model for a narrower workflow.

That flexibility does not guarantee a lower bill or a better result. Costs depend on the model, workload, usage pattern, hosting arrangement, and engineering effort. A careful evaluation should compare the total cost of operating the application, rather than token prices alone.

Where could this be useful?

Consider a company building an AI assistant for customer service. A pilot with a small number of conversations may work well, but production introduces harder questions:

  • Can the service handle spikes in demand?
  • How quickly does it respond during busy periods?
  • What does each resolved conversation cost?
  • Which customer data is sent to the model, and how is it governed?
  • Can the team change models without rebuilding the whole application?

The IBM and Together AI arrangement is aimed at the computing side of this challenge: providing large-scale capacity for serving models. The application team still needs to design retrieval, access controls, monitoring, human escalation, and quality checks around the model.

What should organizations do now?

If your team is planning a production AI workload, this announcement is a reason to evaluate options. You do not need to wait until 2027 to define and test a useful application.

Start with a measurable use case. Test representative data and define acceptable answer quality, response time, and cost. Then compare models and providers under the same conditions. For sensitive workflows, check the proposed service terms, data handling, location, and security controls before sending real customer information.

When the new cluster becomes available, teams can assess its actual pricing, supported models, performance, regions, and contractual terms. Those details matter more than infrastructure specifications by themselves.

The Broaddigix view

The market is moving from AI demonstrations to applications that must work every day. That shift makes inference infrastructure an important part of the architecture, alongside the business process, data, and safeguards that make an AI system useful.

At Broaddigix, we help businesses turn AI ideas into practical workflows and integrations. If you are considering an AI assistant, document automation, or a custom AI application, contact Broaddigix to discuss the use case and the architecture that fits it.