Together AI 携手 IBM Cloud 与 NVIDIA 扩展企业级推理算力
Expanding our enterprise inference capacity with IBM Cloud and NVIDIA
Together AI 与 IBM、NVIDIA 合作,在 IBM Cloud 上部署大型 NVIDIA B300 GPU 集群,并采用 NVIDIA Spectrum-X 以太网组网,用于扩展企业级 AI 推理。
We’re excited to share that Together AI is working with IBM and NVIDIA to scale enterprise-grade AI inference, starting with a large cluster of NVIDIA B300 GPUs on IBM Cloud backed by NVIDIA Spectrum-X Ethernet networking. It's the first dedicated, large-scale inference cluster of its kind on IBM Cloud, and we're the first customer running on it.
What's actually happening
- The infrastructure: A dedicated NVIDIA B300 GPU cluster, purpose-built for inference, running on IBM Cloud.
- The model: We operate the inference layer, IBM provides the cloud, NVIDIA delivers the silicon and networking.
- The trajectory: We're planning for the future as token demand goes parabolic
Why it matters
We started Together AI because we believe the future of AI shouldn't be owned by a handful of closed labs. That bet is paying off faster than even we expected: Hundreds of trillions of tokens served per month to over a million developers, and demand keeps climbing.
Enterprises and AI-native companies choose open models for two simple reasons: their sovereign data stays theirs and they get frontier-level performance at a fraction of closed-model cost. The bigger the usage, the better that math looks.
With this collaboration:
- NVIDIA brings the silicon — B300 GPUs and Spectrum-X Ethernet networking engineered for high-throughput inference.
- IBM brings decades of running mission-critical infrastructure for the world's largest enterprises.
- Together AI brings the inference platform — the fastest, most efficient way to run open models in production, hardened by some of the most demanding AI workloads on the planet.
Put those together and the result is simple: enterprise-grade inference at massive scale – more production grade tokens, with the reliability, security and guardrails enterprises have come to expect.
Open-source AI has to run everywhere, at scale, as fast and reliably as anything closed. We have spent the last few years making sure it can and this collaboration is a big step toward exactly that.
来源:Together AI 研究与产品博客 · together.ai