Together AI Secures $800 Million Series C to Scale Open-Source AI Infrastructure
Together AI closed an $800 million Series C funding round on July 1, 2026, at an $8.3 billion post-money valuation. The round was led by Aramco Ventures, with participation from Vista Equity Partners, General Catalyst, Emergence Capital, Nvidia, March Capital, Pegatron, and SentinelOne's S Ventures. The company, which reports over $1.15 billion in annual bookings, will use the capital to scale its computing capacity and expand its open-source model hosting platform.
Together AI has raised $800 million in a Series C funding round led by Aramco Ventures, valuing the company at $8.3 billion post-money. The financing, closed on July 1, 2026, also drew capital from Vista Equity Partners, General Catalyst, Emergence Capital, Nvidia, Salesforce Ventures, March Capital, Pegatron, SE Ventures, and SentinelOne’s venture arm, S Ventures. The company plans to deploy the capital to scale its hardware infrastructure and expand its specialized cloud platform designed for hosting open-source artificial intelligence models.
Moving Enterprise AI From Closed APIs to Open-Source Infrastructure
The funding marks a significant shift in how enterprises deploy artificial intelligence workloads. Historically, companies integrated generative AI into their products by connecting to proprietary, closed models via cloud-hosted application programming interfaces (APIs). Under this framework, developers pay a fee per request, leaving them reliant on the pricing, uptime, and data privacy policies of a single external vendor.
The alternative is using open-source models, which distribute their underlying weights publicly. This model allows developers to inspect, modify, and host the software themselves.
However, running these massive machine learning architectures requires immense raw computing power, specifically specialized graphics processing units (GPUs). Many enterprises lack the physical hardware, network engineering, and software optimization teams necessary to deploy and maintain these networks of chips at scale.
Together AI aims to solve this operational bottleneck by offering an optimized middle-tier platform. Alongside building and releasing its own open-source foundational models—such as the RedPajama dataset and models, StripedHyena, and the biological foundation model Evo—the company operates primarily as a specialized cloud provider.
The platform rents GPU clusters and runs serverless API endpoints optimized specifically for running leading open-source architectures like DeepSeek, Qwen, and Llama. This approach allows companies to transition their software workloads away from closed systems without the overhead of purchasing and managing their own physical data centers.
Optimizing GPU Utilization to Lower Run Costs
To lower the cost of running these open-source models, the platform relies heavily on software-level optimization. In artificial intelligence operations, running a finished model to answer queries—a process known as inference—can become highly inefficient if the underlying software does not manage how data flows through the physical silicon of the GPU.
The platform bundles its raw GPU compute with a proprietary inference optimization engine. According to the company, this specialized software stack accelerates model performance, which helps maximize the number of queries a single chip can handle per second.
For teams running high-volume production systems, the company provides dedicated endpoints running on high-end hardware like H100 and H200 GPUs. It also offers discounted rates for non-real-time batch inference workloads, such as document processing pipelines or large-scale data classification jobs that can run overnight rather than requiring instant responses.
Efficiency research remains a core pillar of the company’s platform development. Tri Dao, Together AI’s Chief Scientist, co-authored and developed FlashAttention-3 alongside researchers from Princeton University, Stanford University, Meta, Nvidia, and Colfax Research. The open-source software library optimizes memory access on GPU hardware, reducing memory bottlenecks so physical chips can run closer to their theoretical processing limits.
Booking Growth and Infrastructure Scaling Plans
According to the company, its annual bookings have crossed $1.15 billion as of last quarter, driven by a sharp rise in the commercial adoption of open-source models. The platform’s paying customer base currently includes AI-powered coding assistant Cursor, enterprise software engineering platform Cognition, and customer service automation firm Decagon.
The transition to optimized open-source hosting has demonstrated substantial cost efficiencies for early adopters. For example, Decagon co-founder Ashwin Sreenivas reported that migrating their customer service workloads to Together AI resulted in a six-fold reduction in operating costs compared to prior closed-model pricing.
With the new $800 million capital infusion, Together AI intends to purchase and rent additional compute chips to scale its physical infrastructure. The company has stated that it expects its global computing capacity and physical footprint to expand roughly 50-fold over the next five years to keep pace with enterprise demand.
- #Startups
Author
Raj M
Contributor
AI Systems Architect is a seasoned technology leader with over 15 years of experience in the IT industry working with Fortune 500 companies. With a solid foundation in multi-agent systems, open-source LLM infrastructure, and enterprise deployment, he excels at building scalable production-grade AI platforms.