We announced that we are partnering with Nvidia to support rack-scale deployments of our 3D-DRAM enabled Raptor accelerators using the NVLink scale-up fabric on identical infrastructure to Nvidia’s Vera Rubin rack-scale systems.
This allows us to accelerate time to market by combining the speed and ultra-low latency of our stacked 3D DRAM-enabled accelerator, Raptor, with Nvidia’s market-leading scale-up fabric. NVLink allows us to massively scale up the number of Raptor accelerators in a rack to support ultra low latency inference for both Raptor-exclusive workloads and integrate with heterogeneous configurations, offering customers maximum optionality for their specific AI inference needs.
Rack-scale Raptor racks make it possible to create snappy user experiences at scale by enabling colossal frontier-level models with the same low-latency performance benefits of the memory bandwidth our 3D-DRAM enables. More importantly, our rack-scale Raptor architecture is a full plug-and-play solution that instantly brings the benefits of Raptor to any workload.
A plug-and-play solution with Raptor fits alongside one of our core guiding principles: meeting companies in their existing infrastructure with seamless implementation of d-Matrix technology to massively accelerate AI inference.
Supporting the next class of massive AI models—and AI inference
Modern high-performance models run into the hundreds of billions to trillions of parameters and are typically mixture of expert models. As a result, the compute overhead of executing AI inference is a small subset of activating the total parameters.

This makes these problems enormously bound by memory bandwidth. Activated experts represent a fraction of the compute requirements compared to a massive dense model. In addition, the majority of the parameters of a model are dedicated to the feed-forward network where the experts are activated.
AI inference decode workloads—especially those in an MoE configuration—are enormously limited by standard memory bandwidth. GPUs offer more than enough firepower to generate tokens, but the constant round trips to and from HBM limits the ability to scale up low-latency experiences. Faster experiences require less traffic in the HBM, but it leaves GPUs sitting idle.
In the case of Kimi K3 (2.8T total parameters), the model activates 18 experts with 2 shared experts for a total of 102B active parameters. This reduces the total compute overhead required to roughly 3.6% of the model versus activating the dense architecture. In this case, memory bandwidth—getting data to and from those experts—is the major blocker.
Raptor is specifically designed with next-generation memory architecture to offer tremendously more bandwidth without needing the firepower of a GPU. Raptor accelerators activate the smaller subset of experts but with substantially higher memory bandwidth. They can also host smaller models (up to around 27B parameters) themselves with ample headroom for KV cache supporting longer context.
Building on a market-leading scale-up fabric with NVLink and NVFusion
Our rack-scale Raptor architecture is a tray-native design built on Nvidia’s rack architecture that the industry is already deploying for Vera Rubin. Our previous generation accelerator, Corsair, was designed to be full plug-and-play by being built on open standards for existing data centers, and the future of our products have to bring that same level of seamless integration for future next-generation data center buildouts.
Rack-level Raptor architecture supports a total of 18 accelerator trays and nine scale-up switch trays. Each tray contains two boards consisting of four Raptor R4 accelerators each for a total of 144 accelerators per rack. Each tray provides a total memory pool of 128 GB of 3D-DRAM. At rack-level, this configuration creates a total available memory pool of roughly 2.3 TB.

Raptor accelerator trays are accompanied with four NVFusion chips—two per board accompanying 4 Raptor R4s—to provide the fabric for NVLink scale-up along with one Vera CPU for configuration, management, and generic compute tasks.
A rack-scale Raptor solution can support models up to 3T parameters in a single rack. It easily supports models under 1T parameters, such as flash-series models like GLM 5.3 and DeepSeek V4.1, with substantially longer context windows. Workflows with models under roughly 20B parameters can fit on a single Raptor board, utilizing PCIe for scale-inside Raptor-to-Raptor communication.
Partnering with Nvidia on an NVLink scale-up fabric offers us the fastest time to market and industry-leading all-to-all scale-up network within a Raptor rack. NVLink scale-up topology enhancements are on the roadmap to extend the scale-out network across 4 raptor racks—576 Raptor accelerators total using NVLink576 scale-up to four Raptor accelerator racks—to support models up to 10T in parameters. Additional racks (Raptor or GPU) can be added to the cluster via ConnectX-9 scale-up – 3.2Tb/s per tray.
Support for these larger models also enables hyper-efficient pipelines by implementing speculative decoding with high-performance models on half-rack and full-rack solutions, as well as smaller draft models that can fit on a single Raptor board. It can seamlessly support prefill-decode disaggregation and attention-FFN disaggregation just like it can in heterogeneous configurations.
Supporting optionality with disaggregated Nvidia-enabled disaggregated pipelines
Our scaled-up tray-native Raptor rack architecture 2 also offers customers maximum optionality to work with what hardware exists in their current datacenters.
In addition to the NVLink scale-up network, the Raptor accelerator trays also support a scale-out network using four Nvidia ConnectX-9 NICs per tray that provide 3.2Tb/sec of scale-out bandwidth (up to 57.6Tb/sec per rack). The GPU racks can connect to the Raptor accelerators racks via this scale-out network for perform attention-FFN disaggregation, speculative decoding, and prefill-decode disaggregation tasks.
In short, we help GPUs handle tasks that they excel at while taking over memory-bound tasks.
In the case of attention-FFN disaggregation, we can do this at rack-scale with Raptor because the FFN represent a fixed, predictable size when compared to the growing KV cache. The GPU manages activations and the growing KV Cache, while Raptor-enabled racks hosts the FFN weights and handles expert activations and generating results.

This enables our rack-scale architecture to support inference for even the largest models, such as Kimi K3 or the newest GLM-5.3-Flash. The Raptor accelerator rack provides a total memory bandwidth of 7.2 PB/s. The accelerator trays are connected in an all-to-all configuration for Raptor-to-Raptor and Rubin-to-Rubin through nine switch trays.
Integrating NVLFusion with Raptor boards allows us to scale up our Raptor deployments with the same elegance that NVLink7 Nvidia scales up its Rubin deployments. This allows us to accelerate our integration with Nvidia buildouts while bringing the seamless scale-up for Raptor to provide the acceleration it brings to rack-level architecture.

Implementing this disaggregation with a combined Vera Rubin and Raptor rack delivers frontier-scale inference that stays at extremely low latencies. Raptor can sustain nearly 1,000 tokens per second per user while serving Kimi K3 at 1M context.
Raptor accelerator racks can also seamlessly connect with other Nvidia-based rack-scale solutions, such as Vera-only racks and storage-only racks.
How stacked DRAM massively accelerates frontier-level AI models
Our previous generation Corsair was built with on the principle that AI applications needed to both scale up elegantly and provide a snappy, low-latency experience without compromises.
Rack-scale Corsair uses an SRAM-based approach that provides significant support GPUs with using techniques like speculative decoding. Corsair can also scale up and out to support models well into the hundreds of billions of parameters.
Corsair servers carry 4GB of SRAM, requiring a very large volume of servers to scale out to support trillion-plus parameter classes. SRAM 6T cells are roughly 10 times the size of a DRAM 1T1C at a cost of 100x that of DRAM. HBM also carries a practical limit—enabling operating larger model sizes but with a significant handicap to memory bandwidth.
Corsair scales elegantly to support larger models thanks to its air-cooled PCIe form factor, which helps us slip directly into existing data centers. New buildouts, however, place a heavy emphasis on liquid cooling—and we needed to carry that level of innovation on the memory front to Nvidia’s rack architecture.
To that extent, Raptor’s 3D-DRAM provides the best of both worlds: an SRAM-like upgrade to HBM4 at a fraction of the power draw.

Raptor cards can both host large draft models on a single card for speculative decoding and scale up to support frontier-level models at a fraction of the energy cost.
New memory approach, better perf/TCO
Implementing disaggregated pipelines provides enterprises with immediate performance benefits. By implementing speculative decoding alone using Corsair’s SRAM-based approach provided an up to 10X performance boost in AI inference.
In a best-case scenario, this extends directly to whole models in an agentic chain. While Corsair by itself can take on smaller agentic tasks that only require small models, it can also accelerate larger models with speculative decoding. The benefit of implementing disaggregated pipelines in an existing GPU fleet is also immediate by unlocking more performance out of GPUs.

GPU-only fleets provide additive benefits to latency—each new card means supporting a higher volume of sessions, or lowering the batch size needed for existing AI inference and thereby making more memory bandwidth available per session. Even though they improve every generation, adding additional GPUs to preserve lower latencies is a costly endeavor.
Augmented fleets with disaggregated pipelines change the calculus entirely. Each additional memory-optimized card frees up memory pressure on GPUs that it can devote to more compute-heavy workloads, and those benefits only increase with newer generations of GPUs and newer generations of memory-optimized accelerators.
Our Raptor rack provides the same benefit as a significantly better scale:
- High memory bandwidth to provide greater support and optimize GPU-based operations.
- A substantially larger memory pool than Corsair to host larger models at rack-scale deployments, going from a single node to 144 Raptor accelerators in a single rack (up to 576 across multiple racks)
- Enables practical attention-FFN disaggregation for frontier-level models in a side-by-side deployment using NVL72 fabric.
New data centers are planned years in advance of actual completion, and at a certain point the infrastructure is completely locked while the racks and plumbing go into deployment. That creates an extraordinarily large barrier for implementing new kinds of rack architecture and infrastructure in data centers, much less ones that already exist to support fleets of GPUs.
That’s exactly why we built tray-based rack-scale Raptor solution: everything a Raptor rack needs is already spec’d into data centers built for Vera Rubin.
Our Vera-enabled Raptor rack meets enterprises and neoclouds where they are: the infrastructure is there, and they just need trays. Rather than wait for the next construction cycle, you can just roll a Raptor accelerator rack into the building and plug it in with everything else.