d-Matrix
  • Technology
  • Product
    • Aviator
    • Corsair at Rack Scale
    • JetStream
    • 3D DRAM
  • Solutions
    • Voice
    • Security
    • Optimization
    • Code Generation
  • Resources
    • Blog
    • Events
  • Company
    • About
    • Careers
    • Newsroom
    • Blog

d-Matrix Blog

Featured

Scaling AI the Right Way: Introducing Our Rack-Level Inference Solution  

Scaling AI the Right Way: Introducing Our Rack-Level Inference Solution  

October 14, 2025
What is AI Inference and why it matters in the age of Generative AI

What is AI Inference and why it matters in the age of Generative AI

June 4, 2025
The Complete Recipe to Unlock AI Reasoning at Enterprise Scale

The Complete Recipe to Unlock AI Reasoning at Enterprise Scale

February 13, 2025
d-Matrix Adopts NVIDIA NVLink Fusion Rackscale Infrastructure for Ultra-Low Latency AI Inference

d-Matrix Adopts NVIDIA NVLink Fusion Rackscale Infrastructure for Ultra-Low Latency AI Inference

d-Matrix will integrate XPUs into NVIDIA MGX rackscale architecture with NVIDIA NVLink Fusion; Collaboration to enable AI labs, hyperscalers, neoclouds to offer premium-level token services  SANTA CLARA, Calif., Sept. 10,… Read More
September 10, 2026
Model sizes are exploding—and new memory architectures are the way forward 

Model sizes are exploding—and new memory architectures are the way forward 

How disaggregated AI inference pipelines generate less heat — and smaller bills 

How disaggregated AI inference pipelines generate less heat — and smaller bills 

Fast token generation emerges as the key differentiator as heterogeneous inference takes hold

Fast token generation emerges as the key differentiator as heterogeneous inference takes hold

d-Matrix Acquires Wallaroo.ai to Speed up Deployment of Heterogeneous AI Inference Workloads

d-Matrix Acquires Wallaroo.ai to Speed up Deployment of Heterogeneous AI Inference Workloads

The new perf/TCO math: maximizing GPU utilization with disaggregated pipelines

The new perf/TCO math: maximizing GPU utilization with disaggregated pipelines

Advancing Full-Model AI Inference on Corsair with Infinity

Advancing Full-Model AI Inference on Corsair with Infinity

View All Posts >

Trending

Blazing the Trail Toward More Scalable, Affordable AI with 3DIMC 

How to Bridge Speed and Scale: Redefining AI Inference with Ultra-Low Latency Batched Throughput

Going Vertical: Why we created a 3D DRAM solution to advance low latency AI inference

Transforming AI: d-Matrix’s Pivotal Moments in Pursuit of Gen AI Inference At Scale

What is AI Inference and why it matters in the age of Generative AI

Featured Video

The DeepSeek Moment

In this short talk, d-Matrix CTO Sudeep Bhoja discusses the release of the Deep Seek R1 model, highlighting its impact on inference compute. He discusses the evolution of reasoning models and the significance of inference time compute in enhancing model performance.

Learn more about d-Matrix

From the Media

d-Matrix Adopts NVIDIA NVLink Fusion Rackscale Infrastructure for Ultra-Low Latency AI Inference

The Cube

Fast token generation emerges as the key differentiator as heterogeneous inference takes hold

d-Matrix | Wallaroo logos

d-Matrix Acquires Wallaroo.ai to Speed up Deployment of Heterogeneous AI Inference Workloads

Parasail to Combine NVIDIA AI Infrastructure with d-Matrix Accelerators to Achieve 10x Faster Token Generation 

LinkedIn AI Breakthrough award

d-Matrix Corsair Inference Accelerator Wins 2026 AI Breakthrough Award

Upstart chipmakers keep challenging Nvidia. This time it’s Microsoft-backed D-Matrix

View all media articles > For all press inquiries, please email pr@d-matrix.ai>
Transforming AI from
unsustainable to attainable.
  • Technology
  • Product
  • Solutions
  • About
  • Careers
  • Blog
  • Newsroom
  • Newsletter
  • Media Kit
  • Contact
  • Privacy Policy
  • Terms of Use
© d-Matrix, Inc. 2026
Follow us: X Twitter Logo Streamline Icon: https://streamlinehq.com