The companies say combining specialized inference accelerators with GPUs could deliver major speed and power efficiency gains for frontier AI workloads.
AI infrastructure startup d-Matrix and applied AI firm Gimlet Labs have teamed up to bring specialized inference hardware into AI cloud environments, aiming to boost the performance and energy efficiency for real-time, agentic workloads.
Under the partnership, Gimlet plans to integrate d-Matrix Corsair accelerators into Gimlet Cloud alongside traditional GPUs. In the hybrid architecture, GPUs handle compute-heavy stages of inference, while memory- and latency-sensitive operations are routed to Corsair. The companies say this split can deliver up to 10x improvements in latency and throughput per watt compared with GPU-only deployments.