Running a single AI operation efficiently is one thing. Enabling a complete generative AI model—from its first computation through stateful, multi-token generation—is a much broader systems challenge.
d-Matrix and Infinity have collaborated to enable and optimize full-model inference on the d-Matrix Corsair™ platform. Together, the teams brought Qwen3 from initial tensor-parallel operations to complete, stateful inference, demonstrating how purpose-built AI infrastructure and advanced software optimization can accelerate the deployment of leading generative AI models.
The collaboration represents an important step toward making performant AI inference more accessible across a rapidly changing model landscape.
Moving beyond individual kernels
AI accelerator performance is often discussed through individual operations, kernels or isolated benchmarks. These measurements are useful, but production inference requires the entire model to operate as a coordinated system.
A complete inference workload includes much more than matrix multiplication. It also depends on attention, normalization, communication between compute resources, memory movement, token sampling and the management of model state across successive tokens.
Every component must work together efficiently. A bottleneck anywhere in the inference pipeline can affect latency, throughput and the overall user experience.
By enabling Qwen3 end to end on Corsair, the d-Matrix and Infinity teams demonstrated complete model execution rather than the acceleration of only selected model components.
Enabling stateful generative AI inference
Generative AI inference is inherently stateful. As a model processes a prompt and generates a response, it must preserve and reuse information from previous tokens.
This ongoing context is fundamental to conversational AI, reasoning, agentic applications, code generation and other interactive workloads. Supporting it efficiently requires close coordination across the underlying compute architecture, memory system, runtime and model implementation.
The work with Infinity extended from initial tensor-parallel operations across Corsair’s compute resources to a complete inference implementation capable of maintaining the state required for multi-token generation.
The result demonstrates Corsair’s ability to support the broader execution requirements of modern large language models—not simply isolated portions of their computation.
Accelerating the path from model to deployment
The pace of AI model development presents an additional challenge for infrastructure providers. New model architectures, operators and optimization opportunities emerge continuously, making rapid model enablement an essential capability.
Infinity is developing a software layer designed to make AI accelerators inference-ready more quickly. Its work combines automated kernel development, model implementation and full-stack optimization to adapt models to different hardware architectures.
Corsair presented a distinctive enablement challenge. It is a purpose-built inference platform with its own architecture, memory-compute approach and instruction set—not a conventional GPU target with an established library of model implementations.
Working with the d-Matrix team, Infinity progressed from initial tensor-parallel matrix operations to multiple end-to-end frontier-model implementations. According to Infinity, tensor-parallel matrix multiplication was running across Corsair’s 32 compute units within hours, followed by complete model execution within days.
This collaboration illustrates how specialized hardware and advanced software automation can work together to shorten the path between the release of a model and its efficient deployment on production infrastructure.
A purpose-built platform for inference
Corsair was designed from the ground up for generative AI inference in data centres.
Its Digital In-Memory Compute architecture brings memory and computation closer together, helping reduce the data movement that contributes significantly to the cost and energy consumption of conventional AI systems. The platform is designed to provide low latency while maintaining the batched throughput required for large-scale deployments.
This combination is increasingly important as AI applications become more interactive. Users expect models to begin responding quickly, continue generating at a high rate and serve many simultaneous requests without creating unsustainable infrastructure costs.
Corsair is now in full production, with d-Matrix scaling the platform for deployments across hyperscalers, neoclouds and frontier AI labs.
The work with Infinity adds another dimension to that production readiness: the ability to bring complete models onto the platform rapidly and optimize them across the full inference stack.
Building an ecosystem for attainable AI
Delivering production-ready AI requires more than silicon alone.
It requires efficient hardware, a capable software stack, optimized model implementations, scalable system infrastructure and an ecosystem of partners that can respond quickly as models and applications evolve.
Collaborations such as this one help create a more flexible path for deploying generative AI beyond conventional accelerator environments. They also support a broader d-Matrix objective: transforming AI inference from an increasingly expensive and energy-intensive undertaking into infrastructure that is performant, commercially viable and sustainable.
By combining Corsair’s purpose-built architecture with Infinity’s model-enablement capabilities, the teams have demonstrated a practical route from a new hardware target to complete, stateful model inference.
And this is only the beginning.
Read Infinity’s technical perspective
Infinity’s article provides a closer look at the enablement process, including how its engineering system approached Corsair’s architecture, developed the required inference components and progressed toward full-model Qwen3 execution.