News analysis

d-Matrix Raptor Joins NVIDIA NVLink Fusion: What the Rack-Scale XPU Integration Means for AI Inference

d-Matrix plans to put its Raptor inference XPUs into NVIDIA MGX racks through NVLink Fusion. Here is what the integration changes, and what remains a vendor projection.

On this page
  1. The announcement is about integrating a non-NVIDIA inference accelerator into NVIDIA rack infrastructure
  2. Raptor is the next d-Matrix inference XPU after Corsair
  3. NVLink Fusion is the scale-up connection inside the rack
  4. MGX supplies more than the accelerator interconnect
  5. Why this can matter for inference deployment
  6. The target is latency-sensitive inference, not a replacement claim for every AI workload
  7. What is confirmed now, and what still needs evidence

The announcement is about integrating a non-NVIDIA inference accelerator into NVIDIA rack infrastructure

On September 10, 2026, d-Matrix and NVIDIA announced that d-Matrix will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA rack-scale AI infrastructure. Raptor is d-Matrix silicon built for AI inference; it is not an NVIDIA GPU, and the collaboration does not mean NVIDIA designed or owns the accelerator.

The significance is the surrounding platform. d-Matrix says Raptor will enter NVIDIA’s MGX rack-scale reference architecture instead of requiring every rack-level building block to be created as an isolated d-Matrix stack. NVIDIA describes the arrangement as a path for custom silicon to use parts of its broader AI infrastructure platform.

Raptor is the next d-Matrix inference XPU after Corsair

d-Matrix describes Raptor as the follow-on to its Corsair XPU platform and says it extends the company’s memory-centric architecture with a 3D DRAM approach that brings a DRAM memory chip and an SRAM compute chip together in one package. The company says Raptor is designed from the ground up for NVLink Fusion and MGX integration.

Raptor remains future silicon. d-Matrix says it expects the chip to tape out before the end of 2026. The announcement therefore establishes an architecture and product roadmap, not independent evidence of shipping-system throughput, power efficiency, reliability, customer-scale deployment or cost.

NVLink Fusion is the scale-up connection inside the rack

NVIDIA characterizes NVLink Fusion as the scale-up fabric in this design. In practical terms, that is the high-bandwidth, low-latency connection used to tie accelerators and other rack-scale compute components together closely enough to behave as a larger compute system. d-Matrix says Raptor will connect through NVIDIA NVLink switches as part of the rack architecture.

This is different from saying every connection in the rack is NVLink. NVIDIA’s own announcement separately describes Raptor using NVLink for scale-up and Spectrum-X for scale-out. Keeping those roles separate matters because communication among tightly coupled accelerators inside a rack and communication across larger clusters are different networking problems.

MGX supplies more than the accelerator interconnect

d-Matrix says the planned rack reference design includes NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking. It is also working with Astera Labs on connectivity solutions for data flow through the system. The d-Matrix rack is intended to use modular trays built around NVIDIA’s MGX ecosystem.

That makes the collaboration broader than adopting one connector. The value proposition is access to a rack architecture, networking components and an established supply-chain model around MGX while d-Matrix supplies the inference accelerator. It does not establish universal software compatibility: model support, compiler/runtime behavior and workload suitability still depend on the d-Matrix software and accelerator stack.

Why this can matter for inference deployment

Custom inference accelerators compete on more than raw silicon capability. Operators also need hosts, switching, network interfaces, rack mechanics, power and cooling integration, software and a deployment path that can scale beyond one card. Joining an existing rack ecosystem can reduce the amount of surrounding infrastructure a custom-accelerator vendor has to establish alone.

NVIDIA frames NVLink Fusion as a lower-risk and faster route from custom silicon to large-scale deployment, while d-Matrix says the collaboration should improve deployment flexibility and scaling for Raptor clusters. Those are strategic and engineering objectives from the participating vendors, not independent measurements of deployment time, total cost of ownership or production reliability.

The target is latency-sensitive inference, not a replacement claim for every AI workload

d-Matrix positions the planned system for low-latency inference and what it calls premium token services for AI labs, hyperscalers and neoclouds. That target is different from claiming Raptor is a universal substitute for NVIDIA GPUs across training, graphics, scientific computing or every inference workload.

The collaboration is strategically notable because NVIDIA is opening parts of its rack platform to a third-party accelerator while retaining NVIDIA CPUs, switches, DPUs, SuperNICs and Ethernet infrastructure around it. For d-Matrix, that can make its accelerator easier to place inside infrastructure operators already understand. For NVIDIA, it extends the relevance of its rack and networking ecosystem even when the main inference compute is not an NVIDIA GPU.

What is confirmed now, and what still needs evidence

Confirmed today are the September 10 collaboration, Raptor’s planned NVLink Fusion and MGX integration, the named NVIDIA infrastructure components, the separate NVLink scale-up and Spectrum-X scale-out roles, Astera Labs connectivity work, and d-Matrix’s expectation that Raptor will tape out before the end of 2026. Reuters independently reported the collaboration and the end-of-2026 tape-out target.

Not established by the announcement are independently measured Raptor performance, power draw, rack efficiency, total cost, customer adoption, shipment volume or a universal production-availability date. Claims about ultra-low latency, faster deployment, efficiency or premium token economics should therefore be read as vendor objectives until shipping hardware and independent measurements can test them.

Sources

Primary and technical sources

These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.

  1. 01 d-Matrix

    d-Matrix Adopts NVIDIA NVLink Fusion Rackscale Infrastructure for Ultra-Low Latency AI Inference
  2. 02 NVIDIA Blog

    d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
  3. 03 Reuters

    Chip startup d-Matrix to use Nvidia chip-linking tech in AI servers