d-Matrix NVLink Fusion Deal Puts Raptor Inside NVIDIA’s Rack Strategy
d-Matrix has adopted NVIDIA NVLink Fusion for Raptor, despite building an inference accelerator intended to offer an alternative to conventional GPU designs. The agreement connects the future processor to NVIDIA’s rack architecture, networking, CPUs, and infrastructure supply chain. It also exposes the central tension behind the d-Matrix NVLink Fusion partnership.
Specialized accelerators can challenge NVIDIA’s compute silicon without displacing the rest of its platform. For d-Matrix, that bargain offers a faster route from an ambitious chip design to deployable infrastructure. For NVIDIA, every new XPU connected through NVLink strengthens its position around the rack.
The decision arrives before Raptor becomes a commercial product. d-Matrix expects the chip to tape out before the end of 2026, with integrated MGX systems initially available in the fourth quarter of 2027. That schedule makes this a roadmap announcement, not evidence of production performance.
It also places d-Matrix beside a widening group of companies using NVIDIA infrastructure for specialized processors. Those partners include established silicon vendors, cloud providers, and inference specialists. The alternative route is an open rack assembled around technologies such as AMD’s Helios architecture, UALink, and Ultra Ethernet.
What the d-Matrix NVLink Fusion Agreement Changes
d-Matrix is no longer designing Raptor as an isolated accelerator card. It is designing the processor as part of an NVIDIA-defined rack.
The companies announced their collaboration on September 10, 2026. According to the Raptor announcement, the arrangement includes a multiyear product roadmap rather than a single interoperability test.
Raptor will connect to NVLink switches for scale-up communication inside a tightly connected system. Scale-up networking lets multiple accelerators behave more like one large computing resource, with lower communication delays than ordinary data-center networking.
The rack will also use Spectrum-X Ethernet for scale-out networking between systems. That layer connects separate racks or clusters while handling congestion and traffic patterns associated with distributed AI workloads.
The proposed design includes NVIDIA Vera CPUs, BlueField-4 data processing units, and ConnectX-9 SuperNICs. It also uses NVIDIA MGX, a modular reference architecture covering physical trays, power delivery, cooling, and system integration.
Astera Labs will provide connectivity technology intended to maintain high-throughput data movement across the system. d-Matrix says the rack will use modular, cable-free trays from the existing MGX manufacturing ecosystem.
This scope matters because a processor alone is not a usable AI service. Buyers need servers, firmware, networking, orchestration, cooling, repair procedures, and a dependable flow of replacement parts.
A startup must usually build those elements or persuade partners to build them. It must then qualify the complete system for customers that cannot tolerate unexpected downtime.
NVLink Fusion changes that sequence. NVIDIA provides access to selected elements of its scale-up fabric and rack designs, allowing another company’s XPU to enter the same infrastructure framework.
The term XPU refers broadly to a specialized processor optimized for workloads that do not fit a general-purpose CPU. In Raptor’s case, the target is generative AI inference, especially latency-sensitive token generation.
The d-Matrix Raptor XPU can also operate beside NVIDIA GPU systems, including Vera Rubin NVL72. NVIDIA describes that arrangement as disaggregated inference, where different processors handle distinct stages or types of serving work.
That coexistence creates the story’s reversal. d-Matrix does not need to replace NVIDIA across an entire data center to win an inference deployment. It can compete for selected workloads while relying on NVIDIA components everywhere around its chip.
NVIDIA gains something equally important. It can accommodate custom accelerators without surrendering its control over the surrounding system architecture. The competitive boundary moves from the chip to the rack.
The deal therefore represents more than another compatibility badge. It tests whether NVIDIA can make third-party accelerators increase demand for its infrastructure, even when those accelerators compete with its GPUs.
Why Rack Integration Has Become the Real Barrier
The scarce capability is no longer designing an impressive accelerator. It is converting that silicon into infrastructure customers can deploy predictably.
AI inference places different demands on hardware than model training. Training emphasizes large parallel calculations across many accelerators. Inference must also manage response time, concurrent users, model memory, and an unpredictable stream of requests.
Reasoning models intensify that pressure because they can generate far more tokens before returning an answer. Long-context applications also maintain large key-value caches, which store information needed during token generation.
Those workloads turn memory movement into a central constraint. An accelerator can offer abundant arithmetic capacity while leaving compute units waiting for model weights or cached context.
d-Matrix attacks that problem through memory-centric computing. Its architecture places matrix operations closer to stored data, reducing the distance information travels during inference.
However, solving the memory bottleneck inside a processor does not solve deployment outside it. Rack-scale systems must coordinate accelerators, hosts, storage, networking, power, and cooling under real operating conditions.
Each component introduces qualification work. Engineers must validate signal integrity, thermal behavior, firmware compatibility, collective communication, failure recovery, and service procedures.
Liquid-cooled racks add another operational layer. A new vendor must fit established facility designs without requiring customers to rebuild power and cooling around one processor.
NVIDIA’s original NVLink Fusion launch addressed that problem directly. The company introduced the platform in May 2025 for semi-custom AI infrastructure using outside CPUs and XPUs.
Its initial partner list covered several parts of the design chain. MediaTek, Marvell, Alchip, Astera Labs, Synopsys, and Cadence supported custom silicon, connectivity, or intellectual property.
Fujitsu and Qualcomm planned custom CPUs that could work with NVIDIA GPUs. Later collaborations extended the platform to additional processors and cloud designs.
That growing roster creates pressure on every independent accelerator vendor. A chipmaker can join an established rack ecosystem, build an equivalent system alone, or align with a competing open standard.
The first route reduces integration risk but creates strategic dependence. The second preserves more control but demands capital, time, and customer confidence. The third depends on another ecosystem reaching comparable maturity.
d-Matrix has chosen speed and deployability for Raptor. Its chief executive, Sid Sheth, framed the constraint plainly: inference demand is rising while capital, time, and energy remain finite.
The company’s decision also reflects its stage of development. d-Matrix began shipping its Corsair platform before introducing Raptor, but it remains much smaller than the infrastructure vendors it hopes to challenge.
Building a new rack platform alongside a new memory architecture would multiply execution risks. Using MGX lets the company concentrate more engineering resources on its processor, compiler, and inference software.
This choice does not eliminate qualification. Raptor must still implement NVLink correctly, work with NVIDIA’s other devices, and meet system-level performance and reliability targets.
It does narrow the number of new elements customers must accept at once. A familiar rack can make an unfamiliar accelerator easier for an infrastructure team to evaluate.
The result pressures rival accelerator companies as much as GPU vendors. A startup offering only a fast chip now competes against processors packaged within validated, serviceable rack designs.
How NVLink Fusion Works Around Raptor
NVLink Fusion separates processor choice from rack construction, but it keeps NVIDIA’s interconnect at the center of the system.
The architecture has two networking levels. NVLink connects accelerators within a scale-up domain, while Spectrum-X carries traffic across a larger scale-out cluster.
Inside the rack, NVLink switches provide high-bandwidth, low-latency communication among Raptor devices. This connection matters when a model or its working data cannot remain on one accelerator.
Outside that domain, ConnectX SuperNICs and Spectrum-X Ethernet link systems across the data center. BlueField DPUs can handle infrastructure tasks such as networking, isolation, and data movement.
Vera CPUs serve as host processors. MGX defines the mechanical and electrical framework that packages those elements into deployable trays and racks.
Understanding how NVLink Fusion works requires distinguishing access from standardization. NVIDIA is opening its fabric to approved third-party silicon, but NVLink remains an NVIDIA-controlled technology.
The design is therefore horizontally inclusive without becoming vendor-neutral. Partners gain access to the platform, while NVIDIA retains influence over interfaces, qualification, roadmaps, and surrounding components.
That model offers practical advantages. A shared rack can support different accelerator types without forcing an operator to create a separate physical design for every processor.
A data-center builder can standardize floor space, cooling connections, power distribution, and maintenance practices. Compute capacity can then vary according to workload requirements.
NVIDIA also brings an established manufacturing network. Original equipment manufacturers and design manufacturers already produce systems derived from MGX and NVIDIA’s rack-scale GPU platforms.
The company’s technical explanation describes access to NVLink interfaces, chiplets, switches, cabling, and rack technology. It also includes power and liquid-cooling designs.
For d-Matrix, this infrastructure addresses a problem that raw accelerator benchmarks cannot capture. Customers buying inference capacity evaluate deployment schedules, serviceability, utilization, and operational risk alongside tokens per second.
A processor that arrives late or requires a unique rack can lose even with favorable laboratory results. Infrastructure consistency can outweigh a narrow performance advantage.
The approach also lets customers mix specialized inference with GPU-based workloads. NVIDIA says Raptor racks can work beside Vera Rubin NVL72 systems rather than replacing them.
One possible deployment would direct latency-sensitive token generation to Raptor while keeping other model stages on GPUs. The companies have not published a complete production configuration proving that workflow.
Software remains essential to any such split. Model runtimes must place work correctly, manage memory, and move data without erasing the gains from specialized hardware.
d-Matrix has developed its own software stack for Corsair and Raptor. Integration with the broader NVIDIA environment will determine whether buyers experience one manageable platform or two adjacent systems.
That distinction will matter to developers. Hardware heterogeneity offers better workload matching, but it can introduce separate compilers, monitoring tools, performance profiles, and debugging paths.
NVLink reduces communication friction between devices. It does not automatically make their programming models identical.
The partnership’s value therefore depends on coordination above the physical link. The companies must turn compatible components into repeatable deployment patterns that cloud operators can expose as services.
Raptor’s Memory Design Is the Bet Inside the Rack
NVIDIA supplies the system foundation, but Raptor still has to justify why customers need another inference processor.
Raptor extends the memory-centric approach used by d-Matrix’s earlier Corsair platform. Its defining feature is a three-dimensional package that places a compute die directly above custom DRAM.
DRAM offers greater density than SRAM, the faster memory used extensively in Corsair. However, conventional DRAM sits farther from compute and usually requires more energy to move each bit.
d-Matrix’s design exposes many small memory banks directly to compute engines through dense vertical connections. The goal is to combine DRAM capacity with far greater local bandwidth.
At Hot Chips 2026, the company presented a package with a TSMC 4-nanometer compute die bonded above a custom DRAM die. The interface uses a 36-micron connection pitch.
The demonstrated design offers 32GB per card and a stated 100 terabytes per second of internal bandwidth. Those figures describe movement inside the package, not network bandwidth between separate accelerators.
Independent reporting on the 3D DRAM design also highlighted an important qualification. Many published performance results remain projections based on early silicon.
d-Matrix measured the vertical interface at 0.37 picojoules per bit. The company compared that figure with roughly 2.4 picojoules for movement into an HBM4 base die.
An associated research paper projected about 4.7 times more throughput per card than HBM-based designs. Neither a projection nor an interface measurement establishes full production-system performance.
Thermal management presents another challenge. The logic die sits on top so a cold plate can contact it directly, while the DRAM underneath also serves as an interposer.
The complete package has a disclosed power budget of 422 watts. The vertical memory interface accounts for 296 watts when operating at full capacity.
Heat affects DRAM retention, which determines how long memory cells preserve data before refresh. At the stated operating temperature, the design requires substantially more frequent refresh operations.
d-Matrix says smaller memory banks limit the resulting bandwidth cost. It also includes spare banks, error correction, and redundancy intended to maintain reliable operation.
These details explain why integration with a mature liquid-cooled rack matters. Raptor’s architecture does not merely require a connector. It requires power, cooling, and validation designed around an unusual package.
The technology targets token generation because that phase often moves model weights repeatedly while performing relatively modest arithmetic per byte. More local memory bandwidth can keep compute units occupied.
Prefill, which processes an incoming prompt, has a different performance profile. It can demand more arithmetic and may favor a different accelerator configuration.
That difference supports the disaggregated inference argument. Operators could assign separate processors to prefill and decoding, provided software and networking keep the handoff efficient.
Yet Raptor’s value cannot be inferred from bandwidth alone. Useful performance depends on model support, numerical formats, scheduling, batch size, context length, and acceptable response latency.
Memory capacity also matters. A 32GB card cannot hold every large model independently, so larger workloads require partitioning across several devices.
That requirement makes the NVLink scale-up domain more consequential. Raptor’s internal memory design and NVIDIA’s external fabric must work together without creating a new bottleneck.
The d-Matrix Raptor XPU is therefore a compound bet. Its 3D package must work at production yield, and the rack must convert that package into dependable application performance.
NVIDIA’s Open Platform Still Has Boundaries
The primary contest is not d-Matrix against NVIDIA GPUs. It is NVIDIA-controlled integration against vendor-neutral rack infrastructure.
NVIDIA describes its AI platform as vertically integrated and horizontally open. The phrase captures its strategy, but buyers should examine what each part means.
Vertical integration joins NVIDIA processors, switches, network adapters, DPUs, software, rack designs, and supply relationships. Horizontal openness allows selected outside CPUs and XPUs to enter that environment.
This structure expands processor choice inside a system whose essential fabric remains controlled by NVIDIA. It is more open than a GPU-only rack, but less neutral than an industry-governed interconnect.
That boundary is commercially useful for NVIDIA. If custom accelerators gain share, the company can still supply high-value networking and infrastructure components around them.
It can also keep NVLink central as AI systems shift from servers toward rack-scale computers. The rack becomes the product, while individual processors become configurable elements.
d-Matrix benefits because it can reach buyers already preparing facilities for NVIDIA equipment. The partnership lowers the organizational cost of testing a less familiar processor.
However, d-Matrix also inherits dependency on NVIDIA’s qualification process and infrastructure roadmap. Changes to switches, CPUs, software, or commercial terms can affect its system plans.
The companies have not disclosed the partnership’s financial structure. They have also not detailed which software layers will be shared or how customers will procure and support complete racks.
Those unanswered questions matter because openness has several dimensions. Hardware can be physically interoperable while procurement, management, and development remain tightly coupled to one vendor.
The competing model emphasizes open industry interfaces. AMD’s Helios rack design uses OCP Open Rack Wide, UALink for scale-up connectivity, and Ultra Ethernet for larger clusters.
Helios combines AMD Instinct accelerators, EPYC processors, and Pensando networking. Its published design includes 72 accelerators, matching the industry’s movement toward dense rack-scale systems.
UALink aims to let multiple vendors connect accelerators through a shared specification. That approach promises broader portability, although an open specification does not guarantee equal product maturity or deployment volume.
NVIDIA’s advantage is that NVLink already operates across several generations of shipping systems. Its manufacturing partners also have practical experience with dense, liquid-cooled racks.
The open route offers greater theoretical independence. The NVIDIA route offers an established integration path under one company’s architectural direction.
Neither choice removes lock-in entirely. A buyer adopting Raptor also depends on d-Matrix software, its 3D memory supply chain, and the startup’s ability to support future products.
The larger competitive field includes Groq, Cerebras, AMD, Intel, and hyperscaler-designed accelerators. An inference market overview previously identified those companies as alternatives targeting workloads dominated by GPUs.
Several have since moved closer to full-system offerings. Groq, for example, has also entered NVIDIA’s expanding inference infrastructure strategy.
That pattern suggests NVIDIA wants to absorb specialization rather than resist it. A successful XPU can become another reason to adopt NVIDIA networking and rack technology.
For d-Matrix, joining that platform is pragmatic. It also means the company’s challenge to NVIDIA is narrower than a simple rival-chip narrative suggests.
The partnership contests which processor performs inference. It does not contest who defines much of the surrounding infrastructure.
Delivery, Benchmarks, and Customers Will Decide the Outcome
The announcement establishes architectural intent. It does not establish manufacturing readiness, commercial demand, or production performance.
The first observation point is Raptor’s planned tape-out before the end of 2026. Tape-out is the point when a chip design is finalized for manufacturing.
Meeting that milestone would support the current schedule. Missing it would compress time for fabrication, packaging, testing, software work, and rack qualification before late 2027.
The second signal is independently reproducible system performance. Buyers need results from complete Raptor racks, not isolated bandwidth figures or projected model runs.
Those evaluations should disclose models, context lengths, batch sizes, latency targets, accuracy settings, and power consumption. Tokens per second without those conditions can hide major tradeoffs.
Performance per user will be especially relevant for the company’s premium token-service pitch. High aggregate throughput matters less if individual requests wait in large batches.
Energy measurements should cover the whole rack. Package-level efficiency can be diluted by CPUs, switches, cooling equipment, networking, and idle capacity.
The third signal is named customer deployment. d-Matrix says Raptor is under evaluation by hyperscalers and frontier laboratories, but it has not identified those organizations.
A committed cloud instance, managed inference service, or announced production cluster would strengthen the partnership’s commercial case. Evaluations alone do not show purchasing intent.
Initial availability remains scheduled for the fourth quarter of 2027. That long interval gives competing systems time to improve their memory, networking, and inference software.
It also exposes d-Matrix to manufacturing uncertainty. The company must produce a custom DRAM die, bond it to advanced logic, achieve acceptable yields, and validate the package under sustained heat.
Its claim of more than 100 patents does not resolve those production questions. Patents protect technical ideas, while customers need reliable volume and predictable service.
NVIDIA also has work to complete. The relevant Vera, BlueField, ConnectX, Spectrum-X, and NVLink components must reach the required maturity on compatible schedules.
Astera Labs must deliver its connectivity pieces as part of the integrated design. Rack manufacturers then need to qualify cable-free trays, cooling, firmware, and system management.
Software presents a parallel schedule. d-Matrix must support current models and frameworks when Raptor ships, not only those available during design.
Model architectures can change quickly. Sparse mixtures of experts, longer context windows, multimodal workloads, and new decoding techniques alter memory and communication patterns.
A specialized processor succeeds when its architecture remains useful across those changes. It also needs compiler updates fast enough to keep pace with widely used models.
NVIDIA’s platform can reduce physical deployment risk, but it cannot guarantee software adoption. Developers and cloud operators still need a reason to direct workloads toward Raptor.
The strongest case would combine low per-user latency, competitive energy use, and straightforward model onboarding. Weakness in any one area can limit utilization.
Buyers should also watch how NVIDIA positions Raptor beside its own GPUs and other inference products. Equal technical access does not necessarily produce equal commercial visibility.
A final concern is platform concentration. Supporting outside XPUs gives customers more processor choices, while moving more rack decisions under NVIDIA’s influence.
That outcome is neither pure openness nor simple foreclosure. It is a layered market where competition survives inside an increasingly common infrastructure boundary.
The d-Matrix NVLink Fusion collaboration will matter if Raptor crosses that boundary as a shipping, measurable, and widely available service. Until then, its significance lies in the strategy.
d-Matrix has accepted that a better inference chip is insufficient without a credible rack. NVIDIA has accepted that specialized processors can enter its platform without weakening its infrastructure position.
Over the coming months, watch the tape-out, complete-system benchmarks, and named deployments in that order. Each milestone will show whether this roadmap is becoming a product.
For infrastructure buyers, the useful question is not whether Raptor defeats every GPU. Ask whether it delivers a valuable inference profile without adding unacceptable operational complexity. That evidence will determine whether NVIDIA’s rack becomes a launchpad for accelerator choice or another durable point of dependency.



