top of page

d-Matrix Joins the NVIDIA NVLink Fusion Platform, but Raptor Must Still Prove Itself

Sep 13
13 min read

d-Matrix joins the NVIDIA NVLink Fusion Platform with a multi-year plan for Raptor, despite its first integrated racks being more than a year away. The inference-chip company expects Raptor to tape out before the end of 2026. Initial NVIDIA MGX rack availability is targeted for the fourth quarter of 2027.

The partnership gives d-Matrix a direct route into NVIDIA’s rack architecture, networking, cooling, and supply chain. It also reveals a harder truth about competing with NVIDIA. A differentiated accelerator is no longer enough. Customers increasingly expect a complete, deployable system with proven networking and software around it.

NVIDIA benefits from that requirement. NVLink Fusion lets custom processors enter its infrastructure without displacing the surrounding NVIDIA components. For d-Matrix, the arrangement reduces the engineering burden of scaling Raptor. For NVIDIA, it extends the reach of its platform into racks containing processors built for alternatives to general-purpose GPUs.

That makes the announcement less about one interconnect license than about control of the AI rack. Raptor promises specialized, memory-centric inference. NVIDIA supplies much of the system that turns the processor into infrastructure. The open question is whether d-Matrix can preserve a meaningful technical and economic advantage inside that environment.

d-Matrix Joins the NVIDIA NVLink Fusion Platform With a 2027 Target

The agreement moves Raptor from a standalone chip roadmap into NVIDIA’s rack-scale deployment model, but it does not make Raptor a shipping product.

d-Matrix announced the collaboration on September 10, 2026. Its multi-year roadmap starts with Raptor, the company’s next-generation inference XPU. An XPU is a workload-specific accelerator designed to handle selected computing tasks more efficiently than a general-purpose processor.

Raptor will connect through NVLink Fusion, NVIDIA’s technology for integrating outside CPUs and custom accelerators with its scale-up fabric. Scale-up networking joins processors inside a tightly connected computing domain. The design will also use Spectrum-X Ethernet for scale-out networking between systems.

The planned rack contains much more than d-Matrix silicon. The announced configuration includes NVIDIA Vera CPUs, NVLink switches, BlueField-4 data-processing units, ConnectX-9 SuperNICs, and Spectrum-X networking. It will follow the NVIDIA MGX reference architecture, including modular, cable-free compute trays.

Astera Labs will provide connectivity work intended to support high-throughput data movement across the system. This matters because accelerator performance can disappear behind communication bottlenecks once a workload spans processors, trays, and racks.

NVIDIA describes the arrangement as a path from custom silicon to full-scale deployment. Its rack-scale plan places Raptor beside NVIDIA infrastructure rather than treating it as an isolated appliance. Operators would gain a common rack design that can support GPUs, CPUs, and specialized XPUs.

The first proposed workload is disaggregated inference, which divides different stages of model execution among different processors. For AI coding, NVIDIA GPUs could process the compute-intensive prefill stage. Raptor could then handle the latency-sensitive decode stage that generates tokens one after another.

That division reflects a real architectural distinction. Prefill processes the user’s input and builds the model’s working state. Decode repeatedly reads model weights and produces the response. Decode performance can be constrained by memory movement, especially when users expect fast, interactive output.

d-Matrix targets that constraint with a memory-centric design. Raptor is planned to combine a DRAM memory chip with an SRAM compute chip in a vertically stacked package. SRAM provides fast access near computing logic, while DRAM offers greater capacity. The company says the two-story configuration shortens data paths and increases usable bandwidth.

However, the announced rack remains a roadmap. d-Matrix expects Raptor to tape out before the end of 2026. Tape-out marks the completion of a chip design before fabrication, not commercial delivery. Manufacturing, packaging, validation, software qualification, and rack integration still follow.

The company expects initial Raptor MGX systems in the fourth quarter of 2027. That timeline defines the central tension. d-Matrix has secured a credible route into NVIDIA-based data centers, but customers cannot yet measure the finished rack.

For buyers, the announcement therefore changes architectural confidence more than immediate capacity. Raptor now has a named interconnect, CPU, networking stack, rack format, and deployment target. The actual product must still cross the gap between design commitment and validated infrastructure.

NVIDIA Is Turning Accelerator Choice Into Platform Demand

d-Matrix gains access to a mature deployment system, while NVIDIA makes its infrastructure harder to avoid when customers choose non-NVIDIA accelerators.

NVIDIA introduced NVLink Fusion in May 2025 as a way to build semi-custom AI infrastructure. Its initial NVLink Fusion launch named MediaTek, Marvell, Alchip Technologies, Astera Labs, Synopsys, Cadence, Fujitsu, and Qualcomm among participating companies.

The program separates the accelerator decision from the rest of the rack. A cloud provider can select a custom XPU while retaining NVIDIA’s processors, interconnects, switches, network adapters, data-processing units, and reference designs. NVIDIA can therefore participate even when its GPU is not the only compute engine.

That is strategically important because hyperscalers and AI laboratories are developing specialized silicon to reduce dependence on a single accelerator supplier. Custom chips also let operators tune hardware for specific workloads, power limits, and service requirements.

NVLink Fusion answers that movement by making specialization compatible with NVIDIA’s platform. It does not require every processor to be an NVIDIA GPU. It encourages those processors to operate inside an NVIDIA-designed system.

For a smaller chip company, this exchange is attractive. Building a competitive accelerator already requires silicon design, compilers, kernels, runtime software, model support, packaging, and manufacturing. Rack-scale deployment adds power delivery, liquid cooling, cables, switches, serviceability, certification, and field support.

Each extra layer creates another reason for a buyer to delay adoption. An accelerator may perform well in a laboratory yet demand unfamiliar servers or a separate network. Data-center operators must then qualify a new stack while maintaining existing NVIDIA infrastructure.

MGX reduces that friction through reusable rack and server specifications. d-Matrix can focus more engineering effort on Raptor and its Aviator software stack. It can also present buyers with components and operational patterns that already appear elsewhere in NVIDIA deployments.

The cost is greater architectural dependence. d-Matrix is not merely connecting to an open network at the edge of the rack. Its proposed system incorporates NVIDIA technologies across the scale-up, scale-out, CPU, network interface, and data-management layers.

An independent assessment would be useful here, but the core tradeoff is visible from the announced bill of materials. d-Matrix gains deployment credibility by accepting NVIDIA’s rack as the operating environment.

That does not make the XPU irrelevant. The accelerator still determines how efficiently targeted model operations run. Yet NVIDIA controls many interfaces that determine whether isolated chip performance becomes sustained application performance.

This is why d-Matrix pressures competing infrastructure routes more than it directly pressures NVIDIA. A startup could build around Ethernet, PCI Express, or its own fabric. It would then need to persuade customers that the alternative delivers enough independence or efficiency to justify additional integration work.

Raptor instead meets buyers where much of their AI infrastructure already lives. That decision can shorten procurement and qualification cycles. It also reinforces NVLink as a common scale-up option for custom processors.

NVIDIA is effectively broadening its addressable platform. If general-purpose GPUs retain every workload, NVIDIA wins. If specialized accelerators take selected workloads but rely on NVLink, MGX, Spectrum-X, and BlueField, NVIDIA still occupies critical parts of the system.

The result is a subtler competitive map. Accelerator choice expands, but infrastructure convergence can increase. NVIDIA’s influence moves from owning every compute device toward defining how heterogeneous devices become one AI factory.

Raptor’s Mechanism Targets the Decode Bottleneck

Raptor matters only if its memory architecture converts lower data movement into better application latency, energy use, and rack economics.

d-Matrix built its market position around inference rather than model training. Training creates or updates model parameters through large parallel computations. Inference applies those parameters when a deployed model responds to users.

Those phases do not always reward identical hardware. GPUs remain highly flexible and benefit from mature software, broad model support, and massive deployment. Specialized inference processors can concentrate silicon and memory resources on repeated operations that dominate production serving.

Raptor extends the architecture behind d-Matrix’s current Corsair product. Corsair uses digital in-memory compute, which places arithmetic near stored model data. The objective is to reduce the energy and delay associated with repeatedly transferring weights between memory and compute units.

Data movement matters during token generation because large models must continually access substantial parameter sets. A processor with abundant arithmetic capacity can still wait on memory. d-Matrix’s thesis is that bringing computation closer to model weights produces faster small-batch inference.

Small batches are important for interactive applications. A provider can raise hardware utilization by combining many requests into a larger batch. Waiting to assemble that batch can add latency, while processing individual requests can leave conventional hardware underused.

Coding assistants, live chatbots, and voice agents expose this tradeoff directly. A user notices the delay before the first useful output and the cadence of following tokens. A platform operator notices the amount of hardware and energy needed to maintain that responsiveness across concurrent sessions.

d-Matrix calls this market premium token services. The phrase describes inference where customers value low latency enough to justify specialized infrastructure. It is a company framing, not an established economic category with independently reported market size.

The proposed Raptor design tries to expand the memory capacity available to this approach. d-Matrix says it will stack DRAM over an SRAM compute chip, creating a short physical path between capacity and processing. That differs from conventional accelerator designs that connect logic to separate high-bandwidth memory stacks.

The company has presented its 3D memory approach publicly, but a technical concept is not a rack benchmark. Thermal behavior, manufacturing yield, memory capacity, compiler support, and interconnect efficiency will all affect the final system.

NVLink Fusion addresses another part of the mechanism. One Raptor device cannot serve every large or distributed workload alone. Connecting multiple XPUs inside a high-bandwidth, low-latency domain lets d-Matrix pursue larger models and more simultaneous requests.

Spectrum-X then connects racks or system domains over Ethernet. This division between scale-up and scale-out gives d-Matrix a route from one specialized device to a data-center deployment. NVIDIA already packages those networking layers as part of its AI infrastructure.

Heterogeneous disaggregation provides the clearest proposed use case. A GPU handles prefill, where parallel computation is valuable. Raptor handles decode, where memory access and per-user latency carry greater weight. The two processor types must exchange model state efficiently enough to preserve the benefit.

That last condition deserves attention. Splitting a workload creates communication and orchestration overhead. If moving state between GPUs and XPUs consumes too much time, the combined system can lose the advantage measured on either processor separately.

Software determines whether the split remains manageable. Operators need scheduling, model support, monitoring, failure handling, and predictable performance across changing request patterns. They also need developers who can deploy models without maintaining separate pipelines for every hardware combination.

d-Matrix’s Aviator software must bridge those requirements. NVIDIA’s system components help with networking and rack operations, but they do not automatically make models perform well on Raptor. Kernel quality, quantization support, framework compatibility, and production observability remain d-Matrix responsibilities.

Corsair offers some evidence that the company has moved beyond simulation. d-Matrix said in June 2026 that Corsair entered full production, with volume shipments planned for selected customers. Its Corsair production update also described air-cooled cards and rack systems.

The company cited testing by Gimlet Labs in which a GPU and Corsair configuration reduced a baseline response from 24 seconds to under two seconds. That result concerned a particular speculative-decoding setup. It should not be generalized to every model, request pattern, or competing GPU system.

Still, Corsair gives Raptor a predecessor with working silicon and a software base. d-Matrix is not proposing its first accelerator. The harder step is translating that experience into a new 3D package and an NVIDIA-integrated rack.

This is the mechanism behind the d-Matrix NVLink Fusion decision. Raptor supplies specialized decode capacity. NVLink and MGX supply the connective and physical system. GPUs remain available for phases where they are better suited.

If the combination works, buyers could allocate different model stages to different processors without constructing a separate data-center architecture. If it fails, the added XPU becomes another component to purchase, program, monitor, and support.

The 2027 Delivery Date Leaves the Main Claims Unproven

The partnership reduces integration risk, but it does not verify Raptor’s performance, manufacturing readiness, or customer economics.

Raptor’s first announced milestone is tape-out before the end of 2026. A successful tape-out would freeze the design for manufacturing. It would not establish production yield, clock speeds, power consumption, cooling requirements, or sustained application performance.

The next challenge is the stacked memory package. Combining different memory technologies and logic dies requires tight control over manufacturing and thermal conditions. A design that shortens electrical paths can also concentrate heat and increase packaging complexity.

d-Matrix says Raptor is being evaluated by AI hyperscalers and frontier laboratories. It has not identified those evaluators in the announcement. An evaluation also differs from a purchase order, production deployment, or recurring workload commitment.

The same caution applies to the company’s reported patent count. d-Matrix says Raptor is backed by more than 100 patents. Patents can protect implementation choices, but they do not establish that a system is manufacturable or economically superior.

Independent testing will need to compare complete systems rather than isolated components. Useful measurements include time to first token, inter-token latency, sustained throughput, tail latency, energy per token, rack power, memory capacity, and utilization.

Model quality must remain constant during those comparisons. Quantization, speculative decoding, batching, and model modifications can improve speed while changing output or moving work elsewhere. Buyers need transparent configurations to understand what produced each result.

The strongest benchmark would reproduce a production service under realistic concurrency. It should include communication between NVIDIA GPUs and Raptor XPUs, not only a decode kernel running on one device. It should also report performance across several model sizes and architectures.

Software maturity creates another uncertainty. NVIDIA GPUs support a broad set of frameworks, libraries, and operational tools. A specialized accelerator does not need identical breadth, but it must support the workloads that justify its deployment.

Rapid model changes can complicate that task. New attention mechanisms, mixture-of-experts designs, longer context windows, and different numerical formats can alter the balance between compute and memory. Raptor’s architecture must remain useful as those workloads evolve.

The 2027 schedule gives competitors time to respond. NVIDIA will continue improving its own inference performance through GPUs, software, memory, and rack designs. Other specialized vendors can pursue lower latency through different processor and memory architectures.

Cerebras uses wafer-scale systems to reduce communication between separate chips. Groq focuses on deterministic execution and low-latency token generation. Hyperscalers including Amazon and Google develop their own accelerators around vertically integrated cloud services.

The broader inference-chip competition shows why deployment matters as much as architecture. Startups must compete with hardware availability, developer familiarity, cloud access, and procurement confidence, not only benchmark results.

NVIDIA’s platform can help d-Matrix address several of those barriers. It cannot eliminate the cost of adding another processor type. Buyers will still ask whether the latency or efficiency gain exceeds integration, inventory, software, and operational expenses.

Commercial terms remain undisclosed. The companies have not published expected rack configurations, available memory, power envelopes, service arrangements, or customer commitments. Those details will shape the actual value of the collaboration.

There is also a strategic dependence risk. d-Matrix positions Raptor as a complement to GPUs, but its rack relies on numerous NVIDIA components. Changes in NVIDIA roadmaps, licensing, supply allocation, or interface priorities can therefore influence d-Matrix deployments.

That dependence does not invalidate the partnership. Every infrastructure company builds around suppliers and standards. It does mean that claims of expanded accelerator choice should be assessed alongside the concentration of surrounding technology.

The key distinction is between architectural validation and market validation. NVIDIA’s participation validates Raptor as a serious candidate for integration. Market validation requires delivered systems, reproducible measurements, supported models, and customers running sustained production traffic.

Until those signals appear, d-Matrix Joins the NVIDIA NVLink Fusion Platform remains a roadmap story. It is a meaningful roadmap because it names the components, workload split, and availability window. It is not evidence that Raptor has won a share of production inference.

Three Signals Will Show Whether the Partnership Matters

Tape-out, rack-level benchmarks, and named production adoption will determine whether the announcement becomes infrastructure or remains an integration plan.

The first signal is Raptor’s tape-out. d-Matrix expects that milestone before the end of 2026. Meeting it would indicate that the architecture, NVLink interface, and physical design are ready to enter manufacturing.

A delay would weaken confidence in the fourth-quarter 2027 availability target. Tape-out is followed by fabrication, initial silicon testing, revisions, packaging qualification, and system validation. Each stage has limited room for slippage when the commercial system is already dated.

The quality of first silicon matters as much as timing. d-Matrix must show that the 3D memory package operates at the intended performance and power levels. Material revisions after fabrication would place additional pressure on the delivery schedule.

The second signal is an end-to-end benchmark of the Raptor MGX rack. It should compare GPU-only serving with heterogeneous prefill and decode under identical models, precision settings, request patterns, and quality targets.

A credible result must include the entire data path. That means GPU processing, state transfer, Raptor decoding, NVLink communication, networking, and scheduling. Component-level bandwidth claims cannot answer whether users receive faster responses.

Tail latency deserves particular attention. Average output speed can hide slow requests that damage interactive services. Operators should look for percentile measurements across changing concurrency, context lengths, and output lengths.

Energy and utilization should appear beside latency. A specialized XPU can deliver an impressive response time while sitting idle during other workload phases. Rack-level measurements show whether the combined system improves overall efficiency.

These benchmarks would strengthen the case if independent evaluators can reproduce them. They would weaken it if gains depend on narrow models, unusually small batches, or configurations that sacrifice output quality.

The third signal is named production adoption. d-Matrix says hyperscalers and frontier laboratories are evaluating Raptor, but it has not disclosed committed customers for the NVLink Fusion rack.

A named operator running a real coding assistant, voice service, or chatbot would provide stronger evidence than another architecture presentation. Production adoption would show that the latency benefit justifies new hardware, software, and operating procedures.

The depth of adoption will matter. A limited trial validates compatibility. A recurring deployment across multiple racks validates service reliability and economics. Expansion after the trial would offer the strongest evidence of customer value.

Buyers should also watch which model stages customers assign to Raptor. Consistent use for decode would support d-Matrix’s memory-centric thesis. Broader inference duties would suggest the architecture has more flexibility than the initial announcement emphasizes.

d-Matrix Joins the NVIDIA NVLink Fusion Platform at a moment when accelerator startups face a systems problem. Designing distinctive silicon is only the beginning. The product must fit the network, rack, software, cooling, and purchasing environment that customers already operate.

NVIDIA is offering that environment to companies whose chips can reduce demand for NVIDIA GPUs in selected workloads. This looks contradictory only if NVIDIA’s business is viewed as a single processor. Viewed as a platform strategy, the logic is direct.

d-Matrix receives a faster route to credible rack-scale deployment. NVIDIA extends its infrastructure into heterogeneous systems. Customers gain another accelerator option without abandoning the NVIDIA stack surrounding it.

The unresolved question is who captures the resulting value. Raptor must deliver enough latency or efficiency improvement to earn space beside GPUs. NVIDIA must keep NVLink Fusion attractive without making partners indistinguishable inside its architecture.

For developers and enterprise buyers, the practical response is to follow evidence, not topology diagrams. Track tape-out, demand complete rack benchmarks, and look for named production deployments. Those three signals will reveal whether the partnership expands real compute choice or mainly expands NVIDIA’s influence over that choice.

The next move belongs to d-Matrix. Can it turn a 2026 design milestone into validated Raptor racks by late 2027, with measurable gains on production models? Until then, the most important result of the deal is clear: alternative AI silicon increasingly reaches the data center through infrastructure that NVIDIA defines.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page