top of page

Frore’s LiquidJet Claim Puts Delidded Rubin GPUs on the Production Agenda

Frore Systems claims its LiquidJet coldplate can lower an Nvidia Rubin GPU’s temperature by 10°C and increase performance by 15%. The nvidia tom story matters because hyperscalers are considering a more aggressive step: removing GPU lids before installing cooling hardware in production servers.

That approach could put LiquidJet closer to the silicon, improving heat transfer where Rubin generates its hottest thermal concentrations. It also challenges the conservative engineering practices that protect expensive accelerators from assembly damage, leaks, uneven mounting pressure, and long-term material failures.

The central contest is therefore not LiquidJet against one competing coldplate. It is maximum direct-die cooling performance against the serviceability and reliability requirements of hyperscale infrastructure. Frore’s numbers describe the upside, but production validation will determine whether operators can safely capture it.

The Nvidia Tom Claim Connects Cooling Directly to Token Output

Frore is presenting the coldplate as a compute-performance component, not simply a way to keep Rubin within its temperature limit.

According to the original cooling report, LiquidJet can reduce Rubin GPU temperatures by 10°C. The company also links that reduction to a claimed 15% performance increase.

These figures remain company claims unless independent laboratories or deploying customers reproduce them under disclosed conditions. Workload selection, coolant temperature, flow rate, power settings, chip variation, and reference coldplate design can all affect the result.

The comparison becomes especially important because GPU temperature does not translate automatically into more useful work. An accelerator already maintaining its target clock might run cooler without producing additional tokens. A thermally constrained GPU, however, can benefit if lower temperatures let it sustain higher frequency or power for longer periods.

Frore’s performance argument appears to depend on that second situation. Better heat removal would reduce local hotspots, preserve operating headroom, and support more consistent clocks. The resulting gain could appear as higher token throughput, shorter training runs, or greater performance within a fixed rack-power envelope.

Token generation is the economically relevant measure for many inference operators. A coldplate that lets the same installed GPUs serve more requests can improve revenue per rack without requiring another building, electrical connection, or cluster expansion.

However, “performance” needs a precise definition. Training throughput, inference tokens per second, latency, and total tokens per megawatt are related metrics, but they are not interchangeable.

The reported 15% gain also needs a clear baseline. A comparison against a conventional coldplate would reveal something different from a comparison against Nvidia’s validated Rubin cooling design. Results from a simulated thermal load would carry less weight than measurements from production Rubin silicon.

Frore has already made similar claims for earlier designs. At Computex 2026, the company showed LiquidJet Nexus, an integrated cooling assembly for dense AI trays. An ODM test reportedly found a temperature reduction near 6°C and a 10% token-generation gain over the tested reference solution.

LiquidJet Nexus combines coldplates for several heat-producing components into one assembly. That arrangement can include GPUs, CPUs, networking devices, power converters, and voltage regulators. A shared structure can reduce fittings and potential leak points, although its full reliability still depends on manufacturing and deployment quality.

The earlier design reportedly measured 17 millimeters thick, compared with 34 millimeters for the competing assembly. Frore also claimed it weighed 65% less. Those dimensions matter in tightly packed systems where cooling hardware competes with power delivery and networking for limited space.

The latest nvidia tom claim raises the ambition from improving a conventional package interface to considering direct contact with delidded silicon. That change is what turns an incremental cooling story into a production engineering test.

A 10°C reduction would be meaningful if it survives rack-scale validation. The harder question is whether hyperscalers can obtain that reduction without introducing failure modes that erase the performance advantage.

Delidding Removes a Thermal Barrier and a Layer of Protection

Removing the lid shortens the path between Rubin’s silicon and its coolant, but it also moves packaging risk into the data center supply chain.

A GPU lid, often called an integrated heat spreader, is the metal surface above the silicon package. It spreads heat across a wider area while protecting sensitive components beneath it.

Heat normally travels from the GPU dies through an interface material, into the lid, and then through another interface into the coldplate. Every material boundary adds thermal resistance, meaning it slows heat movement across the stack.

Delidding removes the heat spreader and lets a cooling assembly contact the exposed package more directly. The shorter path can lower the temperature difference between the silicon and the coolant, especially around concentrated hotspots.

Direct-die cooling is not a new idea. Enthusiasts have removed processor lids for years, while specialized high-performance systems have used carefully controlled direct-contact assemblies. Hyperscale production creates a different risk profile because operators must deploy and maintain thousands of nearly identical systems.

A Rubin package combines high-power compute dies with high-bandwidth memory and other package components. Their heights, thermal loads, and mechanical tolerances do not necessarily match. A coldplate must contact the intended surfaces without damaging neighboring structures.

Mounting pressure becomes critical. Too little pressure can leave gaps that trap heat. Too much pressure can crack silicon, deform the package, or damage solder connections beneath it.

Pressure must also remain even as materials expand and contract. GPUs repeatedly move between lower idle temperatures and hotter operating conditions. Those thermal cycles can shift interface materials and stress mechanical connections over time.

The cooling assembly therefore needs more than a low initial temperature reading. It must retain that performance after transportation, installation, maintenance, and extended cycling across realistic power states.

Liquid exposure creates another concern. A leak above a lidded package is already serious. A leak or condensation event around exposed silicon can be even less forgiving.

Frore’s technical answer is a coldplate built around three-dimensional, short-loop jet channels. Coolant travels through small paths aimed at specific regions instead of following long, uniform channels across the entire plate.

Conventional microchannel coldplates often move liquid through longer machined passages. Those passages create friction and pressure loss, which increases the pumping work needed to maintain sufficient flow.

Frore says its architecture directs coolant toward a processor’s thermal map, meaning the known distribution of heat across the package. The company adapts semiconductor-style manufacturing processes to bonded metal structures, enabling internal shapes that are difficult to machine conventionally.

The company previously said LiquidJet could handle hotspot densities reaching 600 watts per square centimeter. It also claimed 50% more heat removal per unit of flow and a fourfold pressure-loss reduction in an earlier design.

Those figures provide a plausible mechanism for better cooling. Shorter channels can reduce hydraulic resistance, while targeted jets can disrupt the warm boundary layer that forms along a cooling surface.

Yet the mechanism does not settle the production question. Small channels can be sensitive to contamination, corrosion products, manufacturing debris, and variations in coolant chemistry. A highly optimized thermal structure must remain usable within real facility-maintenance practices.

Delidding also changes responsibility. Nvidia normally validates a complete package and cooling interface with defined mechanical tolerances. Removing the lid would require the system builder, cooling supplier, or hyperscaler to assume more of that integration burden.

That shift may be acceptable to organizations already designing custom accelerators and entire racks. It is harder for smaller operators that depend on standard warranties, replaceable parts, and vendor-qualified service procedures.

The attraction is nevertheless clear. If the extra interface layer accounts for a meaningful portion of Rubin’s thermal resistance, removing it gives coldplate designers direct access to the limiting surface.

The 10°C claim is therefore technically understandable. Its commercial value depends on whether Frore and its partners can package direct-die cooling as a repeatable manufacturing process rather than a specialized modification.

Rubin Makes Thermal Headroom an Economic Resource

Rubin’s density turns every degree of temperature and every watt of pumping overhead into a capacity-planning decision.

Nvidia designed Vera Rubin as a rack-scale platform rather than a loose collection of accelerator cards. The Vera Rubin NVL72 combines 72 Rubin GPUs, 36 Vera CPUs, networking, memory, power delivery, and cooling into a coordinated system.

Nvidia says Rubin infrastructure is its first generation with every major component cooled by liquid. That includes compute chips, networking equipment, and other hardware that previously depended partly on server fans.

The company’s liquid-cooling design supports coolant entering the rack at up to 45°C. Warmer facility water can make chiller-free heat rejection practical in suitable climates, reducing the energy required to cool the coolant itself.

This creates two distinct ways to use a better coldplate. An operator can hold coolant conditions steady and pursue more accelerator performance. It can also raise the inlet temperature or reduce flow while preserving the required silicon temperature.

The first choice targets tokens per GPU. The second targets tokens per megawatt by lowering cooling overhead. Power-constrained data centers can value either outcome, depending on electrical limits and workload demand.

Frore has explicitly described this flexibility. Its LiquidJet roadmap says operators can direct additional thermal headroom toward higher GPU performance or warmer coolant operation.

That framing explains why a cooling supplier would use token generation as a headline metric. Hyperscalers do not earn a return from lower die temperature alone. They care about completed training work, served inference requests, uptime, and facility capacity.

Nvidia is pursuing the same system-level economics through its own architecture. The company says Rubin’s 45°C design enables dry-cooler operation and avoids energy-intensive chilling in appropriate environments.

Its Vera Rubin architecture also uses common manifolds, liquid-cooled power components, quick disconnects, and leak-containment features. Nvidia argues that these improvements create enough power savings to support up to 10% more NVL72 racks within the same power budget.

Frore is effectively arguing that the coldplate can move those economics further. If LiquidJet removes more heat with less pressure loss, the facility could spend less energy on pumps or allocate additional thermal headroom to compute.

This is the primary competitive tension. Nvidia’s integrated cooling path prioritizes a validated rack and partner ecosystem. Frore’s approach promises additional efficiency by optimizing the surface closest to each chip.

The routes are not completely incompatible. Frore describes LiquidJet as a drop-in upgrade that can work with existing manifolds and liquid loops. A hyperscaler could use Nvidia’s broader rack architecture while replacing part of the thermal stack.

However, “drop-in” becomes a complicated description if deployment requires delidding. The external plumbing might remain unchanged, but package handling, assembly, testing, warranty coverage, and repair procedures would all change.

The economic calculation must include those operational effects. A 15% throughput gain could justify significant integration work, especially across a large inference fleet. A small increase in early hardware failures could offset part of that value.

Uptime matters as much as peak throughput. A server producing 15% more tokens when active offers little advantage if maintenance requirements materially reduce availability.

The same principle applies to yield. Hyperscalers would need to know how many accelerators survive delidding, coldplate installation, system assembly, shipping, and field service without damage.

Those costs do not make the idea impractical. Large cloud operators routinely co-design servers, cooling loops, networking, and custom silicon. They can establish controlled assembly lines and automated inspection processes unavailable to ordinary buyers.

Frore says it is working with a majority of hyperscalers on LiquidJet solutions, including designs for custom silicon. The company has not publicly identified production customers for the reported delidded Rubin configuration.

That distinction matters. Technical discussions, evaluation programs, qualification samples, and committed production deployments represent different levels of adoption.

The nvidia tom report suggests the industry is crossing from theoretical interest toward serious production evaluation. Public evidence must still catch up with that interpretation.

A 15% Gain Must Survive Reliability and Benchmark Scrutiny

The strongest case for LiquidJet requires independent throughput results and long-duration reliability data from complete Rubin racks.

Thermal demonstrations typically isolate a favorable part of the system. They can show that a coldplate removes heat effectively without proving that the entire data center benefits by the same percentage.

A production benchmark needs to disclose the GPU model, power limit, clock behavior, coolant inlet temperature, flow rate, pressure, workload, software stack, and comparison hardware. Without those details, readers cannot identify the source of a reported gain.

Chip-to-chip variation matters too. Two nominally identical GPUs can differ in leakage, temperature, and frequency behavior. A credible comparison would test enough devices to separate the cooling effect from ordinary silicon variation.

The 10°C result and 15% performance claim also need to be connected experimentally. Lower temperature can reduce leakage power and support higher sustained clocks, but the relationship varies across operating points.

A GPU running below its thermal limit might show little performance change. A GPU configured to exploit the added headroom could draw more power, complicating claims about facility efficiency.

The most useful metric would combine completed work with total energy. Tokens per joule, tokens per rack, and tokens per megawatt-hour would reveal whether LiquidJet improves economic output rather than only device frequency.

Inference latency should appear beside throughput. Some configurations increase total tokens per second by batching more requests while doing less for the response time experienced by an individual user.

Training presents another measurement challenge. A cooler GPU can maintain clocks, but overall training speed also depends on memory, networking, collective communication, checkpointing, and software efficiency.

Reliability evidence needs an equally broad scope. Coldplate vendors commonly perform pressure, corrosion, thermal-cycling, vibration, shock, and leak tests. Delidded production hardware adds mechanical alignment and exposed-package concerns.

Operators should look for accelerated-life testing across representative Rubin packages. They also need field-replaceable procedures that technicians can execute without turning routine maintenance into laboratory work.

Coolant cleanliness presents a long-term test. Frore’s internal structures depend on small, carefully shaped passages. Qualification should measure performance after extended exposure to the selected coolant, seals, tubing, and facility loop.

Material compatibility matters because mixed metals can drive galvanic corrosion. Biological growth, dissolved oxygen, particles, and chemical additives can also change cooling performance over time.

A data center can filter and monitor its coolant, but tighter requirements increase operational complexity. That complexity must be weighed against any reduction in pump energy or GPU temperature.

Supply capacity creates another uncertainty. Frore says it uses semiconductor manufacturing concepts on metal wafers. That approach enables detailed structures, but hyperscalers need consistent production across large volumes.

A successful qualification would therefore measure dimensional variation and flow balance across many coldplates. One exceptional sample does not establish fleet-wide performance.

The competitive response will also shape adoption. Traditional coldplate suppliers can improve channel geometry, interface materials, mounting systems, and manufacturing methods. Nvidia can revise package and rack designs to reduce thermal resistance without requiring customers to delid GPUs.

Nvidia’s decision to standardize much of its MGX rack architecture through the Open Compute Project creates room for component innovation. It also establishes mechanical and operational expectations that alternative cooling assemblies must satisfy.

Frore’s advantage may rest on customization. Its manufacturing method is intended to match the channel structure to each processor’s hotspot map. That can be valuable for Rubin, Rubin Ultra, and hyperscaler-designed accelerators with different heat distributions.

Customization also creates version-management work. Package revisions, power profiles, and chip configurations can require distinct coldplate designs. Operators must ensure that the correct cooling assembly follows each accelerator through manufacturing and service.

The reported interest in delidded GPUs suggests some hyperscalers accept this complexity when the potential fleet-level return is large enough. It does not mean the practice has become standard.

The exact terms of Nvidia’s support will be decisive. Operators will want clarity about warranty coverage, approved assembly partners, acceptable mounting methods, telemetry, and replacement procedures.

Until those questions receive public answers, the 15% figure should be treated as a promising test result rather than a bankable fleet improvement.

This cautious framing does not dismiss LiquidJet. It identifies the evidence required to move the technology from an impressive thermal comparison into production infrastructure.

Three Signals Will Show Whether LiquidJet Reaches Production

Customer-backed benchmarks, Nvidia qualification, and rack-scale reliability results will determine whether this cooling gain becomes deployable capacity.

The first signal is a benchmark from an identified hyperscaler, server maker, or independent testing organization. It should compare LiquidJet with a disclosed reference coldplate on real Rubin hardware.

That test should report coolant conditions, power settings, sustained frequency, workload throughput, and total cooling energy. It should also separate lidded and delidded configurations.

Independent confirmation near the reported 10°C and 15% results would strengthen Frore’s argument considerably. A smaller gain would not invalidate the design, but it would change the integration economics.

The second signal is formal platform support. Nvidia, an established MGX manufacturer, or a major server vendor would need to qualify the relevant coldplate and assembly process.

Qualification would show that direct-die cooling can fit within defined manufacturing tolerances and support procedures. It would also clarify whether delidding happens before delivery, at a system integrator, or under another controlled arrangement.

Nvidia has already made liquid cooling central to Vera Rubin systems. That commitment creates an addressable market for thermal innovation, but it does not automatically validate every cooling design.

An approved LiquidJet option would strengthen the claim that coldplates can become selectable performance components within the Rubin ecosystem. Continued silence would keep deployments limited to custom programs with greater integration responsibility.

The third signal is durability evidence from complete racks. Buyers need failure rates, thermal drift, leak performance, coolant-maintenance requirements, and service outcomes over sustained operation.

This evidence matters more than another short demonstration. AI infrastructure operates continuously, and small reliability differences multiply across thousands of accelerators.

A useful deployment report would compare uptime and maintenance against standard Rubin cooling. It would also show whether the initial thermal benefit remains stable after repeated power cycles.

These signals should arrive before readers treat delidded Rubin GPUs as a normal production practice. Hyperscalers often evaluate several engineering routes in parallel, and many never advance beyond qualification.

The broader trend is already settled. Cooling has moved into the performance and capacity equation for AI data centers. Rubin’s fully liquid-cooled architecture makes that relationship explicit.

What remains unsettled is how much risk operators will accept to remove the final thermal barriers around the package. Frore believes its targeted channels and direct-die approach justify a more aggressive design.

The nvidia tom claim gives that argument concrete numbers: 10°C lower temperature and 15% more performance. Those figures are large enough to demand attention, but not yet documented well enough to guide a fleet purchase.

Infrastructure teams should now ask for workload-level evidence rather than colder demonstration charts. They should compare total energy, uptime, serviceability, and hardware yield alongside peak throughput.

For developers and AI product teams, the consequence appears further downstream. Better cooling can mean more available inference capacity, steadier latency, and lower infrastructure cost per generated token.

Those benefits only materialize when the cooling system survives production. Watch for a named customer, a qualified Rubin configuration, and long-duration rack data. Together, those three signals will show whether LiquidJet is a compelling experiment or the next step in AI thermal design.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

For the best experience, remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page