top of page

NVIDIA DGX Spark 64GB Arrives as the 128GB Model Gets a Major Price Hike

4 days ago
13 min read

NVIDIA has announced the NVIDIA DGX Spark 64GB, while the original 128GB system is receiving a price increase of nearly 50 percent.

The new configuration preserves the GB10 Grace Blackwell processor, NVIDIA’s software stack, and high-speed networking. However, it halves the unified memory and reportedly reduces local storage.

That trade changes the meaning of NVIDIA’s desktop AI system. DGX Spark originally promised unusually large unified memory in a compact machine. The new entry model makes that defining resource scarcer while charging more than the original 128GB launch price.

NVIDIA says the 64GB systems will become available through Acer, ASUS, Dell, Gigabyte, HP, and MSI on October 23. The company will not sell a NVIDIA-branded 64GB Founders Edition.

The 64GB launch details emphasize private inference, persistent AI agents, and an easier path to two-node clusters. Those uses are credible, but capacity limits will shape which models and workflows remain practical.

The NVIDIA DGX Spark price increase creates the central conflict. Buyers can accept less memory, pay substantially more for 128GB, or evaluate competing systems based on AMD and Apple silicon.

NVIDIA DGX Spark 64GB Keeps the GB10 but Halves Its Memory

The new system retains NVIDIA’s computing platform, but it cuts the resource that made DGX Spark especially valuable.

DGX Spark combines a Blackwell-generation GPU and a 20-core Arm CPU in the GB10 Grace Blackwell Superchip. MediaTek collaborated with NVIDIA on the processor design.

The CPU and GPU share one coherent memory pool. Coherent unified memory lets both processors access the same data without maintaining separate system and graphics memory allocations.

That arrangement matters for local artificial intelligence. Large models must fit their weights, runtime state, prompts, and generated data inside available memory.

Traditional desktop GPUs often have far less dedicated memory than DGX Spark. Developers can add system RAM, but a discrete GPU cannot use it as efficiently as local video memory.

The original 128GB machine offered a different balance. NVIDIA said it could run inference on models containing as many as 200 billion parameters and fine-tune models with up to 70 billion parameters.

Those are vendor limits, not guarantees for every model. Precision, quantization, context length, cache requirements, and framework overhead all change actual memory consumption.

The NVIDIA DGX Spark 64GB reduces NVIDIA’s stated single-system model limit to 100 billion parameters. That remains substantial, particularly when models use compressed weights.

However, fitting a model is not the same as running it comfortably. A workload also needs space for its key-value cache, which stores attention data generated during inference.

Longer prompts consume more cache. Multiple users or agents increase the requirement again. A model that barely loads can leave too little headroom for useful work.

NVIDIA has not cut memory bandwidth alongside capacity. Reports indicate that the 64GB systems retain bandwidth of 273GB per second.

That suggests manufacturers are using lower-density LPDDR5X packages without narrowing the memory interface. The design should therefore avoid an obvious bandwidth penalty when workloads fit.

The new machines also retain DGX OS, CUDA libraries, PyTorch support, NVIDIA’s agent tools, and popular inference runtimes. Supported options include Ollama, llama.cpp, vLLM, and LM Studio.

ConnectX-7 networking remains part of the platform. Its 200Gbps connection allows two systems to exchange data far faster than ordinary consumer Ethernet.

The 64GB version is therefore not a slower processor sold under the same name. It is a capacity-constrained version of the existing platform.

That distinction favors developers running compact models repeatedly. It is less attractive for buyers who chose DGX Spark specifically to avoid memory limits.

Storage reportedly falls alongside system memory. NVIDIA’s public announcement concentrates on memory and clustering rather than providing a complete partner-by-partner storage specification.

Buyers should check each manufacturer’s configuration before ordering. The partner-only release means storage, ports, support, and warranty terms can vary between systems.

NVIDIA’s first DGX Spark launch positioned the machine as a personal AI supercomputer. The 64GB edition narrows that promise to more targeted local workloads.

That is still useful. It is simply a different proposition from placing an unusually large model-development environment on one desk.

The NVIDIA DGX Spark Price Increase Rewrites the Value Argument

The lower-capacity model exists because the original system has moved far beyond its initial market position.

DGX Spark began as Project Digits, a compact desktop system announced at CES 2025. Early positioning suggested a lower entry point than the machine eventually received.

The 128GB Founders Edition later launched at a higher figure. NVIDIA increased that price again during 2026, citing constrained supplies of LPDDR5X memory and NAND flash.

The latest reported adjustment raises the 128GB model by almost half from its preceding list price. It stands nearly three-quarters above the original launch level.

The precise transaction price can vary by seller and availability. However, the direction is clear across NVIDIA’s product line and several partner systems.

According to the original price report, the new 64GB configuration now serves as the less expensive entrance to GB10.

Less expensive is relative here. The reduced model reportedly costs one-quarter more than the 128GB Founders Edition did at launch.

That comparison exposes the reversal. Buyers are not receiving a normal lower-memory discount against a stable product.

Instead, they are receiving half the memory at a higher entry cost than last year’s full-capacity model. The 128GB alternative has moved into another spending category.

NVIDIA presents the 64GB edition as an accessible starting point. That description makes sense only within the current GB10 market, not across the product’s full history.

The company has a reasonable supply-side explanation. LPDDR5X packages are soldered into these compact systems, and high-capacity configurations require many high-density components.

NAND storage also affects costs. Unlike a conventional workstation, DGX Spark does not let buyers install cheaper memory modules after purchase.

Component prices do not fully explain every retail movement. Manufacturer margins, inventory timing, demand, and channel availability also influence listed prices.

NVIDIA has not publicly provided a bill-of-materials breakdown for either capacity. Buyers therefore cannot isolate how much of the increase comes directly from memory.

Memory suppliers nevertheless describe a tight market. Micron has warned that demand will continue exceeding supply through 2027 and 2028.

That outlook matters because AI infrastructure competes for manufacturing capacity across several memory categories. Suppliers naturally prioritize products and customers offering better returns.

Compact AI systems occupy an uncomfortable position. They need enough memory to differentiate themselves from ordinary PCs, but they lack data-center purchasing scale.

The NVIDIA DGX Spark price increase is therefore more than a temporary promotion ending. It reflects pressure on the component that defines the product.

This pressure also changes purchasing calculations. A team considering several Sparks must now compare them against workstations, rented accelerators, and centralized inference servers.

Cloud computing remains expensive for persistent workloads. Yet it avoids the risk of buying a fixed-capacity system shortly before models or requirements grow.

Local hardware offers privacy, predictable availability, and direct control. Its financial advantage depends on utilization, model fit, energy, support, and the buyer’s development schedule.

A workstation sitting idle has poor economics. A Spark running an internal coding or research agent around the clock has a stronger case.

The new price structure makes workload definition essential. Buyers can no longer justify the system mainly because it offers an unusually large memory pool at an aggressive launch price.

DGX Spark 64GB vs 128GB Is Really a Workload Decision

The correct capacity depends on context length, concurrency, fine-tuning needs, and model choice, not parameter count alone.

NVIDIA argues that smaller open models have become capable enough for valuable local work. It points to models in the 26-billion to 35-billion parameter range.

Those models can support coding, document analysis, research assistance, and other agent tasks. Quantization can reduce their memory footprint further by storing weights at lower precision.

The company says one 64GB Spark can support models with up to 100 billion parameters. That figure describes an upper boundary under selected conditions.

A smaller model often produces a better working experience. It can leave memory for long prompts, concurrent requests, tools, embeddings, and operating-system processes.

Consider a developer running a coding agent continuously. The system must hold the model while indexing repositories, processing documentation, and preserving an expanding conversation.

A 64GB pool can handle that workflow with a suitable model. It becomes restrictive when the developer increases context length or launches multiple agents.

Document analysis creates similar pressure. A research agent may need to process thousands of tokens from reports, notes, and source files.

The model weights are only one part of the allocation. The inference engine, cache, document-processing tools, and application services all compete for memory.

Image generation and multimodal workloads add another layer. Models may combine text encoders, image decoders, vision components, and temporary tensors.

Fine-tuning is even more demanding. Training or adapting a model requires memory for parameters, gradients, optimizer state, and intermediate activations.

Techniques such as low-rank adaptation reduce that burden. They do not eliminate it, especially when teams want larger models, longer sequences, or bigger batches.

That is why the DGX Spark 64GB vs 128GB choice cannot rely on NVIDIA’s maximum model labels. Buyers need measurements from their intended software.

The 128GB configuration offers more room for experimentation. It can accommodate larger models, wider context windows, and concurrent workloads without immediate clustering.

Capacity also reduces operational friction. Developers spend less time changing quantization levels, unloading services, or splitting work between machines.

The 64GB system fits a more controlled environment. It suits teams that standardize on a known model and understand the model’s runtime footprint.

Inference-only deployments can be particularly suitable. Once a team selects and tests a model, it can optimize the serving stack around predictable traffic.

Fine-tuning teams face a harder decision. They should treat 64GB as a limit to validate, not as a general substitute for the original Spark.

NVIDIA’s application examples include coding agents, research agents, image generation, and remote model serving. These are categories, not standardized workloads.

A coding agent reading a small repository differs from one analyzing millions of lines. A single-user research assistant differs from a shared departmental service.

The useful question is not whether 64GB runs local AI. It clearly does.

The question is whether the complete target workload runs with enough spare capacity for growth. That answer needs testing with actual models and data.

Teams should record peak memory use, prompt-processing speed, generation speed, and behavior under concurrent requests. They should also test their longest realistic context.

Software compatibility strengthens NVIDIA’s position. CUDA remains the preferred platform for many AI libraries, and developers frequently optimize new tools for NVIDIA first.

That advantage can compensate for lower memory or a higher acquisition cost. It matters most when a competing system requires unsupported kernels or manual fixes.

Yet CUDA compatibility cannot create missing capacity. Once a workload exceeds physical memory, the buyer must use a smaller model, distribute the workload, or choose another system.

NVIDIA Turns Two Smaller Systems Into Its Upgrade Path

Clustering offers more compute and pooled memory, but it does not recreate a single 128GB machine at a lower cost.

Every DGX Spark includes a ConnectX-7 network interface. Two systems can connect directly through a QSFP cable and form a high-speed cluster.

NVIDIA Sync Cluster Assistant detects the machines, validates their configurations, and prepares the ConnectX-7 connection. That removes several manual networking steps.

A forthcoming Model Launcher will simplify deploying supported models across one or two nodes. NVIDIA says Qwen3.8 27B will be among the initial options.

This software work addresses a real weakness. Distributed local inference previously required command-line configuration and changes to vLLM or TensorRT-LLM launch settings.

NVIDIA says two 64GB systems pool their memory to support models containing up to 200 billion parameters. They also provide twice the aggregate memory bandwidth.

In NVIDIA’s Qwen3.8 27B test, a two-node cluster delivered up to 1.7 times the performance of one 64GB system.

That result is vendor-generated and has not received broad independent validation. It should not be interpreted as a universal scaling ratio.

Distributed performance depends on how frequently a workload moves data between nodes. Network transfers add latency even when the interconnect is fast.

Some models divide cleanly across two processors. Others lose more performance to synchronization, communication, or software overhead.

The cluster also doubles compute, storage, networking hardware, and physical systems. It is not merely a method for combining two memory banks.

That makes the design valuable for throughput. A team might serve more requests, run several agents, or assign separate tasks to each device.

However, buying two 64GB systems costs much more than purchasing one newly priced 128GB system. The cluster provides additional compute, but not a cheaper path to equal capacity.

The upgrade pitch is strongest when needs grow gradually. A developer can begin with one node, confirm demand, and add another without replacing the original machine.

It is weaker when the workload already requires more than 64GB. That buyer would accept immediate complexity and higher total spending to avoid the full-capacity model.

Power use and administration also grow with node count. Teams must monitor two operating systems, two storage devices, network health, and distributed software behavior.

NVIDIA Sync reduces setup work. It does not erase the operational differences between one computer and a cluster.

The company says applications can access the clustered model from ordinary laptops and desktops. This creates a small shared inference appliance for a team.

One node could host a private coding model while employees use browsers or familiar development tools. Sensitive data could remain within the organization’s network.

This scenario shows why DGX Spark is not simply a premium mini PC. NVIDIA is packaging data-center software, networking, and deployment patterns for smaller environments.

The approach also protects NVIDIA’s platform strategy. A customer who adds a second Spark increases dependence on DGX OS, CUDA, ConnectX, and NVIDIA’s management tools.

The NVIDIA DGX Spark 64GB therefore works as both a product and an expansion unit. Its long-term value depends on how well the cluster behaves outside curated demonstrations.

Independent tests should measure first-token latency, sustained generation, concurrent users, power draw, and failure recovery across two machines.

Until those tests arrive, NVIDIA’s 1.7-times result is a useful ceiling rather than an expected outcome for every model.

AMD and Apple Put Pressure on NVIDIA’s Memory Tradeoff

NVIDIA leads with CUDA integration and AI networking, while rivals can offer more memory or broader desktop flexibility.

AMD has expanded its response to compact local AI through Ryzen AI Halo systems. These machines use unified memory and integrated Radeon graphics.

An AMD developer system announced in 2026 included 128GB of LPDDR5X memory and support for Linux or Windows. Its initial positioning undercut NVIDIA’s preceding Spark price.

The AMD system comparison illustrates the choice. AMD offers familiar PC flexibility, while NVIDIA provides a more established AI software environment.

Ryzen AI Max systems from several manufacturers also target local model users. Some offer 128GB configurations, standard PC features, and replaceable storage.

Higher-capacity AMD platforms extend that competition further. They can prioritize large memory pools for users willing to trade some AI performance or software convenience.

Memory capacity alone does not settle the comparison. Reviews have often found GB10 faster in GPU-based AI work than contemporary integrated Radeon alternatives.

CUDA remains another advantage. Many frameworks, model packages, and optimization projects publish NVIDIA support before equivalent AMD paths mature.

Developers who need one specific CUDA-dependent library may have no practical substitute. Saving money on hardware brings little value if the required software cannot run correctly.

AMD’s Windows support can matter just as much for another group. Creative workers and developers may want local AI without maintaining a dedicated Linux appliance.

Apple provides a third route. Mac Studio systems combine high memory capacities, efficient processors, and a mature desktop operating system.

Apple’s unified memory allows large models to load without a discrete GPU’s usual video-memory ceiling. Tools such as MLX target Apple silicon directly.

Yet software availability and model performance vary. CUDA-focused workflows do not move automatically to Apple’s Metal or MLX environments.

Mac systems also lack DGX Spark’s built-in ConnectX-7 fabric. They are not designed around NVIDIA’s direct two-node local inference path.

The competitive decision therefore has several dimensions.

Memory capacity

  • NVIDIA 64GB systems: Better suited to defined models and controlled inference workloads.

  • NVIDIA 128GB systems: More headroom, but the latest increase weakens their original value.

  • AMD and Apple systems: Selected configurations can provide larger memory pools with different performance and software tradeoffs.

Software compatibility

  • NVIDIA: Broad CUDA support and a preconfigured AI stack.

  • AMD: Improving ROCm support, plus conventional Windows and Linux options.

  • Apple: Strong native applications and MLX tooling, but limited CUDA compatibility.

Expansion strategy

  • NVIDIA: Direct ConnectX-7 clustering with assisted configuration.

  • AMD: More conventional workstation and network deployment choices.

  • Apple: High single-system capacity, without an equivalent Spark clustering design.

General computing

  • NVIDIA: DGX OS focuses the system on AI development and inference.

  • AMD: Functions more like a standard high-end PC.

  • Apple: Combines local AI with an established creative and productivity platform.

The DGX Spark 64GB vs 128GB decision must therefore include alternatives. Buyers should avoid treating GB10 as the only path to private local models.

NVIDIA’s strongest defense is integration. It combines silicon, CUDA, optimized runtimes, networking, and system management under one supported design.

Its vulnerability is the value of memory. If buyers primarily need capacity, competitors can challenge NVIDIA without matching every GB10 performance result.

The latest pricing widens that opening. A buyer can tolerate slower inference if an alternative holds the required model on one machine.

Conversely, a smaller model running faster through CUDA may outperform a larger but slower deployment in daily use. Real workloads should decide the winner.

What Buyers Should Watch After the 64GB Launch

Three signals will show whether the smaller Spark becomes a practical platform or an expensive response to memory scarcity.

The first signal is partner availability on and after October 23. NVIDIA has named six manufacturers, but their exact configurations and shipping volumes will matter.

A nominal starting configuration offers little value if it remains unavailable. Buyers should track delivered systems, not announcement pages or backordered listings.

Partner specifications also need careful comparison. Storage capacity, port selection, regional availability, warranty service, and bundled software can change the real value.

If several manufacturers maintain reliable stock, the 64GB launch will strengthen NVIDIA’s claim that it created a usable entry point.

If inventory remains scarce, the product will look more like a pricing reference than a broadly accessible development machine.

The second signal is independent workload testing. Reviews should compare one 64GB Spark, one 128GB Spark, and a two-node 64GB cluster.

Model loading is not enough. Tests need long contexts, simultaneous users, agent tool calls, fine-tuning workloads, and sustained operation.

The most important measurements include usable memory, time to first token, output speed, energy consumption, and cluster scaling efficiency.

NVIDIA’s maximum parameter counts should receive special scrutiny. A model that technically fits can still deliver an impractical experience.

Independent testing can also reveal whether unchanged bandwidth preserves most single-node performance. Capacity-constrained workloads should behave differently from bandwidth-constrained ones.

If the smaller machine performs predictably with popular 27-billion to 35-billion parameter models, NVIDIA’s positioning will gain support.

If ordinary contexts or concurrent agents exhaust memory, the 64GB version will serve a much narrower audience than its marketing suggests.

The third signal is competitive response. AMD system builders and Apple can pressure NVIDIA through memory capacity, software support, or lower total ownership costs.

AMD has room to improve ROCm compatibility and promote Windows-based local AI. Better support from model runners would weaken CUDA’s practical lock-in.

Apple can emphasize large unified-memory configurations and efficient local inference. Its opportunity is strongest among developers already using macOS.

NVIDIA can respond with better software rather than another hardware change. Sync Cluster Assistant and Model Launcher are early examples of that strategy.

Watch how many models receive one-click deployment profiles. Also watch whether developers can customize those profiles without returning to complex manual configurations.

The memory market remains the external variable. Continued shortages would normalize higher prices and encourage manufacturers to favor smaller configurations.

Improved supply would test whether recent increases are temporary. It could also restore pressure on NVIDIA to offer more memory at the entry level.

For enterprise buyers, the immediate action is straightforward. Define the workload before choosing the machine.

Test the intended model, context length, user count, and fine-tuning method. Include framework compatibility and operating costs in the comparison.

The NVIDIA DGX Spark 64GB offers the same GB10 foundation with a tighter memory ceiling. That can work for focused inference and agent deployments.

It should not be treated as a universal replacement for the original 128GB system. The 128GB model remains more flexible, but its higher price demands stronger utilization.

Will your workload benefit more from CUDA performance and NVIDIA’s managed stack, or from the largest memory pool one system can provide? Run that test before committing to either Spark configuration.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page