top of page

GLM-5.3 Sparse Attention Cuts Compute, but HBM Demand Persists

14 hours ago
12 min read

GLM-5.3 sparse attention reduces how much context each attention operation reads, yet it does not automatically remove the model’s HBM capacity problem. A September 28 analysis found that selecting relevant tokens can still require access to the complete context history. The result challenges a tempting assumption: fewer attended tokens must mean proportionally less GPU memory.

That distinction matters because GLM-5.3 targets long-running coding and agent tasks, where context accumulates across tool calls, files, test results, and revisions. Sparse attention lowers work inside the main attention calculation. The KV cache, which stores reusable key and value representations from earlier tokens, can continue growing with the full sequence.

DeepSeek Sparse Attention provides the architectural reference point. Its indexer identifies a limited set of useful tokens before the model performs its main attention calculation. GLM-5.3 combines that method with compressed cache representations, cross-layer index reuse, and serving software that can move inactive cache entries into host memory.

The central contest is therefore not sparse attention versus dense attention alone. It is algorithmic sparsity versus the physical requirement to keep a long conversation available. That contest determines whether memory savings appear as lower HBM capacity, lower bandwidth demand, higher concurrency, or simply a different balance between GPUs and system memory.

What GLM-5.3 Sparse Attention Actually Changes

GLM-5.3 reduces the amount of expensive attention work per token, but the model still needs a way to locate relevant information across its history.

GLM-5.3 belongs to Z.ai’s GLM-5 family, which uses a mixture-of-experts architecture. A mixture-of-experts model activates only part of its total parameter set for each token. According to the GLM-5 report, the family combines this routed computation with a long-context attention design derived from DeepSeek’s work.

Z.ai says GLM-5.3 uses the same base model as GLM-5.2. Its reported capability gains come from post-training rather than another round of base-model pretraining. That distinction means the current memory discussion concerns how the existing architecture behaves under real serving loads, not a newly invented GLM-5.3 attention layer.

The relevant mechanism is DeepSeek Sparse Attention, or DSA. Sparse attention limits full attention to a selected subset of earlier tokens instead of processing every previous token at equal cost. DeepSeek introduced the production-oriented design in its V3.2 research.

DSA first runs a lightweight component called a lightning indexer. The indexer scores earlier positions and chooses the top-k tokens, meaning the limited set judged most relevant to the current query. The main Multi-Head Latent Attention calculation then operates on those selected positions.

Multi-Head Latent Attention, or MLA, stores compressed latent representations instead of a separate full key and value vector for every attention head. This compression reduces the cache stored per token. Sparse selection then reduces how much of that cache the main attention operation reads.

These are two different savings. MLA targets the size of each token’s cached state. DSA targets the number of cached positions consumed by the expensive attention calculation.

The combination changes the compute and bandwidth profile substantially. Core attention can move from processing relationships across the full context toward processing a fixed top-k selection. At long sequence lengths, this restrains the growth of the main attention workload.

However, the lightning indexer still needs enough information to score candidates across the retained history. The model cannot select an old token if the serving system has discarded every usable representation of that token.

The SemiAnalysis examination identifies this as the crucial limit. Sparse attention reduces memory traffic during the core scaled dot-product attention operation. It does not necessarily reduce the total memory capacity required to preserve selectable context.

That limit becomes clearer below the sparse threshold. DSA uses a top-k setting of 2,048 positions in the configuration discussed by SemiAnalysis. A sequence containing fewer positions has no larger pool to prune, so attention remains dense.

Serving engines also choose different execution modes according to sequence length and deployment topology. An implementation can favor a lower-compute mode for shorter contexts, then shift toward a lower-memory mode as memory traffic becomes dominant. Sparse attention is therefore not one fixed speedup across every request.

The meaningful change is narrower and more useful. GLM-5.3 sparse attention reduces the recurring cost of consulting a long history. It does not make that history cease to exist.

Why Lower Attention Traffic Does Not Equal Lower HBM Capacity

HBM pressure comes from retained context, while sparse attention primarily changes which retained entries the GPU reads during each operation.

High-bandwidth memory, or HBM, is the fast memory attached directly to an accelerator. Its bandwidth helps GPUs feed large matrix operations, while its limited capacity constrains how many models and active requests fit on each device.

During autoregressive generation, a model produces one token after another. It reuses the keys and values calculated for earlier tokens through the KV cache. Without that cache, the server would repeatedly recompute the entire preceding sequence.

Each active request therefore reserves memory for its context. A long coding session might include repository files, command output, patch attempts, test logs, and earlier reasoning. An agent can produce many more tokens than an ordinary question-and-answer exchange.

Sparse attention changes the read pattern. Instead of loading every historical position into the main attention operation, the model loads the selected top-k set. That can reduce memory bandwidth consumption and the compute performed after selection.

Capacity follows a different rule. If any earlier position remains eligible for selection, its representation must remain accessible somewhere. A conventional serving design keeps the complete KV history in HBM, even when the main attention kernel reads only a small subset.

The resulting system can become capacity-bound before it becomes compute-bound. Each request may perform less attention work, yet still occupy memory proportional to its context length. Increasing concurrency then places more complete histories on the same device.

This explains why sparse attention does not translate directly into an equivalent reduction in HBM demand. The system saves active traffic without necessarily reducing resident state. The model’s logical view of its history remains complete, even when each attention step is selective.

The difference resembles a large archive with a fast retrieval system. Faster retrieval reduces how many documents someone reads for each question. It does not shrink the archive unless older documents move elsewhere or disappear.

GLM-5.3 KV cache compression still matters. Smaller per-token representations let more context fit within a given memory budget. They also reduce the bytes transferred when selected entries enter an attention operation.

However, compressed state continues to accumulate with sequence length. A smaller linear memory curve is still a linear memory curve. Long contexts and many simultaneous requests can eventually consume the saved capacity.

Concurrency exposes the tradeoff quickly. SemiAnalysis reported results where increasing concurrent requests from eight to sixteen reduced prompt-token reuse from GPU memory. The GPU reuse share fell from 90.3 percent to 54.8 percent.

Host-memory reuse rose from 6.0 percent to 40.3 percent in the same comparison. The combined cache hit rate remained above 95 percent at every reported concurrency level. These results show that useful cache capacity can extend beyond the accelerator.

They do not mean host memory matches HBM latency. Moving data across the CPU-GPU connection creates an I/O cost, and cache misses can interrupt an otherwise efficient decode path. The serving system must predict, fetch, and evict data without letting transfers dominate generation time.

The memory market implication is also more nuanced than a simple fall in demand. Sparse attention can reduce HBM traffic per attention step. At the same time, cheaper long-context inference can encourage longer sessions and greater request concurrency.

That rebound matters for infrastructure planning. When each request becomes less expensive to process, operators often admit more simultaneous work. Saved memory bandwidth can become additional throughput instead of unused hardware.

HBM demand can therefore persist even as attention becomes more selective. Host DRAM demand can also increase because complete histories move into a larger, slower memory tier. At still larger scales, storage systems may absorb reusable prefixes or inactive cache data.

The practical question is no longer whether sparse attention saves memory in the abstract. It is which memory tier holds each part of the GLM-5.3 KV cache, and how frequently the serving engine moves it.

HiSparse Moves the Full History Out of the GPU

HiSparse converts sparse attention’s selective reads into actual HBM capacity savings by separating logical cache availability from physical GPU residency.

The SGLang team designed HiSparse as a hierarchical KV cache for sparse-attention serving. It keeps a small working set on the GPU while storing the complete KV history in pinned host memory. Pinned memory is CPU memory prepared for predictable transfers to an accelerator.

Under this design, old cache entries remain logically available to GLM-5.3. They do not all remain physically resident in HBM. The indexer can select a position, and the serving system can retrieve that position when the GPU lacks it.

HiSparse uses a least-recently-used policy for its device cache. When selected tokens are absent from HBM, the system loads them from host memory. It evicts less recently used entries to keep the GPU working set bounded.

This architecture turns a model-level property into a system-level saving. Sparse attention identifies the small set required by the current operation. HiSparse ensures that only a limited selection and working buffer must occupy HBM during decoding.

The HiSparse paper describes the system as exact and indexer-agnostic. Exact means cache placement changes without intentionally approximating the model’s selected attention output. Indexer-agnostic means the memory manager does not depend on one selection algorithm.

Its evaluations cover DSA, Native Sparse Attention, and Quest on H200, B200, and GH200 platforms. The authors report up to 4.7 times higher peak generation throughput on long-context workloads.

That is a system result under tested configurations, not a guaranteed GLM-5.3 speed multiplier. Workload length, request concurrency, interconnect bandwidth, selection locality, and cache miss rates all affect the outcome.

HiSparse also overlaps transfers with useful computation. While one layer executes, the system can prepare selected cache entries for a later layer. This layer-wise overlap hides part of the latency created by host-to-device movement.

Cross-layer reuse makes that scheduling easier. If adjacent layers select many of the same positions, the system has advance knowledge about likely cache demand. Entries fetched for one layer can remain useful for subsequent layers.

The remaining price is I/O. A selection miss requires data to travel from CPU memory into HBM. Frequent misses, scattered selections, or limited host-device bandwidth can erase part of the throughput gain.

That risk separates theoretical sparsity from production efficiency. A sparse kernel may read fewer entries once they arrive. The complete system still has to find those entries, transfer them, map them into usable pages, and coordinate their lifetime.

Time to first token creates another constraint. Prefill, which processes the initial prompt, has different characteristics from token-by-token decoding. HiSparse primarily targets the decode side, where the cache already exists and grows with continued generation.

SGLang’s implementation pairs HiSparse with prefill-decode disaggregation. That architecture assigns prompt processing and token generation to different workers. Each phase can then use a memory layout and hardware allocation suited to its workload.

The design also changes infrastructure demand. HBM becomes a hot cache rather than the only store for the active conversation. Host DRAM holds the larger history, while the interconnect becomes part of the critical path.

This can reduce the HBM capacity needed for each decoding request. It does not eliminate the bytes representing the conversation. It relocates many of them and adds software responsible for keeping the right subset close to the GPU.

For operators, the relevant metric is therefore not just model size or maximum context length. They need the per-request HBM footprint, host-memory allocation, miss rate, transfer volume, and output-token latency at realistic concurrency.

Sparse attention makes that tiered design possible. HiSparse makes it operational. Neither makes memory management free.

IndexShare Cuts the Cost of Finding Relevant Tokens

Once full attention becomes sparse, the indexer itself becomes a visible bottleneck, so GLM’s next optimization reuses selection decisions across layers.

A standard DSA layer has its own lightning indexer. That component scores historical tokens before the main attention calculation selects its top-k set. The indexer is lighter than full attention, but it still examines the context.

As context grows, repeatedly scoring every historical position across every layer becomes expensive. The main attention path has been reduced, so work that once looked minor represents a larger share of total latency.

Z.ai addresses this issue with IndexShare, also described publicly as IndexCache. Instead of running an independent indexer in every sparse-attention layer, groups of layers reuse a shared selection.

The approach relies on an observed pattern: neighboring layers often choose many of the same historical tokens. The IndexCache study reports 70 to 100 percent overlap between top-k selections from adjacent layers in its analysis.

That overlap creates redundancy. A designated full layer can compute an index, while following shared layers reuse the selected positions. The production pattern discussed for GLM assigns one indexer across groups of four DSA layers.

On a 30-billion-parameter DSA model, the researchers removed up to 75 percent of indexer computations with negligible reported quality degradation. They measured up to 1.82 times faster prefill and 1.48 times faster decoding against standard DSA.

The paper also reports preliminary production-scale GLM-5 results. Those findings support the mechanism, but they do not replace broad independent testing across GLM-5.3 workloads and serving stacks.

Selection reuse introduces its own training requirement. A shared indexer must identify tokens that serve several layers, not simply match one layer’s attention distribution. IndexCache trains retained indexers against an average of the attention distributions they support.

That adjustment matters because consecutive layers are related but not identical. An early layer might prioritize lexical details, while a later layer might favor a dependency formed during intermediate processing. Reuse becomes harmful if it removes a token needed only by one member of the group.

The method therefore exposes a second tradeoff. More sharing removes additional indexer work. Less sharing preserves more layer-specific selection behavior.

IndexShare also interacts with HiSparse. When layers share an index, the serving engine can reuse retrieved cache entries across those layers. Shared selections reduce repeated top-k computation and can make host-to-device fetching more predictable.

This combination attacks three different costs:

  • MLA compresses the representation stored for each token.

  • DSA restricts full attention to selected historical positions.

  • IndexShare avoids recomputing similar selections in every layer.

  • HiSparse moves inactive KV entries from HBM into host memory.

Those components should not be collapsed into one memory claim. Compression affects bytes per token. Sparse attention affects active reads. Index sharing affects selection overhead. Offloading affects physical placement.

Each layer of optimization can move the bottleneck elsewhere. Smaller caches can expose compute overhead. Cheaper main attention can expose indexer latency. Offloading can expose transfer bandwidth. Higher concurrency can expose host-memory capacity.

Hardware characteristics determine which bottleneck appears first. SemiAnalysis estimated an arithmetic-intensity profile suggesting that GLM’s attention configuration differs from DeepSeek’s H800-oriented balance. It also connected GLM’s design to support from Chinese accelerator vendor Moore Threads.

That hardware interpretation remains an inference, not a disclosed Z.ai design target. GLM-5.3 supports multiple serving frameworks and accelerator platforms, so operators should measure the model on their own deployment path.

The wider lesson is that GLM-5.3 sparse attention cannot be evaluated through a single FLOP count. Serving performance comes from the joint behavior of its indexer, compressed cache, memory hierarchy, kernels, and workload.

The Real Test Is Production Memory Efficiency

GLM-5.3 will validate its memory design only if operators can sustain long agent sessions without shifting unacceptable costs into latency, DRAM, or operational complexity.

The first signal to watch is independent GLM-5.3 benchmarking under long contexts and high concurrency. Peak single-request speed reveals little about a service handling many persistent agents. Tests should report HBM use, host DRAM use, cache misses, and latency distributions together.

A convincing result would show that GLM-5.3 KV cache offloading admits more concurrent requests while keeping per-token latency stable. If throughput rises only after accepting large latency spikes, the memory saving has limited value for interactive coding agents.

The second signal is broader deployment support for HiSparse and similar memory managers. SGLang has integrated HiSparse, and vLLM has also documented work around the architecture. Consistent behavior across engines would strengthen the case that sparse models can use bounded HBM residency in production.

Fragmented kernel support would weaken that case. Sparse attention depends on specialized selection, page management, cache formats, and attention kernels. A model can be open-weight while remaining difficult to serve efficiently outside a narrow software stack.

The third signal is evidence about long-horizon agent quality. Memory optimization matters only if the model reliably retrieves earlier requirements, code decisions, and tool results. Selection errors that appear late in a session can be difficult to diagnose.

GLM-5.3’s post-training strategy makes this especially relevant. Z.ai says the model improved coding ability by 50 percent over GLM-5.2 on its internal Z.ai Code Bench. That remains a company-reported comparison.

Z.ai also reports an 84.5 percent result on CyberGym, compared with 77.2 percent for GLM-5.2. CyberGym measures whether a model can find and validate software vulnerabilities from source code. The GLM-5.3 release presents these gains as evidence of stronger agentic and cybersecurity capabilities.

Those capabilities increase both usefulness and risk. Longer tool-driven sessions can support repository analysis, testing, and vulnerability research. The same persistence can help automate exploitation steps or preserve sensitive material inside a serving cache.

Memory placement consequently has a security dimension. Host DRAM, shared prefix caches, and distributed cache layers expand the locations where conversation state can reside. Operators need isolation, eviction, access control, and observability across every tier.

The model’s Single-Rollout Asynchronous Optimization work belongs in this context, although it does not directly reduce inference memory. SAO trains on one rollout per prompt and uses a separate value model to estimate token-level returns.

The SAO paper says the method addresses instability and off-policy effects in asynchronous agent training. It was deployed in GLM-5.2’s agentic training pipeline and informs the post-training lineage behind GLM-5.3.

SAO can improve training efficiency for long, uneven agent trajectories. It also carries extra training overhead because the value model runs alongside the policy model. That is another example of reducing one bottleneck by accepting cost elsewhere.

For enterprise teams, the immediate task is disciplined evaluation. Track the full prompt, selected context, cache placement, miss behavior, output latency, and task success under the same workload. Aggregate tokens-per-second numbers hide too much.

Teams also need durable records of model configurations and serving experiments. A searchable knowledge base can connect benchmark results with kernel versions, cache settings, and deployment incidents.

GLM-5.3 sparse attention changes the economics of reading long contexts. It does not repeal the requirement to preserve them. The architecture lowers active attention traffic, while IndexShare reduces selection overhead and HiSparse limits GPU residency.

The question for the next wave of benchmarks is concrete: can GLM-5.3 turn those savings into sustained concurrency without moving the bottleneck into host transfers or retrieval quality? Watch measured HBM occupancy, cache-miss latency, and long-agent accuracy. Together, those signals will show whether sparse attention delivers a better serving system rather than a better kernel in isolation.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page