NVIDIA cuObject Opens AI Storage Beyond Files, but Interoperability Is Still Unfinished
NVIDIA made its cuObject client and server libraries generally available on September 30, extending accelerated AI storage access beyond conventional file systems. The NVIDIA cuObject release adds standardized APIs and an RDMA wire protocol for moving object data without routing payloads through a server CPU.
The company also introduced a SCADA Server SDK for handling storage requests initiated by GPUs. SCADA means Scaled Accelerated Data Access, a framework designed for high volumes of fine-grained storage operations. IBM has already built an early Storage Scale prototype with the SDK.
The announcement targets a persistent divide in AI infrastructure. Object storage offers capacity and familiar S3-compatible interfaces, while high-performance training and inference pipelines often rely on faster file systems or local scratch storage. NVIDIA wants cuObject to narrow that divide, but its larger interoperability plan remains unfinished.
NVIDIA cuObject Moves From a Product Library to an Industry Proposal
The important change is not only that NVIDIA released another storage library. It is asking cloud and storage providers to converge on a common accelerated object protocol.
According to the NVIDIA announcement, both cuObject client and server libraries are now generally available. NVIDIA also released cuObject Server 2.0.0 for storage providers evaluating the server-side integration.
The client library sits inside a GPU application, data loader, or middleware layer. It connects application-level object operations to an RDMA data path. Remote direct memory access, or RDMA, lets network hardware transfer data directly between registered memory regions.
The server library integrates with an object storage service. It manages registered buffers and performs the RDMA operations that move data to or from the client. The system can target either GPU memory or system memory.
This design addresses a problem that GPUDirect Storage did not fully solve. GPUDirect Storage already provided a direct path between storage and GPU memory, but its best-known interface, cuFile, centered on file access. Many AI datasets live behind object interfaces instead.
That difference matters for organizations keeping training samples, documents, images, checkpoints, and generated artifacts in S3-compatible repositories. Those systems are attractive because they scale across large namespaces and separate storage from compute. However, conventional object access usually travels through a TCP stack and CPU-managed buffers.
NVIDIA cuObject keeps object-oriented control operations while changing the payload path. A client can still issue familiar GET or PUT requests through an adapted S3 software development kit. The object data can then move over RDMA instead of following the ordinary TCP data path.
General availability gives developers a supported starting point, but NVIDIA is pursuing a broader objective through xio-sig. The group is expanding from file access into object access, with separate work planned for client APIs, a wire protocol, and conformance testing.
Google Cloud is evaluating expanded participation around cuObject. Microsoft has also said it plans to join the xio-sig board. IBM’s SCADA prototype adds a storage vendor to the early group, although a prototype is not equivalent to production support.
The distinction is central to the story. NVIDIA now has downloadable components, documentation, and partners discussing interoperability. It does not yet have a mature, multi-vendor object-storage standard with several proven production implementations.
That makes the release both concrete and provisional. Developers can begin evaluating the libraries now. The larger promise depends on cloud platforms, storage vendors, frameworks, and application developers adopting the same interfaces.
How NVIDIA cuObject Separates Control From Data
NVIDIA cuObject keeps S3-style control messages on a familiar path while moving object payloads through RDMA.
The cuObject architecture divides each operation into a control plane and a data plane. Standard S3 GET and PUT requests remain part of the control flow. An adapted S3 SDK adds metadata that describes the RDMA transfer.
The object payload follows a different route. The client registers a memory region and creates an RDMA token containing the information needed for the transfer. That token is attached to the HTTP request through a custom header.
The storage gateway parses the request and directs an appropriate data node to perform the operation. That node registers its local buffer through the cuObject server APIs. It then pushes or pulls the payload using an RDMA write or read.
A successful gateway response completes the control transaction after the RDMA operation finishes. This split lets the application retain object semantics while removing the server CPU from the primary payload path.
The approach targets more than raw bandwidth. CPU processing can become a constraint when many accelerators generate concurrent storage operations. Avoiding TCP processing for each payload can preserve CPU capacity for metadata, coordination, security, and other services.
The current implementation uses Dynamically Connected transport, commonly called DC. DC avoids maintaining a reliable connection between every client and storage server pair. That characteristic is useful when a large compute cluster reaches many storage nodes.
NVIDIA documents DC support over InfiniBand and RoCEv2. RoCEv2 carries RDMA traffic over Ethernet networks configured to support the required behavior. Both options demand infrastructure planning beyond installing a library.
The client also does not interact with an unchanged S3 endpoint. NVIDIA’s documentation says integration requires modifications to the client-side S3 SDK and the server-side storage software. Those changes create the accelerated path and exchange RDMA metadata.
Supported operations include GET, PUT, multipart upload operations, and byte-range reads. Range access is relevant when an application needs only part of a large object, rather than transferring the entire object into GPU memory.
This architecture can remove a common staging step. Traditional AI pipelines often copy data from an object repository into a local or distributed scratch file system. Compute nodes then read from that intermediate layer.
Staging can provide predictable performance, but it consumes capacity and adds operational work. Teams must schedule copies, monitor synchronization, clear stale data, and decide which datasets deserve fast storage.
NVIDIA cuObject proposes a flatter path from object storage to accelerator memory. If supported end to end, an application can read from its primary object repository without first creating another full copy.
That benefit does not eliminate every intermediate operation. Data may still require decoding, decompression, validation, batching, or transformation. Storage software may also use internal buffers before delivering a payload through RDMA.
The release therefore changes the transport opportunity, not the entire data pipeline. Applications still need to manage formats, access permissions, object metadata, and failure recovery. Storage vendors must connect the accelerated path to their own placement and durability systems.
NVIDIA cuObject works best as an integration layer between those components. It does not replace the object store, S3 control plane, data loader, or networking fabric.
The SCADA Server SDK Shifts Request Control Toward GPUs
The SCADA Server SDK addresses a different bottleneck: many small requests whose coordination can overwhelm a CPU before storage reaches its limits.
Large sequential transfers make bandwidth the obvious metric. Training jobs reading sizable samples or checkpoints can amortize request overhead across each transfer. The control work becomes a smaller part of the total operation.
Inference and retrieval systems create another pattern. Semantic search, recommendation, fraud detection, and agent memory can generate many smaller lookups. A GPU may have enough parallel work to issue these operations at high concurrency.
Conventional storage software expects a CPU to construct, submit, and complete those requests. That design becomes less efficient as request sizes shrink and operation counts rise. Fixed processing costs consume a larger share of every transaction.
SCADA lets GPU-based clients initiate storage requests through a common interface. The new SDK gives third-party storage providers a way to build servers that receive those requests. A server can fulfill them from local or remote storage before returning data through RDMA.
The SDK does not require every storage system to expose its internal design directly to the GPU. Instead, a SCADA server acts as the bridge between a shared client interface and vendor-specific storage logic.
IBM has demonstrated an initial server based on the SDK for IBM Storage Scale. In that prototype, a SCADA client sends requests to an IBM-built server. The server connects the common request model to Storage Scale.
This is an early interoperability signal, but it remains limited. NVIDIA has not published independent production results for the IBM prototype. The announcement also does not provide comparative latency, throughput, or CPU-utilization measurements.
SCADA includes supporting work beyond the server SDK. NVIDIA has published a Storage Lender Service for provisioning access to NVMe queues. A command-line utility is also intended to help configure and deploy SCADA components.
The Storage-Next initiative supplies the wider industry context. NVIDIA says the group includes more than 40 vendors and customers across flash, controllers, storage systems, cloud infrastructure, and application development.
That group is attempting to define how GPU-driven storage should behave. Its work addresses both large data movement and small, fine-grained operations. The goal is to turn vendor collaboration into interoperable interfaces and standards.
The timing reflects changing AI data access patterns. Training remains important, but production inference introduces context retrieval, tool calls, database lookups, and persistent memory. Each user request can trigger several downstream data operations.
A long-context service also creates pressure outside GPU memory. Key-value cache data records attention state from previously processed tokens. When that state cannot remain in accelerator memory, systems need another tier that can return it efficiently.
Storage is cheaper and more capacious than accelerator memory, but it has different latency characteristics. SCADA attempts to let GPUs tolerate that gap through parallelism. Many GPU threads can keep operations in flight rather than waiting for a CPU-driven request sequence.
The idea is complementary to cuObject rather than a replacement for it. NVIDIA cuObject focuses on accelerated access to S3-compatible object storage. SCADA focuses on GPU-initiated, fine-grained access across storage implementations.
Together, they extend NVIDIA’s influence from compute into the protocols connecting accelerators and stored data. That expansion creates the main competitive pressure behind the announcement.
Open Interfaces Compete With Provider-Specific Acceleration
The primary contest is between shared accelerator-storage interfaces and separate integrations built for each cloud or storage provider.
Object storage over RDMA has lacked a broadly adopted common wire protocol. A developer seeking direct GPU access could rely on traditional S3 transfers, use an intermediate file system, or build around a specific provider’s accelerated path.
Each option imposes a different cost. Traditional access preserves compatibility but retains CPU and TCP overhead. Staging can improve locality while duplicating data. A provider-specific integration can perform well but creates another dependency.
NVIDIA wants xio-sig to reduce that fragmentation. The xio-sig organization describes separate cuObject efforts for a client API, an accelerated wire protocol, and a conformance suite. The same organization also houses related cuFile work.
Conformance is critical because publishing an interface does not guarantee compatible behavior. Implementations must agree on request semantics, memory registration, error handling, security boundaries, and fallback behavior. They must also behave consistently under failure and concurrency.
At publication time, the public organization says code will appear after the founding participants integrate and validate the layers. Its status page also says the stacks must pass conformance tests before that publication occurs.
That leaves a meaningful gap between NVIDIA’s announcement and the finished interoperability layer. The repository structure exists, and the intended components are named. Much of the production-ready implementation is still pending public release.
Google Cloud and Microsoft bring important credibility because cloud providers operate large object-storage platforms. Their participation also tests whether xio-sig can accommodate infrastructure not controlled by NVIDIA.
Yet the partners have made different commitments. Google Cloud is evaluating wider participation for cuObject. Microsoft has indicated an intention to join the board. Neither statement alone confirms general customer availability in their object services.
IBM’s prototype provides a more tangible implementation, but it concerns SCADA and Storage Scale. It does not prove that several independent S3-compatible platforms can exchange cuObject traffic through a single production-tested protocol.
Storage companies also have reasons to preserve their own acceleration methods. Vendor-specific paths can expose differentiated caching, placement, security, or data services. A shared protocol must remain broad enough for interoperability without erasing those features.
NVIDIA has its own strategic interest. A common path from storage to GPU memory can make NVIDIA accelerators easier to use with larger datasets. It can also pull networking products, DPUs, CUDA software, and storage partners into a coordinated architecture.
That does not make the interoperability push inherently closed. The published organization uses an Apache-2.0 license for its current repository, and conformance work can lower integration risk. However, governance and implementation details will determine how open the result becomes.
Other accelerator vendors create another test. NVIDIA’s post refers to GPUs, TPUs, and XPUs when describing compute demand. A genuinely portable storage interface should not depend on one accelerator architecture at every layer.
The clearest evidence will come from non-NVIDIA implementations passing shared tests. Support inside common frameworks would matter as well. Developers care less about organizational membership than whether an existing application can change backends without a rewrite.
For storage buyers, the practical question is portability. An interface has value when it preserves application behavior across several supported products. One fast path tied to one validated combination remains an integration, not an industry standard.
General Availability Does Not Remove Deployment Risk
The libraries are available, but production adoption still requires specialized networks, modified software, careful memory handling, and credible benchmarks.
The cuObject release notes show that the client reached version 1.3.0 in August 2026. Earlier releases added multipath failover, failback, and IPv6 support. Version 1.3.0 added a method for invalidating stale RDMA tokens.
Those additions address operational concerns, but the same documentation lists important limitations. A single memory-registration call has a maximum below 4 GiB. Concurrent GET and PUT operations are not supported on the same registered buffer.
Host-memory transfers require registered buffers. Some configuration behavior differs from cuFile, including the absence of a thread pool for issuing cuObject I/O. Applications must understand these constraints before adopting the path.
Memory lifetimes require particular care. A client must not reuse or deregister a buffer while an operation remains outstanding. Error handling must also prevent stale requests from accessing a memory region after its key is reused.
These requirements are not unusual for high-performance RDMA software. They still shift responsibility toward application, framework, and storage developers. An incorrect integration can produce failures that are difficult to diagnose.
The network also matters. DC transport over InfiniBand or RoCEv2 assumes suitable adapters, switches, routing, and configuration. Organizations cannot expect the accelerated path to appear across an ordinary network without infrastructure work.
RoCE deployments can be sensitive to congestion and fabric design. Multipath behavior, failure recovery, and telemetry must be tested under realistic traffic. A successful laboratory transfer does not establish predictable cluster-wide performance.
Security deserves equal attention. Direct data movement reduces CPU involvement in the payload path, but it cannot bypass authorization. Systems still need a trusted control layer that decides which process can access each registered region and stored object.
NVIDIA describes SCADA as separating unprivileged application work from a privileged setup component. That structure can protect the data path when implemented correctly. Storage providers must still connect it to tenant isolation, auditing, credential management, and revocation.
The announcement contains no standardized benchmark comparing cuObject with conventional S3, staged file access, or provider-specific alternatives. It also does not quantify gains for the IBM prototype.
That omission prevents broad performance conclusions. RDMA can reduce copies and CPU processing, but application outcomes depend on object sizes, access patterns, storage media, network topology, concurrency, and preprocessing.
Large sequential reads may already perform adequately through optimized file systems. Very small requests may expose limits elsewhere, including flash translation layers, metadata services, or application synchronization.
The independent storage analysis surrounding SCADA emphasizes that distinction. Bulk transfers and fine-grained reads impose different requirements, so no single headline throughput figure can describe both.
Developers should therefore treat NVIDIA cuObject as a path to evaluate, not a guaranteed performance result. Tests should use real object sizes, representative concurrency, existing transformations, and expected failure scenarios.
A credible proof of concept should measure more than bandwidth. It should track tail latency, server CPU consumption, GPU utilization, memory-registration overhead, recovery time, and behavior when the RDMA path becomes unavailable.
Teams should also verify fallback behavior. A production application needs a defined response when a server lacks RDMA support or a fabric component fails. Compatibility with ordinary object access can be as important as peak accelerated performance.
The strongest skepticism concerns adoption rather than technical possibility. NVIDIA has shown that the components can be built. It has not yet shown that a wide range of providers will maintain compatible implementations through several release cycles.
What to Watch After the NVIDIA SCADA Server SDK Release
Three signals will show whether NVIDIA cuObject becomes shared infrastructure or remains a set of optimized partner integrations.
The first signal is public xio-sig code accompanied by working conformance tests. The organization has identified repositories for the cuObject client API, wire protocol, and conformance suite. Those repositories need substantive implementations rather than interface descriptions alone.
Passing tests across independently maintained clients and servers would strengthen NVIDIA’s interoperability claim. Delays, narrow test coverage, or dependencies on one hardware combination would weaken it.
The second signal is production support from cloud and storage providers. Google Cloud’s evaluation and Microsoft’s planned board participation are meaningful, but customer-facing availability would matter more.
Buyers should look for supported service combinations, documented deployment requirements, and clear compatibility matrices. More SCADA server implementations beyond IBM Storage Scale would also test whether the SDK generalizes across storage designs.
The third signal is workload-specific evidence. Vendors need to publish reproducible results for training ingestion, checkpoint operations, retrieval, semantic search, and inference cache access. Those tests should compare standard object transfers, staging workflows, and RDMA-enabled paths.
Results should include CPU consumption and tail latency, not only peak throughput. They should also disclose object sizes, network configuration, storage media, and failure behavior. Without that context, performance numbers will be difficult to apply.
Developers do not need to wait before studying the architecture. They can identify where their applications stage object data, profile CPU time in storage transfers, and measure the request-size distribution. That work reveals whether cuObject or SCADA addresses a real bottleneck.
The final question is not whether RDMA can move data faster. It is whether several providers can expose one dependable path without trapping applications inside a narrow stack. Watch the conformance repositories, supported vendor products, and reproducible workload tests. Those signals will determine whether NVIDIA cuObject becomes a portable AI storage layer or another specialized acceleration option.



