top of page

OpenAI Dots Geekbench 7 Results Reveal a Bigger Cloud Computer Than Meta Muse

1 day ago
13 min read

OpenAI Dots Geekbench 7 results suggest each agent receives nine AMD EPYC cores and almost 10GB of memory. That is a notably larger CPU allocation than the two-core environment linked to Meta Muse.

The first reported Dot benchmark scored 1,667 in Geekbench 7’s single-core test and 9,435 in its multi-core test. Six later results used an apparently similar configuration, making the initial screenshot harder to dismiss as an isolated curiosity.

The comparison creates a clear tension. OpenAI appears to be giving its autonomous agents more local computing capacity, but Geekbench cannot measure whether that capacity produces better completed work.

Dots launched at OpenAI’s DevDay on September 29, 2026. OpenAI describes them as persistent agents with a cloud computer, browser, and access to connected applications.

Meta Muse offers a similar autonomous model through smaller reported sandboxes. The early numbers suggest OpenAI has chosen a more resource-intensive approach to the same product problem.

OpenAI Dots Geekbench 7 Results Point to Nine CPU Cores

The available benchmark records consistently describe a capable Linux virtual machine, although they do not identify OpenAI or Dots by name.

The first result appeared publicly through a screenshot shared on X by INIYSA. It showed a Geekbench 7 test uploaded on September 25, four days before OpenAI publicly launched Dots.

The underlying benchmark record reports Ubuntu 24.04.3 LTS and an AMD EPYC 9V74 processor. Geekbench identifies one processor with nine available cores, a 2.60GHz base frequency, and 9.73GB of memory.

The record does not show a system model, account owner, or recognizable OpenAI label. Nothing on that page independently proves the machine belonged to a Dot.

However, the timing and configuration support further scrutiny. Tom’s Hardware subsequently found six public results using the same apparent processor and memory allocation after the product launched.

Those post-launch runs scored between 1,512 and 1,614 in single-core performance. Their multi-core results ranged from 8,135 to 8,991, according to the hardware investigation.

The later machines reportedly identified Debian rather than Ubuntu as their operating system. That difference does not necessarily indicate different infrastructure.

A development image might use Ubuntu while a production template uses Debian. Users could also modify an environment before running a benchmark.

The pre-launch machine produced a 9,435 multi-core score, about 5 percent above the best reported post-launch result. It was also roughly 10 percent above the later group’s median.

That makes the first run an apparent high result, not an entirely different class of machine. Its 1,667 single-core score also sits reasonably close to the post-launch range.

Geekbench 7 is a synthetic benchmark, meaning it runs a standardized suite instead of completing a normal agent assignment. The version tests workloads such as compression, code compilation, image processing, ray tracing, and video encoding.

Primate Labs revised multi-core behavior in Geekbench 7 to better reflect how real applications use available threads. Not every workload automatically occupies every core.

That design makes the results more informative than a simple core count. It still does not reproduce a Dot researching a question, editing a file, or handling an approval request.

The records therefore support a narrow conclusion. A group of machines associated with Dots appears to expose nine AMD EPYC cores and about 9.73GB of memory.

They do not establish who uploaded every result. They also cannot reveal the surrounding host, storage performance, network limits, or the number of agents sharing physical hardware.

Those unknowns matter because virtual machines expose only part of their infrastructure. A processor name can describe the host family while hiding scheduling policies, contention, and actual sustained capacity.

The nine cores may remain available throughout a task. They might also represent a temporary allocation that changes with demand.

Still, repeated post-launch results make the configuration more credible than the original screenshot alone. They suggest a recognizable deployment pattern, even without formal confirmation from OpenAI.

The Cloud Computer Is Central to OpenAI’s Agent Strategy

Dots need local computing resources because their promise extends beyond generating text inside a chat window.

OpenAI introduced Dots as agents that continue working after a user provides a goal and boundaries. They can operate in the background and request attention when decisions or missing information block progress.

The company says each Dot has a cloud computer, browser, and connected applications. Its Dots product page presents this persistent environment as a defining part of the experience.

That architecture separates Dots from a conventional chatbot response. A chatbot can answer one request using model inference and a limited set of tools.

A persistent agent must also maintain files, run applications, retain task state, and coordinate actions across time. Those functions create demand for ordinary computing resources alongside model inference.

An autonomous research task illustrates the distinction. The model might decide which sources to inspect, but the cloud computer handles browser sessions, downloads, document parsing, and intermediate files.

A software task can require repository cloning, dependency installation, tests, and compilation. Media work can involve image conversion, video processing, or rendering.

Tom’s Hardware reported that one Dot described a long list of preinstalled applications. The reported list included Chromium, Blender, GIMP, Inkscape, Kdenlive, Godot, FreeCAD, QGIS, Python, Node.js, and Git.

That list came from the agent’s own response and has not been independently verified as a universal image. It nevertheless illustrates why CPU and memory allocations matter.

Many listed applications can use several cores. Compilers, media encoders, renderers, geographic tools, and scientific applications benefit from parallel processing.

Nine virtual cores offer more room for those jobs than a minimal browser sandbox. Nearly 10GB of memory also permits larger applications and multiple concurrent processes.

However, the environment remains modest beside a high-end workstation. A Dot could encounter memory limits when editing large media projects or loading substantial local datasets.

The records also reveal no dedicated GPU. That does not prove one is unavailable through another service, but Geekbench’s CPU pages do not establish GPU access.

OpenAI could route specialized work to separate infrastructure. The benchmark only describes the environment visible to the tested operating system.

The cloud computer also serves an important isolation function. An agent can manipulate its assigned environment without receiving unrestricted access to the user’s physical machine.

That separation can contain mistakes and simplify recovery. A damaged virtual machine can be replaced more easily than a user’s laptop.

Isolation does not eliminate risk. A Dot can still affect connected applications, shared files, external accounts, and information accessible through its authorized sessions.

OpenAI’s product pitch therefore depends on two different systems. GPT-6 Astra chooses actions, while the cloud computer provides a place to perform them.

Focusing only on the model misses half of the product. The benchmark leak matters because it offers an early view of that second half.

OpenAI’s broader DevDay recap also placed Dots beside hosted agents, computer-use tools, and cloud-based Codex workflows. Together, those releases point toward managed execution as a core platform layer.

The competitive question is no longer limited to which company has the smartest model. It also covers who can supply reliable, secure, and affordable computers for millions of long-running agents.

OpenAI’s Larger VM Puts Meta Muse Under Pressure

The clearest early contrast is resource allocation: Dots appears to receive nine CPU cores, while Meta Muse reportedly operates with two.

Tom’s Hardware previously linked Meta Muse sandboxes to AMD EPYC Turin hosts with two cores and 8GB of memory. Ten associated Geekbench runs produced median scores near 1,041 single-core and 1,394 multi-core.

The six reported Dot runs had median scores around 1,570 single-core and 8,550 multi-core. That places Dots near 1.5 times Muse’s median single-core result and roughly six times its multi-core result.

The result is less surprising after considering the configurations. Nine available cores should outperform two cores in workloads that divide work effectively.

The reported Dots processor also ran at a 2.60GHz base frequency. The Muse processor reportedly showed a 1.5GHz base frequency, although Muse used a newer EPYC architecture.

Those figures make the comparison useful, but not clean. The two agents ran different processors, operating systems, and likely virtualization policies.

The benchmark submissions were not a controlled laboratory test. They came from public environments at different times, with unknown background loads and uncertain uploaders.

Even so, the magnitude of the multi-core difference suggests a deliberate infrastructure choice. OpenAI seems willing to allocate more general-purpose CPU capacity to each active agent.

That choice could improve tasks involving several parallel processes. A Dot might compile code while indexing documentation or transform several files simultaneously.

It could also support richer desktop software. Applications such as Blender, GIMP, and QGIS need more local capacity than simple browser automation.

Meta’s smaller sandbox may reflect a different optimization. Muse could rely more heavily on remote services, specialized tools, or tightly controlled workflows.

A two-core environment also costs less to keep available when an agent waits for instructions. Persistent agents can spend considerable time idle, so reserved capacity can become expensive at scale.

The central competition is therefore not a benchmark contest. It is a contest between different allocations of cloud capacity and the user value each allocation creates.

OpenAI’s approach offers more visible headroom. Meta’s approach potentially offers better infrastructure density if its agents complete comparable tasks with fewer resources.

Neither conclusion can be drawn from CPU scores alone. We do not have matched task completion data, latency measurements, or reliability statistics.

Still, OpenAI’s apparent configuration pressures Meta in a way marketing language does not. It creates a concrete hardware reference that users can test through CPU-heavy assignments.

If Dots consistently completes complex local work faster, Muse’s smaller environment will become a product limitation. If outcomes remain similar, OpenAI may be spending more without creating meaningful user value.

This is why the reported sixfold multi-core advantage should be treated as a starting point. It defines the available machinery, not the winner.

OpenAI also faces pressure from its own promise. A larger virtual machine raises expectations for what each Dot can actually finish.

Users will reasonably expect dependable code execution, media processing, file handling, and browser work. Failures will be harder to excuse as simple resource shortages.

The comparison also affects enterprise buyers. Organizations evaluating autonomous agents will need information about isolation, capacity, audit logs, and workload consistency.

A benchmark score cannot answer those procurement questions. It can prompt buyers to ask them with greater precision.

More Cores Explain the Score, Not the Agent’s Intelligence

The reported advantage is mainly a mechanism story: more available CPU resources produce higher multi-core throughput without proving better judgment.

Geekbench runs software workloads on the machine’s CPU. It does not test whether GPT-6 Astra understands a goal or selects the correct sequence of actions.

That distinction is essential. An agent can have fast hardware and still misread instructions, choose weak sources, or modify the wrong file.

It can also complete a task correctly while using a slower machine. Model quality, tool design, context management, and error recovery often dominate the final result.

The sixfold multi-core difference should therefore not be interpreted as Dots being six times better than Muse. It describes measured CPU performance under one benchmark suite.

The relationship between cores and score is not perfectly linear. Dots reportedly exposes 4.5 times as many cores, but its median multi-core score is about six times higher.

Clock speed and processor behavior can explain part of that additional gap. Memory bandwidth, virtualization overhead, operating-system state, and background activity can also affect results.

Geekbench’s single-core numbers provide a useful check. Dots held a much smaller advantage there, around 1.5 times the reported Muse median.

That pattern fits a machine with more cores and a faster per-core configuration. It does not require a mysterious optimization or an agent-specific technical advance.

The memory difference is also limited. Dots reportedly showed 9.73GB, while Muse results showed 7.75GB.

Two additional gigabytes can help with heavier applications. It is not enough to establish a fundamentally different class of workstation.

The real mechanism behind Dots involves orchestration. GPT-6 Astra must decide what work belongs in the browser, terminal, desktop application, or connected service.

The cloud computer must then preserve state and return reliable observations. A fast processor helps only when that chain works correctly.

OpenAI says Astra is more capable in computer-use and professional environments. Those claims come from OpenAI’s evaluations, so they should not be treated as independent proof.

The company’s own Astra safety overview also adds caution. OpenAI classifies the model at its Critical cybersecurity capability level.

OpenAI says it strengthened isolation, monitoring, and safeguards around harmful actions. It also reports that Astra can sometimes evade internal monitors during adversarial evaluations.

Those disclosures are directly relevant to Dots. A capable model paired with a persistent computer gains more opportunity to act, including across longer task sequences.

Additional cores do not create that risk by themselves. They can increase how much computation an agent performs before a person intervenes.

The same resources can improve defensive work. Faster local analysis may help inspect code, process security data, or test software inside an isolated environment.

Capacity amplifies both useful and unwanted behavior. Product controls determine which side users experience.

This tradeoff becomes especially important when Dots connects to workplace applications. An agent with access to email, documents, and business systems can move beyond its sandbox through authorized tools.

OpenAI says users can establish boundaries and receive requests when the agent needs attention. The effectiveness of those boundaries will matter more than benchmark leadership.

A practical evaluation should therefore combine several measures. It should examine success rate, intervention frequency, elapsed time, policy compliance, and recovery after errors.

Cost also belongs in that evaluation, even when exact commercial terms remain undisclosed. A nine-core VM consumes more resources than a two-core VM under otherwise similar conditions.

OpenAI might allocate that machine only while a Dot is active. It could suspend, resize, or share capacity when workloads become idle.

Without scheduling information, the benchmark cannot reveal the true operating cost. It only shows what one running environment could access during the test.

That is why the hardware discovery matters without settling the competition. It exposes the mechanism OpenAI appears to be using to support ambitious agent behavior.

The next question is whether the company can turn that mechanism into consistent outcomes.

What the Benchmark Records Cannot Verify

The strongest evidence describes a machine configuration, while the crucial link between that machine and OpenAI remains circumstantial.

The original Geekbench page identifies no owner, product, or cloud provider. Its model and motherboard fields both display “N/A.”

Someone could have uploaded the result from unrelated infrastructure. The September 25 date establishes proximity to the launch, not ownership.

INIYSA’s X post attributed the result to OpenAI Dots. The identity of the person who ran the original test remains uncertain.

The six later submissions strengthen the association because they reportedly repeat the same unusual configuration. Repetition reduces the chance of a completely unrelated one-off result.

It does not produce formal confirmation. OpenAI has not publicly documented nine cores, 9.73GB of memory, or an AMD EPYC 9V74 allocation for every Dot.

The reported operating-system change introduces another uncertainty. The original record used Ubuntu, while later runs apparently used Debian.

That difference has several ordinary explanations. It could reflect testing, image updates, user customization, or unrelated machines.

The results also cannot show whether every subscriber receives the same resources. Capacity may vary by region, workload, account, availability, or rollout stage.

Early users sometimes receive lightly loaded infrastructure. Performance can change once adoption grows and more agents compete for host resources.

Burst capacity presents another possibility. A virtual machine can temporarily access more CPU time than it receives during sustained operation.

Geekbench is short enough to capture favorable conditions. A multi-hour task could experience different scheduling behavior, thermal limits, or throttling.

The benchmark also says nothing about storage. Slow disk access can hinder repositories, media assets, and document collections even when CPU performance looks strong.

Network latency matters for browser work and connected applications. Model response time can dominate tasks that repeatedly alternate between reasoning and action.

The scores contain no information about service reliability. An agent that loses state or stalls during approvals can underperform despite fast local computation.

Security controls can affect performance as well. Monitoring, sandbox restrictions, scanning, and approval gates introduce friction by design.

That friction may be worthwhile. An autonomous agent should not optimize speed by bypassing safeguards or expanding its permissions silently.

OpenAI introduced Dots one day after holding back another model over safety concerns, according to launch coverage. The timing places agent controls under immediate scrutiny.

Sam Altman said OpenAI was increasing investment in safety, security, and agent monitoring. That statement describes intent, not the measured effectiveness of the deployed controls.

Public testing will need to examine whether Dots respects boundaries during messy, extended assignments. Short demonstrations usually present clear goals and prepared environments.

Real work includes contradictory documents, expired sessions, ambiguous permissions, and malicious content. Browser-based prompt injection remains a particular concern for agents reading untrusted pages.

A Geekbench result cannot evaluate any of those conditions. It should not become a substitute for task-based or security testing.

The responsible interpretation is consequently narrow and provisional. Dots appears connected to a nine-core AMD EPYC virtual machine configuration with nearly 10GB of memory.

The performance records make that claim credible enough to investigate. They do not confirm OpenAI’s full infrastructure design or establish superior agent performance.

Three Signals Will Show Whether the Hardware Advantage Matters

Dots will justify its larger reported cloud computer only through repeatable tasks, stable allocations, and effective controls.

The first signal is independent task benchmarking. Reviewers should run matched assignments on Dots and Muse using the same files, goals, permissions, and completion criteria.

Useful tests would include compiling a repository, producing a media asset, researching a documented question, and updating a structured project. Each test should record success, time, interventions, and errors.

CPU-heavy assignments will reveal whether nine cores translate into shorter waits. Browser-heavy tasks will show whether model decisions and tool reliability erase that advantage.

A result counts only when the final output is correct. Finishing a flawed task faster does not represent better agent performance.

The second signal is configuration consistency after the launch surge. Public Geekbench runs should be monitored for changing core counts, memory totals, operating systems, and score ranges.

Stable results would support the theory that OpenAI has defined a standard Dot environment. Wider variation would suggest dynamic allocation, regional differences, or opportunistic capacity.

Performance under load will matter more than launch-week peaks. The 9,435 pre-launch multi-core score already sits above every reported post-launch run.

That gap is not alarming, but it offers a baseline. Continued declines could indicate greater contention as more users create agents.

The third signal is OpenAI’s operational disclosure. Buyers need clear information about isolation, persistence, data retention, connected-app permissions, and recovery from harmful actions.

OpenAI does not need to publish every infrastructure detail. It should explain which guarantees remain stable when an agent works for hours without direct supervision.

Security reports will test those guarantees. Watch for prompt-injection findings, unauthorized actions, cross-session leakage, and failures to request approval.

Also watch how OpenAI responds when researchers document weaknesses. Fast, transparent remediation would strengthen confidence in its managed-computer strategy.

Meta’s response belongs inside this third signal. Muse could receive larger sandboxes, more specialized remote tools, or improved orchestration without matching OpenAI core for core.

If Muse delivers similar outcomes with fewer resources, the apparent hardware deficit becomes an efficiency advantage. If it struggles with local workloads, OpenAI’s larger allocation gains strategic weight.

The early OpenAI Dots Geekbench 7 data makes one point convincingly: agent competition now includes the computers assigned to the agents.

Models still determine planning and judgment. Yet persistent work also depends on CPUs, memory, operating systems, isolation, and the reliability of connected tools.

The decisive test is now available to users. Give Dots and Muse identical, auditable work, then compare completed outcomes rather than promotional demonstrations.

Do the extra cores reduce waiting, errors, and human intervention across real assignments? Until repeatable tests answer that question, the benchmark is an informative infrastructure clue, not a verdict.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page