top of page

Holo4: Powering Generalist Computer-Use Agents, but the Benchmark Gap Still Matters

4 days ago
12 min read

Holo4 arrived on September 28 with two models, four interaction modes, and a direct challenge to specialized computer-use systems. H Company describes Holo4: powering generalist computer-use agents as one model family that can navigate screens, execute code, and call software tools.

The release matters because computer automation rarely stays inside one interface. A business process might begin in a browser, continue through an API, and end inside desktop software without modern integrations. Most agent systems handle that transition by combining different models, tools, and control loops.

Holo4 proposes a simpler route. The same model can choose between graphical interfaces, code, Model Context Protocol tools, and APIs. MCP is a standard that lets AI systems access external tools and data through structured connections.

That promise puts Holo4 against a specialist architecture, not merely another model vendor. The specialist approach assigns different models or policies to visual navigation, coding, and tool calling. H Company argues that one trained generalist can coordinate those surfaces more efficiently.

The company has also made thousands of benchmark trajectories available for inspection. That transparency gives developers more evidence than a leaderboard score alone. It does not settle questions about reliability, safety, or performance inside real organizations.

Holo4: Powering Generalist Computer-Use Agents Across Four Interfaces

The central change is architectural: Holo4 treats the interface as a choice within the task, rather than a fixed boundary around the agent.

According to the Holo4 release, the family includes a dense 27-billion-parameter model and a 35-billion-parameter mixture-of-experts model. The latter activates about three billion parameters for each inference step.

A mixture-of-experts model routes inputs through selected internal components instead of activating every parameter. This design can reduce computation, although actual speed depends on hardware, software, and deployment choices.

Both Holo4 models can interact with graphical user interfaces, write and execute code, and call MCP or API tools. H Company says the same model can operate on desktops, websites, Android devices, coding sandboxes, and business systems.

That differs from an agent that only predicts mouse clicks from screenshots. It also differs from a tool-calling model that becomes ineffective when an application lacks an API. Holo4 is designed to switch methods as a workflow changes.

Consider a routine finance operation. An agent might extract fields from a document, normalize them with code, submit them through an API, and verify the result onscreen. Older enterprise software could force another transition into mouse and keyboard control.

A generalist model could preserve one decision process across those stages. A specialist stack would usually route each stage to a separate model, policy, or service. That routing can improve control, but it also introduces more handoffs and failure points.

H Company says it trained Holo4 through supervised learning and reinforcement learning across generated interactive environments. Its internal task factory has reportedly created about 10,000 tasks covering web applications, desktops, MCP servers, and hybrid environments.

These generated tasks are important because static examples cannot reproduce the consequences of an agent’s actions. An interactive environment can test whether a click changed state, whether code executed, or whether an API call produced the intended record.

The approach also lets H Company generate tasks from documentation and screenshots. That could widen training coverage without manually designing every workflow. However, generated environments can still differ from messy production systems with permissions, delays, and unexpected state.

The release includes an updated Holotron4 Nano model alongside the two main Holo4 variants. It also offers model weights in several formats, including BF16, FP8, NVFP4, and four-bit GGUF.

Availability across these formats gives developers several deployment options. Yet the more consequential claim remains that one model can coordinate multiple interfaces without an external model-selection layer.

That makes Holo4: powering generalist computer-use agents a test of whether interface generality can reduce system complexity without sacrificing the precision that specialist agents offer.

Why Long Workflows Put Specialist Agent Stacks Under Pressure

Holo4 pressures specialist stacks because long workflows multiply the cost of every routing decision, context handoff, and recovery step.

A short browser task can hide architectural weaknesses. An agent might open one page, enter a value, and submit a form. Even a brittle system sometimes completes that sequence.

Professional work looks different. It involves several applications, persistent state, ambiguous instructions, and information that appears during execution. An agent must remember earlier constraints while adapting to later events.

OSWorld 2.0 was designed around that harder setting. Its researchers assembled 108 long-horizon workflows spanning everyday and professional work. A skilled human requires a median of about 1.6 hours to complete each task.

The benchmark reports that leading agents can average more than 300 steps per workflow. OSWorld 1.0 tasks required roughly 30 steps, making the newer benchmark a much stricter test of context management.

The failures also extend beyond inaccurate clicking. The researchers observed agents losing constraints, missing incoming information, guessing when clarification was necessary, and skipping verification. Those weaknesses can compound across a lengthy process.

H Company says it rebuilt the Holo4 agent harness in response to these problems. A harness is the execution loop that supplies observations, manages context, runs actions, and returns results to the model.

Its two most notable additions were persistent memory for hundreds of steps and a shell running on the desktop machine. The shell gives the agent a code-based route when direct GUI interaction becomes inefficient.

This is where the generalist design becomes more than a feature list. The model can decide that parsing a local file with code is preferable to reading it visually. It can then return to the interface for actions that require visual confirmation.

A specialist stack can perform the same sequence. However, it must decide when to transfer control and how much context accompanies each transfer. An incorrect routing choice can waste steps or discard information.

Holo4 tries to place that decision inside the trained model. If the approach works consistently, developers could reduce the logic required to coordinate browser control, desktop control, code execution, and structured tools.

That does not eliminate orchestration. Production systems still need credential management, sandboxing, retries, logging, and approval gates. They also need a reliable way to stop an agent before an uncertain action causes damage.

The shift is narrower but still meaningful. Developers could spend less effort deciding which model should handle each interface. They could focus more on defining permissions, validating outputs, and measuring complete workflows.

This distinction matters for teams building a searchable knowledge base. Their workflows often cross local documents, internal search, browser tools, and structured company systems.

Holo4 does not prove that generalists will replace every specialist. It makes specialist routing a design choice that developers must justify, rather than an unavoidable foundation.

One Agent Model Is Simpler, but Specialists Still Set the Reliability Bar

The primary contest is one generalist model against a coordinated stack of specialists, with reliability deciding which architecture wins.

Specialists offer an intuitive advantage. A model trained narrowly for visual grounding can focus on locating controls. A coding model can focus on syntax, execution, and debugging without interpreting every screenshot.

Tool-calling models also benefit from structured schemas. An API exposes permitted actions and predictable fields. A graphical interface offers more flexibility, but its buttons, layouts, and transient states create ambiguity.

The specialist approach lets engineers select the best model for each surface. It can also isolate risky capabilities. A visual agent might receive screen access without gaining arbitrary shell execution.

However, specialization shifts complexity into the surrounding system. A router must classify each stage, choose a component, and preserve the user’s intent across handoffs. The stack must reconcile different context formats and failure signals.

Holo4’s generalist route moves some of that coordination into the model. The agent can look at a screen, recognize that direct manipulation is inefficient, and use code or a structured tool instead.

H Company illustrates this approach with professional software tasks. In one example, Holo4 27B reportedly used 68 calls and 2.4 million tokens to build an autonomous game in Godot. Its Qwen base used 197 calls and 11.4 million tokens under the same prompt and harness.

Those figures come from H Company’s own evaluation, not an independent laboratory. They describe one task rather than average production performance. Still, they show the type of efficiency the company wants Holo4 to deliver.

Other examples involve constructing detailed objects inside FreeCAD. These workflows combine spatial interpretation, software control, and code generation. They are more demanding than filling a single web form.

The examples also reveal a limitation. Holo4’s Eiffel Tower task reportedly required 84 calls and 1.3 million tokens. Long computer-use sessions can remain computationally heavy even when the final result succeeds.

Specialists retain another advantage when the workflow is predictable. A deterministic script or narrow API integration can be faster and easier to audit than an agent choosing among several possible actions.

The generalist case becomes stronger when workflows vary, interfaces change, or legacy systems lack integrations. The specialist case remains stronger when organizations need repeatability and can define the process precisely.

This means Holo4 is unlikely to erase conventional automation. It instead competes for the uncertain middle ground where fixed scripts break but unrestricted frontier agents remain too expensive or difficult to govern.

Developers should therefore evaluate complete tasks, not isolated clicks. The relevant question is whether Holo4 reduces failures and engineering overhead across real workflows.

A model that reaches the correct final state through fewer handoffs can justify lower raw precision on one narrow skill. A generalist that changes methods unpredictably can create a larger debugging burden.

The outcome will depend on trajectory quality, reproducibility, and recovery behavior. Those factors matter more than whether one architecture appears cleaner in a diagram.

Holo4 Benchmark Results Need Their Harnesses and Footnotes

Holo4’s results are notable, but the release itself explains why several headline comparisons are not directly equivalent.

H Company reports that Holo4 27B scored 61.7 percent on OSWorld 2.0. Its 35B-A3B model reached 30.9 percent. The company compares those results with 81.8 percent for Opus 5.5.

The 20.1-point gap between Holo4 27B and Opus 5.5 is substantial. Holo4 does not lead the strongest closed model on this reported measure. Its argument centers on model size, deployment flexibility, and estimated task cost.

H Company also cites scores of 70.2 percent for Opus 5 and 66.2 percent for GPT-5.6 Sol. Those references use maximum-effort partial rewards on an offline OSWorld 2.0 set dated August 8, 2026.

The release cautions that model releases, harnesses, and task subsets vary. That warning should stay attached to every comparison. Agent performance depends on far more than the model checkpoint.

The harness controls memory, tool access, observation formatting, retry behavior, and maximum steps. Changing any of those variables can alter the result, even when the underlying model remains unchanged.

The cost charts require similar care. H Company estimated expenses from the input and output tokens used during each run. It priced Holo4 using its own API rates and used outside list prices for other models.

Such estimates can support internal planning, but they are not controlled economic measurements. Caching assumptions, inference infrastructure, retries, and volume discounts can change actual deployment costs.

AutomationBench introduces another comparability issue. H Company evaluated Holo4 and its Qwen baselines using version 1.0.6 in its internal harness. Other models’ scores came from the benchmark’s public set.

The referenced cost figures for other models came from a leaderboard running a private set. H Company says it plans to report Holo4 on that private evaluation after testing occurs.

Until then, readers should not treat all AutomationBench points as results from one controlled experiment. They represent related measurements produced under different conditions.

Even benchmark definitions can shape a narrative. OSWorld 2.0 supports binary completion and partial-credit scoring. A model can receive meaningful partial credit while failing to deliver a finished workflow.

That does not make partial scoring useless. It can identify progress on long tasks where binary success would hide improvements. Yet buyers care whether the final record, file, or transaction is correct.

Efficiency also needs more than token counts. Research on agent efficiency found that leading computer-use agents took between 1.4 and 2.7 times more steps than necessary in its evaluation.

The same research found that later steps can take much longer than early ones. Planning and reflection calls accounted for much of the latency. A long agent trace can therefore magnify delays beyond its visible action count.

These findings reinforce H Company’s focus on memory and shell access. They also show why a successful benchmark score does not automatically produce an acceptable user experience.

The fair reading is neither dismissal nor acceptance. Holo4 posts a competitive company-reported result for a relatively compact model, while trailing the leading closed system.

The practical test is whether those results persist in independent harnesses, private evaluations, and workflows containing organization-specific permissions and data.

Open Trajectories Improve Verification, Not Safety

H Company’s strongest credibility move is releasing the traces behind its scores, although inspectable behavior is not automatically safe behavior.

The trajectory dataset contains 7,366 runs from Holo4 27B and Holo4 35B-A3B. Each trace can include the task, reasoning, actions, tool results, screenshots, duration, steps, and final score.

The collection includes more than 2,100 OSWorld runs across the two models. It also contains 212 OSWorld 2.0 runs and nearly 3,200 AutomationBench runs.

Additional traces cover AndroidWorld, PinchBench, and Agents’ Last Exam. H Company lets users download the dataset or replay traces through a dedicated viewer.

This disclosure gives researchers several ways to challenge the company’s conclusions. They can inspect whether a successful run used a reasonable path, repeated unnecessary actions, or benefited from task-specific shortcuts.

They can also examine failure patterns. An aggregate score cannot show whether the agent misunderstood the instruction, clicked the wrong target, lost context, or stopped before verification.

Open trajectories can reveal whether performance improvements come from better reasoning or a more forgiving harness. They can also help teams estimate how frequently human intervention might be required.

However, transparency after execution is different from control before execution. A trace helps investigators understand what happened. It does not prevent an agent from sending data, deleting files, or following malicious instructions.

Computer-use agents face risks that conventional chat systems avoid. They operate in environments containing untrusted content and valuable credentials. A webpage can place adversarial text directly inside the model’s observation.

The OS-Harm benchmark tests deliberate misuse, prompt injection, and unintended model behavior across 150 tasks. Its researchers found meaningful unsafe behavior across several frontier systems.

That study did not evaluate Holo4, so its results cannot establish Holo4’s safety. It does demonstrate that competent computer control and safe computer control are separate evaluation problems.

The risk becomes sharper when one model has broad interface access. A generalist can move from reading a webpage to running code or calling an API. That flexibility increases usefulness and the possible impact of a mistake.

Enterprises will need layered controls regardless of benchmark scores. Those controls include scoped credentials, isolated execution, restricted network access, reversible actions, and human approval for consequential steps.

They will also need logs that connect each action to the user’s instruction and the state observed at that moment. Holo4’s trajectory format offers a useful model for such audit records.

Licensing deserves attention as well. Open weights do not guarantee identical commercial rights across every checkpoint or component. Teams should review each model card and dependency before deployment.

The same applies to data governance. Screenshots and agent traces can capture personal information, customer records, or confidential documents. Logging everything can improve debugging while creating another sensitive dataset.

H Company says credentials, internal hosts, and personal data were masked in its public trajectory release. Production operators must build equivalent redaction and retention controls for their own traces.

The open dataset raises the standard for future launches. Vendors claiming superior computer-use performance now have a clearer example of what reproducible evidence can look like.

Still, the most valuable external work will involve adversarial replay, independent scoring, and tests outside H Company’s harness. Transparency opens that process; it does not complete it.

Three Signals Will Show Whether Holo4’s Generalist Bet Holds

The next phase should be judged through independent reproduction, private benchmark results, and evidence from production workflows.

The first signal is independent reproduction of Holo4’s OSWorld 2.0 performance. Researchers need to run the released weights with documented infrastructure, prompts, step limits, and scoring rules.

Matching the reported 61.7 percent would strengthen H Company’s claim that the model itself carries the capability. Large differences would suggest that the company’s harness contributes more than the headline implies.

Reproduction should also compare binary completion, partial credit, steps, latency, and intervention rates. One score cannot capture whether an agent reaches a usable outcome within practical limits.

The released trajectories make this work more achievable. Researchers can start from known runs, examine failure boundaries, and compare alternate harnesses against the same tasks.

The second signal is Holo4’s private AutomationBench result. The release acknowledges that its current scores come from an internal run on the public set, while comparison costs reference a private-set leaderboard.

A private evaluation would create a cleaner comparison and reduce concerns about tuning against visible tasks. It would also test whether Holo4 generalizes across unfamiliar API workflows.

The result should include more than a success rate. Developers need costs, token use, retries, latency, and failure categories under one documented evaluation setup.

The third signal is credible production evidence from mixed-interface workflows. The best cases would involve tasks that genuinely require GUIs, code, and structured tools within one session.

Useful reporting would show completion without human correction, recovery from interface changes, and performance under restricted permissions. It should also count irreversible errors, not only successful runs.

An expense-processing workflow offers a representative test. The agent must read documents, validate fields, interact with business software, and confirm that records reached the intended state.

Another strong test would involve engineering operations across local files, issue trackers, browser consoles, and command-line tools. Such workflows expose whether shared context is an advantage or a source of uncontrolled behavior.

Model updates will provide a related signal. H Company says optimized drafter checkpoints are planned to accelerate inference. Measured latency reductions would strengthen the economic case for the generalist architecture.

Competitor responses matter too, but they remain supporting evidence. Closed model providers can improve computer use, while open-model developers can add broader tool training to their own releases.

The central question is not whether Holo4 remains ahead of every alternative. It is whether a single generalist model delivers a better reliability-to-complexity ratio than a specialist stack.

For developers, the immediate opportunity is controlled evaluation. Use representative tasks, restricted credentials, and reversible environments. Measure completed outcomes rather than isolated model actions.

For enterprise buyers, the procurement question should include the harness. Ask which component manages memory, approvals, retries, secrets, audit logs, and recovery after partial execution.

For knowledge workers, the release suggests that agents will cross more application boundaries. That convenience also makes permission design and visible confirmation more important.

Holo4: powering generalist computer-use agents is therefore less a declaration of victory than a concrete architectural challenge. H Company has supplied models, claims, and unusually detailed traces.

The next move belongs to independent evaluators and deployment teams. Can Holo4 reproduce its reported results, survive unfamiliar tasks, and complete real work without expanding operational risk?

Those three tests will determine whether generalist computer-use agents simplify automation or merely relocate its hardest problems.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page