top of page

Gecko Robotics AI Safety Puts Human Control Ahead of Full Autonomy

Sep 30
13 min read

Gecko Robotics AI safety now has a physical test: keep an autonomous inspection robot inside human-defined limits on a U.S. Navy ship deck. On September 28, Gecko announced work with NVIDIA on an open safety platform for AI agents. The central conflict is immediate. Greater autonomy can increase inspection capacity, but one unsafe command can damage equipment or injure someone.

Gecko CEO Jake Loosararian presented that conflict during a broadcast discussion with Bloomberg Technology. His position challenges a common framing of AI safety as a brake on deployment. Gecko instead argues that enforceable controls can let companies move faster while keeping consequential decisions under human authority.

That argument now faces a harder standard than safety claims about chatbots or office agents. A software agent might expose data, delete a file, or contact the wrong service. A robot can cross a physical boundary, strike a person, or compromise critical infrastructure.

NVIDIA’s Open Agent Safety Platform supplies the technical foundation for Gecko’s experiment. Its OpenShell runtime separates an agent’s planning from the permissions governing its actions. Gecko is testing that separation on Komodo, a robot used to inspect ship decks for corrosion beneath nonskid coatings.

The primary contest is therefore not Gecko against another robotics company. It is policy-enforced autonomy against autonomy that relies mainly on the model obeying instructions. The first approach assumes agents will sometimes make mistakes. It tries to limit what those mistakes can affect.

Gecko’s work offers a specific case for that architecture, but it does not settle the question. The system remains outside production, and a controlled Pittsburgh test cannot represent every shipyard or industrial site. What matters next is whether these boundaries remain dependable when conditions, equipment, and human decisions become less predictable.

Gecko Robotics AI Safety Moves From Promises to Robot Controls

The collaboration turns AI safety from a model behavior problem into an operational control problem.

Gecko’s robotics collaboration with NVIDIA covers the full path between an AI agent and a machine carrying out its instructions. The companies are exploring how OpenShell can place enforceable boundaries around robot actions. Those boundaries define what an agent can access, which commands it can issue, and when a human must approve a change.

OpenShell is an open-source secure runtime, meaning it controls the environment in which an agent operates. The agent sits inside a sandbox with restricted access to files, networks, credentials, and machine interfaces. A separate supervisory layer evaluates requests against defined policies.

That separation matters because an AI model cannot be expected to police itself reliably. Instructions written inside a prompt remain part of the reasoning context that an agent interprets. Runtime controls operate outside that context. The agent cannot simply reason its way around a denied permission.

Gecko is applying this design to Komodo, an electromagnetic acoustic transducer robot used for ship-deck inspections. The robot identifies corrosion under nonskid paint by collecting material-thickness measurements. Its field software tracks position, reads the inspection probe, and controls scanning across the deck.

According to Gecko, Komodo has completed more than a dozen paid inspections and scanned more than 100,000 square feet of deck. Those figures describe the established inspection system, not the autonomous OpenShell configuration. Gecko says the agent-controlled version has been tested on a live robot at its Pittsburgh facility but has not entered production.

The distinction is essential. An operating inspection robot and an experimental autonomous control layer are not the same product state. Gecko has experience with the physical task, while the safety architecture remains under evaluation.

The pilot gives the collaboration a clear objective. Gecko wants one operator to supervise multiple Komodo robots instead of directly controlling one machine throughout an inspection. That arrangement can increase coverage, but it also divides the operator’s attention.

The safety system must therefore do more than reject obviously invalid commands. It needs to preserve local operating rules while the human focuses elsewhere. A technically possible movement can still be dangerous near a deck edge. A faster scan can reduce the measurement density needed for a useful inspection.

Gecko’s announcement does not promise a robot that independently decides what is safe. It describes a system in which people define acceptable operating boundaries before and during a task. The agent plans within those constraints, while external controls intercept actions that exceed them.

That is the most important change in the story. Human control becomes part of the execution architecture, rather than a general promise that an operator remains involved.

Why Physical AI Needs External Guardrails

Physical AI raises the cost of an agent mistake because software decisions become movements, forces, and changes in the real world.

Physical AI refers to systems that perceive an environment, make decisions, and act through machines. The category includes industrial robots, autonomous vehicles, drones, and other equipment operating beyond a computer screen. Its safety requirements extend well beyond generating accurate answers.

Research has already shown why model-level refusals are insufficient. A 2024 study on robot jailbreak research tested attacks against three LLM-controlled robotic systems. The researchers elicited harmful physical actions in white-box, gray-box, and black-box settings.

That paper did not test Gecko’s robot or OpenShell. It nevertheless demonstrates the underlying risk. A language model’s behavioral safeguards can fail when an attacker crafts instructions designed to bypass them. Connecting that model to a mobile machine gives the failure a physical path.

Gecko’s design assumes that the model is nondeterministic, meaning the same situation can produce different outputs. The company also assumes agents can make mistakes. OpenShell limits the consequences by deciding which resources and machine capabilities the agent can reach.

The agent safety platform extends that idea across software and hardware. NVIDIA describes OpenShell as the CPU-level runtime boundary. Its Sentry component provides separate monitoring through BlueField-4 data processing units and can quarantine agents that exceed policy.

NVIDIA says Sentry can stop or isolate an agent within milliseconds. That remains a vendor claim until independent testing establishes performance under varied workloads and failure conditions. Response speed alone also cannot guarantee safety if sensors, policies, or environmental assumptions are wrong.

Gecko’s physical implementation is easier to understand through its deck-edge scenario. The robot’s control interface exposes movement capabilities. OpenShell monitors the commands produced by the agent and compares them with the permitted operating area.

If a proposed command would drive Komodo outside its safe envelope, the middleware can intercept and modify it. The system then informs the agent why the action changed, allowing the agent to produce another plan. This structure preserves useful autonomy without granting unconditional control.

A person entering the work area creates another test. Gecko says its system detects the dynamic object, stops the robot, and alerts both the agent and operator. Work resumes only after the operator confirms that the environment is safe.

These examples reveal a practical definition of human control. It does not require a person to issue every movement command. It requires humans to determine the boundaries, approve consequential exceptions, and retain authority to stop or resume the machine.

The approach also addresses a weakness in prompt-based safety. A prompt can tell an agent not to cross a boundary. An external controller can prevent the command from reaching the robot. One asks for compliance, while the other restricts capability.

That difference is why the project matters beyond one inspection robot. Industrial autonomy will depend on whether companies can translate site knowledge into machine-enforceable rules. Those rules must remain effective even when the AI misunderstands a goal or receives a hostile instruction.

The Real Tradeoff Is Speed Versus Authority

Gecko’s case is that companies can deploy autonomy quickly without surrendering authority, but only if enforcement remains separate from agent planning.

Loosararian rejects the idea that losing control of AI is an unavoidable cost of progress. His argument places responsibility on engineers to design systems that keep agents inside limits set by people. It also reframes safety work as infrastructure that enables deployment.

The commercial pressure behind that view is visible in Gecko’s pilot. Demand for deck inspections is rising, according to the company. Allowing one operator to oversee several robots would increase the amount of deck inspected simultaneously.

That operating model creates an apparent choice. Gecko can preserve direct human attention for every robot, limiting scale. Alternatively, it can grant agents more responsibility and accept that supervisors cannot watch every action in real time.

External enforcement offers a third path. The agent handles routine planning and movement, while policies reserve specific choices for the operator. Human attention shifts from continuous control toward exception handling and authorization.

Gecko’s Komodo pilot illustrates the distinction with inspection speed. A schedule change might produce an instruction to finish the job in half the time. The agent can respond by increasing the probe’s raster speed.

Faster movement, however, can reduce data density. That tradeoff affects the value of the inspection, even if the robot remains mechanically safe. OpenShell can intercept the proposed change and require human approval before altering the scanning rate.

This example expands AI safety beyond collision avoidance. The system must protect the purpose of the work, not only people and equipment. A robot that completes an inspection quickly but collects inadequate measurements has failed its mission.

Human control therefore includes quality thresholds, access permissions, and operational priorities. Each category requires a different policy. A movement boundary can use location data, while an inspection-quality rule may depend on speed, sensor readings, and site requirements.

The design also creates new work. Operators and engineers must convert practical knowledge into explicit constraints. They must identify which actions can proceed automatically, which require escalation, and which should remain prohibited.

That process can expose disagreements that were previously handled informally. A field operator may understand that weather, surface conditions, or nearby activity changes the acceptable risk. A static rule may not capture that judgment without additional sensors and context.

Rapid deployment and safety are therefore compatible only under specific conditions. The relevant hazards must be understood. Policies must represent them accurately, and enforcement must occur outside the agent’s control.

The architecture cannot eliminate uncertainty. It can make uncertainty more manageable by narrowing what the agent can do before a person intervenes. That is a more credible claim than promising that a sufficiently capable model will always choose correctly.

Gecko Robotics AI safety is strongest where the company can define precise physical and operational boundaries. It becomes harder when safety depends on ambiguous context or competing goals. The pilot’s value will come from revealing where that line falls.

NVIDIA Is Building a Control Layer Across the Agent Market

NVIDIA wants agent safety to become a shared infrastructure layer, not a collection of safeguards built separately into every model and application.

The Open Agent Safety Platform is broader than Gecko’s robotics test. NVIDIA describes it as an open reference design covering agent testing, deployment, monitoring, and hardware enforcement. Organizations can use individual components based on their requirements.

OpenShell follows a deny-by-default model. An agent starts without broad access, and policies grant only the permissions needed for its task. The runtime filters system calls, limits reachable files, and brokers network requests through a supervisor.

The supervisor operates outside the agent sandbox. It evaluates network access by software binary, destination, method, and path. NVIDIA says policy changes can apply while an agent is running, with allow and deny decisions recorded for auditing.

A policy prover adds another layer. It uses formal verification, a mathematical method for checking whether a system satisfies defined properties. NVIDIA says the tool can evaluate whether proposed rules remain inside an approved access boundary.

These mechanisms target several agent risks at once. Isolation can limit damage from compromised code. Narrow permissions can protect credentials and files. Audit records can help investigators reconstruct what an agent attempted.

The platform also gives NVIDIA a strategic position between models and the infrastructure where agents operate. It is designed to support open or closed models and different agent frameworks. That model-agnostic stance can make the control layer useful across a fragmented market.

NVIDIA says more than 100 organizations are working with technologies in the platform. Its announced participants span AI companies, infrastructure providers, security vendors, banks, industrial operators, and government-related customers. Figure, Gecko, and Skild AI are among the robotics developers named by NVIDIA.

Those partnerships show interest, not proof of broad adoption. “Working with” can cover integrations, evaluations, contributions, or production deployments. Buyers will need more precise disclosure before treating the participant count as evidence of operational maturity.

NVIDIA is also pursuing a related physical safety strategy through Halos for Robotics. That system combines computing hardware, operating software, external perception, and inspection resources. OpenShell focuses more directly on controlling agent access and behavior.

The two efforts reflect a layered view of safety. A secure runtime can restrict commands, while a robotics safety stack addresses sensing, computing, and machine behavior. Neither layer can substitute for reliable hardware, sensors, maintenance, or site procedures.

This matters for competitors and enterprise buyers. Robotics companies must decide whether to adopt a shared NVIDIA control layer, build proprietary safeguards, or combine both approaches. Industrial customers must determine how those controls fit existing safety and cybersecurity systems.

An open foundation can reduce duplicated work and allow outside review. It can also concentrate architectural influence around NVIDIA’s software and hardware stack. Open source availability does not automatically remove integration costs or dependence on adjacent components.

The platform’s success will depend on portability and verification. Developers need policies that work across models and deployment environments. Security teams need evidence that enforcement holds under adversarial conditions, not only standard demonstrations.

Gecko gives NVIDIA a tangible physical use case. A robot approaching a ship-deck boundary is easier to evaluate than a broad promise about responsible agents. The test is whether that clarity survives deployment beyond a controlled facility.

The Pittsburgh Test Leaves Production Questions Open

The main uncertainty is not whether OpenShell can stop a prepared demonstration, but whether its policies remain dependable in changing industrial environments.

Gecko says all described OpenShell features have been implemented and tested on a live robot in its Pittsburgh testing environment. The company also states that the autonomous configuration has not been deployed into production. That gap should shape every assessment of the project.

A test facility lets engineers control the deck layout, safety zones, network conditions, and people entering the work area. A Navy ship or industrial plant introduces changing equipment, restricted spaces, unusual surfaces, active crews, and site-specific procedures.

The enforcement layer is only as accurate as the information it receives. A geographic boundary cannot protect the robot if localization data drifts. A person-detection rule can fail if sensors miss someone or classify an object incorrectly.

Policies can also conflict. A robot might receive a direction to finish quickly while maintaining measurement quality and avoiding a temporary obstruction. The system needs a clear hierarchy among those goals and a safe response when no permitted plan succeeds.

Human escalation creates its own constraints. A single operator supervising several robots may face simultaneous requests. If every unusual condition triggers approval, the system can lose the productivity advantage that justified greater autonomy.

The opposite failure is more serious. A permissive policy can let an agent proceed without review when the context requires human judgment. Designing escalation thresholds will be as important as the underlying sandbox.

Cybersecurity adds another challenge. Separating enforcement from the agent reduces the chance that a prompt injection can override a rule. It does not automatically secure sensors, robot firmware, operator accounts, policy updates, or communication links.

NVIDIA’s OpenShell architecture addresses files, networks, credentials, sandboxing, and policy enforcement. Gecko’s implementation extends controls to robot commands. Independent assessments must examine how the entire chain behaves when one component is compromised.

Military and critical-infrastructure settings raise the evidence standard further. Customers will want repeatable tests, failure logs, recovery procedures, and clear responsibility when automated decisions cause harm. Vendor demonstrations cannot replace those processes.

The platform’s openness can support scrutiny because researchers and customers can inspect parts of the software. Yet a deployed system includes configuration, sensors, hardware, networking, and local integrations. Reviewing source code alone cannot validate the final installation.

There is also a risk of confusing human oversight with human control. An operator who receives an alert after an unsafe action has begun may only be observing the failure. Meaningful control requires sufficient information, decision time, and authority before consequences become irreversible.

Gecko’s architecture addresses that concern by intercepting certain commands before execution. The production question is how completely engineers can identify those consequential commands. Unknown hazards will not arrive with labels indicating which policy should stop them.

None of these issues invalidate the pilot. They explain why a live-robot test is the beginning of validation rather than its conclusion. The experiment becomes valuable when it produces evidence about failures, ambiguous cases, and operator workload.

Three Signals Will Show Whether Human Control Scales

The next stage must prove that bounded autonomy works across real operations, independent evaluation, and multi-robot supervision.

The first signal is a production deployment with disclosed operating limits. Gecko has described an on-site test in Pittsburgh, but not active autonomous inspections aboard Navy ships. A field deployment would test the architecture against less predictable conditions.

The most useful disclosure would define exactly what the agent controls. Readers should watch for details about movement, scan planning, speed changes, emergency stops, and operator approvals. Clear limits would strengthen Gecko’s claim that authority remains with people.

A production announcement without those details would provide weaker evidence. “AI-assisted” can describe many arrangements, from route suggestions to direct machine control. The degree of autonomy determines which safety claims matter.

The second signal is independent technical evaluation. Researchers or customers should test whether the runtime blocks unauthorized actions, preserves inspection requirements, and fails safely when sensors or policies are incomplete. Adversarial testing should include hostile prompts and compromised components.

Results should distinguish between model failures and enforcement failures. An agent proposing an unsafe action is concerning, but a boundary that blocks it is functioning as designed. An unsafe command reaching the robot indicates a deeper control problem.

Independent evaluation would also clarify NVIDIA’s performance claims. Millisecond quarantine sounds reassuring, but physical safety depends on total response time. Sensors, networking, policy evaluation, robot controllers, and mechanical stopping distance all contribute to the outcome.

The third signal is evidence from one-operator, multi-robot work. This is Gecko’s stated first milestone for broader autonomy. It directly tests whether external controls reduce workload or simply replace manual driving with repeated approval requests.

Useful metrics would include interventions, blocked commands, false alarms, inspection coverage, data quality, and operator attention. Gecko has not published those measurements for the autonomous pilot. Their absence limits comparisons with direct teleoperation.

If one operator can oversee several Komodos without losing situational awareness, the case for policy-enforced autonomy becomes stronger. If escalations overwhelm the operator, the system may need better planning, more conservative policies, or fewer robots per supervisor.

These signals matter beyond Gecko. Industrial buyers need a method for separating deployable safeguards from polished demonstrations. Developers need patterns for keeping models productive without giving them unrestricted access to physical systems.

The emerging lesson is not that humans must manually control every robotic action. It is that autonomy needs an authority structure. Models can propose and execute routine steps, while external systems constrain their reach and people govern consequential exceptions.

That approach resembles mature safety practices in other technical fields. High-risk systems use layered controls because no single component is perfectly reliable. Physical AI will require similar discipline, adapted to agents that can plan, communicate, and change tactics.

Gecko Robotics AI safety provides a concrete test of that principle. Its ship-deck pilot connects abstract agent governance to a machine operating near people and valuable infrastructure. It also exposes the limits of claims based on controlled trials.

The question for developers and enterprise buyers is practical: can their operating knowledge be converted into rules that machines cannot bypass? Teams evaluating physical AI should identify prohibited actions, required approvals, and acceptable failure states before increasing autonomy.

Over the next several months, production evidence should carry more weight than partnership counts. Independent tests should matter more than vendor response-time claims. Operator workload should matter as much as robot capability.

If Gecko publishes credible results across those areas, human-controlled physical AI will look like an implementable architecture rather than a slogan. If it cannot, the industry will still face the original tension: faster autonomous deployment without proven control over what machines do.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page