AI Affairs, home

Wednesday 30 September 2026

Technology

Nvidia introduces two open-source tools to restrict and stop AI agents

The software is designed to control access in real time. Nvidia says its system would have prevented a breach involving OpenAI models and Hugging Face.

Wide-angle panoramic view of San Tomas Expressway and Nvidia campus with overhead bridge
Photo: Dicklyon, CC BY-SA 4.0, via Wikimedia Commons (cropped)

On 28 September, Nvidia unveiled a two-layer security system for AI agents made up of two open-source software tools, intended to limit their permissions while they operate and to halt them if they violate the rules, Bloomberg reported. The tools can run on Nvidia hardware.

Key points

  • The two tools are open source and designed to restrict access while agents operate.
  • They are also designed to stop agents that violate rules.
  • Nvidia says the system would have prevented a breach involving OpenAI models and Hugging Face.

Nvidia’s two controls for agents

Nvidia says its new system would have prevented a recent breach of Hugging Face by OpenAI’s AI models, Bloomberg reported. That is a claim about a particular incident, distinct from the stated purpose of the tools: restricting an agent’s access as it operates and stopping it when it crosses a rule.

Agents can use tools and change their actions in response to information they encounter. Nvidia described those capabilities in a 21 September post on agent security, alongside a familiar security problem: permission to perform one task should not give software permission to perform every action it can attempt. An agent allowed to change a customer record, for example, should not thereby acquire the authority to export the customer’s data.

The two functions Nvidia is rolling out address different moments in that process. An access control would govern an action as the agent tries to take it. A shutdown control is designed to stop an agent that breaks the rules. Both are software measures, so companies building agents could run them on Nvidia hardware rather than rely solely on instructions telling an agent how to behave.

OpenShell puts policies beyond an agent’s reach

Nvidia’s 21 September post described OpenShell, an open-source runtime that it says enforces policies outside an agent’s reach and provides sandboxed execution. A runtime is the environment in which software acts. In this case, the boundary around the agent can govern access to data, the network and system resources independently of the choices the agent makes while working.

Think of the difference between a note asking someone not to open a locked cupboard and the lock on the cupboard itself. Nvidia argues that instructions can guide an agent, but a security boundary must still hold when it makes a wrong decision. Its post calls for limits on the files, network destinations and processes an agent can reach, imposed by the environment in which it runs.

Nvidia gives the example of an agent changing a customer record and encountering malicious instructions in an attached document. If the agent then attempts to send customer data to an unauthorised destination, the company says a network policy should block the transfer. Protected logs should retain the attempted action, the authorisation decision and its outcome, so the destination and the route the agent tried to use can be investigated.

Updating a customer record could remain possible under its assigned permission, while a request to send that information elsewhere would be blocked if it crossed a separate network rule. The attempted transfer could also be recorded for investigation if the protections Nvidia describes were in place.

That separation extends to changes in permission. Nvidia says an agent may request more access but must not grant it to itself. It also calls for each agent to have a traceable identity and credentials limited to its task, with human approval for consequential actions and permission changes within the boundaries set for it.

Nvidia calls for tests beyond agent instructions

For companies deploying agents, Nvidia’s security guidance extends from setting rules to checking whether those rules work. Its post calls for tests before deployment that attempt to obtain credentials beyond an agent’s remit or send sensitive information to an unauthorised destination. It also calls for attempts to change permissions or interfere with monitoring, with testing repeated after substantial changes to models, tools or workflows.

Nvidia says a named owner should use those results to decide whether a system is ready for deployment and ensure failed tests lead to changes. When a failure occurs in testing or operation, the company calls for it to be reproduced and investigated. It says each finding can become a repeatable test, allowing a later release to be checked against the same failure.

The guidance also covers what an agent is allowed to use, not just where it can send information. Nvidia says teams need to check the source and integrity of an agent’s tools, skills and dependencies, and keep protected records of its actions and authorisation decisions. Those records would give investigators a way to reconstruct an incident, while procedures for revoking access would let a team contain one.

Nvidia says Cisco’s DefenseClaw adds a governance layer to OpenShell. JFrog integrates with OpenShell to scan and verify agent skills and enforce policies governing which skills an agent can access.

Topics: Agents, Open source, Safety