📊 Key Data
  • 88 rules benchmarked: Sondera's system successfully autoformalized 23 of 88 clinical AI agent rules in MedAgentBench.
  • 100% success rate: Blocked all 99 adversarial attempts to make agents perform unsafe actions.
  • Industry recognition: Research accepted at ICML 2026 and FLoC 2026, with open-source tool demo at Black Hat Arsenal.
🎯 Expert Consensus

Experts would likely conclude that Sondera's autoformalization technology represents a significant advancement in AI governance, offering mathematically verifiable control over autonomous agents—a critical step toward safer deployment in high-stakes sectors.

20 days ago
Beyond the Prompt: Can Provable Rules Finally Tame Autonomous AI Agents?

Beyond the Prompt: Can Provable Rules Finally Tame Autonomous AI Agents?

NEW YORK, NY – June 30, 2026 – The race to deploy autonomous AI agents is accelerating, promising to reshape industries from finance to defense. These powerful tools can analyze intelligence, manage infrastructure, and execute complex multi-step tasks at machine speed. Yet, beneath the surface of this rapid adoption lies a critical vulnerability: how do we ensure these agents, operating with increasing independence, actually follow our rules? The risk isn't just about a chatbot generating incorrect information; it's about an agent misinterpreting a command and leaking sensitive financial data, violating compliance mandates like HIPAA, or misconfiguring a critical server. This governance gap has become the single greatest barrier to trusting AI with high-stakes work.

Today, a company named Sondera has emerged from stealth with a compelling claim: they can translate natural-language rules—the kind found in thick compliance manuals and standard operating procedures (SOPs)—directly into formally verified, machine-enforceable code. This process, called “autoformalization,” promises a new paradigm of provable, deterministic control over AI behavior, a development that has already earned recognition from top AI and cybersecurity conferences.

The Governance Gap: From Paper Policies to Unpredictable Code

Every organization runs on a complex web of rules. For decades, these policies have lived in documents, their enforcement dependent on human training, oversight, and review. As AI agents are integrated into workflows, the challenge has been to make them understand and adhere to these same constraints. The dominant approaches to AI safety have so far proven insufficient for this task.

Prompt engineering and constitutional AI attempt to steer a model’s behavior from within its context window, but these probabilistic guardrails are notoriously brittle. They can be bypassed by clever prompt injection attacks, and their reliability degrades as models drift or encounter novel situations—so-called “emergent behavior.” They offer no formal guarantees, only a statistical likelihood of compliance. For a hospital deploying a clinical agent or a bank using an agent to handle financial data, “likely” is not good enough.

Manually coding policies, the traditional alternative, is a Sisyphean task. It’s slow, expensive, and fails to scale with the speed of AI deployment and the constant evolution of business logic. This has left many enterprises in a state of “fragmented governance,” where AI innovation outpaces their ability to control it, creating a shadow risk that grows with every new agent deployed. The core problem, as Sondera frames it, is that the real incidents aren't just from malicious outsiders. “The agent incidents we see today aren't from prompt injection and hijacking,” said Josh Devon, co-founder and CEO of Sondera. “They're from authorized humans asking authorized agents to do legitimate tasks... Along the way, the agent reaches the goal with unintended behavior, like leaking or destroying data.”

A Neurosymbolic Blueprint for Control

Sondera’s answer to this challenge is a novel neurosymbolic architecture that combines the strengths of neural networks and classical symbolic logic. At its heart is the autoformalization pipeline, a process that reads natural-language policy and compiles it into formally verified code written in Cedar, an open-source policy language. This approach fundamentally shifts enforcement from a probabilistic suggestion inside the model to a deterministic certainty outside of it.

The process works through a sophisticated “Verification Sandwich.” First, a large language model (LLM) acts as a policy generator, interpreting the intent of a text like a HIPAA manual and producing candidate rules in Cedar code. This is the “neuro” part—using AI’s strength in understanding unstructured language. Then comes the critical “symbolic” layer: a two-part critic loop. A “hard critic” uses a theorem prover and static analysis to mathematically verify the code, checking for syntax errors, logical contradictions, and schema compliance. Simultaneously, a “soft critic,” another LLM acting as a judge, evaluates whether the formalized rule accurately reflects the semantic spirit of the original document. This ensures the rules are not only technically correct but also contextually appropriate.

This method of formal verification provides a mathematical guarantee that the policy behaves as specified. At runtime, every action an agent attempts is checked against these verified rules by a deterministic engine that operates completely outside the agent’s context window. Because this enforcement layer cannot be influenced by the agent’s internal state or manipulated by prompt engineering, it acts as an incorruptible referee. Furthermore, the system is stateful, tracking the agent's entire sequence of actions. This allows it to make context-aware decisions, permitting an action in one scenario while denying the exact same action later if it violates a rule based on the agent's prior behavior.

The academic and security communities are taking note. Sondera’s research paper has been accepted at prestigious workshops at ICML 2026 and the Federated Logic Conference (FLoC) 2026, while a related open-source tool, “GolemHalt,” is slated for demonstration at the influential Black Hat Arsenal showcase.

From the Lab to the Real World: Proving Safety

The true test of any security technology is its performance against real-world challenges. To this end, Sondera benchmarked its system using MedAgentBench, an independent testbed for clinical AI agents published in the peer-reviewed journal NEJM AI. The results are striking. Sondera’s pipeline successfully autoformalized 23 of the 88 rules in the benchmark's policy—more than had previously been accomplished through painstaking manual coding.

More importantly, when subjected to 99 different adversarial attempts to make the agent perform an unsafe action (such as writing unauthorized data), the Sondera-generated rules blocked every single one. A 100% success rate in a high-stakes medical benchmark is a powerful proof point. It suggests this technology could be a critical enabler for deploying AI in sectors where the cost of failure is unacceptably high, from patient data management and clinical decision support to financial compliance and the control of autonomous defense systems.

This capability moves the conversation on AI safety from abstract principles to concrete, provable implementation. It provides a blueprint for building trust in systems that will increasingly operate on the front lines of our economic and national security. By creating a verifiable link between human-readable policy and machine-executable commands, this approach offers a path to deploying powerful, long-running agents with the confidence that they will operate safely and predictably within the boundaries we define. The era of simply hoping our AI agents do the right thing may finally be giving way to an era where we can prove it.

Topics & Related

Sector:
AI & Machine Learning
Cybersecurity
Theme:
AI Governance
Agentic AI
Event:
Product Launch
UAID: 40890