SAS: On governing autonomous AI agents

Autonomous AI agents require controls over their authority, data access, and actions throughout deployment.

Marinela Profi, SAS Global Market Strategy Lead for AI Agents and Generative AI

Ahead of AI & Big Data Expo Europe in Amsterdam, Marinela Profi, Global Market Strategy Lead for AI Agents and Generative AI at SAS, discussed how enterprises should approach systems that execute decisions.

Marinela’s answers examine the architectural choices behind agent deployment, from recording which tools an agent invokes, to establishing when human approval is required. She draws on SAS’ Data & AI Impact (PDF) research and examples from banking and life sciences, distinguishing reported results from work still being explored.

Our Q&A covers why demonstrations provide limited evidence of operational readiness, how governance policies become executable controls, and where engineering teams should place boundaries around machine autonomy.

AI News: Why does the Silicon Valley deployment playbook fail when applied to autonomous software systems that make automated decisions?

Marinela Profi: The traditional Silicon Valley playbook is essentially: ship quickly, observe how users behave, learn, and iterate. That works reasonably well when the cost of failure is a user encountering a bad feature. It becomes much more problematic when software has the authority to act. With autonomous AI, the blast radius changes. An error is no longer necessarily a bad answer on someone’s screen. An agent could trigger a workflow, interact with another system, make a decision, or execute an action before a human ever sees it.

With autonomous AI, the blast radius changes. An error is no longer necessarily a bad answer on someone’s screen. An agent could trigger a workflow, interact with another system, make a decision, or execute an action before a human ever sees it.

The latest SAS Data & AI Impact Report captures this transition very clearly. Trust in generative AI is 76%, but falls to 66% for agentic AI. Yet 89% of agents deployed in production are already acting rather than merely assisting, and more than half operate with limited or no human approval. So I don’t think the lesson is “move slowly.” The lesson is to change what moving fast means.

In autonomous systems, speed has to include the ability to observe, constrain, interrupt, and recover. Before asking, How quickly can we deploy this agent?, organisations should ask: What authority are we giving it, what is the blast radius if it fails, and can we reverse what it does? Autonomy changes failure from something you observe into something the system can propagate.Autonomy changes failure from something you observe into something the system can propagate.

AN: How do early architectural shortcuts translate into compliance liabilities and technical debt over a multi-year horizon?

MP: The dangerous shortcuts often aren’t visible in the AI model itself. They’re in everything around it. Maybe nobody created reliable lineage between the data, model, decision, and resulting action. Maybe permissions were granted at the application level rather than at the individual agent or task level. Maybe logs capture outputs but not which tools were invoked. Maybe governance exists as documentation rather than executable controls.

Those choices feel inexpensive during a pilot. At scale, they become architectural debt. Imagine that two years later a regulator, auditor, or customer asks: Why did this system make this decision, what data did it use, what policy governed it, and who was accountable for it? If your architecture cannot reconstruct that chain, you’re not going to fix the problem by writing another governance policy.

That’s why I increasingly think about governance as architecture, not documentation. The SAS Data & AI Impact Report reinforces this. Only 17.5% of organisations report fully optimised data infrastructure that includes the lineage, governance, validation, and explainability capabilities required for agentic AI.

Technical debt in traditional software makes future development expensive. Trust debt in autonomous systems can make future autonomy impossible, because organisations eventually discover they cannot safely expand what their systems are allowed to do.

AN: Why do standard proofs of concept fail to expose the systemic hazards of agents executing unmonitored backend actions?

MP: Because a proof of concept usually tests whether the AI can perform the task. Production tests whether the entire system can operate reliably under real conditions. A sandbox gives you controlled data, predictable permissions, and a limited set of interactions. Enterprise production gives you none of those things.

And agents introduce another important dimension: state and action. If a chatbot gives me a bad response, I can ignore it. If an agent gives the wrong instruction to another system, I’ve got a much bigger problem on my hands. So I think we need to evolve from model testing toward system testing.

Don’t just ask whether the model produced the correct answer. Test what happens when a tool fails, data is stale, or permissions change. That’s especially important because the research shows that users frequently challenge AI not simply to check accuracy but to probe the system’s capabilities and reasoning. A successful demo proves capability. It doesn’t prove operational readiness.

AN: What common design pattern appears viable on an executive roadmap but causes catastrophic failures in enterprise deployments?

MP: One pattern I worry about is the “one intelligent layer over everything” architecture. It looks fantastic on a roadmap: connect an LLM or agent to enterprise data, give it access to tools, and watch it orchestrate everything.

But intelligence without boundaries creates enormous risk. The agent needs to know not only how to accomplish something, but what it is authorised to do, which data it may touch, and which systems it may affect.

I saw the opposite approach with a major global bank we work with. They are using generative and agentic AI extensively, but the important part is how they architected it. But the insight I love came from their CIO: technology alone doesn’t solve the problem; it’s how everyone works together.

That’s the architectural lesson. The LLM shouldn’t become your business architecture. It should operate inside an architecture of data, decisioning, and control. Or put differently: don’t make the probabilistic component the control plane for the entire enterprise.

AN: How does system verification change when autonomous agents execute actions without human review?

MP: Verification moves from asking “Was the prediction correct?” to asking “Was the entire decision-and-action chain trustworthy?” With predictive analytics, we historically validated things like accuracy, bias, drift, and model performance.

With autonomous agents, we need another layer. Was the agent authorised to take that action? Did it use the correct data? Did it invoke the correct tools? Did it stay within its boundaries?

So verification becomes continuous rather than something that happens primarily before deployment. The research shows a huge difference between mature and immature organisations here: 66% of trustworthy AI organisations have formal verification and validation processes, compared to 15% of those at earlier stages.

I think the conceptual shift is from model assurance to system assurance. And autonomous systems also need verification of the action itself. Sometimes the most important question isn’t whether the model was accurate but whether the action was appropriate.

AN: How do deficits in data quality and system fragmentation restrict an enterprise’s ability to run automated decision pipelines?

MP: I would slightly challenge the premise here: autonomous systems don’t always require deterministic inputs. An agent can only make decisions based on the world it can see.

If customer information is fragmented across six systems, definitions conflict, data is stale, lineage is unclear, and no one owns the decision logic, the agent is effectively flying blind. That’s why I often tell organisations that their AI problem may actually be a data and decision architecture problem.

The SAS Data & AI Impact study found that only 17.5% of organisations have reached fully optimised data and AI readiness. A life sciences example makes this tangible. At one major biopharmaceutical company, as much as 80% of scientist time was spent simply finding, cleaning, and reconciling data before any analysis could begin. That’s the reality behind autonomous AI: you cannot automate your way out of an unreliable data foundation.

AN: How do programmatic controls and policy enforcement accelerate deployments rather than delay them?

MP: I think we’ve made a branding mistake with governance. We’ve spent years describing it primarily as a restriction. Good governance is actually infrastructure for speed.

Imagine two companies. In the first, every new AI deployment triggers a bespoke debate between engineering, legal, security, risk, and compliance about what is permitted. In the second, policies have already been translated into executable controls: what data can be accessed, what actions require approval, what thresholds trigger escalation, what must be logged, and which models are permitted for which use cases. Which organisation deploys the tenth agent faster? Almost certainly the second.

And that’s consistent with the SAS Data & AI Impact research. Organisations with high trustworthiness scores are 15 times more likely to report strong or high AI ROI than low-scoring organisations. Leaders are also much more likely to operationalise policy, continuously validate systems, and conduct governance and explainability reviews. So I don’t see governance as a gate at the end of development. I see it as a reusable operating system for innovation.

A good control doesn’t simply tell an agent “no.” It tells the agent what it can do autonomously, and therefore allows the organisation to delegate more confidently. Brakes don’t exist to make cars slow. They help make speed safe.

AN: Where should engineering teams enforce boundaries on machine autonomy, and what architectures keep humans integrated into supervisory roles?

MP: I don’t think there should be one universal boundary between human and machine autonomy. The boundary should move according to risk, reversibility, uncertainty, and consequence. A low-risk, highly reversible action can tolerate significantly more autonomy than a decision involving someone’s health, credit, employment, or access to public services.

And I think the architecture should reflect that. Some decisions can be fully automated within predefined boundaries. Some can operate autonomously unless an exception threshold is reached. Others should require explicit human approval. And in certain high-consequence situations, humans should remain the final decision-maker.

Interestingly, the SAS Data & AI Impact research suggests that mature organisations don’t simply add more human review. In life sciences, less mature organisations are substantially more likely to require full human review, while leaders are more capable of allowing AI to act autonomously within defined boundaries or escalating exceptions. That’s important because human-in-the-loop should not mean human-in-every-loop. If humans approve every transaction, they become a bottleneck and eventually a rubber stamp.

The architecture I favour is closer to human-on-the-loop: humans establish authority, boundaries, and escalation criteria; machines operate inside them; and humans retain visibility and the ability to intervene. We can delegate execution to AI. We should never delegate accountability to it.

SAS is a key sponsor of this year’s AI & Big Data Expo Europe. Check out Marinela Profi’s presentation ‘Why the leaders who win the Agentic AI race won’t be the ones who moved fastest’ during the event. Swing by SAS’ booth at stand #304 to hear more directly from the company’s experts.

Source: https://www.artificialintelligence-news.com/news/marinela-profi-sas-governing-autonomous-ai-agents/

Leave a Reply

Your email address will not be published. Required fields are marked *