Something shifted in enterprise AI around the middle of this year that I do not think enough people are naming clearly. For the first two years of the generative AI boom, most organizations were using AI as a fancy search engine with typing. You asked it a question, it gave you an answer, and a human decided what to do with that answer. The human was always the last link in the chain. The AI could be wrong, and the only consequence was a wasted few minutes and a mildly embarrassing memo. That model is quietly being retired. The new model has AI taking actions rather than generating text, and the question nobody has cleanly answered yet is the one that will matter most when something goes wrong: who is responsible?
I want to write about this today because I am watching organizations deploy agentic AI at a speed that is clearly outrunning their thinking about governance. I run AI at Mahindra Group, and a meaningful portion of my time goes into conversations that are really accountability conversations dressed up as technology conversations. The excitement about agents is legitimate. The gap in organizational thinking around what happens when they fail is not something that will stay invisible for much longer.
What agentic AI actually looks like in production right now
The term "AI agent" has been used so loosely for so long that it is worth being exact about what it means in a mid-sized to large enterprise today, deployed in a real production system rather than a sandbox.
An agent, in the current usage, is an AI system that takes a goal rather than a prompt. You do not tell it "write me a summary of this customer complaint." You tell it "resolve this customer complaint." Then it figures out the steps: reading the complaint, checking order history, looking up the applicable return policy, deciding whether the case qualifies, initiating the refund in the payments system, and sending the customer a confirmation. The human involvement is at the goal-setting stage, not at each individual step. The agent acts autonomously across multiple connected systems to reach an outcome.
That capability is real today and it is in production. ServiceNow's Now Assist is deploying multi-step agents across IT, HR, and customer service workflows. Salesforce Agentforce went general availability late last year and has clients running agents across their sales and service pipelines. Microsoft's Copilot agents are wired into Microsoft 365 tenants across hundreds of thousands of organizations. The scale is not hypothetical. Gartner has estimated that up to 234 billion dollars of enterprise software spend is at risk from agentic disruption by 2030, as agents complete tasks that currently require a human to navigate between systems. That is about 20 percent of the entire enterprise SaaS market being structurally repriced because agents can do the navigation themselves.
I find the technology compelling. What I want to focus on today is not the opportunity. Enough coverage handles that already. I want to focus on the accountability gap, which is the question that lives between "what can it do?" and "should we let it do that without supervision?"
The accountability vacuum
Here is the problem in plain terms. When a human employee makes a mistake, accountability is reasonably clear. The employee may have been wrong. Their manager may have failed to check their work. The process may have been badly designed. The policy they were following may have been ambiguous. But somewhere in that chain, there is a person. Organizations have spent decades building processes around that assumption, and the assumption is so baked in that most people do not realize it is an assumption until it stops being true.
When an AI agent makes a mistake, the chain becomes murky fast. The agent was trained by one company. It was fine-tuned or configured by another. It was deployed and integrated by an internal team. It was granted access to the payments system by someone in IT. It operated within a policy definition it was told to follow, which was written by someone in legal or operations. The decision it reached drew on data from a third-party CRM, a home-built inventory database, and a lookup table last updated by someone who left the organization eighteen months ago. When the outcome goes wrong, every party in that chain has a defensible argument that the fault originated somewhere else.
I have a name for this in my own head: accountability laundering. The error is real. The harm is real. The cost lands on someone. But responsibility has been distributed so thinly across so many layers that it becomes invisible. The agent is the visible actor. Nobody owns the agent.
The danger of "the AI made that call"
There is a phrase I have started hearing in meetings that I want to flag as a warning sign. When an automated AI process produces a bad outcome, someone in the room says "the AI made that call." Sometimes this is said with genuine confusion about where the problem lies. Sometimes, if I am being direct, it is said with a certain relief, because if the AI made the call then nobody made the call, and if nobody made the call then nobody needs to answer for it.
This is a category error, and I think it is one of the more dangerous habits forming in enterprise AI right now. "The AI decided" is a description of a mechanism. It is not a description of responsibility. A vending machine dispenses a product that injures someone. The vending machine did not decide to injure them. The manufacturer, the operator, and the maintenance company are still responsible. The degree to which we allow "the AI made that call" to function as a statement that terminates the accountability conversation is the degree to which we are building organizations that will be genuinely ungovernable when something goes seriously wrong.
This matters especially because agents are being deployed in the exact domains where the cost of a wrong call is highest. Credit approvals. Contract terms. Medical appointment routing. Supplier payment authorization. Hiring shortlists. These are not low-stakes autocomplete tasks. When an agent makes a wrong call in any of these areas, the consequence lands on a real person, and sometimes on a large number of real people simultaneously, because the agent applies the same flawed logic at scale before anyone notices.
What happened recently is worth taking seriously
The example that stayed with me most from the last few weeks came out of the GPT-5.6 release, which I covered in a broader piece on the chaos of that particular week in mid-July. In early testing, the model exhibited file deletion behaviour that OpenAI's own system card had flagged as a known risk. The system card is the vendor's document for disclosing what a model might do that it should not. The fact that a documented risk shipped into a widely deployed model is not what worried me most. Models have always had edge cases, and no system card covers every one. What stayed with me was the gap between "we know this risk exists" and "someone in a deploying organisation has been named as accountable for preventing it in production." The vendor wrote down the risk. The deploying organisations, in many cases, did not write down the accountability.
This is not a story about one company or one model. It is a structural observation about where the industry sits right now. The people building these systems are, in many cases, more rigorous about documenting risks than the organizations deploying them are about assigning human responsibility for managing those risks. The system card is the vendor doing their job. The deployment accountability framework is the deploying organization doing theirs. The second document is not getting written at the same rate as the first.
The regulator is closer than most Indian enterprises realize
The European Union's AI Act is entering its most consequential enforcement phase. The provisions covering high-risk AI systems, which include AI used in employment decisions, credit, and access to essential services, come into full effect in August 2026. That is next month. Organizations deploying AI agents in these categories in EU markets are legally required to have documented accountability chains, human oversight mechanisms, and audit logs that can reconstruct what the agent did and why. Non-compliance carries fines of up to three percent of global annual revenue under the high-risk provisions.
India's own Digital Personal Data Protection Act is maturing, and the AI governance framework that the Ministry of Electronics and Information Technology put out for consultation this year is moving toward formal status. India tends to follow global regulatory direction on technology with a lag, but the direction is clear and the lag is shrinking. Organizations that treat AI governance as a 2028 problem are already behind.
I am not making a compliance argument here, because compliance is the floor of what you must do rather than the ceiling of what you should do. My argument is that ungoverned agents create organizational liabilities that exceed what any regulator can find or fine. The regulatory risk is the visible part. The operational risk, the reputational risk, and the trust erosion that comes when a large agent deployment makes a high-visibility mistake, those are larger than the fine and harder to recover from.
India is deploying fast and governing slowly
I want to be direct about the Indian enterprise context because it is the one I know most closely, and because there is a specific pattern that is easy to miss from the outside.
India's large enterprises, and particularly the IT services industry, have genuine deployment velocity. The instinct to ship fast, learn fast, and iterate is a real competitive advantage in most contexts. The AI talent density is real. The cost structure that allows rapid experimentation is real. The risk is that the same organizational culture that is excellent at deploying can be slow at governing, because governance feels like it costs speed and nobody gets promoted for writing an accountability framework. Speed is visible. Governance is visible only when it is absent.
The IT services firms deploying agentic solutions for their global clients are operating in a particularly complex accountability picture. When a large Indian IT services firm deploys an agent for a global enterprise client, who owns the agent's decisions? The services firm that built and configured it? The client that wrote the requirements? The model vendor whose API sits underneath the whole stack? The contractual frameworks have not caught up with the technology. I have been in rooms where this question surfaces and the honest answer from everyone present is "we are going to sort that out." Sorting it out after the agent has been in production for six months is harder than sorting it out before deployment. The accountability is not retroactively clarifiable once a claim has been made.
There is also a specific India-context risk around scale. One of India's genuine strengths is the ability to deploy solutions at enormous scale quickly. That same speed means that when an agent has a flaw in its logic or its policy definitions, the flaw propagates across millions of interactions before anyone catches it. The volume that makes Indian enterprise AI impressive in success is the same volume that makes a misconfigured agent extremely costly in failure. Governance at this scale is not bureaucracy. It is risk management.
A tiered autonomy model: what I actually do at Mahindra
I want to offer something practical here, because articles that identify problems without describing how to address them are not very useful and I try not to write them.
When we think about deploying agents at Mahindra, the question is not "should we use agents." We are past that. The question is "what tier of autonomy is appropriate for this task, and who is named as accountable for each tier?" I think about it in three levels.
The first level is advisory autonomy. The agent does the analytical work and generates a recommendation, and a human reviews before anything is actioned. This is appropriate when the consequence of an error is high and the cost of human review is low relative to that consequence. An agent that drafts a supplier contract clause belongs here. It can do ninety percent of the drafting time. The human who approves before signature takes legal responsibility. The accountability chain is clean: the person who signed off owns the outcome. This tier often feels slow to teams that are excited about automation. I accept that feeling. The alternative is faster deployment with murkier accountability, which is a bad trade in high-stakes domains.
The second level is supervised autonomy. The agent acts within defined guardrails, with automatic alerts when those guardrails are being approached, and with a human in the exception path. An agent handling routine customer service cases is a reasonable candidate. It can resolve a large percentage of straightforward cases automatically. When a case is ambiguous, when the proposed resolution cost exceeds a defined threshold, or when sentiment signals a customer who is likely to escalate, a human gets the case. The accountability question here is not just "who owns the agent" but specifically "who owns the guardrail definitions, and how often are those definitions reviewed?" Those people need to be named and they need a review cadence. Guardrails that were appropriate at launch may be wrong six months later as the case mix shifts, and if nobody owns reviewing them, they will stay wrong silently.
The third level is full autonomy. The agent acts without a human in the loop for any individual decision. This level is appropriate for a narrower set of tasks than most deployment roadmaps currently assume. Scheduling a meeting is a reasonable candidate. Approving a supplier payment above a defined value threshold is not, regardless of what any vendor's accuracy benchmarks say about the model. My rule of thumb is that full autonomy is appropriate when the worst-case error is recoverable, cheap to catch, and cheap to reverse. When those three conditions do not hold simultaneously, supervised autonomy is the right call, even if it feels like leaving efficiency on the table.
The practical failure mode I see most often is organizations deploying to the second or third tier for tasks that should be in the first tier, because the first tier feels like it negates the value of the agent. It does not. An agent in the first tier still delivers most of the productivity gain. It compresses the drafting time, the research time, the data retrieval time. The human review at the end is cheap if the output is good, and it catches errors before they reach the customer. The efficiency loss from having a human in the loop at the final stage is usually smaller than it looks from the outside. The accountability gain is much larger than it looks from the outside.
Three questions that need a named answer before any agent ships
Naming accountability is uncomfortable, which is probably why organizations avoid it. It is easier to say "the system owns that" or "governance comes later" than to write a specific person's name next to a specific risk. The discomfort is the point. Accountability that nobody will write down is accountability that nobody actually holds.
For every agent deployment, I think three questions need written answers before the agent goes into production.
Who is responsible for what the agent is permitted to do? This is the policy owner. When the agent's instructions are wrong, too broad, too narrow, or fail to account for a category of case that real usage surfaces, this person answers for the gap. They also own the schedule for reviewing and updating those instructions as the deployment matures.
Who is responsible for what the agent actually does? This is the operational owner. When the agent takes an action that was technically within its instructions but produced a harmful outcome in a specific case, this person answers for that outcome. They own the monitoring, the incident response when something goes wrong, and the post-incident review.
Who is responsible for the effect of the agent's actions on the people they touch? This is harder to name because it involves a judgment about organizational values, not just operational ownership. In organizations that have ethics review committees or AI governance boards, this function exists on paper. In most organizations, it does not exist in any practical sense until an incident makes it necessary. Building it before the incident is not idealism. It is the difference between an organization that learns from a mistake and one that does the same mistake at larger scale the next time.
These three people may be the same person in a smaller organization, or they may be three different people with overlapping mandates in a large one. The specific structure matters less than the act of naming them before the agent ships. The accountability framework does not prevent agents from making mistakes. No framework does. What it does is ensure that when a mistake happens, there is an organization capable of understanding it, owning it, and making it less likely to recur.
My take
I am not arguing against agents. I want to say that very clearly because this kind of piece gets read as technological conservatism and it is not. I think agentic AI is the most consequential shift in enterprise technology since cloud computing moved workloads off owned hardware, and that is a statement I do not make lightly. I have built agentic systems, I am deploying them, and the productivity gains are real in a way that the first generation of generative AI tools often was not.
What I am arguing is that the governance architecture needs to be built at roughly the same pace as the deployment architecture. Right now, those two curves are not close to each other. The deployment curve is steep and the governance curve is nearly flat, and every week that gap widens is a week in which organizations are accumulating what I think of as accountability debt. Accountability debt has a specific repayment structure: it accrues silently and becomes visible all at once when a high-profile agent makes a high-consequences mistake in public.
The organizations that build the accountability structure now, while it is still optional and before any specific incident forces their hand, will have a real advantage when it becomes mandatory or when something goes wrong that requires a credible response. The organizations that treat governance as something to add on after the deployment is already running at scale will find that retrofitting accountability into a live agentic system is roughly as tractable as retrofitting a sprinkler system into a building that is already occupied and on fire.
The question worth asking your leadership team this week is not "how many agents are we deploying." It is "for every agent we are deploying, can we name the three people I described above?" If the answer for even one of those agents is no, the problem is not that you are behind on AI. The problem is that you are ahead on deployment and behind on the part that will protect you when the deployment does something you did not expect.
Agents will do something you did not expect. That is not a reason to stop. It is a reason to know, in advance, whose job it is when they do.