Service management is where a lot of enterprise AI gets tried first, and it’s easy to see why. The tickets are already in a system. The volumes are large. The value case looks like arithmetic.
What AI use cases actually work in ITSM?
AI in enterprise service management works where the ticket is repetitive, well-documented and low-consequence. It stalls where the underlying process is undocumented, because a model cannot resolve what the organisation has never written down. Gartner’s 2026 survey found only 17% of organisations have deployed AI agents, while more than 60% expect to within two years.
It’s also where a lot of it quietly stops working after the pilot — for reasons that have almost nothing to do with the models.
Where the market actually is
The gap between intention and deployment here’s the widest I’ve seen in any technology category.
Gartner’s 2026 CIO and Technology Executive Survey found that only 17% of organisations have deployed AI agents so far, while more than 60% expect to within the next two years — the most aggressive adoption curve among all emerging technologies measured. Gartner places agentic AI at the Peak of Inflated Expectations on its 2026 hype cycle, published in April 2026 by Rajesh Kandaswamy.
Their own reading of that gap is blunter than most of what gets written about it. Most deployments remain narrowly scoped, and fully autonomous agents aren’t ready for the majority of enterprise use cases.
Their ITSM-specific read is more sober. Agentic ITSM is attracting serious market attention, but most organisations are holding off on large-scale commitment until there are proven, practical use cases, and infrastructure and operations leaders are prioritising solutions with measurable return while staying wary of vendor claims.
I’d take that last sentence as the honest state of the field. Interest is near-universal, deployment isn’t, and the people running these functions have learned to want evidence first.
The prediction I’d argue with
One Gartner strategic planning assumption deserves more scrutiny than it usually gets. In IT operations workflows, human-in-the-loop is projected to fall to 40% by 2028, down from 95% in 2025.
As a forecast of where the market is heading? Fine. As a target it’s dangerous, and I’ve watched it quoted as one.
Removing the human from 95% of workflows to 40% is only an improvement if the remaining oversight is real. My experience is that oversight tends to become nominal long before it’s formally removed — the approve button is still there, someone still clicks it, and nobody is actually checking. That’s the failure I wrote about in human oversight becoming a fig leaf, and service management is where I have seen it most often, because the volumes make genuine review impossible long before anyone admits it.
Gartner’s own hype cycle supports the caution better than the projection does. One of its defining signals for 2026 is that governance, security and cost profiles — agentic AI governance, agentic AI security, FinOps for agentic AI — have appeared across the curve rather than clustering at the mature end. Their reading is that the need for oversight and discipline is becoming evident early in the adoption cycle rather than only after large-scale deployment.
So the same research house is saying oversight matters earlier than expected and that human involvement will fall sharply. Both can be true, and they’re only compatible if the oversight that remains gets much better rather than merely thinner.
If you’re planning against that curve, the question is not how fast you can remove humans. It’s which decisions you’re prepared to have made without one, named specifically, in advance.
What works
High-volume, well-documented, low-consequence requests. Password resets, access requests, licence provisioning, status lookups. The pattern is that the correct answer already exists in writing and being wrong is cheap to correct.
Drafting rather than deciding. Suggesting a category, a priority, an assignment group or a first response, with a person accepting or overriding. This preserves genuine oversight because the human is doing something rather than approving something.
Search over your own knowledge base. Often the highest-value and least glamorous option. Most large organisations have written the answer down somewhere and nobody can find it. Fixing retrieval delivers value without any autonomous action at all.
Summarising long incident histories. A major incident with two hundred updates is hard for a human arriving at 3am. Summarisation is low risk because the underlying record stays available.
What stalls
Anything resting on an undocumented process. This is the single biggest cause of failure and it isn’t a technology problem. If the real resolution path lives in the head of one person in a particular team, no model will discover it. You aren’t implementing AI, you’re discovering that your process documentation is fiction — which is useful, but it’s a different project with a different timeline.
Deflection targets set before a baseline exists. A deflection number with nothing to compare it against is unfalsifiable, which is the sizing problem in its purest form. Measure the current resolution path first.
Anything where the cost of a wrong answer was never priced. In service management wrong answers are rarely dramatic. They are a reopened ticket, a second contact, an escalation, a user who stops trusting the tool and goes back to phoning. Each has a cost, and if your case counts only successful deflections it’s a sales deck.
Categorisation as an end in itself. Automatically categorising every incoming ticket is easy to demonstrate and often changes nothing, because the categories were already unreliable and nobody downstream was using them. Coverage of a thing that does not drive a decision isn’t value.
The measurement problem, specifically
Service management has better data than most functions, which produces a particular trap: everything is measurable, so people measure the easy things.
Deflection rate goes up. Handle time goes down. Both can improve while the user experience gets worse, because a deflected ticket includes the user who gave up, and a faster handle time includes the incident that was closed prematurely and reopened next week.
The numbers I’d insist on are reopen rate, second-contact rate, and end-to-end time from first contact to actual resolution rather than to ticket closure. Those are harder to move and they’re the ones that reflect whether anything real changed.
What I’d tell someone starting this
Pick one workflow, not a platform strategy. Establish the baseline before anything is switched on. Choose something where being wrong is recoverable. Name in advance which decisions require a person and what that person is actually expected to check. And write down what result would make you stop.
The vendors will tell you the platform can do considerably more than that. It probably can. The constraint is almost never the platform, and treating it as though it were is how organisations end up with a capable tool sitting on top of a process nobody has ever mapped.
Common questions
What AI use cases actually work in ITSM?
High-volume, well-documented, low-consequence requests such as password resets, access requests and status lookups, where the correct answer already exists in writing and errors are cheap to fix. Drafting suggestions for a human to accept or override, retrieval over an existing knowledge base, and summarising long incident histories also work reliably.
Why do AI service desk projects fail?
Most often because the underlying process was never documented. If the real resolution path exists only in one person's head, no model can discover it. Other common causes are deflection targets set without a baseline to compare against, and business cases that count successful deflections while ignoring the cost of wrong answers.
How many organisations have actually deployed AI agents?
According to Gartner's 2026 CIO and Technology Executive Survey, 17% have deployed AI agents to date while more than 60% expect to within two years, the most aggressive adoption curve among emerging technologies measured. Gartner places agentic AI at the Peak of Inflated Expectations on its 2026 hype cycle.
Is deflection rate a good measure of AI in service management?
On its own, no. Deflection counts include users who gave up rather than got an answer. Reopen rate, second-contact rate and end-to-end time from first contact to actual resolution are harder to move and better reflect whether anything genuinely improved.
Will AI remove humans from IT operations workflows?
Gartner projects human-in-the-loop falling to 40% of IT operations workflows by 2028, from 95% in 2025. Treating that as a forecast is reasonable; treating it as a target isn't. Reducing oversight only helps if the oversight that remains is genuine, and in high-volume environments review tends to become nominal well before it's formally removed.
What should you measure before starting an AI ITSM pilot?
The current baseline for whichever unit the value is claimed in: resolution time, contact volume, reopen rate and the real cost per contact. Without it, any improvement claimed afterwards cannot be verified, and a deflection figure with nothing to compare it against is unfalsifiable.
Does automatic ticket categorisation deliver value?
Rarely on its own. It demonstrates well but often changes nothing, because the categories were already unreliable and no downstream decision depended on them. Categorisation is worth automating only where a specific routing or prioritisation decision actually consumes the category.
Are fully autonomous AI agents ready for enterprise use?
Not for most cases. Gartner's 2026 hype cycle states that deployments remain narrowly scoped and that fully autonomous agents are not ready for the majority of enterprise use cases, while placing agentic AI at the Peak of Inflated Expectations. The same analysis notes governance, security and cost concerns appearing early in the adoption cycle rather than after scale.
Should you set an AI deflection target before starting?
Not before establishing a baseline. A deflection figure with nothing to compare against cannot be proved or disproved, and deflection counts include users who abandoned the channel rather than got an answer. Measure the current resolution path, cost per contact and reopen rate first, then set a target against them.
I’m an AI product manager working across fintech, SaaS, and regulated enterprise — currently leading AI and workflow product at T-Systems International. If you’re building AI governance into a product right now and want to compare notes, I’m at csincsakf@gmail.com or on LinkedIn.