Ask a South African CIO how their AI pilot is performing, and you'll often get a variation of the same answer: "It's going well, the team likes it." That's a feeling, not a number. And when budget season arrives and the CFO wants to know why a chunk of the IT spend went to an AI tool nobody can put a rand figure against, "the team likes it" doesn't survive the conversation.
This is the quiet failure mode of most AI initiatives. Not that the technology doesn't work — it usually does — but that nobody set up the measurement scaffolding before deployment. By the time someone asks for ROI, the baseline is gone. You can't measure improvement against a state you never recorded.
At NewGenIT.ai, we treat ROI measurement as a day-one activity, not a retrospective one. And we've found that credible AI ROI lives across four distinct categories. Track only one and you'll either oversell or undersell the impact. Track all four and you get a defensible, honest picture that survives scrutiny from both the server room and the boardroom.
Why "Surface Metrics" Quietly Sabotage AI Projects
The temptation is to measure the thing that's easiest to count. Tickets closed. Queries answered. Model accuracy on a benchmark. These are real numbers, but they're often vanity numbers — impressive in a slide deck, meaningless in a business case.
Consider a support automation tool that "resolves 2,000 tickets a month." Sounds great. But if 40% of those were tickets the tool created noise around, or if the human escalations now take longer because the AI muddied the context, your headline metric is hiding a net loss. The number went up. The value went down.
Good AI measurement isn't about finding a big number. It's about building a picture honest enough that you'd stake your budget on it.
The four categories below force that honesty because they pull in different directions. Efficiency and cost savings can be gamed in isolation. Pair them with accuracy and user experience, and the gaming becomes visible. A tool that's fast but wrong shows up. A tool that's cheap but hated shows up. That tension is the point.
Category 1: Efficiency Gains — Measuring Time You Get Back
Efficiency is where most teams start, and rightly so — it's the most tangible. But "efficiency" is too vague to track. Break it into specifics you can baseline before deployment and re-measure after:
- Mean time to resolution (MTTR) for incidents, split by severity. AI triage should move the needle most on the routine, high-volume tickets — not the gnarly ones.
- Manual touch-hours per process. How many human hours went into log analysis, patch scheduling, or report generation before, versus after?
- Automation coverage. What percentage of a given workflow now runs without human intervention, and how reliably?
Here's the South African wrinkle: efficiency gains matter more in a constrained labour market. Skilled IT operations staff are expensive and hard to retain, and the strongest ones get poached — often overseas, paid in dollars or pounds. When AI takes routine monitoring off a senior engineer's plate, the real win isn't the hours saved on paper. It's that you've freed a scarce, expensive resource to do work that actually justifies their salary. Quantify efficiency in terms of where the reclaimed time goes, not just how much of it there is.
Category 2: Accuracy and Reliability — The Trust Multiplier
An AI system that's frequently wrong doesn't just fail to add value — it actively destroys it, because every false alarm trains your team to ignore the tool. Once trust is gone, adoption collapses regardless of how clever the underlying model is.
The metrics that matter here:
- False positive rate in anomaly and threat detection. This is often the single most important number for an IT Ops AI. A model that flags 100 anomalies where 90 are noise will get muted within a fortnight.
- Predictive accuracy for capacity, failures, and demand — measured against what actually happened, not against a training set.
- Outage prevention. How many incidents were caught before they became downtime? This is harder to measure because you're counting things that didn't happen, but tracking near-misses and pre-emptive interventions gives you a proxy.
Reliability carries extra weight for South African operations because of load-shedding. Infrastructure here already absorbs power interruptions, failover events, and UPS cycling that teams in more stable grids never think about. An AI monitoring system needs to distinguish a genuine anomaly from the daily reality of a Stage 4 evening. If it can't, it'll drown your team in false positives every time the lights go out. Accuracy under local conditions isn't a nice-to-have — it's the difference between a useful tool and abandoned software.
Category 3: Cost Savings — The Number the CFO Actually Reads
This is where AI ROI has to translate into rands. Efficiency and accuracy are inputs; cost savings are the output your finance team will hold you to. The categories to track:
- Direct operational cost reduction — fewer manual hours, consolidated tooling, reduced overtime during incidents.
- Downtime cost avoided. Calculate your cost-per-hour of downtime (revenue impact plus recovery labour) and multiply by hours prevented. This is usually the biggest number and the most persuasive.
- Resource optimisation — right-sized cloud spend, better capacity planning, fewer over-provisioned instances sitting idle.
Rand economics make this category sharper here than elsewhere. Most serious AI tooling — cloud compute, foundation model APIs, enterprise platforms — is priced in dollars and billed monthly. A weak rand means your AI running costs climb even when your usage stays flat. That cuts both ways: it raises the bar for justifying the spend, but it also raises the value of every optimisation, because dollar-denominated waste hurts more. Track your AI's cost in the currency you're billed in, and be brutally honest about it. An initiative that looks profitable at R18 to the dollar can go underwater at R20 if you never modelled the exposure.
Category 4: User Experience Impact — The Metric Everyone Forgets
The hardest category to quantify, and the one that most often decides whether an AI deployment survives its first year. You can have brilliant efficiency, tight accuracy, and real cost savings, and still watch a tool wither because people quietly route around it.
What to measure:
- End-user satisfaction — short, regular pulse surveys tied to specific AI-touched interactions, not a once-a-year generic survey.
- System responsiveness as experienced by users, not as reported by dashboards. Perceived speed and actual speed diverge more than teams expect.
- Service quality — first-contact resolution, escalation rates, and whether users feel the AI helped or got in the way.
- Adoption and retention. Are people still using the tool three months in, or has usage quietly decayed? Sustained adoption is the truest signal that value is real.
User experience is also where POPIA lives. If your AI tools process personal information — and most IT Ops and service-desk tools touch it somehow — the experience has to be trustworthy in a compliance sense, not just a convenience sense. Users and regulators alike need confidence that data is handled lawfully. A slick AI assistant that mishandles personal information isn't a UX win; it's a liability that a single complaint to the Information Regulator can turn into a very expensive lesson. Build compliance into your experience metrics from the start.
Bringing the Four Together
The reason these four categories work is that they cross-check each other. An initiative that scores well on all four is genuinely delivering. One that spikes on efficiency and cost but tanks on accuracy and experience is borrowing against future trust — and that debt always comes due.
Here's what we'd urge any IT Ops team to do before their next AI deployment:
- Baseline before you build. Record where you stand across all four categories now. You cannot reconstruct this later.
- Pick two or three metrics per category — not twenty. Measurement you can't sustain is measurement you'll abandon.
- Model your currency and load-shedding exposure explicitly. These aren't edge cases in South Africa; they're the operating environment.
- Review on a rhythm. ROI tracking isn't a launch checklist, it's a quarterly discipline that lets you iterate the strategy instead of defending it.
The teams that win with AI aren't the ones with the most advanced models. They're the ones who can tell you, in plain numbers, exactly what their AI is worth — and who started counting on day one.
So the honest question worth sitting with: if your CFO asked tomorrow what your AI initiative has returned, could you answer across all four categories — or just say the team likes it? We'd genuinely like to hear how your team is measuring it.