AI

Transforming Operations with AI: Beyond Chatbots and Customer Support

Almost every AI conversation I am invited into starts in the same place: someone wants a chatbot.

It is the most visible thing you can build with AI, and usually the least valuable. It sits on the website where the board can see it, it demos well in a meeting, and it changes almost nothing about how the business actually runs. Meanwhile the work that quietly costs the most — the invoices being keyed in twice, the stock that is ordered on instinct, the report that takes two days to assemble every month — carries on untouched.

This is where AI earns its keep. Not at the front desk, but in operations.

The numbers everyone quotes, and the one that matters

McKinsey’s State of AI, published in November 2025, found that 88% of organisations now use AI in at least one business function. That is the number that gets quoted in every presentation.

Two others from the same study matter far more. Nearly two-thirds of those organisations have not begun scaling AI across the enterprise. And only 39% can attribute any EBIT impact to it at all — with roughly 6% seeing more than 5%.

Almost everyone is using AI. Almost nobody is getting paid for it.

The organisations in that top group are not distinguished by better models. Everyone rents the same models. What separates them is that they redesigned the workflow around the tool, rather than bolting the tool onto a workflow that was designed for people.

That is the whole argument of this article, and it is why chatbots disappoint. A chatbot is bolted on by definition. It sits in front of a process nobody changed.

Diagram contrasting front-office AI such as chatbots with operational AI such as document processing, forecasting and anomaly detection, showing visibility against business value
The most visible AI projects and the most valuable ones are rarely the same projects.

Why operations is the better place to start

Three reasons, in order of how much they matter.

  • The data already exists. Operational work leaves a trail — invoices, tickets, timesheets, delivery records, purchase orders. Customer-facing work often does not. You cannot train or evaluate anything on conversations nobody recorded.
  • The measurement is honest. If invoice processing took four days and now takes four hours, that is a fact. Chatbot success gets measured in deflection rates, which is a polite way of counting people who gave up.
  • The risk is contained. A mistake in a back-office workflow is caught by a human before it reaches a customer. A mistake in a public chatbot is a screenshot on social media.

How to choose the first use case

Most failed AI projects were doomed at selection, not at implementation. Before anyone talks about models or vendors, put the candidate process through four tests. It needs to pass all four.

Test The question Fails when
Volume Does this happen hundreds of times a month? It is important but rare — the payback never arrives
Rules Could you write down how a good decision is made? Every case is judgement, and the judgement is not written anywhere
Trail Is there a record of past cases and their outcomes? The knowledge lives in one experienced person’s head
Tolerance Is a 5% error rate survivable with a human check? One wrong output is a legal or safety problem
Diagram showing the four tests for choosing a first operational AI use case: volume, rules, data trail and error tolerance
A process has to pass all four. Most of the ideas that arrive in the room fail on the third.

In my experience the third test kills the most proposals, and it is the one nobody expects. A company will describe a process confidently, then discover that no record exists of what was decided or why — only the outcome. That is a data problem, and it has to be fixed before it is an AI problem.

Where it actually works

Forget the sector labels. Operational AI is worth considering when work takes one of five shapes. All five are variations on automating the repetitive middle of a process rather than the judgement at either end of it.

Document-heavy work. Invoices, purchase orders, permits, contracts, HR files, insurance claims. Anything where a human reads a document and types its contents into a system. This is the single most common win in the GCC because so much of the region’s back office still runs on PDFs and email. It is also the work that connecting AI to your own systems makes practical rather than theoretical.

Forecasting. Demand, stock levels, staffing, cash flow, energy load. Anywhere a person currently guesses from last year’s spreadsheet. The bar is not perfection — it is beating the guess, which is usually a low bar.

Scheduling and routing. Fleet movements, field engineers, appointment slots, shift rotas. Constraint problems where a small percentage improvement compounds daily.

Anomaly detection. Fraud, duplicate payments, quality defects, unusual access patterns. Machines are better than people at noticing that one row in fifty thousand looks wrong.

Knowledge retrieval. Policy documents, procedures, past projects, regulatory guidance. The value here is not clever answers — it is that a new employee stops interrupting a senior one twelve times a day. If you want the detail on how the underlying models behave, I covered that in using large language models in a Kuwaiti business.

The three ways it goes wrong

Pilot purgatory. A successful proof of concept that never becomes a production system, because nobody budgeted for integration, security review or the boring work of connecting it to the ERP. The pilot proves the model works. It does not prove the organisation can absorb it. Budget the second half before you start the first.

No owner. The IT department owns the tooling, the operations manager owns the process, and the AI vendor owns neither. When the model’s output disagrees with the old process, no one has the authority to decide which changes. The project stalls in a meeting. This is the same gap I wrote about in the decision nobody in the room was qualified to make, and it needs the same fix: one person with authority across the whole chain.

Automating the wrong version of the process. Most processes have accumulated steps that exist because of a system limitation from 2015 or a manager who has since left. Automate that faithfully and you have made the wrong thing faster and much harder to change. Map the process first, delete what should not exist, then automate what remains.

The Kuwait context

This is not happening in isolation. Kuwait’s National AI Strategy 2025–2028 sets out a phased plan: a centre of excellence and pilot projects first, then wider sectoral deployment, with full integration targeted by 2028. Across the border, Abu Dhabi has committed to becoming the world’s first fully AI-native government by 2027.

For a private business the consequence is procurement, not policy. Government and semi-government buyers whose own operations are being rebuilt around structured data and APIs will expect the same from their suppliers. Being able to explain how your systems handle data is becoming part of qualifying for work, whether or not you ever deploy AI yourself.

What it costs and how to sequence it

A first operational use case is not a transformation programme, and it should not be priced like one — the wider systems-and-data work is a separate decision, taken later and on its own evidence. The realistic shape is a short diagnostic, one narrow process, a working system in production, and a measured before-and-after. If a proposal starts with a platform licence and an eighteen-month roadmap, someone is selling you the second project before the first has proved anything.

  1. Pick one process that passes all four tests, and write down what it costs today in hours and errors.
  2. Fix the data trail if there is not one. Frequently this alone pays for itself.
  3. Redesign the process assuming the machine does the first pass and a person reviews exceptions — not the other way around.
  4. Build it small and put it in production. A pilot that never ships teaches you nothing about your organisation.
  5. Measure the same numbers you wrote down in step one, then decide whether there is a second use case worth doing.

Done in that order, the first project is usually cheap enough that the argument about AI strategy stops being theoretical. Done in the wrong order, you get a chatbot.

Frequently asked questions

Is a chatbot ever the right first AI project?

Sometimes — if your support volume is genuinely high, your answers are well documented, and deflection saves measurable staffing cost. What makes it the wrong default is that it is chosen for visibility rather than value. If the honest reason for building one is that the board wants to see AI on the website, the money is better spent elsewhere.

How long does a first operational AI project take?

For a single well-chosen process, expect weeks rather than quarters to reach a working system, then a period of running it alongside the existing process to compare results. Projects that stretch beyond that are usually blocked on data access or on a decision nobody is authorised to make, not on the technology.

Do we need our own AI model or data scientists?

Almost never for a first project. The models are rented, and the work is integration, data quality and process design rather than research. Hiring data scientists before you have a use case with a data trail is a common and expensive way to start.

What if our data is messy?

It will be — everyone’s is. The question is whether a record exists at all. Messy but present data can be cleaned as part of the project. Absent data cannot, which is why the trail test matters more than the volume test.

Does this work in Arabic?

For document processing and knowledge retrieval, yes, and noticeably better than two years ago — though Arabic still needs more evaluation than English, particularly with mixed-language documents and scanned handwriting. Test on your own documents rather than trusting a vendor’s demo, which will have been built in English.

Where to start

If you want to know whether AI can do anything useful in your operations, the answer starts with a process map and a look at what data you are already sitting on — not with a vendor demo. That is usually a short piece of work, and it either finds a candidate that passes all four tests or it saves you from an expensive project that was never going to pay.

If you would like an independent read on that, see how I work on AI or tell me what your operation looks like and I will tell you where I would look first.

 

Have a project, problem or idea?

Let's discuss what you're trying to build, improve or grow — and whether I can help.

Discuss Your Project