There are two positions on this and neither helps. One says everything has changed and you are behind. The other says it is a toy that makes things up. Both are arguments about the technology, and you have a business rather than an opinion to run.
The question that actually sorts the work is narrower: how expensive is a wrong answer here, and can you check the output faster than producing it yourself? Two questions, and between them they decide almost every case.
Where it saves real hours today
- First drafts of things you were going to rewrite anyway. Product descriptions from a specification, a first pass at a policy, the boring middle of a long email. You were always going to edit; now you edit instead of starting.
- Sorting what arrives. Classifying incoming enquiries, tagging support messages by cause, routing. A mistake is cheap and visible, which is exactly the profile that suits it.
- Pulling data out of documents. Invoices, supplier lists, spreadsheets somebody built in 2014. Tedious, checkable at a glance, and previously an intern's week.
- Translation as a starting point. Good enough to be edited by someone who speaks the language, which is a different job from translating from scratch — and much cheaper.
- Summarising long threads so a person can decide. The summary is wrong occasionally, and the person is reading the thread anyway when it matters.
Notice the pattern. In every case a human still signs, and checking takes less time than doing.
Where it creates work instead of removing it
Anything you must verify end to end. If confirming the answer means redoing the task, you have added a step. This is where most disappointment comes from: the output looks finished, so nobody checks properly, and the errors surface later at a worse moment.
Anything where being subtly wrong is worse than being absent. Prices, stock, legal wording, medical or financial specifics, anything a customer will act on. A confident wrong number is more expensive than no number.
Anything customer-facing without supervision. The company owns what the machine said. That is not a hypothetical: businesses have been held to promises their automated assistant invented.
The question to ask about any tool being sold to you
Vendors now attach the word to everything, so a single question cuts through: what does this do when it is wrong, and how will I find out?
A good answer sounds like: it flags low confidence, a person reviews the queue, mistakes are visible within a day. A bad answer changes the subject to accuracy percentages. Accuracy is not the issue; what happens in the remaining share is the issue, and that is a process question rather than a technology one.
The three costs nobody counts
Verification. It is real work, it lands on your most senior person, and it is rarely in the business case. Count it, or the saving is imaginary.
Accountability. When the output is wrong and it reached a customer, the answer to “who decided this” has to be a person. Design that in — a name, a step, a signature — or discover the gap during the incident.
The thinking that stops happening. The subtlest and the most expensive. A team that stops drafting stops arguing about what it means; a developer who stops reading stops noticing. Not an argument against the tools — an argument for keeping the judgement work with people on purpose rather than by accident.
The rule we use
Put it where a human still signs, and where checking is faster than doing. That single line settles almost every proposal that arrives, including the ones we make to ourselves.
It also explains why the best results are unglamorous. Nobody demonstrates “we cut two hours a week off tagging support tickets” at a conference, and that is precisely the shape of automation that survives contact with a real week — the same argument as deleting back-office jobs one at a time.
How to find out in two weeks
You do not need a project or a platform. Pick one task that repeats weekly and annoys someone. Write down how long it currently takes, honestly, for two weeks. Then do it with assistance for two weeks and record both the time spent and the time spent checking.
Keep it if the total dropped and quality held. Drop it if it did not, and tell the team you dropped it — a business that can abandon something publicly will try more things than one where every initiative has to be a success.
What we do and do not automate
Ours goes into the repetitive middle: drafts we rewrite, classifying what lands in the inbox, converting messy data into a shape we can work with, the mechanical parts of moving content between systems. Every one of those has a person at the end who signs.
What we keep human: what to build, what to leave out, what to tell a client when the honest answer costs us the project, and the final read of anything that goes out with our name on it. Those are the decisions people pay a studio for, and they are also the ones where a confident wrong answer does the most damage.
