[ Where agents pay, and where they do not ]
A plain guide to the work AI agents do well, the work they do badly, and how to tell which is which before you commission anything. Written from what we have shipped and what we have watched fail.
A plain guide to the work AI agents do well, the work they do badly, and how to tell which is which before you commission anything. Written from what we have shipped and what we have watched fail.
Bring us a task and we will tell you honestlyThe best candidates are tasks a careful new joiner could learn in a week from a written procedure. Agents are good at following a procedure across systems.
Drafting, gathering, sorting and preparing are safe to automate because a mistake costs a correction. Sending, paying and deleting are not, and stay behind approval.
The arithmetic only works when the task happens hundreds of times. A monthly job is cheaper to keep doing by hand.
If nobody can tell whether the result was right, nobody can tell when it stops being right. Tasks with a checkable answer come first.
Asked
What is the refund window on a bulk order?
Retrieved from your documents
Answered
Bulk orders over ₹50,000 can be returned
within 21 days of delivery, against the
standard 7 days. The goods must be unopened
and in original packaging.
The model arranged the sentence. Every fact in it — the amount, the window, the condition — came from a document you own, and the source is attached so a wrong answer can be traced rather than argued about.
Enquiry triage and routing, invoice and document extraction, CRM hygiene and enrichment, first-draft replies, report preparation, reconciliation of two lists, and chasing missing documents.
Anything contractual or regulated without review, judgement calls with no right answer, one-off tasks, work needing context that exists only in someone's head, and anything where a wrong action cannot be undone.
Four questions that sort almost any task: how often does it happen, what does a mistake cost, can the result be checked, and how many systems does it touch.
Volume per month and minutes per instance. Below a few hundred instances a month an agent rarely repays the engineering and the ongoing supervision.
A wrong draft costs a minute. A wrong payment costs a lot more. This single question sorts most candidates into automate, assist, or leave alone.
Is there a right answer someone could verify. Without that there is no evaluation set, and without an evaluation set you cannot tell degradation from noise.
Each one needs an API, credentials and a permission model. Three systems is a project; seven is usually a different, longer project.
Fully automatic, approve-before-commit, or draft-only. Almost everything worth doing starts at draft-only and moves inward on evidence.
One workflow, one team, full approvals, four weeks. The completion and intervention rates from that decide whether it widens.
Agents are genuinely useful for a narrower band of work than the current enthusiasm suggests, and the band has clear edges. They are good at following a written procedure across several systems, hundreds of times a month, where the output can be checked and a mistake can be undone. They are poor at judgement, at anything irreversible, and at tasks whose real rules live in an experienced person's head and have never been written down.
The four questions on this page — volume, reversibility, verifiability, and how many systems are involved — will sort most of your candidate list in an afternoon, before anyone writes a brief. We would rather you arrived with a good candidate than commissioned a poor one.
Almost every worthwhile agent begins by preparing work for a person to approve, and earns each removed gate with measured evidence.
Effort scales with integrations, not with cleverness. Two systems is a sensible first project; six is a programme.
No verifiable answer means no evaluation set, and no way to notice the day it quietly starts being wrong.
The general-purpose autonomous assistant has largely given way to narrow, single-workflow agents with approvals — because those are the ones still running after six months.
Draft-and-approve is now the standard starting point rather than a cautious variant, with autonomy granted per action on evidence from the traces.
The model is a small part of the bill. The number of systems and the state of their APIs is what the estimate actually tracks.
Someone has to read the exception queue every week. Projects that did not budget for that quietly stopped being trustworthy.
High-volume, rule-shaped work across two or three systems where the output can be checked and a mistake can be undone: enquiry triage and routing, invoice and document extraction, CRM enrichment and hygiene, first-draft replies, reconciliation, and chasing missing paperwork. The common thread is that a careful new joiner could learn the task from a written procedure.
Let’s talk about your agent use cases project. No obligation, just a conversation.
Next service
Custom Ecommerce Development