Agentic AI · 29 September 2026
Do you really need an AI agent?
You need an AI agent only when a repeated job has enough judgement to earn one. Start with the part you can take back, then put a person in front of anything that sends, spends, deletes, or changes a record.
When somebody offers an AI agent, do not begin by asking which agent to buy. Put the job on the table first.
Start with the job, not the product
the first questionChoose one job that returns every week. It might be sorting enquiries, checking invoices, or pulling the same facts from a stack of files. Name the input, the output, and the moment where a wrong answer reaches another person.
If the job is still "make the business smarter," there is no job to hand over yet. You have a vague project, and vague projects are where expensive promises grow.
The useful question is smaller. Does the job need somebody to read what arrives and choose the next step? Or does it follow a fixed path every time?
The reversibility test
write this firstI use one line before I let a tool touch work. If I can take the result back, it can run. If another person, their money, or their record feels the result, it waits for me.
Writing a draft into a local folder is usually reversible. A weekly list of missing details is reversible. A message in an unsent drafts folder is still reversible because you can read it and remove it.
Sending the message is different. Moving money, deleting a file, and changing a booking also need an approval step that is visible before the action happens.
| Job outcome | Can it run on its own? | What makes it safe |
|---|---|---|
| A local review draft | Usually | It stays in a named folder for you to read |
| A list of missing information | Usually | It points out gaps without changing the source files |
| A customer email, payment, deletion, or booking change | Wait | A person sees the exact action and approves it first |
I ran the same safe shape two ways
saved evidenceI kept the test deliberately boring. A practice folder held 6 fictional enquiries and a response guide. The job was to refresh one local action list, show which items need a person, and leave the input files alone.
In Claude Code Desktop, I saved that as a local routine. I ran it once by hand, then a later scheduled run appeared without another Run now click. It needed one-time permission to replace the existing report. After that approval, it completed.
The raw scheduled report was useful, and it was not clean enough to trust without reading. It counted the days before a course incorrectly and added a few actions that the practice files did not support. I kept that raw output, checked it against the guide, and wrote a corrected report beside it. That is the part many demos leave out.
For the plain route, I copied the same 6 fictional enquiries and response guide into a small fixed script. The copied practice files stayed the same for both routes. The script wrote the same local report filename. It has no instruction to send a message, change a booking, or edit an input file.
| Route | What it reads | What it writes | Where judgement lives | Cost evidence |
|---|---|---|---|---|
| Claude Code local routine | The same 6 copied practice enquiries and response guide | A refreshed local action list | The task can read the case and sort it, then waits when permission is needed | It ran under the saved account, but the retained record has no per-run dollar receipt |
| Plain automation | The same 6 copied practice enquiries and response guide | A local action list with fixed cases | The rules sit in the script, so a changed case needs the script changed | The recorded run used no paid API, model, cloud service, or network request: zero external service charge only |
What the test does and does not prove
keep the boundary honestThis test shows a useful boundary. A scheduled agent can handle a report that needs it to read the current case, while a fixed automation can handle a report whose cases and rules are already known.
My recorded routine needed permission before it replaced its existing output. I only observed one scheduled firing, so I do not treat it as evidence of unattended work over time. The saved routine records its model and effort setting, but no dollar cost reading. The fixed script used no paid API, model, cloud service, or network request in its recorded run. That proves zero external service charge for that run, not zero electricity, computer time, or maintenance work.
Claude Code's current documentation makes another difference clear. A local Desktop task runs on the machine and needs the app open and the computer awake. A cloud Routine can run when the computer is off, but that is a different setup with its own access decisions.
When an agent earns the extra moving parts
the buying decisionUse ordinary automation when the job follows fixed rules and the input stays predictable. A calendar reminder, a spreadsheet formula, or a short script can be the honest answer. You do not need a model to rename a file the same way every time.
Use an agent when the same job arrives in many shapes and somebody normally has to read, compare, or decide what matters next. The agent earns its place when it reduces that reading while leaving the irreversible part for you.
Pay a builder when they can name your real exceptions and the approval boundary. They are not selling a word when they can show what the tool reads, what it writes, what it will never do alone, and where you inspect it.
Leave the project alone until the repeated job has a name. A tool cannot repair a process before that work has been named.
Your first reversibility test
try it this weekPick one recurring job. Choose a job you understand well and can discard after the test.
Write the boundary. Say which files it may read, which local output it may create, and which actions always need you.
Run it once. Keep the input, prompt, output, permission choice, and date together.
Check one source line. Find one claim in the output and trace it back to the source file before you trust the rest.
Only then schedule it. A saved task, a manual test, and a later scheduled run are different pieces of evidence.
The whole thing on one card
if you read nothing elseYou do not need an AI agent because a job repeats. You need one when the job still needs judgement after you have named the rules.
Keep it reversible
Drafts and reports stay in a folder where you can read and remove them.
Wait for a person
Messages, payments, deletion, and record changes stop before they affect somebody else.
Use the simpler route
Fixed path, fixed rules may only need ordinary automation.
The job arrives in different shapes, and somebody must still decide what matters next.
If another person, their money, or their record feels it, put a person before it.
Keep the input, prompt, output, and permission choice together. A schedule is only proved by its later run.
Where do you stand?
15 short situations about handing a job to an AI tool. In each one you pick what you would do next.
- ▸Whether you sized the job right
- ▸What you handed over
- ▸How you checked what came back
- ▸When you stopped
About 5 minutes. You see your score, your weakest area and the reasoning behind every answer before anybody asks you for an address.
Questions people ask next
answered in one line eachDo I really need an AI agent?
You need one when the same job keeps arriving in different shapes and a person normally has to read the case before choosing the next step. A fixed repeated task may need only a script, formula, or ordinary automation.
What should an AI agent never do on its own?
Put a person before actions that send a message, spend money, delete a file, or change a record. The exact boundary depends on your work, but it should be written before the first run.
Can a scheduled Claude Code task run without my computer?
A local Claude Code Desktop task runs on the machine, with the app open and the computer awake. Claude Code also documents cloud Routines, which are a separate setup that can run when the computer is off.
Is a script cheaper than an AI agent?
The fixed script in this test used no paid external service. Somebody still writes and maintains its rules, so compare the real setup, maintenance, and review work for one job before deciding.