Skip to content

Agentic AI · 29 September 2026

Do you really need an AI agent?

You need an AI agent only when a repeated job has enough judgement to earn one. Start with the part you can take back, then put a person in front of anything that sends, spends, deletes, or changes a record.

When somebody offers an AI agent, do not begin by asking which agent to buy. Put the job on the table first.

Start with the job, not the product

the first question

Choose one job that returns every week. It might be sorting enquiries, checking invoices, or pulling the same facts from a stack of files. Name the input, the output, and the moment where a wrong answer reaches another person.

If the job is still "make the business smarter," there is no job to hand over yet. You have a vague project, and vague projects are where expensive promises grow.

The useful question is smaller. Does the job need somebody to read what arrives and choose the next step? Or does it follow a fixed path every time?

One recurring job produces a removable local result while outward actions remain closed.
Start with the small pile that returns to your desk, not the product name in the sales call.

The reversibility test

write this first

I use one line before I let a tool touch work. If I can take the result back, it can run. If another person, their money, or their record feels the result, it waits for me.

Writing a draft into a local folder is usually reversible. A weekly list of missing details is reversible. A message in an unsent drafts folder is still reversible because you can read it and remove it.

Sending the message is different. Moving money, deleting a file, and changing a booking also need an approval step that is visible before the action happens.

The reversibility test
Job outcomeCan it run on its own?What makes it safe
A local review draftUsuallyIt stays in a named folder for you to read
A list of missing informationUsuallyIt points out gaps without changing the source files
A customer email, payment, deletion, or booking changeWaitA person sees the exact action and approves it first

I ran the same safe shape two ways

saved evidence

I kept the test deliberately boring. A practice folder held 6 fictional enquiries and a response guide. The job was to refresh one local action list, show which items need a person, and leave the input files alone.

In Claude Code Desktop, I saved that as a local routine. I ran it once by hand, then a later scheduled run appeared without another Run now click. It needed one-time permission to replace the existing report. After that approval, it completed.

The raw scheduled report was useful, and it was not clean enough to trust without reading. It counted the days before a course incorrectly and added a few actions that the practice files did not support. I kept that raw output, checked it against the guide, and wrote a corrected report beside it. That is the part many demos leave out.

For the plain route, I copied the same 6 fictional enquiries and response guide into a small fixed script. The copied practice files stayed the same for both routes. The script wrote the same local report filename. It has no instruction to send a message, change a booking, or edit an input file.

Same safe job, different route
RouteWhat it readsWhat it writesWhere judgement livesCost evidence
Claude Code local routineThe same 6 copied practice enquiries and response guideA refreshed local action listThe task can read the case and sort it, then waits when permission is neededIt ran under the saved account, but the retained record has no per-run dollar receipt
Plain automationThe same 6 copied practice enquiries and response guideA local action list with fixed casesThe rules sit in the script, so a changed case needs the script changedThe recorded run used no paid API, model, cloud service, or network request: zero external service charge only
Two routes receive the same copied notes and each leaves a report for a person to inspect.
Both routes leave a file for a person to inspect. The difference is whether the next case is decided from the files or from rules somebody already wrote.

What the test does and does not prove

keep the boundary honest

This test shows a useful boundary. A scheduled agent can handle a report that needs it to read the current case, while a fixed automation can handle a report whose cases and rules are already known.

My recorded routine needed permission before it replaced its existing output. I only observed one scheduled firing, so I do not treat it as evidence of unattended work over time. The saved routine records its model and effort setting, but no dollar cost reading. The fixed script used no paid API, model, cloud service, or network request in its recorded run. That proves zero external service charge for that run, not zero electricity, computer time, or maintenance work.

Claude Code's current documentation makes another difference clear. A local Desktop task runs on the machine and needs the app open and the computer awake. A cloud Routine can run when the computer is off, but that is a different setup with its own access decisions.

When an agent earns the extra moving parts

the buying decision

Use ordinary automation when the job follows fixed rules and the input stays predictable. A calendar reminder, a spreadsheet formula, or a short script can be the honest answer. You do not need a model to rename a file the same way every time.

Use an agent when the same job arrives in many shapes and somebody normally has to read, compare, or decide what matters next. The agent earns its place when it reduces that reading while leaving the irreversible part for you.

Pay a builder when they can name your real exceptions and the approval boundary. They are not selling a word when they can show what the tool reads, what it writes, what it will never do alone, and where you inspect it.

Leave the project alone until the repeated job has a name. A tool cannot repair a process before that work has been named.

Your first reversibility test

try it this week
  1. Pick one recurring job. Choose a job you understand well and can discard after the test.

  2. Write the boundary. Say which files it may read, which local output it may create, and which actions always need you.

  3. Run it once. Keep the input, prompt, output, permission choice, and date together.

  4. Check one source line. Find one claim in the output and trace it back to the source file before you trust the rest.

  5. Only then schedule it. A saved task, a manual test, and a later scheduled run are different pieces of evidence.

A checked report can move to review; mail, payment and deletion wait beyond the approval line.
The report can move to your review folder. The envelope, payment card, and file bin stop at the approval line.

The whole thing on one card

if you read nothing else
The answer

You do not need an AI agent because a job repeats. You need one when the job still needs judgement after you have named the rules.

Keep it reversible

Drafts and reports stay in a folder where you can read and remove them.

Wait for a person

Messages, payments, deletion, and record changes stop before they affect somebody else.

Use the simpler route

Fixed path, fixed rules may only need ordinary automation.

The agent test

The job arrives in different shapes, and somebody must still decide what matters next.

The approval test

If another person, their money, or their record feels it, put a person before it.

The proof

Keep the input, prompt, output, and permission choice together. A schedule is only proved by its later run.

Free · no sign-up to see your result

Where do you stand?

15 short situations about handing a job to an AI tool. In each one you pick what you would do next.

  • ▸Whether you sized the job right
  • ▸What you handed over
  • ▸How you checked what came back
  • ▸When you stopped

About 5 minutes. You see your score, your weakest area and the reasoning behind every answer before anybody asks you for an address.

Take the readiness check

Questions people ask next

answered in one line each
Do I really need an AI agent?

You need one when the same job keeps arriving in different shapes and a person normally has to read the case before choosing the next step. A fixed repeated task may need only a script, formula, or ordinary automation.

What should an AI agent never do on its own?

Put a person before actions that send a message, spend money, delete a file, or change a record. The exact boundary depends on your work, but it should be written before the first run.

Can a scheduled Claude Code task run without my computer?

A local Claude Code Desktop task runs on the machine, with the app open and the computer awake. Claude Code also documents cloud Routines, which are a separate setup that can run when the computer is off.

Is a script cheaper than an AI agent?

The fixed script in this test used no paid external service. Somebody still writes and maintains its rules, so compare the real setup, maintenance, and review work for one job before deciding.