Independent field guidance. Some links may earn us a commission, at no added cost to you.
Field guide · 9 minute read

How to test an AI agent before you trust it.

Do not start with your biggest workflow. Start with one recurring annoyance, a stack of known source material, and a finish line you can inspect.

Published September 20, 2026 · No affiliate links in this guide

Most AI agent demos begin after the hard decisions have already been made. The data is clean. The request is polished. The expected answer is known. Real work rarely arrives that way.

Pick a task that is boring, not important

Your first test should happen often enough to matter and be safe enough to fail. A weekly status summary is better than a customer refund. Drafting a meeting recap is better than changing production data.

Write the job in one sentence. If it needs a paragraph full of exceptions, the process probably needs cleanup before it needs an agent.

A useful first test“Every Friday, read these three project documents and draft a status note that lists completed work, blockers, and the owner of each next action.”

Give it a finish line

“Help with operations” is not a task. A finished report with five named fields is. Decide what must be present, where the result should appear, and what would make you reject it.

Keep a copy of the expected structure beside the test. You are not trying to see whether the agent can surprise you. You are trying to see whether it can follow a repeatable definition of done.

Control the source material

For the first run, use a closed set of documents. Ask the agent to cite the file or message behind each factual claim. Then add one deliberate trap: an old file, a blank field, or a contradiction that should trigger a question.

If the agent smooths over the gap instead of flagging it, you learned something useful before connecting a larger data set.

Draw the approval line in plain English

Drafting and acting are different permissions. An agent may be allowed to prepare an email while being forbidden to send it. It can identify stale records without deleting them. It can build a campaign plan without purchasing an ad.

List the actions that always require a person: sending, publishing, deleting, spending, changing access, signing an agreement, or moving money. Test at least one of those boundaries on purpose.

Make failure visible

A dependable system does not merely succeed. It fails in a way you can understand. Disconnect a source or remove a required field. The agent should name the missing input, preserve the work it completed, and stop before guessing.

“Done” is the wrong message when half the job failed. Look for a short receipt: what ran, which sources were used, what changed, and what still needs attention.

Count rework

Timing the first answer is easy. Timing the cleanup is more honest. Record how long it takes to prepare the task, review the result, correct mistakes, and run it again. Include credits or usage fees.

If a ten-minute task becomes a three-minute prompt followed by twelve minutes of checking, the agent did not save time. It moved the work around.

Pass standard

Three clean runs before expansion

Do not connect more systems after one good result. Run the same job three times with slightly different source material. A pass means the output is complete, the evidence is traceable, the approval boundary holds, and failures are reported without guessing.

What to write down

  • The exact task and finish line
  • The allowed source material
  • Actions that require approval
  • Expected output location and format
  • Time spent preparing, reviewing, and correcting
  • Credits or usage cost per run
  • What happened when a required input was missing

A note about our own evaluations

Product Field Guide separates product claims, public documentation, and direct use. If we have not run a feature ourselves, we will not write as if we did. An affiliate relationship may support the work, but it will be labeled and will not turn a product into a winner by default.

Have a task in mind?

Use the worksheet to see whether it is specific, inspectable, and safe enough for a first run.

Open the first-task worksheet →