Skip to content

AI for business · Long read

AI agents for small businesses: from first process to measurable pilot

A long-form guide with a complete fictional case: scope, documents, permissions, evaluation, costs and maintenance of an AI workflow.

By Eric Muriel15 min read
An input, review and output pipeline with action boundaries.
Original Eric Muriel illustration for this topic; not a product screenshot.Illustration: Eric Muriel · © Eric Muriel

01Define the job before choosing an agent

An AI agent can retrieve information and use tools to complete a task. Anthropic’s December 2024 article makes a useful distinction: workflows follow predefined steps, while agents decide parts of their own path. We use that distinction as a conceptual starting point. Greater autonomy is not assumed to produce better business outcomes, and this guide does not reproduce a vendor implementation.

The remaining sections develop an original proposal around a fictional company, North Workshop. It receives maintenance requests by email and turns them into draft work orders. A request may contain a machine, a symptom and a location. The first objective is to prepare a record for human review. Sending quotations, assigning technicians and promising dates are outside this version.

Write a testable sentence: “Given a request, prepare a record containing the machine, symptom, location and missing information, linked to the original message.” Agree on this before choosing tools. Save two examples of acceptable records and one dangerous example with an invented location. Those examples will become the beginning of your test collection.

Reference [1]: Anthropic

02Observe the manual process and its exceptions

Before automating, observe the person doing the work. Record what they read first, where they look up a reference and when they ask someone else. A general interview often produces “we read the email and create an order.” Concrete examples reveal details: one machine has two names, the sender is not necessarily at the affected location, and an attachment may correct the message.

For this exercise, prepare twenty fictional requests: ten complete, four missing a location, three with ambiguous identifiers and three unrelated to maintenance. This is a learning sample selected for the guide, not an estimate of real company traffic. Describe the expected result for every case before trying a model. Otherwise, it is tempting to redefine correctness around whatever the system generates.

Measure complete manual handling time for each class. Reading a draft is not equivalent to correcting and saving it. Comparing the full manual process with generation time alone exaggerates savings. Keep the exceptions in the sample; they may reveal more about the required design than the straightforward requests.

03Decide what a rule can resolve

Separate the process into decisions. Checking whether an identifier exists, validating a date and preventing duplicate work orders can be expressed deterministically. Turning “it has made a strange noise since yesterday” into a readable summary needs different treatment. Reserve the model for language interpretation and keep explicit checks in code or visible rules.

North Workshop’s first proposal has four steps: receive the request, extract fields, validate the result and save a draft. The model cannot select arbitrary tools or search every folder. Whatever label a vendor uses, understand the actual behaviour. Draw the path and mark every decision that depends on the model. Those decisions require examples and review criteria.

Try this limited version first. Additional autonomy should address a demonstrated need, such as choosing between two authorised sources based on equipment type. Even then, specify which decisions are allowed and how many attempts are available. “Investigate until solved” is not an operational boundary. A short path that stops when information is missing can be more useful than an extended, convincing conversation.

04Design a record that allows unknown values

Define the output before writing instructions. Our record contains request_id, machine_id, symptom, location, missing_information, evidence and status. Specify a type and acceptance condition for each field. Location can be text or an explicit empty value; missing_information is a list; status permits only draft or needs_review. The request identifier comes from the input system rather than the model.

Specify conflict handling. If the message says building A and its attachment says building B, the record must expose that disagreement. Selecting whichever text appeared last is not a reliable resolution. Preserve both pieces of evidence and identify what needs confirmation. An incomplete but faithful record is preferable to a seemingly complete record that hides an assumption.

Prepare examples with both sufficient and missing information. Require extracted values to be linked to a passage or an authorised lookup. Where possible, verify the relationship outside the model: does the identifier exist in the permitted catalogue, and does that location belong to the equipment? This separates formatting failures, extraction errors and genuine business ambiguity.

05Prepare a small document collection

If the record needs internal instructions, start with a small reviewed collection: an equipment catalogue, a location table and a classification procedure. Record the owner, current version and permitted audience for each document. A folder containing duplicated files does not become reliable simply because semantic search is added.

Include an outdated document deliberately. The previous catalogue places machine M-17 in building A; the current one places it in building B. Define which date or version governs the lookup. If the business has not established that rule, the system cannot independently determine which relocation was authorised. A data owner must resolve it.

Separate reference material from third-party submissions. A sentence in an attachment has no authority to change an export destination or expand permissions. Treat incoming content as information to analyse. Tool access, destinations and permissions belong in the application. The companion RAG guide provides a further exercise for testing retrieval, currency and evidence separately. Record the chosen source version alongside each result so reviewers can reconstruct what the system actually saw.

06Turn permissions into a testable list

Write a simple table containing tool, operation, scope and accountable person. In the first version, North Workshop’s account can read a test folder and create drafts in one list. It cannot delete documents, send email or modify the catalogue. Enforce those boundaries through credentials and server checks as well as instructions. Text saying “do not delete” does not remove deletion permission.

Test two fictional users. Alice may access building A equipment; Bruno may access building B. Ask the same question under both identities and confirm that records from the other scope are excluded. Repeat with indirect questions, such as the latest incident involving a machine. Hiding a row in the interface is not equivalent to controlling access at the source.

Document how access will be revoked after the pilot. Avoid reliance on a personal account whose owner may change roles. Store credentials in server configuration and document which connection uses them without copying secrets into the workbook. The deliverable should be a short list of verifiable capabilities rather than a general claim of security.

07Write instructions with decision examples

Useful instructions specify the task, permitted sources, output format and stopping conditions. For the exercise: “Prepare a maintenance draft using the request and authorised catalogue. Do not invent location or priority. Preserve conflicting evidence and mark needs_review.” Add the examples prepared when defining the record. Express decisions instead of accumulating adjectives such as expert, precise or infallible.

Include an irrelevant request: “Could you send your commercial catalogue?” The expected result is to classify it outside the process without inventing an incident. Add a message that attempts to instruct the system to change its configuration. Verify that it remains request content rather than becoming an authoritative instruction. These checks do not establish universal protection, but they expose specific failures before real data is connected.

Version instructions alongside examples. When changing a sentence, rerun affected cases and compare results. Avoid changing the model, catalogue and instructions simultaneously because the cause of a regression becomes unclear. Record a short reason for each change, especially when it alters handling of missing information or out-of-scope requests.

08Build a proposal-only version first

The first version can use local files and a draft screen. It does not need to send messages or write to the final system. Make progress visible: input received, extraction complete, validation passed and draft available. Preserve the request identifier when a stage fails. “Location absent from catalogue v3” points towards a next action; “something went wrong” does not.

In the fictional pilot, a reviewer sees the original message beside the proposed fields. They can accept, correct or return the record for clarification. Corrections include a reason. This feedback can improve evaluation examples; it does not automatically retrain the model or make every edit a permanent rule.

Avoid features that do not test the initial hypothesis. A dashboard with twenty charts does not establish whether drafts save work. At this stage, capture what arrived, what was proposed, what a person corrected and how long the complete process took. Once those questions are answered, decide whether integrating the work-order system is worthwhile. The prototype produces useful evidence even if the decision is to stop.

09Review the exact action before execution

If real work orders are later enabled, approval must show the exact machine, location, description, destination and effect. A button labelled “approve agent” is too broad for a specific action. Reviewers should compare the draft with the original request and see whether anything changed after review. A relevant change requires renewed approval before execution.

In this design, approval is bound to the request identifier and draft version. The server checks that relationship when creating the order. A model response containing “approved” is insufficient: the decision must come from an authenticated interaction by an authorised person. Keep the approver distinct from the component that generated the proposal, and record the destination system’s returned identifier.

Define what happens if nobody responds. During the pilot, the draft remains pending in a review queue; a timer does not send it automatically. Reminders or internal escalation can be explicit rules. Measure the queue’s size and age. Otherwise, automation may simply transfer effort into an accumulating review backlog.

10Prevent duplicates when a connection fails

Imagine the work-order system saves a record but the connection drops before confirmation arrives. Retrying without checking may create a second order. This problem is independent of generated text quality. Define a stable operation key, such as request plus action type, and preserve its relationship with the destination identifier. The precise strategy depends on the API’s guarantees.

In a local test, receive request S-17 twice and expect one associated order. Then simulate failure after saving but before responding. Recovery should inspect recorded state or reconcile with the destination instead of assuming every error means nothing happened. If the outcome cannot be established, mark the operation uncertain and route it for review.

Do not identify requests solely by message text: two customers may use identical wording for different jobs. Do not generate a new key on every retry either. Document which operations can be repeated safely and which require checking state. The site’s idempotency tutorial offers practice with fictional data before adapting the pattern to a real integration.

11Evaluate quality, omissions and behaviour

A single score hides important distinctions. Separate correct fields, invented information and correct review decisions. For a request without a location, leaving the field empty and asking for clarification is successful behaviour. Filling it with a plausible location is an error even when the remaining text reads well. Also count unnecessary rejections that force manual rework.

For the original twenty examples, record expected outcome, observed outcome, evidence and decision. Review failures with someone who understands the process. An apparent error may reveal an outdated catalogue. Disagreement about priority may expose the absence of a shared rule. Resolve those ambiguities instead of adjusting a score until it looks satisfactory.

Keep additional examples out of instruction development. Include short, long, contradictory and out-of-scope requests, and repeat selected cases to observe variation. A small collection cannot guarantee future performance; it helps identify problems and select the next experiment. Define stopping conditions, such as unauthorised access or duplicate creation, separately from minor wording improvements.

12Calculate complete cost with explicit numbers

These invented numbers teach the calculation; they are not vendor prices or a commercial forecast. Assume 600 monthly requests taking six minutes each manually: 60 hours. If the new process needs two minutes of review per request, that is 20 hours. Add eight hours for exceptions and maintenance, leaving 32 hours of released capacity if quality remains acceptable.

At a hypothetical internal value of €20 per hour, those hours represent €640 of capacity. If tools, infrastructure and calls cost €90 monthly, €550 remains before amortising implementation and accounting for other costs. Released capacity is not automatically cash income. Its value depends on whether the team uses the time productively and how actual costs behave. Include initial data preparation and testing too.

Now test a less favourable assumption. Four minutes of review means 40 hours plus eight maintenance hours, releasing only twelve hours. At the same hypothetical valuation, €240 less €90 leaves €150. This sensitivity tells you what to measure carefully. Track per-request usage, retries, storage and human effort, then consult current provider prices for a real budget.

13Launch a pilot with a defined group and exit

A pilot should fit into one sentence: one request type, one location and two reviewers over an agreed period. Duration depends on the volume needed to encounter relevant cases. A week with no requests provides different evidence from a week with fifty. Define results that justify expansion, require changes or stop the integration.

Keep the manual procedure available. Requests must remain findable when an external service fails, and someone must know how to process them. Test that exit beforehand by disabling the integration in a test environment and handling a pending request. If only the developer knows how to recover work, there is still a significant operational dependency.

Explain which responsibilities change and which still require judgement. Here, the system proposes fields and evidence; people resolve conflicts and authorise orders. Ask for specific reports containing an identifier and expected result rather than general opinions about AI. Concrete observations support improvements and expose interface steps that add unnecessary work. Record both successful cases and interruptions so the review is not based only on memorable failures.

14Observe the system without retaining everything

Design operational records around specific questions: which request arrived, which version ran, which stage failed and what result was saved. You do not need to copy complete documents into every log. In this example, source references, version identifiers, timings, status and review reasons are sufficient. If debugging requires retaining a sample, define its access and retention according to the project’s actual needs and obligations.

Distinguish transient failures from content problems. A dropped connection may justify a bounded retry; an unknown identifier needs correction or clarification. Repeating extraction ten times cannot recover a location that was never provided. Tell operators what kind of failure occurred and what they can do without exposing credentials or other users’ information.

Prepare periodic counts of received requests, reviewed drafts, pending items, prevented duplicates, corrections and handling time. Explain every percentage’s denominator. If twenty requests arrived but only ten were reviewed, acceptance among those ten does not yet describe the complete set. Keep a path from each metric back to its contributing cases to investigate unexpected changes.

15Maintain versions and check changes

A working pilot can degrade when the catalogue, model, API or request-writing habits change. Assign responsibility for the business process and for the integration, even if one person fills both roles. List changes that require retesting: a new field, expanded permission, a different source or an update that alters expected output.

Run saved cases before deploying a change and inspect differences. If a version improves long messages but invents locations in short ones, an average must not hide that regression. Compare by case type and preserve the previous version where technically possible. Document a return to manual handling when an external dependency cannot restore its previous behaviour.

Review reference content with its owner. Correct extraction can still produce an incorrect order if it consults an obsolete table. Maintenance must examine data, permissions and review workload as well as server availability. This project belongs to a business process. If nobody can explain why a rule remains valid, review it before extending automation to more teams.

16Decide the next step with a short evidence record

At the end of the pilot, prepare a decision page containing the initial problem, tested cases, results, complete cost and known limits. Include a meaningful failure and explain how it was resolved or why it remains open. A record containing only successful examples supports a demonstration but offers little basis for deciding whether people should depend on the system daily.

Three outcomes are reasonable: expand the same process, correct a specific limitation or stop. Expansion does not mean connecting every application. Add a second location while retaining the same operations, repeat permission checks and compare review effort. Improvement may involve repairing the catalogue or removing an interface step. Stopping may be best when volume is low or the manual task already works well.

Use Eric Muriel’s downloadable workbook to record decisions and the related guides to explore data, documents and validation. Practise with local tutorial examples before connecting production systems. The goal is a process people can understand, verify and maintain. Technology should be replaceable without losing the job definition, acceptance criteria or knowledge gained during the pilot.

Sources and further reading

These references expand on the concepts indicated. The examples and exercises are original editorial material.

  1. [1] Anthropic

    Building effective agents ↗

    Conceptual reference published 19 December 2024. The case, exercises and calculations in this guide are original editorial material.

    Back to the related section
How we use sources, quotes and images

Frequently asked questions

Do I need an autonomous agent to start?

Not necessarily. This guide starts with extraction, validation and human review along a defined path. Expand autonomy only to address a demonstrated need.

Are the example costs real prices?

No. They are hypothetical figures for practising a complete calculation. Replace them with your volume, review time and current provider prices before making a decision.