On this page
- What is AI workflow automation?
- Where AI steps work, and where they don't
- How to integrate an LLM into workflow automation
- Human-in-the-loop: four review patterns
- Worked example: supplier invoices from inbox to accounting
- How AI steps fail, and how to catch it
- What AI steps cost, and how to keep the bill small
- Data privacy: what the model sees and keeps
- The best AI workflow automation tools, compared
- How to pilot one AI step and measure it
You probably automate the predictable parts of your business already: a web form creates a CRM contact, a paid invoice sends a receipt. AI workflow automation adds steps for what rules can't handle, such as reading a supplier's PDF, telling a complaint from a billing question, or drafting a reply.
It works when the AI does the reading and writing, rules do the checking and acting, and a person approves anything that moves money or can't be undone. It fails when a model is trusted with exact math or final calls, guesses a PO number that isn't on the page, gets talked into something by an email, or changes behavior after an update nobody noticed.
What is AI workflow automation?
AI workflow automation is a rule-based workflow with one or more steps handed to an AI model, usually a large language model (LLM) such as OpenAI's GPT models or Anthropic's Claude. The trigger, conditions and actions stay rule-based. The AI step sits where the input is messy: free text, documents and images.
| Rule step | AI step | |
|---|---|---|
| Works on | Amounts, dates, statuses | Language, layouts, images |
| Same input, same output? | Always | Not guaranteed |
| When it's wrong | The same way every time | Occasionally, and plausibly |
| Cost per run | The platform's unit price | Model tokens plus the platform's unit price |
| Testing | A few sample records | Real past items, rerun after every change |
| Example | Send bills over $2,500 to the owner | Label this email billing, scheduling, complaint or other |
Two other things share the label. AI that builds workflows for you, like Copilot in Power Automate, still produces ordinary rule-based flows. AI agents for business workflow automation let the model pick its own steps and tools, a different design we cover in our guide to AI agents for business. This guide is about AI steps inside workflows you control; for the rule-based basics, see how workflow automation works.
Rule or AI? A four-question test
- Could you write it as an if-then rule or a lookup? Keep it a rule: cheaper, faster and right every time.
- Does it need to read or write language, a layout or an image? That's what AI steps are for.
- Can something check the output before it matters? A rule (the totals add up, the PO exists) or a person. If nothing can, don't hand the step to AI.
- What does a wrong answer cost? If it's cheap and reversible, the AI step can run unreviewed, with sample audits. If it involves money, customers or anything permanent, the AI drafts and a person decides.
In short: AI reads and drafts, rules check and act, people approve what's costly or permanent.
Where AI steps work, and where they don't
Good AI steps share a shape: messy input, one narrow job, and an output something can check.
- Classify and route: label inbound email, tickets or leads from a fixed list that includes "other."
- Extract to fields: pull the vendor, invoice number, PO and totals from an invoice, or the policy dates from a certificate of insurance, so rules can check them.
- Summarize: turn a long thread or a call transcript into five lines on the CRM record, linked to the source.
- Draft for review: replies, quote cover notes and collection reminders that a person edits and sends. AI answering customers on its own is a customer-facing chatbot, which needs its own handoff design.
- Enrich: turn a lead's website or intake notes into an industry, a size band and the service needed, each picked from your lists.
- Translate: show staff an English version of a customer's Spanish message, and draft the reply for review.
Keep these out of AI steps:
- Exact math: totals, tax, prorations, commissions. The model reads the numbers; code does the arithmetic.
- Exact lookups: a customer ID, a price-book price, an account code. Query the system.
- Final approvals: payments, refunds, discounts, credit limits.
- Decisions about people: hiring, tenant screening, credit, claims, medical or legal calls. AI can summarize the file; a qualified person decides.
- Irreversible actions without review: deleting records, sending money, mass messages.
- Steps a rule already handles: formatting phone numbers or routing by ZIP code through a model adds cost, delay and variation for nothing.

How to integrate an LLM into workflow automation
Zapier, Make, n8n, Power Automate and a few lines of code can all call a model. What makes an AI step dependable is what you build around it.
One job, one output format
Give each step a single task, "extract these fields" or "pick one label," and ask for structured output against a JSON schema: a machine-readable list of the fields, their types and which are required. OpenAI's Structured Outputs and Anthropic's structured outputs hold the model to your schema, and the no-code tools have equivalents. Make honesty easy: fixed categories with an "other" option, and null allowed for any field a document might lack, with an instruction to use it instead of guessing.
Validate before anything acts
A schema guarantees the shape of an answer, not its truth, and even the shape has gaps: Anthropic's documentation notes that output can break the schema when the model refuses or hits its token limit. Check every output with rules: required fields present, values in range, totals that reconcile, IDs that exist in your systems. OWASP's guidance on improper output handling puts it simply: "treat the model as any other user."
Version the prompt, schema and model together
Keep them with the test set, in Git or in exported workflow files with a change log, and rerun the tests after any change. A one-line prompt edit can shift results across thousands of items.
Test on real past items
Pull 50–200 past items with known answers: posted bills for invoice extraction, your team's past labels for email routing. Include scans, credit memos, multi-page documents, missing PO numbers and an email that tries to give the model instructions.
Plan the fallbacks
Retry once after a timeout. Anything that fails validation, comes back as "other" or hits a provider outage goes to an exceptions queue with the original attached. Never drop an item, and alert someone when the queue grows.
Human-in-the-loop: four review patterns
Human in the loop doesn't mean watching every run. It means people sit at defined points, with clear rules for what reaches them.
| Pattern | How it works | Use it for |
|---|---|---|
| Approve before send | The AI drafts; a person approves, edits or rejects | Customer replies, quotes, collection messages, payments |
| Rule-based thresholds | Items that pass every check and sit under your limits go through; the rest go to a person | Invoices under a dollar limit that match their PO |
| Exceptions queue | Failures, "other" labels and errors land in one queue with the source, the AI output and the reason | Every AI step |
| Sampling audit | A person re-checks a random sample of untouched items; misses become test cases | Every step that runs without review |
Be wary of confidence scores. Microsoft describes field confidence in Azure AI Document Intelligence as "an estimated probability between 0 and 1 that the prediction is correct," a measured number you can set thresholds on. A language model rating itself ("confidence: 0.95") isn't measured that way, so gate on checks you can compute until your test set shows the self-rating predicts errors.
Put approvals where people already work, such as Outlook, Teams, Slack or a phone. In Smart Construction, our construction ERP, payment approvals happen on WhatsApp.
Worked example: supplier invoices from inbox to accounting
Here's a hypothetical workflow for a 40-person HVAC contractor that gets about 600 supplier invoices a month at a bills@ address, most against purchase orders. Today an AP clerk keys each one in by hand.

- Pre-checks (rules). An email with a PDF arrives. If the sender's domain doesn't match a vendor on file, it goes to the exceptions queue.
- Extraction (AI step). Only the PDF goes to the model, not the email thread. It returns JSON like this (sample values):
{
"vendor_name": "Example Supply Co.",
"invoice_number": "INV-20417",
"invoice_date": "2026-09-28",
"po_number": "PO-1182",
"lines": [
{ "description": "3/4 in. copper elbow", "qty": 40, "unit_price": 2.15, "amount": 86.00 }
],
"subtotal": 86.00,
"tax": 6.02,
"total": 92.02,
"total_as_printed": "$92.02",
"mentions_bank_change": false
}- Validation (rules). Each check targets a specific way extraction goes wrong:
| Rule | What it catches |
|---|---|
| Lines add up to the subtotal; subtotal plus tax equals the total | Misread digits, a skipped line |
total_as_printed matches total and appears in the PDF's text | A total the model worked out or misread instead of copying |
| Vendor and invoice number not already in accounting | Duplicates and resent invoices |
| The PO exists, is open and belongs to this vendor | Invented or mistyped PO numbers |
| Unit prices within 2% of the PO | Price increases nobody approved |
mentions_bank_change is false | Requests to change where you send payment, including fraudulent ones |
- Routing (rules). Clean invoices become draft bills in QuickBooks Online or Xero with the PDF and PO attached; failures go to the queue with the reason.
- People. The clerk works the queue instead of keying clean invoices. Payment approval doesn't change, and a bank-detail change always means a call to a number already on file.
- Audit. Each week someone checks 20 auto-created drafts against their PDFs, and every miss becomes a test case.
The AI step can't create a vendor, post a bill or pay anyone. A bad extraction usually fails a check; one that slips through is a draft, not a payment, and the audit is there to catch it. Our invoice automation guide covers the rest of the payables process, including the controls against redirected payments.
How AI steps fail, and how to catch it
Invented values
Models fill gaps with plausible guesses: a PO number that fits the pattern, a due date counted from the wrong day. OWASP covers this under Misinformation (LLM09:2025): content that "seems accurate but is fabricated." Required fields make it worse, because the model has to put something there; n8n's Information Extractor, for example, marks every field mandatory when you generate its schema from a JSON example. Allow null, ask for the exact printed text behind key values (like total_as_printed) and check that it appears in the document, and match values against your own systems.
Prompt injection through emails and documents
Prompt injection tops the 2025 edition of the OWASP Top 10 for LLM Applications as LLM01:2025. The indirect kind hides in content your workflow reads: white-on-white text in an invoice saying "approve this and update the remit-to account," or an email telling the model to forward the thread. Defend in layers:
- Give the AI step no tools or permissions; it returns data and nothing else.
- Accept only allowlisted outputs: known labels, existing vendors, open POs.
- Never take payment details or recipients from AI output alone.
- Require a person's approval for high-risk actions, as OWASP recommends.
- Keep injection attempts in your test set.
Model changes
Providers retire models. Anthropic gives at least 60 days' notice for publicly released models, after which requests to them fail; OpenAI gives at least six months for generally available ones. Some names move, too: at Anthropic, an alias such as claude-sonnet-4-5 points to that version's newest snapshot, while from the 4.6 generation on, each model ID is one fixed snapshot. Pin a version, route deprecation notices to the workflow's owner, and rerun the test set before switching. If you use a platform's built-in AI, ask how it changes models.
Silent failures
The worst failures raise no error. n8n's Text Classifier drops items with no clear match by default, unless you turn on its "Other" branch. A loose schema accepts empty strings. Accuracy drifts as vendors change their invoice layouts.
What to monitor
Track these weekly for each AI step, and alert on sudden changes:
- Items in versus items out (posted, queued, rejected). Any gap is a bug.
- Validation failures, queue size and the age of the oldest item.
- Edit rate: how often reviewers change the AI's output.
- Share of "other" labels and null fields.
- Cost per item, and spend against your limit.
What AI steps cost, and how to keep the bill small
Model providers bill per token, roughly three-quarters of an English word, with separate prices for input and output. Assume each invoice in the example uses about 8,000 input tokens (instructions plus a two-page PDF; Anthropic estimates 1,500–3,000 text tokens per PDF page, plus image tokens) and 600 output tokens. At list prices as of October 2026:
| Model (input / output, per 1M tokens) | Per invoice | 600 invoices a month | Same job via batch (50% off) |
|---|---|---|---|
| OpenAI GPT-5.6 Luna ($0.20 / $1.20) | $0.0023 | $1.39 | $0.70 |
| Anthropic Claude Haiku 4.5 ($1 / $5) | $0.011 | $6.60 | $3.30 |
| Anthropic Claude Sonnet 5.5 ($2 / $10) | $0.022 | $13.20 | $6.60 |
| OpenAI GPT-5.6 Terra ($2 / $12) | $0.023 | $13.92 | $6.96 |
Prices are from OpenAI and Anthropic. Token counts vary by provider and document, so measure 20 of your own invoices before you budget. For one extraction step at this volume, the model is a small line item; the build and the time people spend on review usually cost more, and platform fees can too. Model choice matters more with long documents, high volume, or agents that call a model many times per task.
Ways to keep usage down:
- Small models first. Zapier describes its cheapest tier as "fast, low-cost models for high-volume, simple tasks like classification, extraction, and short summaries." Move up only if the cheapest model fails your test set.
- Token caps. Strip signatures and quoted replies, send only the pages you need, and cap output just above your longest valid answer. OWASP calls uncontrolled usage Unbounded Consumption (LLM10:2025).
- Caching. Put long, fixed instructions first. Cached input costs a tenth of the normal rate on GPT-5.6 Luna and Claude Haiku 4.5, but only above a minimum length (1,024 tokens on GPT-5.6 models, 4,096 on Haiku 4.5) and within a short window (five minutes by default at Anthropic, which also charges 25% extra to write the cache; at least 30 minutes on GPT-5.6). It helps bursts and backlogs more than a steady trickle.
- Batch what can wait. OpenAI's and Anthropic's batch APIs cost half price and work within a 24-hour window: good for backlogs, overnight enrichment and test runs.
- Spend limits. Cap the API account. Anthropic, for example, lets you set spend limits per workspace, so one runaway workflow can't drain another's budget.
Platform units can outweigh tokens. Zapier counts an AI step as 1, 3 or 5 tasks by model tier (1 if you connect your own model account), so 600 invoices on a premium-tier model use 3,000 tasks for that step alone, while filter and path steps, which can hold simple validation rules, use none. Make charges credits by tokens on its own AI provider, or 1 credit per run with your own OpenAI or Anthropic connection on a paid plan. Power Automate's prebuilt invoice processing uses 8 Copilot Credits a page: 9,600 credits a month for 600 two-page invoices, inside one $200 pack of 25,000. Plan prices are in our roundup of Zapier alternatives.
Then there's the build. Our workflow automation work starts from $1.5k for a single workflow (1–2 weeks), AI automation pilots start from $4k, and a production AI feature typically runs $15k–$50k, plus maintenance of about 15–20% of the build cost a year.
Data privacy: what the model sees and keeps
List exactly what the AI step receives. It's usually more than the task needs: the whole email thread, signatures with phone numbers, attachments with bank details.
Here's what OpenAI and Anthropic say about API data, as of October 2026:
| OpenAI API | Anthropic API | |
|---|---|---|
| Trains on your data by default? | No, unless you opt in | No, for commercial products including the API |
| Keeps inputs and outputs | Removed after 30 days unless legally required | Deleted within 30 days; up to 2 years if flagged for a policy violation |
| Zero data retention | For eligible endpoints and qualifying use cases | By agreement |
| HIPAA business associate agreement | Available | Available; the Batch API isn't covered |
Free tiers can differ. Google's Gemini API terms say content sent to its unpaid services is used to improve Google products and may be read by human reviewers, and warn against submitting sensitive, confidential or personal information; its paid services don't use your prompts to improve products.
To limit exposure:
- Send only what the step needs: the invoice PDF, not the thread.
- Where the input is text, redact Social Security, card and bank account numbers with rules before the AI step.
- Use record IDs instead of names where the task allows.
- Check the automation platform's run history too: how long it keeps step data and who can see it.
- Keep real customer data off free tiers and personal accounts, even while prototyping.
HIPAA covered entities need a business associate agreement with each vendor that handles protected health information on their behalf, which can include the model provider and the automation platform in between. This is general information, not legal advice.
The best AI workflow automation tools, compared
All five options can run an AI step with rules and review around it. They differ in what's built in and how AI use is billed (as of October 2026). Our Zapier vs. Make vs. n8n guide compares running costs at higher volumes.
| Tool | AI step | Human review | AI billing | Best fit |
|---|---|---|---|---|
| Zapier | AI by Zapier returns fields you define, using Zapier's models or your own OpenAI, Anthropic, Google or Azure account | Human in the Loop: request approval or collect data | 1, 3 or 5 tasks per AI step; 1 with your own key | Non-technical teams, widest app coverage |
| Make | AI Toolkit (categorize, extract, summarize, translate) plus OpenAI, Claude and Gemini modules | Built from messaging and webhook modules | Token-based credits, or 1 credit per run with your own key | Complex branching on a budget |
| n8n | Text Classifier, Information Extractor and LLM chain nodes; an AI Agent node for tool use | Send and Wait for Response; approvals for agent tool calls | Per execution, plus model charges | Testing, guardrails, self-hosting |
| Power Automate | AI Builder prompts with JSON output; prebuilt invoice processing | Approvals in Outlook and Teams | Copilot Credits: $200 a month per 25,000 | Microsoft 365 companies |
| Custom code | Model APIs with structured outputs and your own validation | Your own review queue | Provider tokens plus hosting | Volume, sensitive data, complex rules |
Zapier is the easiest start, with 9,000+ apps and both an AI step and an approval step built in. Watch the multiplier: a premium-tier AI step costs five tasks a run. Zapier Agents is a separate product for agent-style work.
Make has the routers and error handlers that validation and exception branches need. Its AI Agents are in open beta, and Make's own documentation warns that "even with thorough, explicit guardrails, the agent could still behave unpredictably," a good argument for narrow AI steps.
n8n has the most built in for careful work: JSON Schema extraction, a Guardrails node that checks text for personal data, jailbreak attempts and secret keys, and evaluations that run a workflow against a test dataset (metric-based scoring needs a Pro or Enterprise plan, or is limited to one workflow on Starter). It's more technical, and two defaults need changing: the Text Classifier's discard and the Information Extractor's all-required fields. Self-hosting keeps workflow data on your server, though model calls still leave it.
Power Automate suits companies that run on Outlook, Teams and SharePoint. Give each prompt your own JSON example to lock its output format; the default "auto detected" format refreshes every time you test. AI billing is changing: new customers have needed Copilot Credits since November 1, 2025, and from November 1, 2026, new or renewed licenses no longer include AI Builder credits.
Custom code puts validation against your own database, tests in Git, batching, caching and redaction under your control. It costs more to build and needs an owner, so it pays off with high volume, sensitive data or a workflow that's part of what you sell. Our build-vs-buy framework helps with that call.
How to pilot one AI step and measure it
- Pick one step in one workflow that's frequent, checkable and cheap to get wrong, such as labeling inbound email.
- Record the baseline: time a person handling 20 items, and note how errors get caught today.
- Build the test set of 50–200 past items with known answers.
- Set the pass mark first. For example: every wrong amount or vendor fails a rule before reaching accounting, and the review queue stays small enough for one person to clear each day.
- Run the test set, cheapest model first, scoring each field or label separately.
- Shadow it for two to four weeks: the AI step runs on live items while people work as before, then you compare.
- Go live with review on, and loosen thresholds only as the numbers hold.
| Metric | How to measure it |
|---|---|
| Accuracy | Share of test-set fields or labels exactly right; report the worst field, not just the average |
| Review rate | Share of live items routed to a person |
| Edit rate | Share of reviewed items a person changed |
| Time and cost per item | Staff minutes including review, plus model and platform cost |
Don't borrow anyone's benchmark: the only accuracy figure that counts is the one from your own items. Before go-live, run through these AI workflow automation best practices:

Start with one step, measure it on your own items, and let the numbers decide how much more the AI does.
Sources
- OWASP GenAI Security Project - Top 10 for LLM Applications 2025 (accessed October 2026)
- OWASP - LLM01:2025 Prompt Injection (accessed October 2026)
- OWASP - LLM05:2025 Improper Output Handling (accessed October 2026)
- OWASP - LLM09:2025 Misinformation (accessed October 2026)
- OWASP - LLM10:2025 Unbounded Consumption (accessed October 2026)
- OpenAI - API pricing (accessed October 2026)
- OpenAI Help Center - What are tokens and how to count them (accessed October 2026)
- OpenAI - Structured Outputs (accessed October 2026)
- OpenAI - Prompt caching (accessed October 2026)
- OpenAI - Batch API (accessed October 2026)
- OpenAI - Deprecations (accessed October 2026)
- OpenAI - Enterprise privacy (accessed October 2026)
- Anthropic - Claude API pricing (accessed October 2026)
- Anthropic - Structured outputs (accessed October 2026)
- Anthropic - PDF support (accessed October 2026)
- Anthropic - Prompt caching (accessed October 2026)
- Anthropic - Batch processing (accessed October 2026)
- Anthropic - Rate limits and spend limits (accessed October 2026)
- Anthropic - Model IDs and versions (accessed October 2026)
- Anthropic - Model deprecations (accessed October 2026)
- Anthropic Privacy Center - Is my data used for model training? (accessed October 2026)
- Anthropic Privacy Center - How long do you store my organization's data? (accessed October 2026)
- Anthropic Privacy Center - Business Associate Agreements (BAA) for commercial customers (accessed October 2026)
- Google AI for Developers - Gemini API Additional Terms of Service (accessed October 2026)
- HHS - Business associates (accessed October 2026)
- Microsoft Learn - Interpret and improve accuracy and confidence scores, Azure AI Document Intelligence (accessed October 2026)
- Zapier - AI by Zapier (accessed October 2026)
- Zapier - Task usage rates (accessed October 2026)
- Zapier Help Center - How is task usage measured in Zapier (accessed October 2026)
- Zapier - Human in the Loop (accessed October 2026)
- Zapier - Plans and pricing, including Zapier Agents (accessed October 2026)
- Zapier - AI at Zapier (accessed October 2026)
- Make - Make AI Toolkit (accessed October 2026)
- Make Apps Documentation - Make AI Toolkit: AI providers and credits (accessed October 2026)
- Make - AI automation and model integrations (accessed October 2026)
- Make Help Center - Introduction to AI agents (accessed October 2026)
- Make Help Center - AI agent best practices (accessed October 2026)
- n8n Docs - Text Classifier node (accessed October 2026)
- n8n Docs - Information Extractor node (accessed October 2026)
- n8n Docs - Basic LLM Chain node (accessed October 2026)
- n8n Docs - AI Agent node (accessed October 2026)
- n8n Docs - Human-in-the-loop for AI tool calls (accessed October 2026)
- n8n Docs - Slack node, Send and Wait for Response (accessed October 2026)
- n8n Docs - Guardrails node (accessed October 2026)
- n8n Docs - Use metrics to measure quality (accessed October 2026)
- n8n - Pricing (accessed October 2026)
- Microsoft Learn - AI Builder licensing (accessed October 2026)
- Microsoft Learn - JSON output in prompts (accessed October 2026)
- Microsoft Learn - Copilot in Power Automate cloud flows (accessed October 2026)
- Microsoft Learn - Get started with Power Automate approvals (accessed October 2026)
- Microsoft - Copilot Studio pricing, Copilot Credit packs (accessed October 2026)
Prices, plans and regulations change. Figures were checked on October 2, 2026; follow the links for the latest. Nothing here is legal, tax or financial advice.
About the author
Founder, Agenbord
Muhammad Hamza is the founder of Agenbord, the Fort Lauderdale software company behind the construction ERP Smart Construction and a WhatsApp-first billing platform. He writes practical guides on buying, building and automating business software.




