AI•10 min read

AI agents in production: what changes beyond the pilot

Pilots make agents look simple. You demo a ChatGPT or Claude workflow in Slack, wire it into n8n, Zapier or Make, and it works well enough to get buy in.

Production is different. The questions stop being about the model and start being about who approved what, what got logged, what happens when a tool call fails, and how you stop costs drifting.

The short version

  • In production, the hard part is control surfaces: approvals, budgets, audit logs, and permissions, not prompts.
  • You need deterministic fallbacks for every external call, because retries can create duplicates and replays can double charge.
  • Treat cost as a budgeted resource per workflow run, not as an invoice surprise at month end.
  • MCP changes your integration boundary: you must log what data leaves your systems, and what tool was called, not just the final outcome.
  • Do not ship an agent without an error route, a dead letter queue, and a clear owner for ongoing tuning.

What actually changes when an agent goes live?

In a pilot, you can accept vague outcomes. In production, you need the same things you would demand from any integration: traceability, predictable failure handling, and bounded blast radius.

A practical definition we use is this: an agent is in production when it can trigger real world changes, without an engineer watching it, and you would notice if it silently stopped. That could be raising supplier POs, updating customer records, sending messages in Slack, or reconciling finance data.

Once you cross that line, four things change.

  1. You need an approval model. “Human in the loop” becomes “who is allowed to approve which action”. OpenAI’s own guidance on workspace agents explicitly calls out admin controls, activity logs, and the ability to require human approval before actions like sending messages or updating records in a business workspace. (OpenAI workspace agents)
  1. You need an audit trail you can trust. Not screenshots in a chat thread. Immutable events, tied to a run id, with inputs, outputs, and approval decisions.
  1. You need cost controls at workflow level. Not “try it and see”. Zapier bills primarily on tasks, Make on operations and credits, and agent steps can multiply quietly when retries or replays kick in. (Zapier pricing, Make operations)
  1. You need explicit failure routes. Retries, rate limits, and partial failures are normal, especially when you involve Slack, finance systems, and shipping systems. Slack documents method rate limits and notes new limits for some endpoints for newly created apps. (Slack API rate limits)

The popular advice that is wrong: “Just add a retry.” Retries are not free, they create duplicates unless you design for idempotency.

Do you need human approval steps, and where should they live?

If an agent can do anything irreversible, approval is not optional. The question is where the approval happens.

There are three common patterns, and each has a different failure mode.

Pattern 1: Approval in the agent tool layer

If you are using ChatGPT agents inside a managed workspace, OpenAI describes the ability to require human approval for certain actions, and to track agent activity and review logs. (OpenAI workspace agents)

This can be a good default for actions like “send this Slack message” or “update this CRM record”. The risk is that you end up with approvals scattered across individual agents, rather than a single business process.

Pattern 2: Approval in your workflow orchestrator (n8n, Make, Zapier)

For many UK teams, the workflow tool is where approvals belong because it is already the system that owns triggers, webhooks, retries and branching.

Make’s docs are explicit that error handling is implemented as routes that define the logic path when a module hits an exception. That same routing concept is useful for approvals: an agent proposes an action, the route pauses and asks for approval, then resumes or diverts. (Make error handlers)

Zapier’s strength here is speed of rollout and governance within the Zapier account. Its weakness is that complex approvals often turn into lots of Zaps, which becomes hard to reason about.

Pattern 3: Approval in a custom “control plane”

For finance and logistics, the approval decision itself often needs context: spend limits, supplier status, contract terms, last three exceptions, and whether the record has already been touched today.

That context rarely lives in Slack. It usually lives across a few systems. The cleanest pattern is often a small internal tool that shows “proposed actions” and writes the approval as an auditable event. The agent does not approve itself, it just queues the work.

If you are weighing custom build versus no code, it can help to separate “orchestration” from “approval UI”. We commonly keep n8n or Make doing orchestration, then build a small approval interface where the business risk sits. Swarm Labs does this kind of work under AI automation and custom software development.

What should you log for audit, and what will you not get “for free”?

Most teams start by logging outcomes. In production you need to log decisions.

At minimum, you want to be able to answer these questions months later:

  • What triggered the run (webhook payload, schedule, Slack command)?
  • What external systems were called, with which parameters (redacted where needed)?
  • What did the model decide to do, and why did it choose that tool?
  • Who approved it, when, and based on which summary?
  • What did the tool return, and what did we do next?

OpenAI’s documentation around agent features and compliance makes an important point: conversation level logs are not the same thing as action level traces. For example, the ChatGPT agent help page notes that conversations involving agent tasks can appear in compliance logs, but individual agent actions may not be fully represented as you might expect. (ChatGPT agent help)

Separately, OpenAI’s own engineering write up on running Codex describes exporting agent aware events (including tool approvals and MCP usage) via OpenTelemetry. That is much closer to what you want for production. (Running Codex safely at OpenAI)

If you do not have that level of telemetry, you can still build a reliable audit trail by treating every tool call as an event that you persist yourself. That means:

  • Generate a run id at the start.
  • Log inputs, tool calls, responses, and approvals keyed to that run id.
  • Log retries as separate events, not overwriting the first attempt.

If you already run n8n or Make heavily, it is worth tracking run level telemetry as a first class dataset. Swarm Labs’ own tool Time Hive is built around logging automation runs and turning them into time saved and ROI. It is useful for the cost and value side of agent operations, not just debugging.

How do you control costs when the agent can loop?

Two things drive surprise bills.

  1. Fan out. One input record becomes 20 tool calls. This is common in logistics (one order triggers label, invoice, customs docs, tracking messages, exception checks).
  1. Replays and retries. The system re runs work you already paid for.

Zapier is explicit that replay attempts can still count tasks. Its replay article states that successful steps count towards task usage even if they were already counted in a previous run. (Zapier replay)

Make uses an operations and credits model, and its docs define operations as the unit consumed by activity. (Make operations)

n8n’s cost model is mainly your hosting and any third party APIs you call, which can make it look “free”. n8n’s own documentation positions it as a fair code tool with a self host option, and n8n also provides a free community edition activation key for self hosted instances. (n8n docs, n8n Community Edition key)

The mechanism that matters for production is not the plan, it is budget enforcement.

A simple budget model that works

Create a “cost envelope” per run. You can implement this even in no code.

  • Decide a max cost per run in pounds, or in internal units like “max 200 Zapier tasks per day”.
  • Estimate per tool call cost. For example, a Slack API post is not billed by Slack, but it can be rate limited, and it does have business impact. A model call has direct cost. A Zapier step has direct cost.
  • Track spend as the workflow executes. When it hits budget, stop and send for approval.

This forces useful behaviour. If the agent starts asking follow up questions or trying different approaches, it runs out of budget and asks for a human.

What are the failure modes you only discover in production?

These are the ones we see repeatedly when teams connect ChatGPT or Claude to workflow tools.

Rate limits and throttling

Slack rate limits are documented per method tier, and some endpoints have specific constraints for certain apps. When you hit these limits, you must decide if you queue, back off, or fail the run. (Slack API rate limits)

Retrying causes duplicates

In n8n, retries are typically node level. Community guidance on partial failures points out a painful truth: retries replay execution data, so if you re run a later node you can re fire earlier side effects unless you designed for it. It also points out that you can set an error workflow for failures that outlive retries. (n8n community discussion on partial failures)

The fix is to make your tool calls idempotent:

  • Use an external idempotency key when the API supports it.
  • Write your own de duplication table keyed on run id plus action type plus target record.
  • Prefer “upsert” style updates over “create” where you can.

Replays hide real incident rates

Zapier’s automatic replay keeps things running, but it can mask a flaky integration until your task usage spikes. You want alerts based on replay frequency, not just outright failures. (Zapier replay)

Error routes do not exist unless you add them

Make is unusually clear that error handling is a separate route, and if a module in the error handling route errors, the scenario ends with an error. You have to design those routes and test them. (Make overview of error handling)

This is where production differs from a pilot. In a pilot, errors show up while you are watching. In production, they show up at 2am unless you engineered for them.

What does MCP change for production integrations?

MCP, the Model Context Protocol, matters because it is now a common way of exposing internal tools and data to an assistant without hard wiring one custom integration per vendor.

Anthropic describes MCP as an open sourced standard for connecting assistants to systems where data lives, via a specification and SDKs. (Anthropic on MCP)

OpenAI’s developer documentation on MCP connectors highlights an operational detail that matters: by default, OpenAI requests approval before data is shared with a connector or remote MCP server, and it recommends carefully reviewing and optionally logging data shared with remote MCP servers. (OpenAI MCP servers guide)

That is the production shift. When you add MCP, you are no longer just “calling an API”. You are opening a surface where the model can choose from many tools. That means:

  • Your audit logs must include which MCP server and which function was invoked.
  • Your approval flows must consider data leaving your boundary, not just actions being taken.
  • Your security model must include tool poisoning and unsafe tool design. If a tool can do too much, the agent will eventually do too much.

A good default is to publish narrowly scoped tools. “Create invoice draft” is safer than “run arbitrary SQL”. “Fetch order by id” is safer than “search all customer notes”.

A production readiness checklist you can run this week (UK)

If you are moving from pilot to production, run this as a gate.

AreaMinimum for productionWhat usually goes wrong
Approval flowsNamed approver groups and action categories, with an escalation pathApprovals are in Slack DMs, nobody can prove who approved
Audit loggingRun id, inputs, tool calls, outputs, approvals, retries, stored immutablyOnly final outcomes are logged, you cannot reconstruct decisions
Cost budgetsPer run and per day budgets, with stop and notifyRetries and replays inflate usage, finance only sees it after billing
Failure handlingError routes, dead letter queue, replay policy, manual re run procedure“Retry on fail” causes duplicates, silent partial failures
SecurityPrinciple of least privilege for tokens, secrets vault, access reviewsShared API keys, broad scopes, unknown data leaving via tools
MCP compatibilityTools are narrow, versioned, and logged, connector approvals reviewedOne MCP server exposes too much, nobody tracks tool invocations

If you want a printable version for your team, this is the lead magnet we use internally: AI Agent Production Readiness Checklist (UK). It covers approval flows, audit logging, cost budgets, fallback routes, security, and MCP compatibility.

Getting an agent from pilot to boring

Boring is the goal. An agent that only works when an engineer watches it is not in production.

At Swarm Labs, we are a UK software studio in Manchester. We build custom software and integrations with n8n, Make, Zapier or custom code. If you are at the point where your pilot worked, but you are now stuck on approvals, logging, budgets, and failure handling, our AI Agent Productionisation and Monitoring service is designed for that. It covers the build, plus monthly optimisation as real world traffic and edge cases arrive. If you want to talk through your workflows, talk to us about your integration.

Sources

  1. OpenAI: Workspace agents for business
  2. OpenAI Help Center: ChatGPT agent
  3. OpenAI Developers: MCP servers guide
  4. Anthropic: Introducing the Model Context Protocol (MCP)
  5. Slack API: Rate limits
  6. Zapier: Plans and Pricing
  7. Zapier Help: What is replay?
  8. Make Help Center: Operations
  9. Make Help Center: Error handlers
  10. Make Help Center: Overview of error handling
  11. n8n Docs: n8n documentation
  12. n8n Help Center: How to activate Community Edition key
  13. n8n Community: Partial failures when an n8n workflow calls multiple APIs

Want this wired up
for you?