An AI agent runtime with a leash
tools, limits and a human approval step built in
An AI agent that can only talk is a chatbot. An agent that can actually update a CRM, check inventory or issue a refund needs a runtime that limits what it can touch and asks a human before anything irreversible. We build that runtime, not just the conversation on top of it.
What it is
An AI agent runtime is the layer that lets a language model actually do something beyond generating text: call a function that checks inventory, update a CRM record, issue a refund, send a message. It is also the layer that decides what the model is allowed to do, what requires a human to say yes first, and what gets logged so a decision can be reviewed later. The conversation people see is the smallest part of this; the runtime around it, permissions, approvals, logging, a kill switch, is what makes an agent safe to actually run in production.
When you need it (and when you do not)
You need this once an AI agent needs to take action, not just answer questions, checking real inventory, updating a real order, issuing a real refund, because that is exactly where a confident but wrong model response stops being an annoying chat answer and starts being a real-world mistake. It is also essential once multiple channels, a website chat, WhatsApp, Telegram, all need the same agent with the same limits, rather than each channel’s integration re-implementing its own, inevitably inconsistent, safety rules.
You do not need a full runtime for an agent that only answers questions from a knowledge base with no ability to change anything, that is a RAG pipeline, a simpler and cheaper build. The runtime earns its cost specifically once “the agent can do something” becomes true.
How we build it
Every tool the agent can call is defined narrowly and explicitly, an agent that can “look up an order” gets exactly that function, not broad database access it could misuse in a way nobody anticipated. We define approval rules with you before launch, not as a generic default: a refund under a set amount might go through automatically, a refund over it waits for a human, a price change always waits, the specific lines depend on your actual risk tolerance, not ours. Every tool call, its input, and its result get logged, so a reviewed decision later has the full trail, not just the final chat transcript. A kill switch stops the agent across every channel at once, built in from day one rather than added after an incident makes the need for one obvious. We built this exact pattern into a seven-channel sales agent now running hundreds of unit tests against its tool-calling logic, and into the agent team running our own marketplace’s day-to-day operations.
What to watch
The biggest risk in an agent runtime is scope creep: a tool added quickly for one use case that turns out to allow something nobody intended once the agent finds a creative way to use it. We review tool scope specifically for this before each addition, not just whether the tool works. The second risk is approval fatigue, if everything requires a human, the agent adds overhead instead of saving time, which means the approval rules need real calibration against actual risk, not a blanket “ask a human for everything” default that defeats the purpose. Cost of ownership is mostly in maintaining the tool definitions and approval rules as your business processes change, a new refund policy or a new product category means revisiting the rule set, which is a normal and expected part of running an agent, not a sign something was built wrong.
We also test the agent against adversarial inputs before launch, a user deliberately trying to get it to do something it should not, because that test tends to surface a permission gap that testing with cooperative users alone never will.
Price and timeline
| Scope | Price | Timeline |
|---|---|---|
| 2-3 tools, basic approval flow | from $2,000 | 2 to 3 weeks |
| Multiple tools, multi-channel, full escalation | from $5,000 | 4 to 6 weeks |
Related
Built as part of AI agents and custom development. Runs on top of an LLM gateway with cost control and benefits from a model evaluation and test harness before launch. See it in production in a seven-channel AI sales agent and ProBay’s AI agent team. Tell us what you want an agent to actually do: get in touch.
FAQ
How much does an AI agent runtime cost?
A runtime with two or three tools and a basic approval flow starts at $2,000. A fuller build with many tools, multiple channels and a detailed escalation workflow runs $4,000 to $10,000, close to what we quote for a full AI agent build since the runtime is most of that work.
How long does it take?
2 to 6 weeks depending on how many tools the agent needs and how much judgment the approval rules require. A narrow agent with one or two well-defined tools ships faster than one touching several systems with different risk levels.
What counts as needing human approval?
Anything costly, irreversible, or outside a normal pattern: a refund over a threshold, a price change, deleting a record, an action the agent has not seen before. We define the exact rule set with you before launch, it is a business decision, not a default we impose.
What stops the agent from doing something it should not?
Scoped tool permissions first, the agent cannot call a tool it was never given access to, no matter what it decides to try. Approval rules second, for the things it is allowed to touch but should not do without a human checking. A kill switch third, for everything else.
Who owns the runtime and the logs?
You do. It runs on your infrastructure, every tool call and its result is logged on your systems, and the agent's permissions are defined in your own configuration, not inside a vendor's black box.