Skip to content
All articles

·3 min read

AI Agents That Call Your APIs: A Practical Guide to Tool Calling

A chatbot answers questions. An agent gets work done: it asks for missing inputs, calls your backend and recovers from errors. How to build one that is safe enough for real business processes.

#ai#agents#architecture

A chatbot answers questions. An agent gets work done. The difference is tool calling: the language model does not only produce text, it decides which of your functions to call, with which arguments, and uses the results to continue. Done well, this turns a multi-step business process (create a customer, check a VAT ID, send a welcome mail) into a short conversation.

Done badly, it turns into a model that guesses parameters and writes to your database. This guide is about the first case.

How tool calling works

You describe your tools to the model: a name, a short description and a JSON schema for the arguments. The model replies either with text or with a request to call a tool. Your code executes the call, sends the result back, and the loop continues until the task is done.

{
  "name": "crm_lookup",
  "description": "Find a customer by email address",
  "parameters": {
    "type": "object",
    "properties": { "email": { "type": "string", "format": "email" } },
    "required": ["email"]
  }
}

The model never touches your systems directly. It only proposes calls; your backend decides whether to run them.

Keep the business logic in the backend

The most important design rule: the model orchestrates, the backend decides. Validation, permissions, prices, stock and every business rule stay in your existing services. Tools should be thin wrappers around APIs you already trust, not new logic written for the model.

This has three advantages. The behaviour stays predictable, you can test the tools without any AI, and you can swap the model (cloud or local) without rewriting rules.

Define workflows, not just tools

Giving a model fifty tools and a vague goal rarely works in production. What works is a workflow definition: the steps of a process, which inputs each step needs, which tool it calls and what happens on failure. The model’s job becomes much narrower:

  • understand what the user wants and map it to a workflow,
  • collect missing inputs by asking the user,
  • fill tool arguments from the conversation,
  • explain results and errors in plain language.

Workflows can live in JSON or a database table, so business teams can change steps without a deployment.

Ask instead of guessing

When a required argument is missing, the agent should ask. A good pattern is a dedicated ask_user tool with a clear question. Combined with schema validation, this removes most hallucinated parameters: if the model cannot fill a field from the conversation, it has to ask for it.

Validate everything, confirm what matters

  • Validate every tool call against its schema before executing it. Reject and return the error to the model, which usually corrects itself.
  • Separate read and write tools. Reads can run immediately; writes that change money, contracts or customer data should show a summary and wait for confirmation.
  • Make tools idempotent where possible, so a retry does not create a second order.
  • Log every step: prompt, tool call, arguments, result. You will need it for debugging and audits.

Recover from errors

APIs fail. The agent should see the error message, decide whether to retry, ask the user for a correction (for example an invalid postcode) or stop and hand over to a human. Limit the number of steps per task so a confused model cannot loop forever.

Measure before you scale

Build a small evaluation set: realistic conversations with the expected tool calls. Run it on every prompt or model change. Track how often the agent picks the right workflow, fills arguments correctly and completes the task. This is what turns a demo into something you can put in front of customers.

Cloud or local model?

Both work. Cloud models are strong at tool calling out of the box. Local models (for example via Ollama) are good enough for well-defined workflows and keep data in-house, which is often the deciding factor in Germany. Because the logic lives in the backend, you can start with one and switch later.

If you have a process that eats hours of manual clicking every week, an agent with a handful of well-defined tools is often the fastest win. Happy to sketch one for your case in a free call.

Get the AI Readiness Checklist

10 questions to answer before you start an AI project, plus occasional notes on applied AI and SaaS. Free.

Double opt-in, unsubscribe anytime with one click. Details in the privacy policy.

More articles