AI· 16 min read

How to Implement AI in Your Business: 99% Is a Workflow

Cris Ugarte

CEO & co-founder, Puente OS

How to Implement AI in Your Business: 99% Is a Workflow

In short: Implementing AI in your business starts with separating rules from judgment. Almost all of your processes have a known right answer and run best as workflows with explicit rules. AI pays off in two places: reading what arrives messy, like an order sent over WhatsApp, and acting as an agent on the steps where the next move depends on what turns up. Workflows for rules, models for reading, agents for judgment and people for exceptions.

A workflow is an automated process that follows steps and rules defined in advance: given the same data, it always reaches the same result. An AI agent is a system in which a language model decides on its own which step to take and which tool to use to reach a goal. Implementing AI well means assigning each step of each process to the piece that fits it.

I'm Cris Ugarte, CEO and co-founder of Puente OS, the platform that builds the systems that run your business. We work with retail, eCommerce and logistics companies across Latin America, in Chile, Peru and Mexico, and since AI agents became the trend, many conversations open with the same question: where do we put an agent?

My answer is a thesis: 99% of a company's processes are deterministic and are solved with workflows. Agents pay off in the rest, where coordination has no fixed order or the input arrives in a thousand forms, and even there they deliver a structured result. Here I review what the companies that build the models recommend, what the benchmarks measure and what it looks like in real operations.

What is the difference between a workflow and an AI agent?

The difference is who decides the next step. In a workflow, the code decides, with rules your company defined in advance; in an agent, a language model decides, case by case. That is why a workflow delivers the same result every time it runs on the same data, and an agent can reach different results.

Anthropic, the company behind Claude, draws the same line: in a workflow, the model and its tools follow predefined code paths, and in an agent the model directs its own process. Its advice is to find "the simplest solution possible" and add complexity only when needed, which can mean building without agents.

Workflow and agent, along four dimensions:

  • Who decides the next step. Workflow: the code, with your rules. Agent: the model, case by case.
  • What happens with the same data. Workflow: the same result every time. Agent: the result can vary from one run to the next.
  • What each case costs. Workflow: a predictable cost. Agent: a variable cost that is usually higher.
  • How it is audited. Workflow: the rule explains each decision. Agent: you have to reconstruct its reasoning case by case.

A workflow can also carry AI inside it. When a model reads a message or classifies a document as one more step, the code still decides what comes next. Microsoft calls that level a "direct model call" and places it as the least complex option, below any agent.

Why are 99% of a company's processes deterministic?

Because almost every process in a company has a known right answer, and your business wants that answer every time: the same carrier for the same package, the same commission for the same sale, the same reconciliation for the same payment. What changes from one case to the next is how the information arrives and the minority of cases that fall outside the rule.

A process is deterministic when, given the same input data, it always produces the same output. Its rules are usually written down somewhere: the carrier contract, the commission table, the returns policy, the electronic invoicing standards. Your team applies them from memory, and your auditor expects to see them applied the same way every time.

The 99% is a thesis, and data from real processes supports its core: the rules are few and stable, and the variation concentrates in the input and in the exceptions.

  • Cases concentrate on a few paths. Wil van der Aalst, a professor at RWTH Aachen University and a leading figure in process mining, observes that 20% of a process's variants often describe 80% of what happens. In the purchasing process of a Dutch multinational paint company, analyzed by Celonis, 87% of the cases in the main flow were resolved right the first time, along 95 path variants; the 13% with rework produced 6,198 variants.
  • Exceptions are the minority, and they concentrate the work. In accounts payable, the rule-based process par excellence, 18.4% of invoices carry an exception on average, and 11.1% at the most efficient companies, according to Ardent Partners. The same study found that 48.6% of invoices still arrive through non-electronic channels.
  • Input arrives through conversational channels. In Chile, chatting on WhatsApp (76.2%) and using email (65%) are among the most common things people do online, according to the Subtel and Cadem survey. Many orders, complaints and delivery confirmations arrive the same way.

Those studies measure cases and exceptions; the exact share of deterministic processes still has no public measurement. Counted in cases, the 99% works as a design goal: your team touches one order in a hundred. The closest measured data point comes from a multichannel electronics seller in Mexico that runs on Puente OS: 97.9% of its orders get their shipping label with no human intervention.

What do OpenAI, Anthropic, Microsoft and Google recommend about agents?

They recommend the same thing: start with the simplest solution, use code or workflows for anything with clear rules, and save agents for problems without a fixed path. That direct competitors agree gives the advice weight, since each of them sells its own agents.

  • Anthropic reserves agents for open-ended problems, where you can't hardcode a fixed path, and warns that their autonomy brings higher costs and errors that compound.
  • OpenAI lists three signals that justify an agent in its practical guide to building agents: complex decisions with nuance and exceptions, rules that have become hard to maintain, and heavy reliance on unstructured data. For everything else, it concludes, a deterministic solution may be enough.
  • Microsoft says it plainly in its Agent Framework documentation: if you can write a function to handle the task, write the function. Its Architecture Center lists among common mistakes "using nondeterministic patterns for workflows that are inherently deterministic," along with the reverse mistake.
  • Google Cloud recommends, in its design pattern guide, exploring non-agentic solutions when a workload is predictable, highly structured or fits in a single model call, such as summarizing a document or classifying customer feedback.
  • AWS marks an agent as not needed for a deterministic workflow in its comparison of orchestration models.

Harrison Chase, CEO of LangChain, one of the most widely used libraries for building agents, sums it up from what he sees in production: nearly all agentic systems combine workflows and agents.

What does the data from companies already using agents show?

It shows that agents in production are still few, that part of what is sold as an agent is automation, and that value appears when the company redesigns its processes. Most of the use cases sold as agents today are solved with a workflow and a good reading step.

  • McKinsey reviewed more than 50 agent implementations it led and concluded that in low-variance, high-standardization processes, agents built on nondeterministic models can add more complexity and uncertainty than value. Its rule of thumb: rule-based automation for repetitive tasks with structured input, generative AI to extract information from unstructured input, and agents for multistep decisions with a long tail of highly variable cases.
  • Gartner warns that "many use cases positioned as agentic today don't require agentic implementations", estimates that only about 130 of the thousands of vendors claiming to sell agents actually have them, and predicts that over 40% of agentic AI projects will be canceled by the end of 2027. In April 2026, 17% of organizations had deployed agents.
  • Deloitte and Stanford see it in adoption: 11% of organizations have agentic systems in production, according to Deloitte's Tech Trends 2026, and Stanford's AI Index 2026 puts agent deployment in the single digits across nearly all business functions.

Value comes from the process side. In McKinsey's 2026 survey of 1,719 respondents, 37% attribute some impact on their company's operating profit to AI, a share similar to the year before. Among the companies capturing the most value, about 6%, nearly three in four redesigned their workflows; among the rest, one in four.

What happens when an agent runs a rule-based process?

The process becomes less reliable, more expensive and harder to audit. An agent decides each case from scratch, and a decision made anew every time can come out differently.

The benchmarks show the gap:

  • Office tasks. In TheAgentCompany, from Carnegie Mellon, agents work in a simulated company with 175 tasks. The best agent in the study completes 30%, and the top score on the public leaderboard reaches 42.9%. Administration and finance are the weakest categories.
  • Customer conversations. In CRMArena-Pro, from Salesforce AI Research, the leading agents solve about 58% of CRM tasks in a single turn and about 35% when the conversation runs several turns.
  • Consistency. In τ-bench, from Sierra, a GPT-4o agent solved fewer than half the tasks, and in retail its success dropped below 25% when it had to solve the same task eight times in a row.

Two effects explain those numbers. The first is arithmetic: if each step is right 95% of the time, a 20-step process ends well only 36% of the time (0.95 to the 20th power), which Demis Hassabis, CEO of Google DeepMind, compared to compound interest. The second is that the same model, with the same instruction, can answer differently: Thinking Machines Lab requested 1,000 responses to the same prompt at temperature zero, the setting meant to make answers repeatable, and got 80 different ones.

When the rule exists and the system improvises another, the cost can end up in court. Air Canada's chatbot invented a bereavement refund policy for a passenger, and in 2024 a British Columbia tribunal held the airline responsible for what its chatbot said, just as for any page on its website.

The workflow contains that risk: it keeps the model only in the steps that need one. If 18 of 20 steps are rules and 2 depend on the model, each right 95% of the time, the whole process is right about 90% of the time, and validating each output plus human review of doubtful cases pushes that figure higher.

Where does AI actually pay off in your operation?

In two places: in reading, where a model turns messy input into structured data, and in judgment, where an agent handles the steps whose path depends on what turns up. In both, the output is structured and a rule validates it before anything is written to your systems.

Each step of a process falls into one of four layers:

  • Workflows for rules. The dispatch cutoff, carrier assignment, exact-match reconciliation, commission calculation.
  • Models for reading. The order sent over WhatsApp, the purchase order in a PDF, the supplier's email, the photo of a shipping label.
  • Agents for judgment. The failed delivery, the ambiguous complaint, the inventory discrepancy that needs investigating.
  • People for exceptions. Anything irreversible, high-value or outside every rule, with the full context in view.

Reading: from a WhatsApp message to structured data

A customer writes on WhatsApp: "send me 20 boxes of the same as last time, to the Quilicura warehouse." A model reads the message and returns fields: customer, product, quantity and address. From there the workflow takes over: it looks up that customer's last order, checks stock and coverage, and creates the sales order in your ERP.

Reading natural language is the new capability generative AI brought, and McKinsey attributes much of the jump in automation potential to it. In Latin America this layer weighs more, because a large share of operations still arrives through WhatsApp, PDFs and email.

Structured output comes with concrete guarantees. When OpenAI launched Structured Outputs in 2024, its model followed complex data schemas 93% of the time, and to reach 100% the company added a deterministic constraint on the output. That guarantee covers the shape of the data: Air Canada's chatbot could have returned a flawless record containing an invented policy. That is why Google recommends always validating values in your application. In an operation, that means checking that the quantity is reasonable, the product exists and the address is within your coverage.

Judgment: coordinating without a fixed order

Coordinating systems in a known order is workflow work: Microsoft recommends a workflow when multiple agents or functions must coordinate. Coordination becomes agent work when the next step depends on what the system finds along the way, which Google calls adaptive routing. Consolidating orders from seven channels and assigning the carrier by rule is coordination in a known order. Resolving a failed delivery, where the customer's reply decides the next step, is adaptive coordination.

A well-designed agent works within three limits:

  • A mandate, a budget and a stopping rule. McKinsey proposes that every autonomous system operate with all three defined.
  • Structured output. The agent returns its decision in a fixed shape, and the workflow executes and logs it.
  • Human approval for anything irreversible. OpenAI gives canceling orders, authorizing large refunds and making payments as examples.

OpenAI adds one more case for agents: the step whose rules have grown so large that maintaining them costs more than delegating the judgment. Further down I explain why AI changes that math too.

What does it look like in a real operation?

It looks like processes where rules do almost all the work and AI shows up at precise points. Three examples, from less judgment to more.

Example 1: shipping labels at a multichannel seller in Mexico

A consumer electronics seller runs seven sales channels in Mexico without an ERP. Every order follows this path:

  1. The workflow reads the order as soon as it enters one of its channels.
  2. It puts the address into Mexican format and validates it against the official postal code catalog. When the address arrives written irregularly, an AI model normalizes it: in July 2026 that step worked on 41.6% of orders, and it used to be fixed by hand.
  3. It picks the carrier by zone, cost and insurance, with explicit rules.
  4. It creates the label and writes the tracking number back to the order.
  5. Whatever the rules don't resolve reaches the team, with its resolution time logged.

Of the 38,628 orders processed between November 2025 and August 2026, 97.9% got their shipping label with no human intervention. Between August 1 and 12, the incident rate was 0.59%, and generating the labels for a cutoff went from 5 to 7 hours to 15 minutes. The company chose auditable rules for its critical logic, and AI stayed in the step that needs it: reading addresses written a thousand different ways.

Example 2: a fashion retailer with an agent where the conversation leads

A Chilean multibrand fashion retailer uses each piece according to what the process calls for. In customer service, where every question arrives differently over WhatsApp, Instagram or email, an AI agent handles the first line, resolves exchanges and returns inside the chat and hands the conversation to a service rep when needed. That agent runs on a CRM with 20 tables and 11 workflows. In alterations, fixing a garment always follows the same path (a ticket in the store, routing to the area's workshop, six statuses and an email to the customer at every change), and that process runs on rules.

Example 3: the failed delivery

The carrier marks "customer not home" for the second time on an order. This is how you design it with an agent inside a workflow:

  1. The workflow detects the status and opens a case.
  2. The agent checks tracking, the customer's history and stock, and messages the customer on WhatsApp.
  3. Depending on the reply, it proposes rescheduling, changing the address or collecting the package at a pickup point.
  4. It returns its decision in a fixed shape, and the workflow executes it in the carrier's portal and in your eCommerce store.
  5. If the customer asks for a refund above the amount you set, the decision waits for a person's approval in an App.

Every action is logged with what triggered it, like every governed action.

Where is the AI if almost everything is a workflow?

In how it is built. The strongest argument for agents is that writing and maintaining hundreds of rules is expensive, and AI makes exactly that cheaper. With Puente OS you describe in natural language what your operation needs and get software running: apps, workflows, databases and agents.

At the electronics seller in Mexico, one person built 20 apps and 75 automated processes in ten months, across seven channels and 19 integrations. At the fashion retailer, the customer service CRM, with 20 tables and 11 workflows, was built in four days. At a pet retailer in Chile, a workflow calculates the margin of every sales line across four channels each day: 247,626 lines so far, 3,830 of them with negative margins that nobody saw before.

The rules stay written out explicitly, with history and rollback: changing one means editing it. The most profitable AI in your operation may be the one that wrote the workflows that run on their own.

How do you implement AI in your business, step by step?

One process at a time: pick one that hurts, separate the rules from the exceptions, assign each step to its layer and measure against the baseline. With Puente OS, each process reaches production in 2 to 4 weeks.

  1. Pick a process that hurts. With volume, manual handoffs and a visible cost when it fails.
  2. Take the baseline. Cases per month, time per case and the share resolved right the first time.
  3. Write down the rules your team applies from memory. For example: "if the package weighs more than 30 kg and ships outside the metro area, it goes with the second carrier."
  4. Put a reading step at every messy entry point. A model extracts the fields from the WhatsApp message, the PDF or the email, and a rule validates them before anything is written to your systems.
  5. Give an agent only the steps with no fixed path, with a mandate, a budget, structured output and human approval for anything irreversible.
  6. Decide where each exception goes. The system resolves it, or it reaches a named person with full context.
  7. Measure from the first week. Cases resolved end to end, cost per case and consistency: the same case has to produce the same result. Every exception that repeats is a rule waiting to be written.

Steps 3 and 6 are the core of autonomous operations design: designing each process so the repeatable work runs on its own, with your rules, and your team decides what takes judgment.

What changes as models improve?

Agents win steps, one at a time. Business rules, approvals and auditing stay in code, because those decisions have to come out the same every time, and a design that lasts lets you move a step from a workflow to an agent once the agent proves it can handle it.

The frontier moves fast. METR, a research organization that measures agent capabilities, estimates that the length of software tasks an agent completes with 50% success doubles every 4 to 7 months. At 80% success, the best model measured was at around three hours in May 2026, on tasks METR itself describes as cleaner than real work. Anthropic expects agent design to leave more and more room to the models, and keeps "do the simplest thing that works" as its best advice.

In customer service, Gartner predicts that by 2029 agentic AI will autonomously resolve 80% of common issues, and also that half of the organizations that planned to cut their service staff will abandon that plan by 2027. Agents take the standard volume, and people keep what takes judgment.

Workflows for rules, models for reading, agents for judgment and people for exceptions. Your ERP records what happened. Your BI shows you what happened. Puente makes it happen.

Key takeaways

  • Almost all of a company's processes have a known right answer, and they run best as workflows with explicit, auditable rules.
  • AI pays off in two places: reading, which turns a WhatsApp message, a PDF or an email into structured data, and judgment, where an agent handles steps with no fixed path. In both, the output is structured and a rule validates it.
  • Anthropic, OpenAI, Microsoft, Google, AWS, McKinsey and Gartner recommend starting with the simplest solution and saving agents for work with no fixed path.
  • AI also makes writing and maintaining rules cheaper: at an electronics seller in Mexico, one person built 20 apps and 75 processes in ten months.
  • It is implemented one process at a time and measured by cases resolved end to end, consistency and cost per case.

Want to know which parts of your operation are workflows and which need an agent? We start with a one-month pilot: in that month we map your processes, separate the rules from the exceptions and put your first process live in production. If you continue, the pilot fee is credited in full toward implementation. Book a demo with our team.

Frequently asked questions

What is the difference between a workflow with AI and an AI agent?

In a workflow with AI, the code decides the next step and the model performs a bounded task inside it, such as reading a message, extracting data from a PDF or classifying a complaint. In an agent, the model decides which step to take and which tool to use. A workflow with AI is repeatable and auditable, and an agent fits the steps whose path depends on what turns up.

When should a company use an AI agent?

When the next step depends on what the system finds along the way, when the input arrives incomplete or ambiguous and someone has to ask or fill in the gaps, or when maintaining a step's rules costs more than delegating the judgment. In retail, eCommerce and logistics, typical cases are a failed delivery, an ambiguous complaint or an inventory discrepancy to investigate. The agent works within limits, returns structured output and leaves irreversible actions for a person to approve.

What is a deterministic process?

A deterministic process always produces the same output from the same input data. In a company, these are the processes with known rules: assigning a carrier by weight and zone, calculating a commission, matching a payment to its invoice or accepting a return within the allowed window. They run best as workflows, because their results can be repeated, audited and explained.

How much does an AI agent cost compared to a workflow?

A workflow has a predictable cost per case, because it runs the same steps every time. An agent consumes more: according to Anthropic, an agent uses about 4 times the tokens of a chat, and a multi-agent system about 15 times. McKinsey reports that the same agentic task can vary up to 30 times in cost from one run to the next. Save agents for the steps where their judgment is worth that cost.

What is the difference between a workflow and RPA?

RPA imitates a person's clicks on a screen and depends on that screen staying the same: EY reported in 2016 that 30% to 50% of initial RPA projects failed. A workflow connects to your systems through APIs and works with the data: it reads the order, applies the rule and writes back. That is why it logs every decision and keeps working when a screen's design changes.

Where should I start implementing AI in my business?

With a process that hurts, with volume and manual handoffs: label generation, payment reconciliation or orders that arrive over WhatsApp. Measure the baseline, write down the rules, add a reading step where information arrives messy and save agents for what has no fixed path. With Puente OS we start with a one-month pilot in which your first process goes live in production, and each process after that arrives in 2 to 4 weeks.