Build

Agentic Workflows: What They Are and When They Survive Production

Gartner expects over 40% of agentic projects to be cancelled by end-2027, naming cost, unclear value and weak risk controls. All three are design decisions.

PLANACTOBSERVEREVIEWHUMANTOOLSSCOPEDAGENTIC LOOPPLAN · ACT · OBSERVE · REVIEW
What this actually means

An agentic workflow is one where a model plans a sequence of steps and executes them against real systems, rather than returning text for a human to act on. The capability jump is real; so is the failure surface, because an agent that misreads a situation now does something rather than saying something.

What the evidence says about agentic workflows

FIG. 1 — AGENTIC INVESTMENT VS EXPECTED SURVIVAL
of organisations had made significant agentic AI investment (Jan 2025 poll, n=3,412) 19%
Source: Gartner press release, 25 June 2025
had made conservative investment 42%
Source: Gartner press release, 25 June 2025
of agentic AI projects expected to be cancelled by end-2027 40%+
Source: Gartner press release, 25 June 2025

Gartner names three causes: escalating costs, unclear business value, and inadequate risk controls. Each is addressable at design time and expensive to retrofit — which is the whole argument for scoping narrowly first.

Which agent workflows survive production

01

Narrow, high-volume first task

The first agent should handle something that runs dozens of times a week with little variation between instances. High volume makes impact visible quickly; low variance means the edge cases that break agents are rare.

The instinct is to pick the interesting problem. Interesting means variable, and variance is where agents fail publicly.

02

Explicit tool boundaries

Define exactly which systems the agent may read from and write to, and which actions require confirmation. An agent with broad write access and narrow testing is an incident waiting for a trigger.

03

Human review at consequence points

Any output that reaches a customer, a financial record or a regulated artefact passes a person first. This is not a maturity stage to graduate out of; it is the design.

04

Consumption budget with an owner

Agent actions typically bill per use, so cost scales with usage and has no natural ceiling. Model spend at three times expected volume before go-live; if that number is uncomfortable, scope more narrowly.

Escalating cost is the first cause Gartner names. It is entirely predictable and therefore entirely preventable.

05

Observability on decisions, not just uptime

Log what the agent decided and why, not merely that it ran. When someone asks why a case was routed a particular way, the answer should be a query rather than a meeting.

06

A defined stop condition

Write down what would make you switch the agent off, before it goes live. Projects without a kill criterion tend to consume budget well past the point the evidence stopped supporting them.

When not to do this

Where an agent is the wrong tool

  • The workflow has high variance or many exceptions. Agents amplify ambiguity — they do not resolve it.
  • The underlying data is unreliable. An agent reasoning over inconsistent records acts on them with more confidence than a person would.
  • No one owns consumption spend. Usage-based pricing punishes the absence of an owner more than any other cost model.
  • The goal is to demonstrate AI capability rather than to move a specific number. That is a demo, and demos do not survive production.

Green light, red light

Agentic workflows: which agents survive production
Dimension Go ahead when Do not when
Task shape High volume, low variance, dozens of runs a week that look alike The workflow has many exceptions — agents amplify ambiguity rather than resolving it
Data reliability The records the agent reads are consistent and owned Underlying data is unreliable; the agent will act on it with more confidence than a person would
Tool boundaries Read and write scopes are explicit, and consequential actions require confirmation The agent has broad write access and narrow testing — an incident waiting for a trigger
Cost ownership Consumption modelled at three times expected volume, with a named owner and a cap Nobody owns spend; usage-based pricing punishes an absent owner more than any other cost model
Stop condition Written down before go-live: what would make you switch it off No kill criterion, so the project consumes budget past the point evidence stopped supporting it
Observability Logs record what the agent decided and why Logging records only that it ran

How to pilot an agent — without hiring anyone

  1. 01
    List candidate workflows with two columns: how often it runs, and how much each instance differs from the last.
  2. 02
    Discard anything running fewer than a few dozen times a week, and anything highly variable.
  3. 03
    From what remains, pick the one where you can already state a number — time taken, error rate, cost per unit.
  4. 04
    Model consumption at three times expected volume and confirm the figure is acceptable.
  5. 05
    Name the owner and write the stop condition before a line of code.

Questions buyers ask about agentic workflows

A chatbot returns text for a human to act on. An agent plans steps and executes them against real systems. That distinction is why agent projects need explicit tool boundaries and human review at consequence points — the failure mode changes from a wrong answer to a wrong action.

Model consumption at three times expected volume before go-live, name one accountable owner, and put a monthly spend review in the calendar. Consumption pricing has no natural ceiling, so the control has to be organisational.

Rarely. Failures are public and variance is highest where humans are involved. Prove the mechanism on an internal, high-volume workflow, then move outward from a success.

Stuck pilots almost always trace to one of three things: data the agent cannot rely on, a workflow with too much variance, or no clear definition of done. Two of the three are fixable without additional licence spend.

Want an honest read on your readiness?

Six evidence tests, four minutes, no email required. It will tell you plainly if the foundations are not there yet.

Run the assessment

Read the rest of the kit.

Five capability guides, every figure cited to a named source and year.

Open the AI Resource Kit
Book a Strategy Session