FreeWhat moves your North Star?
AI Agents

The Agent Doesn't Own the Outcome

An AI agent can count its own work. The value it creates lands in systems it does not own. We looked at about 240 agent companies to see where their proof of AI agent ROI stops.

SS

Swarnim Shrey

Founder, MindPalace

September 24, 202616 min read

An AI receptionist answers the phone for a plumbing company. Three months in, its dashboard looks good. Thousands of calls answered, none missed after hours, a booking rate well above what the front desk used to manage.

Then the owner asks the question the dashboard was not built for.

"Am I making more money because of this?"

That is the AI agent ROI question, and the answer lives in the owner's systems, not the agent's.

Answering it means following each call past the point where the agent stops seeing it. The call becomes an appointment. Some appointments get cancelled. Some become a forty-dollar drain clearing, others a water heater replacement. Some never happen because there was no technician free that week. The revenue shows up days later, as an invoice, in a different system.

The agent saw the call. The business lives in what came after.

Four levels of proof

There is a ladder here. Naming the rungs helps, because most arguments about agent value are two people standing on different ones.

LevelWhat it showsFor the plumbing company
1. ActivityWhat the agent didCalls answered
2. Workflow outcomeWhat the agent achieved, counted from its own logsAppointments booked
3. Business outcomeWhat happened in the business, from the business's own systemsCompleted jobs, invoices, revenue
4. Incremental impactAn estimate of what the agent caused, from a comparison against a baseline or holdoutEstimated revenue that would not exist without it

Level 2 feels like proof. It is the agent grading its own homework. Level 3 is what the owner is actually asking about. Level 4 is what the owner's accountant asks next, and the gap between 3 and 4 matters: an agent can be associated with $100,000 of revenue without having caused $100,000 of new revenue.

Where the dashboards stop

I wanted to know where real agent companies land on this ladder. So we read the public job boards and product pages of about 240 companies building AI agents: support and sales agents, voice receptionists, procurement agents, and vertical agents in insurance, healthcare, manufacturing, freight, legal, real estate and retail. For each one we asked two things. What does the company publicly show customers about value? And is it hiring people to prove it?

How we counted

Sample. Known AI agent companies in eight segments (support, sales and marketing, voice and booking, procurement and back office, financial services, healthcare, industrial, and a group covering legal, real estate, retail, hospitality, education, government and recruiting), plus companies we found along the way.

Method. Sources checked as of 24 September 2026, with the help of AI research agents. Each company's level comes from what it documents publicly: product pages, docs, help centers and changelogs. We did not test the products.

What counts. Level 3 means a public page shows business outcomes drawn from the customer's own system, not estimated from the agent's logs. Level 4 means it describes a comparison group, holdout or baseline. About a third of the level 3 count rests on weaker evidence (a case study, a marketing claim or a third-party summary), and we marked those.

Hiring. From public job boards read the same day. About 90 companies had no board we could reach, so the hiring count is a floor.

What this is not. A product can do more than its public pages say, so these are preliminary observations, not a census. We are not publishing the company-by-company list, but if you want to know how a particular company was classified and why, ask us. Every quote here was checked against its live source on 24 September 2026.

The pattern was more consistent than I expected.

Of the roughly 150 companies whose public pages show any customer reporting at all, about three in four stop at level 1 or level 2. Resolutions, bookings, hours saved, "estimated revenue." All of it computed from the agent's own logs. Where the agent is priced per resolution, that self-counted number is also the one on the invoice.

We found about 37 that document level 3, and nearly all of them own the system where the outcome is recorded, or read it directly through an integration. The CRM, the store, the maintenance system, the claims platform. ServiceTitan's own voice agent can show "revenue from booked jobs" because the jobs live in ServiceTitan. An independent receptionist on top of it has to build that link itself. A recruiting screener put the boundary in one line on its website: "Your ATS stays the source of truth." (HeyMilo)

We found three that describe a comparison group or holdout in the product, and all three decide who the agent acts on and who it does not. One decides which customers get a message. One tests versions of its own conversations. One controls ad spend and holds some back. The closest thing to comparing having the agent against not having it, in the customer's own outcomes, was Microsoft's Copilot impact report, and it says itself that the differences have not been tested for statistical significance.

Ownership is the easy case. What really matters is access: reliable, authorized access to the systems where the outcome is recorded, and a way to connect those records to the agent's own work. The independent receptionist on top of ServiceTitan could show the same revenue as ServiceTitan's own agent, if it had access to the jobs and a way to match them to its calls. For most agents, that access stops at their own logs.

What the agentrecordedWhat the businessrecordedWhere valuecan be proven
The agent's logs and the business's systems each hold half the story. Value can only be proven where the two are connected.

The work moved to people

The other half of the search was hiring. About 30 of the roughly 150 companies whose job boards we could read are hiring right now to prove value to their customers. Roughly a third are building it into the product. The rest are hiring people to do it by hand.

The job descriptions are specific, and they describe the gap better than I can.

A support agent company wants someone to "work directly with customers to define the metrics that will be used to measure performance, then operationalize them so that both sides use and trust them." (Decagon)

A logistics agent company wants someone to own "what the deployment is worth, whether it delivered, and what happens next commercially." (Pallet)

A healthcare intake company wants someone who will "track outcomes against what was sold, run the underlying ROI analysis." (Tennr)

A legal agent company wants to "prove the result in numbers the firm already cares about: cases per attorney, cycle time, revenue per matter." (Eve)

A prior authorization company is hiring actuaries to measure savings with "pre and post, difference-in-differences, quasi-experimental" methods, "including sensitivity analyses that isolate true Cohere impact and validate savings." (Cohere Health) That is level 4, done by a statistician, one customer at a time.

And an AI receptionist company for home-service contractors is hiring its first data engineer, whose work "powers customer-facing dashboards, cross-customer benchmarks, and eventually the business intelligence we ship inside the product." (We have left this one unnamed.)

None of this is a mistake. These are good companies making a rational choice. But look at what the choice implies. A company that built an agent to automate one workflow is now building a data team to understand each customer's business.

The cost of proving AI agent ROI

Now imagine doing it for a hundred customers.

Each one runs a different mix of systems. Each one calculates the same metric a little differently. One marks a job "Completed," another "Closed," a third only counts it once the invoice is paid. Records are incomplete, and the link between a call and a job is not always obvious. Someone on your team has to work out what this customer means by a completed job, where that lives, and how to connect it to what your agent did. Then someone has to check the result. And when the customer changes how they work, some of that has to be done again.

How much of it carries over to the next customer?

If every new customer needs another round of data engineering, metric definitions, analysis and validation, the cost of proving AI agent ROI grows with the customer base. It shows up in onboarding time, in margins, and in how many customers one solutions engineer can carry.

So the question for an agent company is not only whether the agent can automate the workflow. It is whether you can deliver, and keep delivering, an understanding of what the agent changed, without rebuilding that understanding for every customer.

The customer's side of the same gap

Now stand where the plumbing company owner stands.

The receptionist's dashboard shows bookings. The field-service system shows completed jobs and technician hours. The accounting system shows revenue and margin. The marketing platform shows what each lead cost. To know whether the agent is working, somebody has to connect all four, agree on what the words mean, and work out why the numbers do not match. That somebody is usually the owner, or an operations manager who already has a full-time job.

And the question does not stop at "is it working." Say bookings are up and profit is down. Is the agent booking the wrong jobs? Are technicians scheduled badly? Is it just January? You need to know which before you change the agent, and after you change it, whether the change helped.

If an agent saves a business time on the work but costs it time understanding the work, how much value was delivered? The agent may still be worth it. But the honest number has to include the effort of running it.

Three agents, three answers

It gets harder when a business runs more than one agent.

A marketing agent generates leads and counts the leads. A receptionist books them and counts the bookings. A follow-up agent chases the quotes and counts the closed jobs. Each one reports its contribution from its own logs, and each one is correct by those logs.

Add the claims together and they can exceed the total increase in revenue. Say revenue rose $40,000 this quarter. The marketing agent claims $40,000 from its leads, the receptionist $40,000 from its bookings, the follow-up agent $40,000 from its closed jobs. That is $120,000 of credit for $40,000 of growth. Nobody is lying. There is simply no shared record of what happened to a customer across all three systems, so every agent takes credit for the same job.

Marketing:the leadReceptionist:the bookingFollow-up:the closed jobOne job,claimed three times
Each agent counts the same customer from its own logs. Without a shared record, one job is counted three times.

The business did not stop needing to understand itself when it automated the work. It now has three partial views of itself, with three sets of definitions, and possibly three different answers to the same question. The work did not disappear. Some of it moved.

What would have to exist

I think this points at a missing layer. Not another dashboard inside each agent, but one place that holds the business's own understanding: what its outcomes are, where they are recorded, what each term means for this particular company, and how the work flows from one step to the next.

Agents would read from it before they act. What counts as a completed job here? What is this owner actually trying to maximize, bookings or profit? The organization would read from it to see what the agents changed. Both need the same thing underneath: a verified map of how the business works, joined across systems that no single agent owns.

This is the problem we have been working on at MindPalace, from the other direction. We started by helping companies map their own metrics from their own data, and by refusing to show a number we could not trace back to a definition someone approved. It turns out that is much of what agent companies are now hiring people to do by hand, which is why I want to test whether it can be shared.

A Living Map for the workflow

Here is what that looks like for the receptionist. We call it a Living Map: a tree of how the business believes value gets created, where every node is tied to real data and every change is a hypothesis you can check.

Solid arrows are arithmetic: each metric is the product of the ones below it. Dashed lines are operational limits: open technician slots cap both bookings and completed jobs. Hypotheses about agent changes are written on these edges and checked later.

Two things make this different from an ordinary KPI tree.

First, every node is tied to the system its data lives in. The top of the funnel lives in the agent's logs. Completion and price live in the field-service system. Add profit and lead cost, and accounting and the marketing platform join in. The map is the join.

Second, every node carries who can move it. The agent controls how it handles the call: how it qualifies the caller and which appointments it offers. Whether that becomes a booking also depends on the customer's decision, the open slots on the schedule, and the shop's own rules. Whether a booking then becomes a completed, well-priced job is almost entirely the business. So the map separates what the agent influences from what the rest of the operation decides.

When bookings rise and revenue does not, it tells you which side to look at before anyone touches the agent's settings. If completion rate fell because technicians were fully booked, retuning the agent's script will not help. Improving an agent means understanding the business around it, not just its instructions.

None of this proves cause on its own. The map keeps four kinds of link apart. How a metric is computed: completed jobs are bookings times completion rate. How the operation works: technicians limit how many jobs can be completed. What the data has shown together: bookings and revenue rose in the same month. And what a change is expected to do: "booking after-hours emergency calls directly will raise completed jobs, not just appointments." That last one is a claim written on an edge before the result arrives. Only a check against a comparison group gives level 4 evidence, and even then it is an estimate, not proof.

The analysis on top of the map is not exotic. A funnel from call to paid invoice, across systems. A drilldown when a number moves: which service type, which hour, which location. A comparison of customers who first came through the agent against those who came through staff. A before-and-after for every change to the agent, with a comparison group wherever one exists. And a read of the conversations themselves, to see why the calls that did not book went wrong.

What MindPalace would take off the plate

The question I am exploring is how much of this work can be made reusable. An agent company already understands the workflow its agent runs. It knows what the agent does, what data it produces, and which outcomes it is meant to move. The shape of the map is the same for every plumbing company it serves. What changes per customer is the mapping: which fields in their system mean "completed," how they price, what they call a callback.

Say the map has been built and checked for the first plumbing company. For the second, nobody should have to rediscover how calls lead to appointments, appointments to completed jobs, and jobs to revenue. What could carry over is the workflow structure, the common business concepts (a call, a booking, a completed job, an invoice), the questions worth asking, and the calculation patterns. What stays theirs is how their data maps onto the structure, what their own terms mean, who is allowed to see what, the rules they run the shop by, and the validation. A calculation checked for the first shop is a starting point for the second, not a result: it still has to be checked against the second shop's own data. The goal is for every customer after the first to need less custom work than the one before, without pretending every business runs the same way.

What if the shared part became a starting point for every customer? MindPalace could build the Living Map above for each customer, keep the plumbing-company shape shared, and version what is theirs: which field means "completed," how they price, what they call a callback. The analysis would run on top, inside the agent company's own product.

The agent company would still own its agent and its customer relationship. The customer would still authorize data access and confirm what their own terms mean. But neither would have to rebuild the analytical foundation from scratch each time the agent goes into a new business.

That is the opportunity I want to test. It is not a finished product today.

What we do not know yet

I want to be careful about the claim here, because it is easy to overstate.

Linking a phone call to an invoice across two systems is genuinely hard. Phone numbers change, addresses get typed differently, and the job might be booked under a spouse's name. Proving what the agent caused, level 4, needs a comparison, and holding some calls back from the agent costs the customer real calls. One healthcare company is paying actuaries to do it, and not by accident.

Inside a single warehouse, we have learned how much work sits behind one number we are willing to certify. On one demo warehouse, "revenue by product category" turned out to be several different numbers: the full value of every order containing a bike, or only the bike line items, before or after discounts, with or without tax and freight. Choosing the one the business meant took a written decision, not a better query. Across an agent's logs and a customer's systems, that work does not get smaller. The question is whether it can be made reusable: the workflow shape shared across customers, the definitions checked once and versioned, the analysis run the same way every time.

If it cannot, every agent company ends up with a data team, and every business running agents ends up with one too. That is a lot of new work created by software sold to remove work.

The value of an agent is not only how much work it takes away. It is also how much work it takes to know what it did.

Read this next