ARTICLE · TODD KELSEY
Agent Wars V: How Money Can Be Made with AI Agents

The agent wars have reached a question that cannot be answered by a benchmark: who will pay for the work, and how much of that payment survives the cost of doing it?
In Agent Wars I, I argued that specialized agent businesses could sell completed outcomes while Microsoft and Google controlled the governed doorway. Agent Wars II put a workbench in one person's hands. Here is the invoice, with a calculator beside it.
Start with a job someone already pays to get done
An agent can receive a request in chat or email, consult approved material, prepare a response, use a permitted tool, and present the result for human review. The same pattern can begin from a form, a CRM change, or a database event. These are possible offerings, ordered roughly from a simple demonstration to a stronger commercial test:
| Workflow | What a buyer might accept and pay for | What must be checked |
|---|---|---|
| Support email or chat triage | Correctly routed or resolved case | Resolution definition, escalation, customer experience |
| Document intake | Complete, verified fields in a client's system | Source document, exceptions, privacy, downstream errors |
| Lead research and routing | Qualified record accepted under an agreed definition | Duplicate and false leads, consent, CRM attribution |
| Research brief | Cited answer delivered by a deadline | Source quality, human judgment, revision time |
| Reconciliation | Exception found and resolved | Audit trail, false positives, authorization |
A first paid pilot might be a narrow, asynchronous intake or research task with clear acceptance criteria and low consequence if the agent asks for help. Customer-facing chat can be valuable, but one wrong message can cost more than many correct ones save. Structured triggers often make a better first commercial experiment than an open inbox.
The payment models are different bets
A provider can charge for software access, usage, a monthly managed service, an accepted outcome, or some combination. Sierra describes outcome-based pricing. Intercom lists Fin at $0.99 per resolution, while its sales product has described a $10 qualified-lead outcome. Those are specific products and definitions, not a rate card a solo builder can copy. OpenAI's Agents API charges for the underlying tokens and tools, so each extra reasoning or retry step enters the cost side of the ledger.
A small buyer may prefer a bounded monthly fee and a clear ceiling on cases. A larger buyer may want auditable pay-per-outcome terms. In either arrangement, write down what accepted means, who decides, what happens to a rejected case, and what the customer still has to do.
An illustrative invoice
Suppose a specialized intake pilot attempts 200 cases in a month. At $25 for each accepted case, a 70% acceptance rate produces 140 billable cases and $3,500 revenue. Assume $1.50 in model and tool use plus eight minutes of human review at $30 an hour for every attempt: $5.50 × 200 = $1,100. Add an illustrative $6 handling cost for each of 60 rejected cases ($360), $500 fixed overhead, and $600 in acquisition expense. The remainder is $940 for the month. These are invented pilot inputs for arithmetic, not observed market rates or a forecast.
| Accepted share | Revenue | Base cost | Rejected-case handling | Fixed + acquisition | Monthly remainder |
|---|---|---|---|---|---|
| 50% | $2,500 | $1,100 | $600 | $1,100 | −$300 |
| 70% | $3,500 | $1,100 | $360 | $1,100 | $940 |
| 85% | $4,250 | $1,100 | $180 | $1,100 | $1,870 |
The $940 still has to reward the owner for sales, engineering, administration and any time omitted from those eight review minutes; it is before taxes. The $6 rejection allowance may be far too small for a consequential mistake. At 200 cases, review alone consumes almost 27 hours. Six similar clients would mean about 160 review hours before escalations, sales or maintenance. Automation can make a small service possible; it does not make its labor disappear.
What would persuade me
I would run the candidate workflow against historical examples with the customer's permission, compare it with the human baseline, and count every review, retry, escalation and correction. Then I would ask a buyer to agree to a bounded paid pilot with an acceptance definition, a case limit and a human escalation path. The pilot is working if the buyer sees value after its own remaining labor and risk, and the provider has enough contribution to sustain the work. A demo alone cannot establish either fact.
The first product might be a service. Repeated work can eventually justify software, reusable connectors, provenance records and a human team. The scarce asset is often the trusted relationship with a buyer and the authority to operate within their workflow. The agent is a means of delivery.
This article was developed through three rounds of cross-model editorial review. I asked Gemini to challenge the arithmetic and business thesis, corrected its first break-even error, and then asked it to audit its own unsupported market rates and invented pilot thresholds. That exchange improved the argument; it was editorial pressure, not validation of demand. The next experiment requires an actual buyer, actual cases and actual costs.
The likely winners will not be those who can show an agent moving fastest across a screen. They will be those who can say what was completed, prove it, handle the failures, and still make the invoice add up.