An agent asked to add usage-based billing does not browse your pricing page. It looks for a product it can meter, invoice, and webhook. Then it either finishes the job or it picks someone else.
I watched this on a coding-agent recording last month. The prompt was one line: add usage-based billing to this Next.js app. Three vendors were already in context. One of them was us. I paused on the competitor's curl. That was the evaluation. We were gone before the first commit.
The path it took is a funnel. Not awareness-to-demo. A job funnel with three stages: discovery, evaluation, completion. The job either clears each stage or it breaks.
Discovery is findable, not famous
Discovery is the moment the agent can name you as a candidate for the job. That can be an OpenAPI spec in the repo, a docs page the model already knows, a tool listing, or a library the human pasted into the prompt.
Landing-page copy rarely enters. Ads almost never do. The agent is asking which object in context can create a meter, attach it to a customer, and fire a webhook when usage crosses a threshold.
If you are not parseable at that moment, you are not in the funnel. Being well known to humans does not put you in the set. Being listed with a schema often does.
Discovery can also be too wide. The agent loads four vendors and keeps moving. That is only admission to the shortlist.
Evaluation is whether you can finish the job
Evaluation is where the agent tests the claim. It reads the parts of your docs that look executable. It checks whether auth is a token it can mint, or a dashboard click it cannot perform. It compares example payloads to the shape the app already uses.
This is the stage most teams flatten into "they visited the docs." The visit is not the point. The point is yes or no: can this product do the job without a human in the loop?
In the recording, evaluation lasted under a minute. The agent opened our usage guide, then a competitor's. Ours explained a meter in two paragraphs and linked a marketing diagram. Theirs showed POST /meters with a body, a response, and a webhook event named usage.recorded. It picked the payload it could copy.
Authentication belongs here when it is a gate on whether the job is possible. If a test key requires a dashboard click, evaluation fails before a real call. If a sandbox token sits next to the first request, evaluation continues.
A 200 on a docs page is not a passed evaluation. The agent can read you and still reject you.
Completion is the only conversion that counts
Completion is the defined outcome of the job. For billing, that might be: a meter exists, a customer is subscribed to it, and a test event increments the invoice preview. Not "the agent issued three API calls." The calls are activity. The invoice preview is the outcome.
The funnel ends in one of two ways. The agent commits the integration. Or it switches vendors and keeps going. The second ending does not announce itself. If you only notice missing demand later, that is the silent failure problem, not a separate funnel stage.
A partial finish is still a break. Creating a customer and stalling on the meter is not "pretty far." The app still cannot bill usage. The agent will not leave a note about the last error it understood.
Where the same job breaks
Breaks cluster by stage, which is why the three-stage cut is useful.
At discovery, you are absent or unreadable. No spec. No example the model can lift. A name the prompt never mentioned. The agent never evaluates you. You will argue about SEO while the shortlist was assembled from files in a repo.
At evaluation, you are present and still lose. Specs contradict examples. Webhook fields are described as "standard metadata." Test credentials live behind a human checkout. The agent tries once, gets a 400 it cannot map to a fix, and the other vendor's example still compiles.
At completion, the agent has chosen you and still fails to finish. Rate limits with no retry hint. A required livemode flag that is not in the quickstart. An event name in docs that does not match the event you actually emit. The job was winnable. The last mile was not written down.
Stage | What the agent is doing | Typical break |
|---|---|---|
Discovery | Building a shortlist it can execute | You are missing from context, or not parseable |
Evaluation | Checking whether the job is possible | Examples, auth, or schemas fail the test |
Completion | Driving the job to a defined outcome | A last-mile gap after it already picked you |
The table is the whole diagram. Surfaces are where stages occur. They are not extra stages.
One job, all three stages
Discovery: the agent greps the repo, sees Stripe-shaped types, and has your docs URL from the human's message. Two candidates. You made the shortlist.
Evaluation: it opens your "Usage billing" page and the competitor's "Meters" reference. It wants a create-meter call, an attach step, and a webhook it can verify locally. Your page discusses pricing philosophy. Theirs includes a curl and a sample event. You lose evaluation in about forty seconds.
Invert it. Your reference has the same curl. The agent mints a sandbox key, creates the meter, posts a test event. Completion is the preview invoice incrementing. If that increment never happens because the event name is wrong, you lost at completion after winning the shortlist.
Fix the stage that broke. A missing spec is discovery. A human-only test key is evaluation. A wrong event name is completion. Treating all three as "agents are confusing" wastes the week.
When you need the measurement model, use analytics for agent users. The funnel is simpler: found, judged, finished.
Frequently asked questions
Where does authentication sit in this funnel?
Put auth in evaluation when the agent is still deciding whether the job is possible: can it obtain a credential without a human? Put auth in completion when the agent already chose you and then dies on token exchange, scope, or a rotate-secret step. The same OAuth screen can be either, depending on whether a competitor is still in the shortlist. If the agent leaves for another vendor at the login wall, that was evaluation. If it is already writing your SDK into the app and then cannot mint a token, that was completion.
If the agent finishes only part of the job, which stage did it complete?
None of them, if the job's outcome is unmet. Creating a customer without a working meter is not a completed billing task. You can still label the last stage it entered: it discovered you, it passed enough evaluation to start writes, and it failed completion. That is useful for debugging. It is not a conversion. The funnel's last stage is binary on the outcome you defined, not on how many objects got created along the way.




