ENGINEERING / 6 MIN READ

An agent that has
to prove it’s done.

The same model that did the work was also grading it. We took the verdict away.

The worker does not grade itself
Concept illustration of separate work and review panels with an inspection lensConcept artwork · Illustration, not product footage
010203
01Worker02Read-only checker03Return or report

We asked Solace to make a Spotify playlist of every song from the Need for Speed: Underground soundtrack. We asked seven times. Seven times it said "Done", and seven times it was wrong.

The failures were different each time:

  • one playlist had ten songs the model made up;
  • one had four songs "as a foundation";
  • one run saved somebody else's playlist instead of making one.

Every run ended with a confident, friendly summary.

Why "done" meant nothing

When we read the run records, the reason was simple. Nothing ever looked at the phone.

In four of the six runs we examined, the agent never set itself a goal. The run ended when the model stopped talking. In the other two, it wrote its own success criteria and then marked them met using its own evidence. "Added two playlists to library" closed "Add all 18 tracks to a playlist". The same model that did the work was also grading it, and it graded generously.

This isn't a Spotify problem, or even a model problem. Any agent that decides for itself when it's finished will sometimes finish early. It will also tell you it succeeded, because the model's last message is the one that sounds most sure.

Moving the grading out of the worker

So we took the verdict away from the model doing the work. When a run has changed something (driven an app, sent, saved), a separate check runs before anything is reported. It has:

  • Your words, not the agent's summary. It judges against the request as you wrote it.
  • Read-only tools. It can open apps, look at screens, scroll and read files. It cannot send, save, delete or change a setting. A checker that could "fix" things would just be a second worker with the same incentive to finish.
  • No stake in the answer. It didn't do the work, and it isn't the one who has to explain failure.

The check returns one of two things. The first is met: the answer goes out. The second is not met, with a list of what is missing, and the run goes back to work with that list.

Two things we got wrong first

"Not there" and "not done" look the same. The playlist check kept sending runs back for five songs Spotify simply doesn't carry. The run searched for them again and again across four checks. The verdict now has a third list: things the request names that don't exist in the app. Those never send the run back. Solace tells you about them instead.

Some requests have nothing to check. "Find me some stays in Amsterdam on Airbnb" leaves nothing on the phone to inspect; the answer is the result. A read-only checker couldn't even redo the search: it couldn't type the city, so it spent its minutes tapping around. Requests that only ask to find or tell something are now judged by their answer. Anything that also asks to keep, send or change something is checked.

What it costs

A check is extra model calls, and it runs on every task that changed something. It makes successful runs slower. We think that's the right trade: a run that reports "done" when it isn't is worse than a slow one, because you stop checking and it eventually costs you something real.

What we haven't measured yet

We haven't yet published a controlled before-and-after rate for the completion check. That belongs in the technical report, as an ablation run with the check on and off over the same tasks, several runs each. When we have it, it'll be here, whichever way it comes out.

Solace AI is in closed beta. Request an invite ↗