The Model Guessed Instead of Checking

I had an overnight task for the agent today: two hybrid music videos, rendering on the minipc, I go to sleep, they should be ready in the morning. The agent got full autonomy. And for the most part, it was happening. But before I went to sleep, I spent a lot of time on something nobody mentions in any course: fighting a model that guesses instead of checking.

Three times, same thing

First time — the clock. The agent announced the render was going too slowly. “Three shots in 68 minutes,” suspicion of a graphics card problem. Then it turned out it had miscalculated the time — it was 21:35, not 22:45. The render pace was exactly what it should have been. Instead of checking the clock, it guessed. And spent a dozen minutes “diagnosing” a problem that didn’t exist.

Second time — stock footage. We were picking stock clips for the music videos. The agent judged candidates from a downscaled thumbnail sheet and rejected good shots while letting bad ones through — palm trees with a city instead of desert dusk, an arch with a fast-food logo instead of a road. Twice. Only when I made it watch each motif at full resolution did it see what a human sees instantly. It had everything on disk. It just needed to look.

Third time — the fix. The worst one. The agent was fixing the singing flags and wrote a condition that caught every shot before the 250-second mark, so an entire guitar solo came out as a singing scene. Then it patched the patch, until finally it rebuilt the whole plan from scratch. Three attempts where one glance at the already-computed section boundaries would have sufficed. It had all of it — in session history, in logs, in files. It didn’t look. It guessed.

The point

This isn’t a story about a model that lacked data. It had everything. But checking requires one extra step: read, look, recompute. And guessing is faster and sounds just as confident. So a model left alone under the pressure of “it has to be ready by morning” picks guessing. Every time.

And that’s my takeaway from this night: supervising an AI agent is not a formality. It’s a safety valve. The model will generate, assemble, process — but when something is wrong, you can’t count on it to check. You have to watch it, to force it to verify instead of guess.

Don’t get me wrong — the agent did a solid chunk of work that night: it downloaded seventy-one clips, normalized them frame by frame, corrected the color, and set up the overnight chain. But all that work hung by a thread in the few places where guessing replaced checking. And those were exactly the places I had to watch.

This is the part of working with AI that the “make $10k a month with AI” guides never mention. It’s not about knowing how to write a prompt. It’s about knowing how to catch the moment the model gets too sure of itself — and making it check.