A Music Video That Doesn't Illustrate the Lyrics

For a while now I’ve been building a music video pipeline with Hermes. The idea is simple: I drop in an audio file with metadata, the agent writes the script, generates frames and clips, and assembles the whole thing. It’s an overnight job — the GPU on the minipc renders, and by morning there’s a finished video.

But today we didn’t render a single frame. We worked on something more important than the render itself: figuring out what a real music video in a given genre actually looks like.

The problem: AI illustrates lyrics literally

Anyone who’s played with AI music video generation knows this: the model takes the lyrics and produces one-to-one images. The line says “driving down the highway at night”? You get a car on a highway. “Fire in my eyes”? You get eyes with fire.

Except real music videos don’t work that way. Nobody shoots literally what the lyrics say. A music video is its own film language — with its own conventions, shots, lighting, and pacing. And those conventions differ by genre.

What Hermes did

Instead of guessing, I had it research the topic properly. The instruction was specific: six calm and narrative genres — country, americana, folk/acoustic, ballads, singer-songwriter, indie folk — and for each one establish:

  • which video type dominates (artist performance or narrative with actors),
  • typical shots and framing (how close to the face, whether you see hands on the guitar, or wide landscapes),
  • lighting and palette,
  • typical locations and motifs,
  • editing pace,
  • how the artist’s face is shown,
  • plus 2-3 specific, real music videos with working links.

The agent searched for sources — articles on conventions, music video rankings, AI generation guides — and assembled them into a document. A few things surprised me.

First, country and folk aren’t just “guy with a guitar”. They’re largely a story told through images — the road, the family home, golden hour, warm sepia. Performance is there, but often intercut with narrative.

Second, no source gives the editing pace in seconds. And here’s the thing I appreciated most: Hermes didn’t write “about 4 seconds per shot” because no source said so. It wrote “not established.” Exactly as I’d instructed — don’t invent data.

What this is all for

This document isn’t for show. It’s the front door to the pipeline: before the GPU gets anything to render, the script has to speak the language of the right genre. A ballad has to look like a ballad, and indie folk like indie folk — otherwise you get the same generic “AI look” everyone recognizes at a glance.

And this is where the agent’s second role shows up, the one worth saying out loud: advisory. I work in energy, I know a fair amount about producing electricity, but I have no idea how country music video framing differs from indie folk framing. Hermes stepped into the topic like someone who knows film conventions, found real sources, and came back with knowledge I’d have spent a long time collecting on my own.

Of course — I still keep watch. I took the “don’t invent data” instruction seriously and check that the links actually work. But most of the work, from finding sources to organizing it into one note, basically did itself.

The music videos will come later. First, the grammar.