MIRCHI is a 2:41 music video generated end to end — the cast, the city, the song. The hard part isn't one good-looking clip. It's twenty-one of them where the same four faces come back every time.
Four characters registered once as reusable identities. Every plate and every clip that follows references these records — never a fresh description.
Every frame below came out of a separate generation. This is the whole claim, and it is the thing generative video normally gets wrong — look at the faces.
Story plates generated before any motion — Mumbai at night, tricolour used as light rather than iconography. These are the source frames every clip animates from.
Generative video has one failure that ruins everything else: faces don't survive between shots. You generate a clip, it looks great. You generate the next one and it's a different woman. Cut them together and you don't have a film, you have a mood board.
Getting past it isn't a prompt. It's a pipeline — characters registered once and reused as identity references, wardrobe locked in text, an anchor phrase carried into every prompt featuring a given character, and a review gate before anything expensive runs. That pipeline is VideoClaw.
VideoClaw is a command-line pipeline for generated video: character and identity locking,
storyboard and plate generation, multi-provider rendering, review gates before spend, and
local assembly. It is on npm as videoclaw. The source is private for now.