Maintaining character identity across AI-generated stories
Continuity is the difference between a reel of nice clips and a story an audience can follow. It is also the hardest thing to get from a generative video model, because models do not remember.
The naive path
Scene 1: Jacob appears. Scene 28: Jacob appears again. A naive system sends two independent prompts — "a young man with dark hair" — and gets two different men. Nothing in the pipeline knows they are supposed to be the same person.
Adding more adjectives does not fix it. Prompt text is a lossy description of a face, and every extra clause competes with the composition you actually asked for.
What Voiceido does instead
- 01Resolve the person, not the phrase
"Jacob," "the schoolmaster," and "he" are resolved to one entity with a stable identifier before any art exists. Narrative presence is resolved at the same time, so a character who is only spoken about is never staged as if he walked into the room.
- 02Generate one canonical reference
That entity gets a character sheet: a canonical portrait plus an identity description, invariant traits that must never change, and explicit forbidden drift.
- 03Lock every downstream frame to it
Scene frames are generated with the approved reference attached, not with a re-typed description. The prompt carries the identity contract along with the shot.
- 04Audit before spending
Before paid video runs, a completeness check confirms every scene participant exists in the registry as an on-screen character. If the plan references someone who does not, production stops instead of rendering a stranger.
Why this is an architecture problem, not a prompt problem
The useful mental model is a database, not a prompt library. A character is a row with an identifier, references, and constraints. A scene is a query against that row. Generation is the last step, and it is the only step allowed to be probabilistic.
That ordering is what makes the system improve when models improve. When a better image or video model appears, the story layer, the registries, and the approvals stay exactly where they are.
Questions engineers ask
Is this just image-to-image with a reference?
Reference conditioning is one mechanism, but it only works if something upstream decided which reference belongs in this shot. The hard part is entity resolution and presence inference across a whole book, before any frame exists.
What happens when a character legitimately changes?
Visual state is part of the record. A character can age, be injured, or change costume as a tracked state change rather than as accidental drift.
How do you know it worked?
Every shot records which identity references it used. Continuity is auditable after the fact instead of being judged by eye.