Story understanding
The book is chunked along narrative boundaries and read as a whole: plot, chronology, relationships, tone, and which passages actually depict on-screen action rather than commentary.
Most generative video systems operate on one prompt at a time. Long-form storytelling needs something else: a persistent representation of the story that every later generation is forced to obey.



A prompt-by-prompt video tool has no memory. Ask it for the same character in scene 1 and scene 28 and you get two different people — different face, different clothes, different age. That is tolerable for a single clip and fatal for a story.
A book contains hundreds of interconnected facts — who is in a room, what they are wearing, which house this is, what happened two chapters ago, who is dead. None of that survives in a prompt box. Voiceido's approach is to extract that structure first and generate second.
Each stage writes a durable artifact that later stages read.
The book is chunked along narrative boundaries and read as a whole: plot, chronology, relationships, tone, and which passages actually depict on-screen action rather than commentary.
Every character gets a stable identifier, a canonical visual description, invariant traits, and forbidden drift. Characters who are only referenced — the dead, the historical, the offstage — are marked so they never get directed as visible cast.
Locations and props are resolved the same way, so the same house is the same house in chapter 2 and chapter 20.
Shots are planned against specific source units, which means every shot can be traced back to the passage it came from and no stretch of the book is silently skipped.
Narration is synthesized and measured before motion is cut. The voice becomes the clock for shot length, transitions, and caption timing.
Image, video, and speech models sit behind one adapter with idempotency keys and content-addressed storage, so a model can be swapped without losing the story layer above it.