VoiceidoAI StudioFree pilot

A story representation, then video

Most generative video systems operate on one prompt at a time. Long-form storytelling needs something else: a persistent representation of the story that every later generation is forced to obey.

A story world mapped into locations and relationships
Story graph
A locked character identity reference
Identity lock
Narration timing aligned to the cut
Voice clock

The problem with clip-first systems

A prompt-by-prompt video tool has no memory. Ask it for the same character in scene 1 and scene 28 and you get two different people — different face, different clothes, different age. That is tolerable for a single clip and fatal for a story.

A book contains hundreds of interconnected facts — who is in a room, what they are wearing, which house this is, what happened two chapters ago, who is dead. None of that survives in a prompt box. Voiceido's approach is to extract that structure first and generate second.

The production graph

Each stage writes a durable artifact that later stages read.

  1. 01BookEPUB, PDF, DOCX, or manuscript text
  2. 02Story understandingPlot, chronology, and relationships
  3. 03Character registryOne canonical identity per character
  4. 04World registryLocations and props that persist
  5. 05Scene graphShots planned against the source text
  6. 06Visual generationFrames locked to the cast references
  7. 07MotionShot-level animation
  8. 08NarrationVoice measured first, then timed to picture
  9. 09VideoCaptioned, assembled, downloadable

The parts that carry the continuity

Stage 02

Story understanding

The book is chunked along narrative boundaries and read as a whole: plot, chronology, relationships, tone, and which passages actually depict on-screen action rather than commentary.

Stage 03

Character registry

Every character gets a stable identifier, a canonical visual description, invariant traits, and forbidden drift. Characters who are only referenced — the dead, the historical, the offstage — are marked so they never get directed as visible cast.

Stage 03

World registry

Locations and props are resolved the same way, so the same house is the same house in chapter 2 and chapter 20.

Stage 04

Scene graph

Shots are planned against specific source units, which means every shot can be traced back to the passage it came from and no stretch of the book is silently skipped.

Stage 05

Voice-first timing

Narration is synthesized and measured before motion is cut. The voice becomes the clock for shot length, transitions, and caption timing.

Stage 06–08

Provider-independent generation

Image, video, and speech models sit behind one adapter with idempotency keys and content-addressed storage, so a model can be swapped without losing the story layer above it.

Engineering constraints we hold ourselves to

  • Paid provider work is never issued twice for the same logical asset; ambiguous submissions stop for human review instead of re-charging.
  • Every asset is content-addressed, so two different generations can never overwrite each other.
  • Jobs are leased, checkpointed, and resumable, so a deploy or a closed laptop does not lose a run.
  • A cast completeness audit blocks production when the scene plan references someone who is not a real on-screen character.
  • Rights are inspected before paid generation, not after.

Go deeper

Keep reading