VoiceidoAI StudioFree pilot

Turn a book into an animated video

Upload the manuscript. Voiceido reads it end to end, locks a cast that stays recognizable, times narration against the picture, and returns captioned video ready to publish.

A book page becoming a structured source
01 · Source
A canonical character reference sheet
03 · Cast locked
A generated scene in motion
07 · Motion

What actually comes out

A finished pack is a vertical video with burned-in captions, generated scene art, shot-level motion, and narration timed to the cut. Long books are split along narrative boundaries rather than page counts, so each video covers a coherent stretch of story instead of an arbitrary ten pages.

You approve the cast before any of it renders. Characters are generated once, reviewed as reference sheets, and then reused as identity references in every downstream frame.

The path a book takes

Nine stages, each one inspectable and resumable.

  1. 01BookEPUB, PDF, DOCX, or manuscript text
  2. 02Story understandingPlot, chronology, and relationships
  3. 03Character registryOne canonical identity per character
  4. 04World registryLocations and props that persist
  5. 05Scene graphShots planned against the source text
  6. 06Visual generationFrames locked to the cast references
  7. 07MotionShot-level animation
  8. 08NarrationVoice measured first, then timed to picture
  9. 09VideoCaptioned, assembled, downloadable

Where authors put the video

Book trailers

A 60–120 second animated trailer built from your own opening chapter rather than stock footage.

Animated excerpts

One scene, fully narrated, for the moment in the book you already read aloud at events.

Character reveals

Canonical character art with narration, released one cast member at a time before launch.

Short-form social

Vertical, captioned, sized for YouTube Shorts, Reels, and TikTok without a second edit.

Narrated adaptations

Serialized episodes across a whole book for a channel that publishes on a schedule.

Backlist revival

An older title gets a visual edition without reprinting or re-recording anything.

What makes this different from prompting a video model

  • The book is read as one story, not as a series of disconnected prompts.
  • Characters, locations, and props are stored in registries and reused by identity, not re-described each time.
  • Narration is measured first, then motion and captions are fit to it, so nothing drifts out of sync.
  • Every generated asset carries a checksum, lineage record, and idempotency key, so a retry never silently pays twice.
  • Work is checkpointed in the cloud, so closing the tab does not lose a production run.

Common questions

What formats can I upload?

EPUB, PDF, DOCX, Markdown, and plain text. Scanned images without a text layer are not supported yet, because the system needs real text to build a story graph.

Do I keep the rights to the video?

You keep the rights to your own book, and the video generated from it is yours to publish. Voiceido asks you to declare the rights status of the source before any paid production starts, and blocks production when the declaration and the edition metadata disagree.

How long is one video?

Most packs land between one and three minutes because they follow a narrative unit. A full book becomes a series of packs rather than one long file.

Can the cast stay consistent across a whole series?

That is the core of the system. Approved character references are stored per book and reused in every later scene, so book two looks like book one.

Keep reading