Case study: a book, an 8-minute animated explainer and 18 Shorts
How our pipeline turned one book summary into a long YouTube video, 18 vertical Shorts, thumbnails and channel art - and what we learned.
Before offering this to anyone else, we ran the full pipeline on a project of our own: BookReel, an app that retells product management books as short stories you swipe through. We wanted video for its YouTube channel, and we wanted to know whether AI agents could produce something we’d actually be proud to publish.
This is what we made, how, and what didn’t work the first time.
The brief
Take the BookReel summary of Inspired by Marty Cagan - an original fictional story about a product manager named Theo, told in 18 parts - and turn it into:
- one long animated video for YouTube,
- a Short for every part,
- thumbnails, a channel avatar and a banner.
The story itself is original; nothing is copied from the book.
What we delivered
| Output | Result |
|---|---|
| Long explainer | 8 min 28 s, 1920×1080 |
| Scenes | 73 illustrated story shots + 18 title cards + 18 lesson cards |
| Shorts | 18 videos, 21-31 seconds each, 1080×1920 |
| Thumbnails | 3 variants for A/B testing |
| Channel art | Avatar + banner, checked against mobile and desktop crops |
| YouTube copy | Titles, description with chapters, tags, a title per Short |
You can see excerpts on our work page.
How it was made
1. Script. The story was already written in short parts, each with a one-line lesson. The script agent split every part into sentences - each sentence became one shot.
2. Voice. A deep, warm narrator voice, generated with word-level timestamps. Those timestamps drive everything else: when each shot starts, when each caption word lights up, when each prop appears.
3. Direction. For each sentence, the director picks a set (office, meeting room, studio, restaurant…), the characters, their poses and moods, and the props - roadmap boards, charts, calendars, video-call grids, sticky notes. Cues are tied to words: when the narrator says “roadmap”, the roadmap appears.
4. Sound. Original background music and a library of sound effects, all composed for the project. The music drops under the voice and comes back up on title and lesson cards.
5. Render. Code-based animation rendered to video: about six minutes for the long version and about five and a half for all 18 Shorts.
6. Review. A human went through a still frame of every shot before the final render, and fixed overlaps, cut-off text and moments where the picture didn’t match the words.
What didn’t work the first time
- The first voice was too flat. We switched to a richer narrator voice and the whole video felt more confident.
- Text-only videos weren’t enough. Our first attempt animated the summary as text on screen. It was clean but forgettable. Characters and scenes made it a story.
- Captions lingered. Early versions left the last caption of a sentence on screen into the next scene. We now group captions per sentence.
- Shorts need hooks, not titles. “Part 3: Two uncomfortable truths” became “Half your ideas won’t work.”
What we learned
- Word timing is everything. Once every animation is tied to the exact word being spoken, the video feels directed rather than generated.
- Review stills, not just video. A contact sheet of every shot reveals problems faster than scrubbing through minutes of footage.
- One story, many formats. The long video took the most thought. The 18 Shorts reused its scenes and narration - no extra recording.
- Humans stay in the loop. The agents did the heavy lifting; the judgment calls - which hook, which voice, what feels off - were human.
That’s the process we now use for client videos. If you’d like to see what it could do with your product, send us a link.