A skill your AI agent reads — Claude Code, Codex, Cursor, Kimi Code — that turns
prompts, scripts, agent-readable sources, and real product flows into narrated, captioned video. 2D HTML/CSS, 3D Three.js/WebGL, and SVG motion all driven by the voice track. Local-first speech and rendering; optional tools only when the project needs them.
MIT licensed · local-first by default · cloud only when you register it
Claude Code·Codex·Cursor·Kimi Code·opencode·any agent that reads skills·
Don't read about it. Watch it.
This 30-second reel was made by narova, about narova.
Speech drives everything
The voice is the clock. Every word is a trigger.
Your voice, your way
Start with local Piper, XTTS, Qwen, or cloned Chatterbox voices. Add an explicitly registered provider only when a project calls for it.
One narrator, a full panel, or no voice at all. Give each speaker a color; the captions follow.
Karaoke captions
Every word lights up exactly as it's spoken, in the speaker's color. No manual timing, ever.
Cue-timed reveals
data-cue="2" keeps an element hidden until turn 2 is spoken. Visuals land on the beat.
Two renderers. Both local. Both free.
Keep the full browser canvas and Studio in HyperFrames. Switch to Narova No-Browser when a machine cannot launch a browser: Skia draws the portable scene tree, FFmpeg plays local media and encodes the result, and unsupported HTML fails clearly instead of degrading silently.
HyperFramesHTML · CSS · WebGL · 3D · Studio
Narova No-Browserbrowserless · images · SVG · video · RTL
Every scene carries a content hash. Both renderers now reuse untouched scenes from cache — changed scenes re-render, the rest are spliced via ffmpeg. "Try five versions of scene 4" costs five scene renders, not five full videos.
Prove the direction before full production.
For ambitious work, creative-brief.md turns medium-neutral intent into 2–3 small proof branches with rationale. The agent renders decisive, project-bound evidence, rejects weak directions, restores one winner with its proof-time overrides, records exact expansion lineage, and expands only that branch. Camera, depth, and lighting stay conditional on the chosen medium.
A raw canvas means the whole frame.
Zero-style scenes have no implicit centering, max-width, gutter, or caption reserve. Patterns, chrome, and safeLayout are independent opt-ins. Captions remain a default-on overlay and can be disabled with captions: false.
Sources, with an agent in the loop
Your agent reads product sites, articles, papers, docs, and repos, then selects the story and evidence. Narova handles the mechanical ingest and video workflow.
Show the product doing the work
Explore an app in an agent-readable browser, then record highlighted semantic clicks, typing, waits, and scrolling on the narration clock. Narova frames the real UI, adds voice and word-synced captions, and refuses stale takes.
83 seconds · one continuous browser take · 24 narration-timed operations ·
voiced click-ripple proof
Plan before rebuilding
narova plan reads the canonical manifest and tells you whether a change needs TTS, alignment, mixing, composition, or rendering.
Urdu-aware dialogue
No-browser ۔ and ؟ punctuation now drive sentence timing correctly. For meaningful Urdu scripts, Narova can hand dialogue polishing to the optional urdu-voice-director skill.
Bring your own narration
Already have a recording? Set narration.file and drop in your own voice. TTS is skipped. Add wordTimings for automatic karaoke overlays — every word lights up without a single caption line in your HTML.
Preview locally. Ship deliberately.
Use Studio or a no-browser draft plus shots --beats to inspect every promised visual beat, --reuse for revisions, and build --release for a preflight before synthesis plus a measured-timing recheck before rendering,
then export presets for TikTok, Reels, Shorts, LinkedIn, X, YouTube, and WhatsApp.
Optional by architecture
Local by default. Extensible on purpose.
Narova ships with four local backends and keeps them as the default. Premium cloud TTS lives in separate companion skills, with no provider SDK, endpoint, model, authentication rule, or dependency inside Narova core.
ElevenLabs
voice library
Account voices, voice IDs, and provider-native controls.
Tell your agent what you want to make. Add a URL, repo, document, script, or just an idea.
02
Prove the creative intent
For ambitious work, your agent saves 2–3 small, orthogonal proofs with rationale, compares the rendered evidence, and expands only one selected branch. For demos, it also explores the product and records declared actions after narration timing is known.
03
Preview, direct, ship
Inspect narration and marker beats, watch Studio, ask for changes in your own words, then run the fail-before-render release gate and export final deliverables.
Your first video starts with one prompt.
Install the skill once. After that, use Narova the way it was designed: talk to your agent. Pick a starting point or say it your way.
pick an example · ask your agent · render locally
# install once$ npx skills add ammar-hasan/narova --skill narova -g
# idea → videoyou →Make a 45-second vertical explainer about why starting
small makes habits easier to keep. Make it warm and practical,
and show me a preview before rendering.agent →I’ll use one warm narrator, a calm visual rhythm, and
simple step-by-step scenes. I’ll keep the language practical
and show you the vertical preview first.