Offline ai Text-to-Animation Pipeline
Budget: ₹12,500 – ₹37,500 INR
I need a self-contained AI pipeline that turns plain sentences into fully rendered 2D animated videos. The entire workflow must run offline on our own GPU servers—no external APIs or cloud calls.
Pipeline goals
• Ingest: clean text input (full sentences)
• Interpret: extract narrative, characters and scene cues with an LLM or rules-based NLP module
• Generate: create on-model storyboards and keyframes, then interpolate to smooth 24-30 fps motion
• Render: output MP4 or MOV at 1080p (minimum) with transparency-capable layers so we can later composite or edit
• Export: knock-out audio placeholders so our sound team can sync voice-overs later
Tech preferences
I’m comfortable with PyTorch, Stable Diffusion/AnimateDiff, ControlNet, and Open-Source motion-transfer libraries, but I’m open to any stack that can be reproduced offline. If you foresee a whiteboard variant down the line, make sure your approach can switch to a modern digital whiteboard look without major re-engineering.
Acceptance criteria
1. One-click script launches the entire pipeline locally (Linux).
2. Demo with at least three sample sentences producing three distinct 10-second 2D clips.
3. Clear README covering installs, model weights location, and how to swap art styles.
4. All code, models, and assets delivered under permissive licenses suitable for commercial use.
If this sounds like your field, tell me briefly which generation models or motion libraries you’d combine and how you’d keep everything strictly offline.
Pipeline goals
• Ingest: clean text input (full sentences)
• Interpret: extract narrative, characters and scene cues with an LLM or rules-based NLP module
• Generate: create on-model storyboards and keyframes, then interpolate to smooth 24-30 fps motion
• Render: output MP4 or MOV at 1080p (minimum) with transparency-capable layers so we can later composite or edit
• Export: knock-out audio placeholders so our sound team can sync voice-overs later
Tech preferences
I’m comfortable with PyTorch, Stable Diffusion/AnimateDiff, ControlNet, and Open-Source motion-transfer libraries, but I’m open to any stack that can be reproduced offline. If you foresee a whiteboard variant down the line, make sure your approach can switch to a modern digital whiteboard look without major re-engineering.
Acceptance criteria
1. One-click script launches the entire pipeline locally (Linux).
2. Demo with at least three sample sentences producing three distinct 10-second 2D clips.
3. Clear README covering installs, model weights location, and how to swap art styles.
4. All code, models, and assets delivered under permissive licenses suitable for commercial use.
If this sounds like your field, tell me briefly which generation models or motion libraries you’d combine and how you’d keep everything strictly offline.