Build Free Real-Time Educational Avatars

Job ID: 40207684

Budget: $15 – $25 USD

I am building a completely free-to-use educational platform where realistic human and animated avatars hold fluid, real-time conversations around video-based lessons. Because the service itself must stay cost-free for users, every element—from data collection to deployment—has to avoid paid or closed platforms.

Here is the core challenge I need solved:

• Source a large, fully royalty-free image set (no watermarks, no usage restrictions) and train a model that can generate and animate photorealistic human avatars.
• Enable two or more avatars to speak with each other live, with no perceptible latency, while they reference or react to video content the learner is watching.
• Achieve all of this on infrastructure we can self-host or run inexpensively on open-source stacks—no reliance on commercial avatar APIs, paid speech engines, or proprietary chat platforms.

The finished solution should let me upload or stream a lesson video and immediately drop in multiple avatars that discuss, explain, or quiz the viewer, all synchronized in real time. Think WebRTC for low-latency audio/video, a conversational engine (e.g., Rasa or LLM-based) for dialogue, and a lightweight front-end that embeds cleanly into any web page.

Acceptance criteria
1. Avatars look lifelike, articulate words accurately, and conform to the legal image dataset you provide.
2. Two-way (or multi-way) conversational flow stays under 300 ms round-trip on a standard broadband connection.
3. No component requires a paid license; every dependency is open-source or self-built.
4. I receive full source code, model weights, and documentation so the system can be redeployed on a fresh server without external fees.

If your team can architect, train, and deliver this stack end-to-end, let’s discuss milestones and timelines.