Multimodal AI Agent Development

Job ID: 39633463

Budget: ₹1,500 – ₹12,500 INR

Title:
"Build an Autonomous Multimodal AI Agent that Understands Video, Audio, and Text"

* Description:
We’re looking for an AI developer to build a multimodal AI agent capable of understanding, reasoning, and acting based on input from video, audio, and text—just like a human assistant. The agent should perceive its environment, understand instructions or events, and respond with logical, goal-oriented behavior.

* Key Requirements:
Multimodal input processing (video + audio + text)

(vision), (speech-to-text/text-to-speech), (image-text), (LLM)

Reasoning with LangChain or AutoGen

Store and retrieve memory

Capable of autonomous task execution and self-correction

* Tech Stack:
Python, PyTorch

LangChain or AutoGen

speech-to-textor text-to-speech / vision / image / llm

store/memory


* Deliverables:
Working prototype (Colab or Jupyter)

Code and documentation

Demo: Upload video → Agent gives summary, insight, or answers

Optionally: Add memory and real-time webcam interface

* Timeline:
2–3 weeks (extendable with more features)