Real-time Meeting Proxy App -- 2

Job ID: 40156831

Budget: ₹12,500 – ₹37,500 INR

I'm looking for a Windows application that acts as a proxy for team meetings. It will listen in on meetings from my office laptop and provide responses in text based on the conversation.

Key Features:
- Real-time transcription of meeting discussions
- Summarized transcripts post-meeting
- Action item extraction during meetings
- Automatic language detection for multilingual meetings

Ideal skills and experience:
- Proficiency in app development for Windows
- Experience with real-time transcription and NLP technologies
- Familiarity with GPT models
- Multilingual processing capabilities

App Requirements – From Venkata
1. Audio Input

The app should listen to live Microsoft Teams meetings directly from my office laptop speaker output (not from microphone).

It should clearly detect each individual speaker and separate voices (Speaker 1, Speaker 2, etc.).

2. Real-Time Response

The app should generate instant text responses within 1–2 seconds.

The reply must sound human-like and not like AI.

It should understand when to give short answers or detailed answers based on the meeting conversation.

3. Pre-Meeting Context

Before the meeting starts, the app should accept:

Meeting context

Summary or background

It should use this information to respond correctly during the meeting.

4. During the Meeting

The app should:

Listen continuously

Identify who is talking

Process the conversation

Provide simple, clear English replies in my tone (as Venkata)

5. Post-Meeting Output

After the meeting ends, the app must automatically generate:

Full meeting summary

Action items / next steps

Key decisions

6. AI Model Requirements

The app must use the latest GPT models (GPT-5 or newer when available).

Should also support switching to other LLMs if required (e.g., Gemini, Claude).

7. Performance Expectations

Delay: 1–2 seconds max for real-time chat replies

Voice separation accuracy: High, with clear labeling

Responses must be:

Short by default

Detailed only when needed or when asked

8. Tone & Style

App responses must represent me (Venkata) as a human.

Should always use simple English, no complex paragraphs unless requested.

9. Additional Notes

The app should not repeat my voice when I speak.

It should only activate after detecting my voice first, then join the meeting.

The UI should include:

Live transcript

Identified speakers

Real-time generated replies

Post-meeting summary and actions