Setup Whisper + pyannote.audio for Meeting Transcriptions on GCP

Job ID: 38165259

Budget: $25 – $50 USD

I'm looking for an experienced freelancer to set up and integrate OpenAI's Whisper and pyannote.audio for high-quality speech-to-text transcription and speaker diarization on Google Cloud Platform (GCP). This setup will support our platform’s meeting feature, providing accurate and efficient transcription and speaker separation, and it must integrate seamlessly with our existing platform built on Django.

Project Scope and Requirements:

1. Set Up GCP Environment:

- Provision and configure a VM instance on GCP to handle audio processing tasks.
- Ensure the VM has sufficient resources (CPU, RAM, and storage) for efficient transcription and diarization.

2. Install and Configure Whisper:

- Install Whisper from OpenAI on the GCP VM.
- Ensure Whisper is correctly set up for high-accuracy transcription of audio files.

3. Install and Configure pyannote.audio:

- Install pyannote.audio on the same VM.
- Configure pyannote.audio for accurate speaker diarization of audio files.

4. Integrate Whisper and pyannote.audio:

- Develop a pipeline to process audio files using Whisper for transcription and pyannote.audio for diarization.
- Ensure the integration handles both batch and real-time processing of meeting audio files.

5. Integrate with Django Platform:

- Develop an API or service that allows our Django platform to send audio files for processing and receive transcriptions with speaker diarization.
- Ensure seamless integration with our existing Django application, including necessary endpoints and authentication.

6. Optimize Performance:

- Optimize the processing pipeline for performance and efficiency.
- Ensure that the solution can handle large audio files (up to several hours of meeting recordings) without significant delays.

7. Documentation and Testing:

- Provide detailed documentation on the setup process, configurations, and usage.
- Conduct thorough testing to ensure the solution meets quality and performance requirements.
- Include test cases and sample audio files to validate the setup.

8. Deliverables:

1. Fully Configured GCP VM:

- A VM instance on GCP with Whisper and pyannote.audio installed and configured.
- Documentation on VM setup and configurations.

2. Integrated Processing Pipeline:

- A Python script or set of scripts to process audio files using Whisper and pyannote.audio.
- Integration details and documentation.

3. Integration with Django Platform:

- API or service for processing audio files and returning transcriptions with speaker diarization.
- Documentation on how to integrate and use the service with our existing Django application.

4. Performance Optimization:

- Optimized configurations for handling large audio files.
- Performance benchmarks and testing results.

5. Documentation and Testing Report:

- Comprehensive documentation on the setup, integration, and usage.
- Testing report with test cases and results.

Skills and Experience Required:
- Proven experience with Google Cloud Platform (GCP).
- Expertise in setting up and configuring virtual machines on GCP.
Strong Python programming skills.
- Experience with Whisper and pyannote.audio for transcription and diarization.
- Familiarity with Django and API integration.
- Familiarity with audio processing and optimization techniques.
- Ability to document processes clearly and provide thorough testing.

How to Apply:
Please provide the following in your application:

- A brief overview of your relevant experience.
- Examples of similar projects you have completed.
- Your approach to setting up and integrating Whisper and pyannote.audio on GCP.
- Your approach to integrating the solution with a Django platform.
- Estimated timeline and cost for completing the project.
Related categories: Python Django Whisper AI