Offline Medical Live Transcription & AI Categorization
Budget: $30 – $250 USD
We are looking to add a new offline streaming and AI categorization feature to our existing Android-based healthcare application.
1. Use case
A doctor interacts with a patient while the app listens in real time and performs **live speech transcription**. Once the interaction ends, the transcription is processed fully offline, without any internet connection, using Whisper. After transcription, a lightweight on-device LLM categorizes the medical text into predefined clinical categories.
2. Key requirements
- Real-time or near real-time streaming transcription
- Offline transcription using Whisper
- Fully offline LLM inference on an Android tablet
- No use of cloud APIs, GPT, Gemini, or any internet-based models
- The Android device has 6GB RAM, so the model must be quantized and optimized
- The LLM must understand medical terminology and perform accurate categorization
3. LLM expectations
We are specifically interested in medical-domain trained models, similar in capability to MedGemma, but runnable locally on-device. The app should support a dropdown selection allowing us to switch between multiple LLMs for evaluation and accuracy comparison.
Preferred candidate models include:
- MedAlpaca 7B (LLaMA-based, medical fine-tuned)
- Meditron 7B
- BiomedLM
- Biomistral 7B
4. The selected developer must:
- Identify suitable quantized versions of these models (or recommend better alternatives)
- Ensure they can run efficiently on an Android tablet with 6GB RAM
- Integrate model selection and inference into the app
- Optimize for performance, stability, and accuracy
5. Project goals
- Accurate streaming Whisper-based transcription
- Lightweight, on-device medical LLM
- Reliable categorization of medical speech
Experience with Android ML deployment, Whisper, quantized LLMs, and medical NLP is strongly preferred.
1. Use case
A doctor interacts with a patient while the app listens in real time and performs **live speech transcription**. Once the interaction ends, the transcription is processed fully offline, without any internet connection, using Whisper. After transcription, a lightweight on-device LLM categorizes the medical text into predefined clinical categories.
2. Key requirements
- Real-time or near real-time streaming transcription
- Offline transcription using Whisper
- Fully offline LLM inference on an Android tablet
- No use of cloud APIs, GPT, Gemini, or any internet-based models
- The Android device has 6GB RAM, so the model must be quantized and optimized
- The LLM must understand medical terminology and perform accurate categorization
3. LLM expectations
We are specifically interested in medical-domain trained models, similar in capability to MedGemma, but runnable locally on-device. The app should support a dropdown selection allowing us to switch between multiple LLMs for evaluation and accuracy comparison.
Preferred candidate models include:
- MedAlpaca 7B (LLaMA-based, medical fine-tuned)
- Meditron 7B
- BiomedLM
- Biomistral 7B
4. The selected developer must:
- Identify suitable quantized versions of these models (or recommend better alternatives)
- Ensure they can run efficiently on an Android tablet with 6GB RAM
- Integrate model selection and inference into the app
- Optimize for performance, stability, and accuracy
5. Project goals
- Accurate streaming Whisper-based transcription
- Lightweight, on-device medical LLM
- Reliable categorization of medical speech
Experience with Android ML deployment, Whisper, quantized LLMs, and medical NLP is strongly preferred.