High-Accuracy Multimodal Cry Classifier
Budget: ₹12,500 – ₹37,500 INR
I need a researcher who can build a production-ready model that listens to a baby’s cry, watches the paired video, and decides—reliably—whether the cause is hunger, discomfort, or simple attention seeking. Audio and video must be fused inside one architecture; running them in parallel but independently will not satisfy our accuracy goals.
You may use the deep-learning stack you trust most (PyTorch, TensorFlow, Keras, OpenCV, torchaudio, etc.) provided the final network can run in real time on an edge device and be exported to ONNX or TFLite. I will share product constraints and a small proprietary data set; you will expand it through public sources or augmentation, perform rigorous cross-validation, and refine the model until we consistently exceed 90 % precision and recall on an unseen hold-out set.
When you apply, show me past work—links to papers, GitHub repos, Kaggle solutions, or shipped features—demonstrating experience with cry detection, sound-event recognition, emotion analysis, or any other multimodal perception problem. A concise paragraph with links is enough; no full proposal is needed at this stage.
Deliverables
• Well-documented training pipeline and source code
• Trained model file(s) plus lightweight export (ONNX/TFLite)
• Inference script or microservice, ready for product integration
• Evaluation report: confusion matrix, per-class metrics, brief methodology
• Integration guide detailing inputs, outputs, and runtime footprint
Payment is released as soon as the artefacts are reviewed and meet the stated accuracy target.
You may use the deep-learning stack you trust most (PyTorch, TensorFlow, Keras, OpenCV, torchaudio, etc.) provided the final network can run in real time on an edge device and be exported to ONNX or TFLite. I will share product constraints and a small proprietary data set; you will expand it through public sources or augmentation, perform rigorous cross-validation, and refine the model until we consistently exceed 90 % precision and recall on an unseen hold-out set.
When you apply, show me past work—links to papers, GitHub repos, Kaggle solutions, or shipped features—demonstrating experience with cry detection, sound-event recognition, emotion analysis, or any other multimodal perception problem. A concise paragraph with links is enough; no full proposal is needed at this stage.
Deliverables
• Well-documented training pipeline and source code
• Trained model file(s) plus lightweight export (ONNX/TFLite)
• Inference script or microservice, ready for product integration
• Evaluation report: confusion matrix, per-class metrics, brief methodology
• Integration guide detailing inputs, outputs, and runtime footprint
Payment is released as soon as the artefacts are reviewed and meet the stated accuracy target.
Related categories:
Python
Data Processing
Algorithm
Data Science
Keras
Computer Vision
Deep Learning
Natural Language Processing