Wristwatch VTO System: Computer Vision Enhancement

Job ID: 38608549

Budget: ₹37,500 – ₹75,000 INR

We are developing a Virtual Try-On (VTO) system for wristwatches and are seeking an experienced computer vision specialist to enhance hand and wrist landmark detection. The project initially used Google’s Mediapipe hand tracking library, but we encountered key limitations:

Issues with Mediapipe:
Lack of Detailed Wrist and Forearm Landmarks: Mediapipe only provides one wrist landmark, which hinders accurate wristwatch alignment during wrist rotations and hand tilts.
Inaccurate Depth Information: The 3D landmarks provided by Mediapipe are not accurate in terms of depth perception, causing misalignment when the hand moves closer to or further from the camera.
We are looking for a freelancer to address these limitations by improving wrist and forearm landmark detection and ensuring the accurate placement of the watch model in 3D space.

Project Scope:
1. Dataset Preprocessing:
Work with FreiHand and Rendered Hand Pose (RHD) datasets.
Extract, preprocess, and augment the datasets to include additional wrist, forearm, and elbow landmarks.
Use semi-automatic or automatic annotation tools to create these additional keypoints.

2. Model Training:
Fine-tune a hand pose model (starting from a pre-trained model like Mediapipe or OpenPose) to improve 2D and 3D wrist and forearm landmark detection.
Pre-train on Rendered Hand Pose (RHD) and fine-tune with FreiHand for improved real-world performance.

3. Integration:
Ensure the model performs in real-time, accurately aligning the wristwatch during hand movements, wrist rotations, and hand tilts.
Optimize the model for lightweight performance (e.g., TensorFlow Lite or ONNX format) to facilitate integration into our VTO system.

4. Testing and Delivery:
Provide a fully tested model with metrics for 2D/3D landmark accuracy, focusing on wrist and forearm movements.
Deliver the model ready for integration, complete with documentation for usage and setup.
Skills Required:
Proficiency in computer vision and machine learning.
Experience with hand pose estimation and landmark detection.
Knowledge of deep learning frameworks (TensorFlow, PyTorch).
Experience with real-time model optimization (e.g., TensorFlow Lite, ONNX).

Preferred Experience:
Previous experience with Mediapipe or OpenPose.
Background in virtual try-on (VTO) systems or augmented reality (AR) applications.

Deliverables:
Preprocessed datasets with additional wrist, forearm, and elbow landmarks.
A trained and optimized hand pose model for real-time wristwatch alignment.
Full documentation and setup instructions.