Machine Learning Engineer - Pose Detection & Motion Analysis

Job ID: 39857619

Budget: $250 – $750 AUD

What I’m building:
I’m developing a system that detects when each foot makes contact with the ground using human pose data, combining pose estimation, temporal filtering, and signal analysis. As a Foley Artist, I’m building this tool to streamline the process of syncing footsteps and movement with sound. The system runs videos through a pose detection model to pinpoint the exact frame of every footstep, then outputs a text file listing those events in precise SMPTE timecode, ready for syncing with my recorded footstep audio. The long term goal is to extend this to multi character videos, accurately separating and labelling footsteps for multiple people within the same scene.

We’re not a big company (barely a company) more like a sound guy with an idea and a recent computer science graduate who’s been building out the core code.
Our goal is high accuracy, frame by frame detection that matches ground truth down to precise heel-strike. Our prototype is functional and demonstrates the core idea. We’re now looking for an experienced engineer to review the current code, help track down bugs, and set a clear roadmap toward a stable and commercially compliant system.
You’ll be giving us guidelines to shape how the system grows, improving pose accuracy, tightening the pipeline, choosing safe models and datasets, and sticking around for ongoing check ins as we keep building.

Parts I Need Help With:
- Review and refine our existing Jupyter based prototype (Python, PyTorch)
Evaluate:
- Pose model integration (RTMPose / WholeBody / MediaPipe)
- Event detection and filtering logic (heel-strike, toe-off)
- Ground truth alignment, metrics, and overall reliability
- Identify any non commercial or research only assets currently in use and suggest safe replacements
Advise on datasets and annotation strategy:
- We’re currently using smaller research databases to debug the system before investing in a large commercial dataset
- We’ll need help determining what kind of annotations (2D vs 3D, keypoint density, heel/toe labels, etc.) will deliver the best results for training and evaluation
- Design a roadmap from quick technical fixes to long term system improvements
- Offer ongoing consultation as the project expands, including dataset planning, model fine tuning, and deployment guidance

You’ll Need:
Solid experience with pose estimation frameworks (MMPose, RTMPose, RTMW, MediaPipe, or OpenPose)
Strong Python / PyTorch background and ability to improve or refactor research style code
Understanding of temporal filtering and kinematic event detection (heel-strike, toe-off)
Familiarity with dataset structure, labelling formats, and annotation quality for 2D and 3D pose data
Knowledge of open-source licensing and ability to flag non commercial code, models, or datasets

Deliverables:
A technical review report covering:
- Code quality, pose accuracy, and pipeline structure
- Non-commercial or restricted assets and suggested replacements
- Dataset recommendations (annotation type, 2D vs 3D, keypoint coverage)
- A roadmap document outlining short term fixes and long term improvements
- Option for ongoing consulting to support implementation, dataset curation, and scaling

How to Apply

Please send:
A short intro and relevant experience (GitHub, portfolio, or prior ML projects)
Experience with pose detection, dataset preparation, or motion analysis
Hourly/day rate and availability for ongoing consulting