Computer Vision R&D Engineer: Benchmark Foot Contact Detection Methods (Moving Camera)

Job ID: 40145291

Budget: $25 – $50 AUD

Project Overview
I'm seeking an experienced Computer Vision Research Engineer to lead a technical benchmarking project for my development team. This is an exploratory R&D role, not a straightforward implementation task.
The Technical Challenge: We need to achieve frame accurate foot contact detection (±1-2 frames) from 2D video footage captured with a moving/panning camera. This is a non trivial computer vision problem because camera ego motion makes it difficult to distinguish actual ground contacts from apparent motion in screen space.

What We Need
You will evaluate multiple state of the art approaches to solve this problem and recommend the best solution for our production pipeline.

Your Responsibilities
This is an exploratory R&D collaboration, not a rigid checklist. The goal is continuous knowledge transfer to our lead developer as you investigate solutions.

Research & Propose

1.Survey current SOTA methods (WHAM, TRAM, SLAHMR, 4D Humans, physics-based approaches, etc.)
2.Propose 3-5 candidate methods worth benchmarking with technical justification
3.Define evaluation metrics (frame accuracy, false positive rate, robustness to camera motion, etc.)

Implement & Test

1.Get multiple approaches running on our test footage
2.Handle the inevitable dependency conflicts and missing implementation details
3.Systematically compare results across consistent test cases
4.Document failure modes and edge cases

Analyze & Recommend

1.Deliver comprehensive technical report comparing all methods
2.Provide clear recommendation for production deployment
3.Include code repositories (Docker containers preferred)

Important: These phases are a flexible guide, not rigid milestones. We value ongoing communication and knowledge sharing over formal deliverables. The real goal is to equip our lead developer with deep understanding of what works, what doesn't, and why, so they can make informed decisions going forward.
Think of this as a technical partnership where you're the domain expert helping us navigate the landscape of foot contact detection methods.


Required Qualifications
Must Have:
- Demonstrable experience with 3D human pose estimation from monocular video
- Familiarity with research codebases (PyTorch, handling dependency issues, adapting paper implementations)
- Understanding of camera calibration, SLAM, or structure-from-motion concepts
- Experience with at least one of: SMPL models, motion capture, biomechanics, or gait analysis
- Strong Python skills and comfort with Docker/conda environments
- Ability to read and implement from recent computer vision papers

Strongly Preferred:
- Published research or contributions to open-source CV projects
- Experience with temporal filtering, Kalman filters, or physics-based constraints
- Previous work on moving camera scenarios (not just static camera pose estimation)
- Familiarity with ground contact detection in sports analytics or motion capture

Application Process

Step 1: Pre Screening Questions

Please answer these questions in your application:

1.What's the most recent computer vision paper you've implemented or experimented with? What challenges did you encounter?
2.Have you worked with monocular 3D human pose estimation? Which specific methods (e.g. HMR, PARE, WHAM, etc.)?
3.In your view, what's the fundamental challenge of detecting foot contacts from a moving camera versus a static camera?
4.If you had to start this project with a 2 week deadline for initial results, what would be your day one strategy?
5.Share a GitHub link to a computer vision project you've worked on (or describe your most relevant technical work).

Applications without answers to these questions will not be considered. Applications using AI generated answers will not be considered.

Step 2: Trial Task

If your screening answers are strong, we'll invite you to a trial:

Trial Deliverables:
- We'll send you a test video (moving camera, multiple footstep events)
- You propose 2-3 approaches you think are worth benchmarking and explain why
- You implement one approach as a proof-of-concept

Deliver:
- Video with contact visualization overlay (red/blue boxes when foot contacts detected)
- CSV/JSON export with frame-by-frame contact predictions
- 1-2 page write up on methodology and limitations
- GitHub repo or Docker container we can run

Timeline: 5-7 days from receiving test video

Evaluation Criteria:
- Technical depth of proposed approaches
- Implementation quality (does it run without manual fixes?)
- Clarity of visualization and documentation
- Realistic assessment of accuracy and failure modes

Step 3: Full Project

Top trial performer will be invited to complete the full benchmarking project (8-12 weeks, part-time).
Time Commitment: ~15-25 hours/week for 8-12 weeks

Start Date: Immediately for qualified candidates

What Success Looks Like

At the end of this project, we will have:
- A comprehensive technical report comparing 3-5 methods
- Working code repositories for each approach (Docker preferred)
- Clear recommendation on which method to deploy in production
- Frame-by-frame accuracy benchmarks on our test dataset
- Documentation our team can use to implement the winning approach

Ideal Candidate Profile

You're a great fit if you:
- Stay current with computer vision research (read recent papers, attend conferences, engage with research community)
- Enjoy the messy work of getting research code to actually run
- Think systematically about evaluation metrics and tradeoffs
- Can explain complex technical concepts clearly to non-CV experts
- Are comfortable saying "this method won't work because..." rather than overselling

You're NOT a great fit if you:
- Only work with commercial tools or pre-built APIs
- Claim every problem can be solved with a custom neural network
- Can't handle ambiguity or exploratory research
- Need hand holding with dependency management or debugging


How to Apply
Your application must include:
1.Answers to all 5 pre screening questions (see above)
2.Links to relevant work (GitHub, papers, portfolio)
3.Availability (hours/week you can commit)
4.Your rate for the trial task and estimated rate for full project
Applications without complete answers will be automatically rejected.