Create Ground-Truth CSV for Traffic-Light Color per Frame (11 driving videos) — Manual Only (No Auto-Labeling/YOLO)

Job ID: 39765701

Budget: $10 – $100 CAD

Project: Create Ground-Truth CSV for Traffic-Light Color per Frame (11 driving videos) — Manual Only (No Auto-Labeling/YOLO)

What I need I have 11 MP4 driving videos. I need a single master CSV (GT_file.csv) + per-video CSVs that mark every visible traffic-light face per frame and label its color as red, yellow, green, or off, with tight bounding boxes.
Important: This ground truth will be used to compute confusion matrices / precision-recall for our improved models. To keep evaluation unbiased, annotation must be 100% human. • Not allowed: any auto-labeling, model-assisted labeling, pre-labeling, or using YOLO/Detectron/etc. • Allowed: standard tool conveniences like manual box interpolation/track only if you visually confirm every frame.
Scope & labeling rules
Annotate every frame (also quote an option for 10 FPS sampling).
One row per traffic-light box per frame (if 3 lights are visible, that’s 3 rows).
Color labels: red, yellow, green, off (green arrows count as green; off = face visible but no bulb lit).
If glare/occlusion makes color unclear: still box it; set quality_flag=uncertain.
Don’t annotate backs of signals or unrecognizable far blobs; if boxed for continuity, set quality_flag=ignore.
Keep boxes stable across adjacent frames (no jitter), tightly enclosing the front lamp face.
Do not rename files; video must match the original filename exactly.
CSV format (exact columns & order) video,frame,t_sec,x1,y1,x2,y2,color,quality_flag
video: exact filename (e.g., VID_01.mp4)
frame: zero-based frame index (int)
t_sec: seconds from start (frame/fps, 3 decimals)
x1,y1,x2,y2: pixel coords at original resolution (ints)
color: red|yellow|green|off
quality_flag: ok|uncertain|ignore (default ok)
Tiny example
VID_01.mp4,0,0.000,812,146,860,260,red,ok VID_01.mp4,1,0.033,812,146,861,260,red,ok VID_01.mp4,2,0.067,812,146,861,260,yellow,uncertain
Deliverables
GT_file.csv (master across all videos)
Per-video CSVs (same schema)
README: tool used, FPS policy, any edge-case notes
(Nice-to-have) YOLO TXT per frame with class map {red:0, yellow:1, green:2, off:3}
Short attestation that all labels were produced manually with no model assistance
Quality bar / acceptance
Tight, consistent boxes on lamp faces; low jitter across frames.
Clear frames labeled correctly ≥ 98% (spot-check).
Reasonable use of uncertain for glare, occlusion, transitions.
I’ll spot-check 5–10% of frames and may request fixes.
For audit, I may ask for: tool export/project file, brief screen recording of a tiny segment, or timing logs on a small sample.
Tools
Use a professional tool: CVAT, Label Studio, Roboflow Annotate, VIA, etc. Do not use auto-labeling / model-assisted features. Interpolation is fine only with per-frame human confirmation.
Files & privacy
I’ll share the 11 MP4 videos via Google Drive.
Confidential/NDA.
Budget & timeline
Please provide a fixed price and timeline for:
Every frame
10 FPS option
If preferred, quote per annotated minute and state your FPS/length assumptions.
How to apply (short answers, please)
Confirm you will do manual-only labeling (no auto-labeling/model assistance) and describe your process.
Tool you’ll use + why.
Link/screenshot of similar work (small objects/traffic signals is a plus).
How you handle glare/occlusion; when you use uncertain.
Quote & timeline for every frame and 10 FPS.
Willing to do a paid 1-minute pilot first? (yes/no)
Availability & time zone.
(Start your bid with “HUMAN ONLY” so I know you read this.)