LLM Engineer Classification Data Labeling

Job ID: 40065619

Budget: $15 – $25 AUD

I have a growing collection of engineer-focused text (linkedin profiles) that I’m using to fine-tune a large language model. I need each snippet labeled with the precise engineering role it reflects and a short, well-reasoned explanation of why that label fits. The immediate focus is on infrastructure-heavy roles—DevOps Engineer, Site Reliability Engineer, and Golang-centric Infrastructure Engineer—but you’ll also come across specialised profiles from big-tech and high-frequency-trading environments, so depth and nuance matter.

Your computer-science background will let you spot the technical signals that separate, say, a Kubernetes-driven SRE from a CI/CD-heavy DevOps professional or a performance-tuned Go engineer. The emphasis is on technical skills rather than soft skills; I want the model to learn the subtle language patterns that reveal tooling, architectures, and scale experience.

Deliverables
• A consistently labeled dataset (JSON or CSV) with one role tag per record.
• A brief, structured justification for each label—bullet-point evidence drawn from the text.
• A concise taxonomy document capturing any edge-case decisions, so future annotators can stay aligned.

Acceptance Criteria
• ≥95 % inter-annotator agreement on a 50-item spot check.
• Justifications reference concrete technical cues (tool names, protocols, architectures).
• Taxonomy remains internally consistent when new roles are introduced.

If you’ve ever interviewed or coached infra engineers—and enjoy turning that expertise into clean, machine-readable annotations—this will be straightforward and engaging work.
Related categories: Computer Science Software Engineering