Audio Transcription Analyst / Data Annotator (Speech Data)

Job ID: 40521310

Budget: $2 – $8 USD

Job Title: Audio Transcription Analyst / Data Annotator (Speech Data)
Overview

This role involves listening to short audio clips (conversational snippets between users and voice agents) and producing detailed, rule-governed transcriptions used to train and evaluate AI speech models. This is not casual transcription — it requires close adherence to a 38-page style guide covering tagging conventions, formatting, and metadata labeling.
Core Responsibilities

For each assigned audio file, the analyst will produce a spoken-form transcription (lowercase, untagged punctuation, tag-based annotation of audio features like overlapping speech, whispering, filled pauses, unintelligible audio, etc.) and a written-form transcription (the same content reformatted with standard capitalization, punctuation, and digit-based numbers). They'll also assign a save state (Good or Discard, with a discard reason when applicable) and tag each identifiable speaker's gender and nativity according to defined categories.
Required Skills and Qualities

A strong fit for this role would have native or near-native fluency in English (specifically familiarity with American spelling conventions, since the American Heritage Dictionary is the spelling authority), excellent listening comprehension including the ability to parse overlapping speech, accents, mumbled or distorted audio, and disfluent speech, meticulous attention to detail since the rules involve dozens of tags with strict positional and formatting requirements that must be applied consistently, comfort with ambiguity and judgment calls (e.g., distinguishing low voice from quiet audio, or deciding when a word is "guessable" versus truly unintelligible), basic research skills for verifying spellings of catalog entities, names, and non-target-language words using specified dictionaries and search strategies, and the discipline to follow a rules-based system precisely rather than transcribing intuitively or relying on ASR shortcuts.
Helpful but Not Strictly Required

Exposure to linguistics, audio annotation, closed captioning, or content moderation work is a plus, as is familiarity with a second language (for handling non-target language speech) and general pop-culture awareness (the catalog entity list includes musicians, YouTubers, actors, and shows, so recognizing names speeds up verification).
Work Style

This is detail-oriented, repetitive, rules-heavy work best suited to someone who's comfortable with structured, almost checklist-like tasks rather than open-ended creative writing. Good candidates tend to be the type who double-check formatting against a style guide rather than going from memory.
If you want, I can turn this into a polished job posting (with a punchier intro, required vs. preferred sections, and a call to action) or a formatted Word doc you could actually post.