Catalan Audio Transcription & Annotation
Budget: $8 – $15 USD
About the Role
AI models learn spoken language from carefully transcribed and labeled audio. That is the work here.
You receive batches of short audio clips in Catalan. For each one you produce an accurate transcript of what is said, then answer a small set of structured questions about the voices in the clip: how many people are speaking, the gender of each speaker, and how fluent each sounds (native versus non-native). The value is your ear as a native speaker, so the labels have to reflect what you actually hear, not a guess. Clear audio takes only a few minutes per clip; harder clips with overlapping or background speech take longer.
No prior AI knowledge is needed, and there is no education requirement. If you are a native Catalan speaker with solid English (C1/C2) and a good ear, you can do this well.
Key Responsibilities
Transcribe Catalan audio clips accurately, following the formatting and spelling conventions in the project guidelines
For each clip, record the number of distinct speakers
For each speaker, label gender and fluency (native or non-native)
Flag clips where the audio is unclear, corrupted, or not actually Catalan
Keep your transcripts and labels consistent across the full batch
Apply reviewer feedback and correct your work where the guidelines call for it
Ideal Qualifications
Native Catalan speaker
English at C1 to C2
A careful ear for accent, dialect, and fluency in spoken Catalan
Comfortable working independently and remotely on task-based work
Reliable internet, good headphones, and a quiet space to listen
Nice to Have
Prior experience on an audio transcription or subtitling project
Familiarity with regional Catalan varieties (Central, Valencian, Balearic, Northern)
Experience with structured labeling or annotation guidelines.
What Success Looks Like
Transcripts that match the audio word for word, with the project's spelling and formatting conventions applied consistently
Speaker counts and gender and fluency labels that hold up when a reviewer checks them against the clip
Consistency across the whole batch rather than clip-by-clip drift
Steady, reliable throughput once you are up to speed
AI models learn spoken language from carefully transcribed and labeled audio. That is the work here.
You receive batches of short audio clips in Catalan. For each one you produce an accurate transcript of what is said, then answer a small set of structured questions about the voices in the clip: how many people are speaking, the gender of each speaker, and how fluent each sounds (native versus non-native). The value is your ear as a native speaker, so the labels have to reflect what you actually hear, not a guess. Clear audio takes only a few minutes per clip; harder clips with overlapping or background speech take longer.
No prior AI knowledge is needed, and there is no education requirement. If you are a native Catalan speaker with solid English (C1/C2) and a good ear, you can do this well.
Key Responsibilities
Transcribe Catalan audio clips accurately, following the formatting and spelling conventions in the project guidelines
For each clip, record the number of distinct speakers
For each speaker, label gender and fluency (native or non-native)
Flag clips where the audio is unclear, corrupted, or not actually Catalan
Keep your transcripts and labels consistent across the full batch
Apply reviewer feedback and correct your work where the guidelines call for it
Ideal Qualifications
Native Catalan speaker
English at C1 to C2
A careful ear for accent, dialect, and fluency in spoken Catalan
Comfortable working independently and remotely on task-based work
Reliable internet, good headphones, and a quiet space to listen
Nice to Have
Prior experience on an audio transcription or subtitling project
Familiarity with regional Catalan varieties (Central, Valencian, Balearic, Northern)
Experience with structured labeling or annotation guidelines.
What Success Looks Like
Transcripts that match the audio word for word, with the project's spelling and formatting conventions applied consistently
Speaker counts and gender and fluency labels that hold up when a reviewer checks them against the clip
Consistency across the whole batch rather than clip-by-clip drift
Steady, reliable throughput once you are up to speed
Related categories:
Audio Services
Transcription
Catalan Translator
AI (Artificial Intelligence) HW/SW
Data Annotation