Corpus coding
Budget: $30 – $250 AUD
The file contains aligned source-text and target-text segments. Each row includes one source segment and its corresponding target-language rendering. The task is to review the aligned texts and create a coded annotation sheet for culture-specific items, CSIs.
The work involves identifying, recording, categorising, and coding CSIs in the source text and comparing them with their renderings in the target text.
For each aligned segment, please do the following:
Read the source-text segment carefully.
Identify any culture-specific item, CSI. A CSI refers to any word or phrase that carries cultural, religious, legal, social, historical, place-related, time-related, behavioural, or formulaic meaning.
Record each CSI occurrence separately. Do not remove repeated items. If the same CSI appears more than once, each occurrence must be recorded because frequency and recurrence will be analysed later.
Record the exact target-language rendering of the CSI. Do not translate, correct, or rewrite the target text. Copy the rendering exactly as it appears.
Keep the original text ID for every entry. Do not change, shorten, or remove the ID. This ID is needed for later analysis.
Keep the theme or general topic label for every entry. This will be used later to compare whether different topics show different CSI patterns or translation techniques.
Group spelling and grammatical variants under one main CSI when they refer to the same concept. For example, different spellings or grammatical forms of one term should be linked to the same main CSI. However, if a similar expression has a different function or meaning, it should be treated as a separate CSI.
Classify each CSI according to the agreed CSI category system.
Code each target-language rendering according to the agreed translation technique system.
Assign only one main category and one main translation technique for each occurrence. If more than one option seems possible, choose the dominant one and add a note in the comments column.
The annotation sheet should include these columns:
Text ID
Date or sequence number, if available
Theme or general topic
Source segment
Target segment
CSI in the source text
Main CSI form
Target-language rendering
CSI category
Translation technique Please keep all repeated items, spelling variants, and different target-language renderings. These differences are important for the later analysis.
The final output should be a clean Excel sheet where each CSI occurrence is linked to its text ID, theme, source segment, target segment, CSI category, and translation technique.
Please treat the material as confidential. Do not share, copy, upload, or discuss the data or the project details with anyone. The task is limited to annotation and coding only.
The work involves identifying, recording, categorising, and coding CSIs in the source text and comparing them with their renderings in the target text.
For each aligned segment, please do the following:
Read the source-text segment carefully.
Identify any culture-specific item, CSI. A CSI refers to any word or phrase that carries cultural, religious, legal, social, historical, place-related, time-related, behavioural, or formulaic meaning.
Record each CSI occurrence separately. Do not remove repeated items. If the same CSI appears more than once, each occurrence must be recorded because frequency and recurrence will be analysed later.
Record the exact target-language rendering of the CSI. Do not translate, correct, or rewrite the target text. Copy the rendering exactly as it appears.
Keep the original text ID for every entry. Do not change, shorten, or remove the ID. This ID is needed for later analysis.
Keep the theme or general topic label for every entry. This will be used later to compare whether different topics show different CSI patterns or translation techniques.
Group spelling and grammatical variants under one main CSI when they refer to the same concept. For example, different spellings or grammatical forms of one term should be linked to the same main CSI. However, if a similar expression has a different function or meaning, it should be treated as a separate CSI.
Classify each CSI according to the agreed CSI category system.
Code each target-language rendering according to the agreed translation technique system.
Assign only one main category and one main translation technique for each occurrence. If more than one option seems possible, choose the dominant one and add a note in the comments column.
The annotation sheet should include these columns:
Text ID
Date or sequence number, if available
Theme or general topic
Source segment
Target segment
CSI in the source text
Main CSI form
Target-language rendering
CSI category
Translation technique Please keep all repeated items, spelling variants, and different target-language renderings. These differences are important for the later analysis.
The final output should be a clean Excel sheet where each CSI occurrence is linked to its text ID, theme, source segment, target segment, CSI category, and translation technique.
Please treat the material as confidential. Do not share, copy, upload, or discuss the data or the project details with anyone. The task is limited to annotation and coding only.
Related categories:
Translation
Data Entry
Research
Proofreading
Excel
Writing
Documentation
Linguistics