AI Prompting & Data Parsing Specialist Needed
Budget: $30 – $250 AUD
We’re seeking an experienced Prompt Engineer or AI/LLM Automation Specialist to help us develop a scalable, structured data parsing workflow for timestamped transcripts (typically property inspection recordings). You’ll be working with GPT (OpenAI) or Gemini (Google) to convert audio-derived transcripts into structured CSV or spreadsheet-ready formats using a defined parsing ruleset.
This role requires both prompt engineering expertise and the ability to help us overcome common LLM issues like memory limits, prompt chaining failures, and systemic parsing inconsistencies.
The task involves converting timestamped transcripts (typically audio/video inspection reports) into structured CSV outputs based on predefined logic.
Key Requirements:
- Extract the following information from transcripts:
- Damage descriptions
- Timestamp details
- Inspector's remarks
- Allocation of the areas mentioned
- Allocation of the items mentioned
- Use a combination of rule-based and machine learning approaches for parsing.
What you'll be doing:
• Reviewing our current prompting methods and parsing rules for transcripts.
• Designing and testing better prompt flows to improve reliability and reduce errors.
• Creating automated workflows, agents, or APIs to repeatedly run parsing tasks at scale.
• Advising us on LLM platform selection (e.g. GPT-4o vs Gemini Advanced) for best results.
• Solving for LLM memory limitations by implementing chunking, memory tokens, or tool integrations.
• (Optional) Writing lightweight scripts or tools to speed up future parsing runs.
Ideal Skills and Experience:
• Expertise in AI and machine learning, particularly with large language models (LLMs).
• Advanced experience in GPT-4, GPT-4o, or Gemini Advanced.
• Proven ability to build structured prompt flows or agents that handle complex logic.
• Familiarity with token limits, memory handling, and LLM automation (e.g., Assistants API, LangChain, or similar).
• Experience parsing or restructuring large text datasets into structured formats (e.g. CSV, Notion, Airtable).
• Strong documentation and communication skills.
• Strong data parsing and structuring skills.
• Proficiency in generating consistent and accurate outputs tailored for property inspection reports or condition assessments.
• Experience with CSV file formatting and export.
Example Input → Output:
• Input: A transcript with time-based entries (e.g., “0:04 The kitchen sink is slightly leaking”).
• Output: A structured table with columns like:
• Timestamp, Area, Item, Comment
• Using a fixed mapping list (e.g., “Kitchen” → Area, “Sink” → Item)
We will provide:
• A fixed list of Areas and Items to guide the parser.
• Sample transcripts and ideal outputs.
• A detailed list of parsing rules (e.g. keyword exceptions, fuzzy matching rules, inheritance logic).
Deliverables:
• A reusable and well-documented LLM-powered parsing agent that:
• Takes transcript + item list input.
• Returns structured output via UI or exportable format.
• Recommendations on:
--> Optimal LLM platform and model choice.
--> How to scale/automate further parsing tasks (batch processing, API flows, etc.).
--> Debugging and enhancement of systemic memory or misclassification issues.
To Apply, Please Include:
• A short intro outlining your experience with LLM prompt engineering and data parsing.
• Any relevant examples of prompt/agent design (especially if transcript or CSV-related).
• Which platform (OpenAI/Gemini/etc.) you’d recommend for this type of job and why.
Looking forward to your expertise to make this project a success!
This role requires both prompt engineering expertise and the ability to help us overcome common LLM issues like memory limits, prompt chaining failures, and systemic parsing inconsistencies.
The task involves converting timestamped transcripts (typically audio/video inspection reports) into structured CSV outputs based on predefined logic.
Key Requirements:
- Extract the following information from transcripts:
- Damage descriptions
- Timestamp details
- Inspector's remarks
- Allocation of the areas mentioned
- Allocation of the items mentioned
- Use a combination of rule-based and machine learning approaches for parsing.
What you'll be doing:
• Reviewing our current prompting methods and parsing rules for transcripts.
• Designing and testing better prompt flows to improve reliability and reduce errors.
• Creating automated workflows, agents, or APIs to repeatedly run parsing tasks at scale.
• Advising us on LLM platform selection (e.g. GPT-4o vs Gemini Advanced) for best results.
• Solving for LLM memory limitations by implementing chunking, memory tokens, or tool integrations.
• (Optional) Writing lightweight scripts or tools to speed up future parsing runs.
Ideal Skills and Experience:
• Expertise in AI and machine learning, particularly with large language models (LLMs).
• Advanced experience in GPT-4, GPT-4o, or Gemini Advanced.
• Proven ability to build structured prompt flows or agents that handle complex logic.
• Familiarity with token limits, memory handling, and LLM automation (e.g., Assistants API, LangChain, or similar).
• Experience parsing or restructuring large text datasets into structured formats (e.g. CSV, Notion, Airtable).
• Strong documentation and communication skills.
• Strong data parsing and structuring skills.
• Proficiency in generating consistent and accurate outputs tailored for property inspection reports or condition assessments.
• Experience with CSV file formatting and export.
Example Input → Output:
• Input: A transcript with time-based entries (e.g., “0:04 The kitchen sink is slightly leaking”).
• Output: A structured table with columns like:
• Timestamp, Area, Item, Comment
• Using a fixed mapping list (e.g., “Kitchen” → Area, “Sink” → Item)
We will provide:
• A fixed list of Areas and Items to guide the parser.
• Sample transcripts and ideal outputs.
• A detailed list of parsing rules (e.g. keyword exceptions, fuzzy matching rules, inheritance logic).
Deliverables:
• A reusable and well-documented LLM-powered parsing agent that:
• Takes transcript + item list input.
• Returns structured output via UI or exportable format.
• Recommendations on:
--> Optimal LLM platform and model choice.
--> How to scale/automate further parsing tasks (batch processing, API flows, etc.).
--> Debugging and enhancement of systemic memory or misclassification issues.
To Apply, Please Include:
• A short intro outlining your experience with LLM prompt engineering and data parsing.
• Any relevant examples of prompt/agent design (especially if transcript or CSV-related).
• Which platform (OpenAI/Gemini/etc.) you’d recommend for this type of job and why.
Looking forward to your expertise to make this project a success!
Related categories:
Data Processing
GPT Agent
ChatGPT Prompt
LLM Prompt Engineering
Large Language Models (LLMs)