Automated Extraction and Structuring of Recipes from PDF Cookbooks
Budget: $250 – $750 USD
I have a collection of approximately 50 PDF cookbooks, containing an estimated total of 13,000 recipes. Some of these cookbooks have already been split into individual recipe files, while others remain in their original format. I am seeking a freelancer to automate the process of separating the remaining recipes, using the most appropriate tools, potentially including DALL-E or other advanced OCR and text recognition software.
Scope of Work:
The selected freelancer will be responsible for the following tasks:
1. Text Recognition:
• Extract text from the recipes, even if they are embedded in images within the PDF files.
• Ensure the recognition process accurately captures all text, including any non-standard fonts or formats used in the cookbooks.
2. Recipe Identification:
• Detect and separate individual recipes, even if they span across multiple pages.
• Ensure that each recipe is fully captured, without splitting across pages unless the recipe itself does.
3. Data Conversion:
• Convert the extracted text from each recipe into a structured JSON format.
• The JSON should include fields such as Title, Title of the Book, Short Description, Ingredients, Cooking Process, Categories, and Tags.
• Categories and Tags will be provided by me.
4. Language Consistency:
• Ensure that all extracted recipes are in English. For any non-English recipes, a translation process may be necessary.
5. Database Creation:
• Input the JSON-formatted recipes into a database platform, such as Airtable, which I will provide access to.
• Each recipe entry in the database must include a unique identification number and all the relevant fields.
Deliverables:
• A fully populated Airtable database containing all the recipes, accurately categorized and tagged.
• JSON files for each recipe, stored in a systematic folder structure.
• A report detailing the process, including any challenges encountered and how they were addressed.
Skills Required:
• Expertise in OCR technology and text extraction from PDFs, especially where text is embedded in images.
• Experience with tools like DALL-E or similar for image and text recognition.
• Strong knowledge of JSON formatting and database management.
• Familiarity with Airtable or similar database platforms.
• Fluency in English, with experience in translation if necessary.
Timeline:
Please provide an estimated timeline for completing this project, considering the volume of work involved.
Budget:
I am open to bids, but please provide a detailed breakdown of costs, including any software licenses or tools that may be required.
Application Requirements:
• Please provide examples of similar projects you have completed.
• A brief outline of the tools and methods you would use to accomplish this project.
• Your proposed timeline and budget.
Scope of Work:
The selected freelancer will be responsible for the following tasks:
1. Text Recognition:
• Extract text from the recipes, even if they are embedded in images within the PDF files.
• Ensure the recognition process accurately captures all text, including any non-standard fonts or formats used in the cookbooks.
2. Recipe Identification:
• Detect and separate individual recipes, even if they span across multiple pages.
• Ensure that each recipe is fully captured, without splitting across pages unless the recipe itself does.
3. Data Conversion:
• Convert the extracted text from each recipe into a structured JSON format.
• The JSON should include fields such as Title, Title of the Book, Short Description, Ingredients, Cooking Process, Categories, and Tags.
• Categories and Tags will be provided by me.
4. Language Consistency:
• Ensure that all extracted recipes are in English. For any non-English recipes, a translation process may be necessary.
5. Database Creation:
• Input the JSON-formatted recipes into a database platform, such as Airtable, which I will provide access to.
• Each recipe entry in the database must include a unique identification number and all the relevant fields.
Deliverables:
• A fully populated Airtable database containing all the recipes, accurately categorized and tagged.
• JSON files for each recipe, stored in a systematic folder structure.
• A report detailing the process, including any challenges encountered and how they were addressed.
Skills Required:
• Expertise in OCR technology and text extraction from PDFs, especially where text is embedded in images.
• Experience with tools like DALL-E or similar for image and text recognition.
• Strong knowledge of JSON formatting and database management.
• Familiarity with Airtable or similar database platforms.
• Fluency in English, with experience in translation if necessary.
Timeline:
Please provide an estimated timeline for completing this project, considering the volume of work involved.
Budget:
I am open to bids, but please provide a detailed breakdown of costs, including any software licenses or tools that may be required.
Application Requirements:
• Please provide examples of similar projects you have completed.
• A brief outline of the tools and methods you would use to accomplish this project.
• Your proposed timeline and budget.