Node.js Parser for Comprehensive Document Extraction

Job ID: 38730401

Budget: $250 – $750 USD

I'm looking for a skilled Node.js developer to create a parser that can extract text from various types of files, including pdf, doc, images, xlsx, elm, md, and pptx. The parser should also be capable of translating tables and extracting embedded images and graphics. The final output of the extracted data should be in plain text format.

Key requirements:
- Proficiency in Node.js
- Experience in developing parsers for document extraction
- Ability to handle various file formats
- Capable of translating tables
- Skilled in extracting embedded images and graphics

Please note, the content from text, tables, and embedded images and graphics needs to be extracted. I am open to suggestions for the prioritization of content in the parsing process.
Related categories: Python Node.js LangChain LLM Prompt Engineering