Data Extraction from CSV and PDF using Langchain for Q&A Chatbot

Job ID: 37747218

Budget: $750 – $1,500 USD

We are seeking someone capable of resolving an issue related to extracting data from CSV and PDF files containing unstructured tables. We have developed a Q&A chatbot utilizing Langchain, Rag, and llm code in the backend, implemented using Plumber PDF. We have attached two sets of documents for you to utilize in testing your ability to extract data accurately.

Below are the questions we need answers to:

1. Does Dairyland accept Matricula?
- Answer: Yes, Dairyland accepts Matricula. The information can be found in the CSV file at row D17, where the corresponding column for Matricula is marked as TRUE.

2. Does Anchor accept California ID?
- Answer: Anchor does not accept California IDs (CA ID) for drivers; it is indicated as FALSE in the guidelines. This information is located in the CSV file at C6.

For the PDF attachment, the question is:
What is the filing fee for an SR filing with Triton Seaside?
- Answer: SR filings are issued in California only, and the initial filing fee is $5. It's important to note that the data in the PDF includes colored text, which is presented as an image, requiring special handling for extraction.

If you can successfully resolve this issue, you will be considered for a role in this project. If you have alternative solutions, we are open to exploring them as well.

*********In your bid please type langchain on top**********
Let me know if you test it if so please write in the bid