Machine Learning expert advice required for detecting similarities in text in uploaded document from Source document database
Budget: ₹37,500 – ₹75,000 INR
This is the brief description of the idea which I want to implement:
1. There should be a **SOURCE DATABASE** which stores the uploaded documents like .pdf, word etc.
2. So, after we have uploaded/stored, say, 10 PDFs or documents in the **SOURCE DATABASE**
3. Then, we should submit a pdf/doc or text document (**UPLOADED DOCUMENT**) containing several paragraphs, and the algorithim should be able to extract the text from the **UPLOADED DOCUMENT** and search the **SOURCE DATABASE** for occurrences of similar content.
4. The reason we are using machine learning here is because it should scan the **SOURCE DATABASE** and **UPLOADED DOCUMENT** sentence by sentence and paragraph by paragraph keeping the context into consideration and only highlight those instances which are really copied. ( after removing the stop words etc)
**#NOTE:** You dont need to make any gui interface, you can just start with lets say 5-6 pdf or word files which act as **source documents** and 1 or 2 files which act as **uploaded document** and start parsing them directly using Python. Please write "I have read and understood the description" in your proposal before placing your bid.
1. There should be a **SOURCE DATABASE** which stores the uploaded documents like .pdf, word etc.
2. So, after we have uploaded/stored, say, 10 PDFs or documents in the **SOURCE DATABASE**
3. Then, we should submit a pdf/doc or text document (**UPLOADED DOCUMENT**) containing several paragraphs, and the algorithim should be able to extract the text from the **UPLOADED DOCUMENT** and search the **SOURCE DATABASE** for occurrences of similar content.
4. The reason we are using machine learning here is because it should scan the **SOURCE DATABASE** and **UPLOADED DOCUMENT** sentence by sentence and paragraph by paragraph keeping the context into consideration and only highlight those instances which are really copied. ( after removing the stop words etc)
**#NOTE:** You dont need to make any gui interface, you can just start with lets say 5-6 pdf or word files which act as **source documents** and 1 or 2 files which act as **uploaded document** and start parsing them directly using Python. Please write "I have read and understood the description" in your proposal before placing your bid.