Extract articles from PDF page -- 2

Job ID: 36397986

Budget: $15 – $25 USD

I need to extract articles from any PDF file like the sample attached.

You can find a sample of how the texts and regions are extracted here:
https://www.pressreader.com/brazil/folha-de-s-paulo/20220720/

Here's a tool that promised to do the same
https://www.pdftron.com/pdf-tools/article-extraction/

Here's an article about it
https://www.cse.iitd.ac.in/~sumantra/publications/das14_fixed_point.pdf


You're supposed to developed an article extraction that generates a JSON or XML file from any newspaper or magazine PDF file.

Technologies accepted: Java, Linux, Kotlin - open source, it can't depend on cloud or any other paid services.

Step1 - Development - You generate a json/xml from a pdf that follows these rules and you win the project.
Step2 - Tests - You send us the JAR (executable) file so we can test with other pdf files
Step3 - Payment - If works, we release you 50% of the payment and you send the sources. If it's everything ok with the source code you'll have the other half released.