Extract articles from PDF page -- 2

Job ID: 34290784

Budget: $250 – $750 USD

I need to extract articles from any PDF file like the sample attached.

You can find a sample of how the texts and regions are extracted here:
https://www.pressreader.com/brazil/folha-de-s-paulo/20220720/

Here's a tool that promised to do the same but it's offline:
https://www.pdftron.com/pdf-tools/article-extraction/

You're supposed to developed an article extraction that generates a JSON or XML file from any newspaper or magazine PDF file. In the image "article-extraction.jpg" you can see how it should be extracted from.

Technologies accepted: Java, Linux, Kotlin - open source, it can't depend on cloud or any other paid services.

Step1 - Development - You generate a json/xml from a pdf that follows these rules and you win the project.
Step2 - Tests - You send us the JAR (executable) file so we can test with other pdf files
Step3 - Payment - If works, we release you 50% of the payment and you send the sources. If it's everything ok with the source code you'll have the other half released.