Python Engineer for Metadata Processing & RAG

Job ID: 40501416

Budget: ₹1,500 – ₹12,500 INR

Title: Python Data Engineer for Anna's Archive Metadata Pipeline, Book Deduplication & RAG Dataset Preparation

Project Overview

We are building a large-scale book and knowledge ingestion pipeline for an AI-powered RAG (Retrieval Augmented Generation) platform.

The freelancer will be responsible for setting up the complete pipeline, starting from Anna's Archive metadata dumps through deduplication, content extraction, chunking, and preparation of a clean corpus ready for embeddings and vector database ingestion.

Scope of Work

1. Download and process Anna's Archive metadata dumps:

* LibGen RS
* LibGen Fiction
* Magazines
* Sci-Hub / Scientific Articles

2. Build metadata ingestion pipeline:

* Parse large JSON/JSONL/GZ/ZST files
* Create efficient storage schema
* Handle incremental updates

3. Deduplication Engine:

* Remove duplicate titles across sources
* Author normalization
* Version selection logic
* Format prioritization (EPUB > PDF > MOBI)

4. Download & Processing Pipeline:

* Metadata filtering
* File retrieval workflow
* EPUB/PDF text extraction
* Data cleaning and normalization

5. RAG Preparation:

* Token-based chunking
* JSONL corpus generation
* Metadata enrichment
* Embedding-ready output

6. Infrastructure Setup:

* Linux server setup (Hetzner/VPS)
* Python environment configuration
* Logging and monitoring
* Resume and recovery mechanisms

Deliverables

* Complete production-ready codebase
* Installation and deployment documentation
* Automated ingestion scripts
* Deduplication pipeline
* Text extraction pipeline
* Chunking pipeline
* Final JSONL corpus generation workflow
* Server setup guide
* Knowledge transfer session

Required Skills

* Advanced Python
* Large-scale data processing
* JSONL / GZIP / Zstandard processing
* Async programming (aiohttp, asyncio)
* Data engineering
* ETL pipeline development
* Linux server administration
* Elasticsearch/OpenSearch
* PostgreSQL or SQLite
* Text extraction from PDF and EPUB
* RAG and vector database fundamentals
* Git and deployment workflows

Preferred Experience

* Anna's Archive datasets
* LibGen metadata processing
* Digital library projects
* Knowledge graph or search systems
* Qdrant, Weaviate, Milvus, or pgvector
* LLM/RAG pipelines
* Large-scale document ingestion systems

When Applying

Please provide:

1. Similar projects completed.
2. Experience processing datasets with millions of records.
3. Experience with PDF/EPUB extraction.
4. Experience building RAG or vector-search systems.
5. Estimated timeline and cost.
6. GitHub profile and relevant portfolio links.
Related categories: Python LaTeX PostgreSQL MariaDB SQLite Elasticsearch Git