Python Engineer for Metadata Processing & RAG
Budget: ₹1,500 – ₹12,500 INR
Title: Python Data Engineer for Anna's Archive Metadata Pipeline, Book Deduplication & RAG Dataset Preparation
Project Overview
We are building a large-scale book and knowledge ingestion pipeline for an AI-powered RAG (Retrieval Augmented Generation) platform.
The freelancer will be responsible for setting up the complete pipeline, starting from Anna's Archive metadata dumps through deduplication, content extraction, chunking, and preparation of a clean corpus ready for embeddings and vector database ingestion.
Scope of Work
1. Download and process Anna's Archive metadata dumps:
* LibGen RS
* LibGen Fiction
* Magazines
* Sci-Hub / Scientific Articles
2. Build metadata ingestion pipeline:
* Parse large JSON/JSONL/GZ/ZST files
* Create efficient storage schema
* Handle incremental updates
3. Deduplication Engine:
* Remove duplicate titles across sources
* Author normalization
* Version selection logic
* Format prioritization (EPUB > PDF > MOBI)
4. Download & Processing Pipeline:
* Metadata filtering
* File retrieval workflow
* EPUB/PDF text extraction
* Data cleaning and normalization
5. RAG Preparation:
* Token-based chunking
* JSONL corpus generation
* Metadata enrichment
* Embedding-ready output
6. Infrastructure Setup:
* Linux server setup (Hetzner/VPS)
* Python environment configuration
* Logging and monitoring
* Resume and recovery mechanisms
Deliverables
* Complete production-ready codebase
* Installation and deployment documentation
* Automated ingestion scripts
* Deduplication pipeline
* Text extraction pipeline
* Chunking pipeline
* Final JSONL corpus generation workflow
* Server setup guide
* Knowledge transfer session
Required Skills
* Advanced Python
* Large-scale data processing
* JSONL / GZIP / Zstandard processing
* Async programming (aiohttp, asyncio)
* Data engineering
* ETL pipeline development
* Linux server administration
* Elasticsearch/OpenSearch
* PostgreSQL or SQLite
* Text extraction from PDF and EPUB
* RAG and vector database fundamentals
* Git and deployment workflows
Preferred Experience
* Anna's Archive datasets
* LibGen metadata processing
* Digital library projects
* Knowledge graph or search systems
* Qdrant, Weaviate, Milvus, or pgvector
* LLM/RAG pipelines
* Large-scale document ingestion systems
When Applying
Please provide:
1. Similar projects completed.
2. Experience processing datasets with millions of records.
3. Experience with PDF/EPUB extraction.
4. Experience building RAG or vector-search systems.
5. Estimated timeline and cost.
6. GitHub profile and relevant portfolio links.
Project Overview
We are building a large-scale book and knowledge ingestion pipeline for an AI-powered RAG (Retrieval Augmented Generation) platform.
The freelancer will be responsible for setting up the complete pipeline, starting from Anna's Archive metadata dumps through deduplication, content extraction, chunking, and preparation of a clean corpus ready for embeddings and vector database ingestion.
Scope of Work
1. Download and process Anna's Archive metadata dumps:
* LibGen RS
* LibGen Fiction
* Magazines
* Sci-Hub / Scientific Articles
2. Build metadata ingestion pipeline:
* Parse large JSON/JSONL/GZ/ZST files
* Create efficient storage schema
* Handle incremental updates
3. Deduplication Engine:
* Remove duplicate titles across sources
* Author normalization
* Version selection logic
* Format prioritization (EPUB > PDF > MOBI)
4. Download & Processing Pipeline:
* Metadata filtering
* File retrieval workflow
* EPUB/PDF text extraction
* Data cleaning and normalization
5. RAG Preparation:
* Token-based chunking
* JSONL corpus generation
* Metadata enrichment
* Embedding-ready output
6. Infrastructure Setup:
* Linux server setup (Hetzner/VPS)
* Python environment configuration
* Logging and monitoring
* Resume and recovery mechanisms
Deliverables
* Complete production-ready codebase
* Installation and deployment documentation
* Automated ingestion scripts
* Deduplication pipeline
* Text extraction pipeline
* Chunking pipeline
* Final JSONL corpus generation workflow
* Server setup guide
* Knowledge transfer session
Required Skills
* Advanced Python
* Large-scale data processing
* JSONL / GZIP / Zstandard processing
* Async programming (aiohttp, asyncio)
* Data engineering
* ETL pipeline development
* Linux server administration
* Elasticsearch/OpenSearch
* PostgreSQL or SQLite
* Text extraction from PDF and EPUB
* RAG and vector database fundamentals
* Git and deployment workflows
Preferred Experience
* Anna's Archive datasets
* LibGen metadata processing
* Digital library projects
* Knowledge graph or search systems
* Qdrant, Weaviate, Milvus, or pgvector
* LLM/RAG pipelines
* Large-scale document ingestion systems
When Applying
Please provide:
1. Similar projects completed.
2. Experience processing datasets with millions of records.
3. Experience with PDF/EPUB extraction.
4. Experience building RAG or vector-search systems.
5. Estimated timeline and cost.
6. GitHub profile and relevant portfolio links.