LLM-Powered RAG Pipeline with Weaviate and MCP Server on AWS
Budget: €18 – €36 EUR
We are looking to build a fully serverless Retrieval-Augmented Generation (RAG) pipeline on AWS, combining LLMs, document embedding, and structured data querying via MCP Server that interacts with Postgres or Redshift database. The system will allow users to chat with their data, leveraging both unstructured documents and relational data from Postgres or Redshift. Note that UI is not needed, only AWS API Gateway Endpoints
System Overview:
- Allow uploading of text into Weaviate Cloud
- Chunk and embed document content using LLM-based embeddings.
- Store and index these embeddings in Weaviate Cloud for fast retrieval.
- Enable user interaction via natural language:
- Retrieve relevant content chunks from Weaviate.
- Combine that with structured queries via MCP Server to Postgres/Redshift.
- Return coherent responses using an LLM (OpenAI, Anthropic, etc.). The system should be model agnostic.
- Be deployed on AWS using a serverless architecture wherever possible.
Technical Requirements:
- Upload → Chunk → Embed → Store in Weaviate.
- Triggered via S3/Lambda/SNS or similar AWS-native flow.
Query Pipeline:
- Accept user prompts.
- Use LangChain (or DSPy) to:
- Query Weaviate for context.
- Query Postgres/Redshift via MCP Server for live data.
- Generate responses via LLM.
Infrastructure as Code:
All infrastructure must be provisioned via Terraform.
Inputs should be parameterised for:
VPC
Route53 DNS
S3 bucket names
Database connection details
Use AWS-native services as much as possible:
Lambda, SNS, SQS, DynamoDB, S3, API Gateway, CloudWatch, IAM, Batch if needed.
Vector Store:
Use Weaviate Cloud (managed) for vector search and context retrieval.
Proper schema design and metadata tagging is expected.
LLM Framework:
Prefer LangChain for chaining logic.
Open to using DSPy where it simplifies prompt optimization or retrieval quality.
Deliverables:
AWS-deployable RAG system, with:
Stateless ingestion pipeline (Lambda or Batch-based).
Stateless query pipeline (Lambda or API Gateway trigger).
LLM orchestration using LangChain (or DSPy).
Terraform templates:
Modular, clean, and parameterized.
Suitable for deploying into any existing AWS environment (accepts VPC/DNS/etc. as input).
MCP Server integration:
Secure, scalable connection to Postgres/Redshift.
Working example queries incorporated into LangChain/DSPy chain.
README and setup guide:
Clear instructions for deployment.
Architecture diagram.
Example usage (upload, chat, response).
Response Requirements:
Only submit a proposal if you have direct experience building RAG pipelines with LLMs.
In your application, only include links or summaries of actual, relevant projects involving:
LLM-based retrieval or document Q&A systems
Integration with Weaviate (or other vector stores)
LangChain / DSPy usage
AWS serverless deployments
Do not submit templated or generic applications. Submissions that don’t follow this will be ignored.ly code commits,
Our expectations: daily updates, daily code commits, usage for freelancer hour tracking system.
System Overview:
- Allow uploading of text into Weaviate Cloud
- Chunk and embed document content using LLM-based embeddings.
- Store and index these embeddings in Weaviate Cloud for fast retrieval.
- Enable user interaction via natural language:
- Retrieve relevant content chunks from Weaviate.
- Combine that with structured queries via MCP Server to Postgres/Redshift.
- Return coherent responses using an LLM (OpenAI, Anthropic, etc.). The system should be model agnostic.
- Be deployed on AWS using a serverless architecture wherever possible.
Technical Requirements:
- Upload → Chunk → Embed → Store in Weaviate.
- Triggered via S3/Lambda/SNS or similar AWS-native flow.
Query Pipeline:
- Accept user prompts.
- Use LangChain (or DSPy) to:
- Query Weaviate for context.
- Query Postgres/Redshift via MCP Server for live data.
- Generate responses via LLM.
Infrastructure as Code:
All infrastructure must be provisioned via Terraform.
Inputs should be parameterised for:
VPC
Route53 DNS
S3 bucket names
Database connection details
Use AWS-native services as much as possible:
Lambda, SNS, SQS, DynamoDB, S3, API Gateway, CloudWatch, IAM, Batch if needed.
Vector Store:
Use Weaviate Cloud (managed) for vector search and context retrieval.
Proper schema design and metadata tagging is expected.
LLM Framework:
Prefer LangChain for chaining logic.
Open to using DSPy where it simplifies prompt optimization or retrieval quality.
Deliverables:
AWS-deployable RAG system, with:
Stateless ingestion pipeline (Lambda or Batch-based).
Stateless query pipeline (Lambda or API Gateway trigger).
LLM orchestration using LangChain (or DSPy).
Terraform templates:
Modular, clean, and parameterized.
Suitable for deploying into any existing AWS environment (accepts VPC/DNS/etc. as input).
MCP Server integration:
Secure, scalable connection to Postgres/Redshift.
Working example queries incorporated into LangChain/DSPy chain.
README and setup guide:
Clear instructions for deployment.
Architecture diagram.
Example usage (upload, chat, response).
Response Requirements:
Only submit a proposal if you have direct experience building RAG pipelines with LLMs.
In your application, only include links or summaries of actual, relevant projects involving:
LLM-based retrieval or document Q&A systems
Integration with Weaviate (or other vector stores)
LangChain / DSPy usage
AWS serverless deployments
Do not submit templated or generic applications. Submissions that don’t follow this will be ignored.ly code commits,
Our expectations: daily updates, daily code commits, usage for freelancer hour tracking system.