Kafka Optimization, Configuration, Performance Tuning & Capacity Planning
Budget: $15 – $25 USD
Objective
We aim to optimize the architecture and fine-tune the configuration of our message middleware (Kafka/MQ). By combining horizontal scaling with high-availability design for service clusters, we want to proactively detect and eliminate potential failures, ensuring system stability under the following workloads:
- Single-node environment: Sustain 10,000 TPS / 100,000 QPS
- Cluster mode: Sustain 500,000 TPS / 2,000,000 QPS
Deliverables
1. Architecture Design Document
- Overall system topology and cluster deployment plan
- High availability and elastic scaling design
- Key architectural details such as partitioning, replication, and load balancing strategies
2. Deployment Document
- Installation and configuration steps for middleware and dependencies
- Core parameter tuning (partitions, replicas, batch sending, flush policies, memory/network parameters, etc.)
- Guidelines for cluster expansion, upgrades, and version migration
3. User Integration Guide
- Producer/consumer access standards
- API usage examples and best practices
- Rate limiting, retry mechanisms, and error handling instructions
4. Operations & Maintenance Document
- Core monitoring metrics (TPS/QPS, message latency, backlog)
- Alerting rules and incident handling process
- Capacity planning methods
- Troubleshooting, recovery, and emergency response plans
We aim to optimize the architecture and fine-tune the configuration of our message middleware (Kafka/MQ). By combining horizontal scaling with high-availability design for service clusters, we want to proactively detect and eliminate potential failures, ensuring system stability under the following workloads:
- Single-node environment: Sustain 10,000 TPS / 100,000 QPS
- Cluster mode: Sustain 500,000 TPS / 2,000,000 QPS
Deliverables
1. Architecture Design Document
- Overall system topology and cluster deployment plan
- High availability and elastic scaling design
- Key architectural details such as partitioning, replication, and load balancing strategies
2. Deployment Document
- Installation and configuration steps for middleware and dependencies
- Core parameter tuning (partitions, replicas, batch sending, flush policies, memory/network parameters, etc.)
- Guidelines for cluster expansion, upgrades, and version migration
3. User Integration Guide
- Producer/consumer access standards
- API usage examples and best practices
- Rate limiting, retry mechanisms, and error handling instructions
4. Operations & Maintenance Document
- Core monitoring metrics (TPS/QPS, message latency, backlog)
- Alerting rules and incident handling process
- Capacity planning methods
- Troubleshooting, recovery, and emergency response plans