Scalable Backend Architecture on AWS

Job ID: 40421964

Budget: $50 – $0 USD

I need a seasoned backend engineer to partner with me in shaping, building, and shipping a production-grade API service that can grow effortlessly while staying rock-solid online. AWS is my chosen cloud, so every architectural decision—from VPC layout to IAM policies—must align with best-practice patterns in that ecosystem.

Scalability and reliability are the two absolutely critical pillars. The system should sustain heavy, spiky traffic without degradation and recover gracefully from failure scenarios. You’ll guide the high-level design (microservices vs. modular monolith, service-to-service communication, fault isolation) and drill down into concrete AWS components such as ALB, ECS / EKS, Lambda, DynamoDB or other NoSQL options, SQS, EventBridge, and appropriate caching layers.

A NoSQL data store will back the core domain. I’m expecting you to propose a clean, future-proof schema (or data model) that supports rapid queries at scale, detail read/write patterns, and outline backup, replication, and disaster-recovery strategies. Observability must be woven in from day one: metrics, traces, and logs surfaced through CloudWatch, OpenTelemetry, or comparable tooling so bottlenecks never hide.

Security, CI/CD, and infrastructure-as-code complete the picture. Please include advice and sample templates (Terraform, CDK, or CloudFormation), plus a deployment pipeline that runs automated tests, vulnerability scans, and zero-downtime rollouts.

Deliverables
• Architecture document and diagram (PDF/Draw.io)
• NoSQL data model with justification and sample migration script
• IaC templates and CI/CD pipeline definition ready for use
• Monitoring & alerting guide with example dashboards
• Short hand-off session or recorded walkthrough explaining key decisions

Acceptance criteria
• Design supports horizontal scaling to 10k+ requests per second with P95 latency below 100 ms under load tests
• All components deploy via a single automated pipeline, passing unit & integration tests
• Critical paths have dashboards and alerts proving MTTR < 15 minutes

If this aligns with your expertise, I’m ready to dive in and iterate quickly.