Fixing AWS production-stack
Budget: $250 – $750 USD
Introduction
We have a production stack on AWS that functions perfectly in our Native app, available on both the Play Store and App Store. However, our development environment cannot access it correctly. We need to identify the cause and implement a solution.
About the Stack
Our backend is hosted on AWS, leveraging a serverless architecture with the following services:
S3: Image storage
DynamoDB: Database
Cognito: User authentication
Amazon SNS: SMS verifications
IAM: Authorization
API Gateway & Lambda: API endpoints
Amplify Framework: API integration
CloudWatch: Logging and monitoring
The backend code is written in Node.js, and we use CloudFormation, YAML, and Serverless to manage the infrastructure.
Objective
Since this environment is live and serves our customers, all changes must be meticulously documented and carefully executed to ensure uninterrupted service. Under no circumstances should customer-facing functionality be disrupted.
During this investigation, it might become apparent that copying the stack (while retaining the Cognito database) or rebuilding it entirely could be a better solution. Therefore, the project is divided into three milestones to allow flexibility in the approach and scope of work.
Milestones
Milestone 1: Initial Investigation (2 Days)
Conduct an in-depth investigation of the current issue.
Provide detailed reports at the end of Day 1 and Day 2.
Milestone 2: Extended Investigation (2 Additional Days)
Continue the investigation based on findings from Milestone 1.
Propose refined solutions or confirm the root cause.
Provide detailed reports at the end of Day 3 and Day 4.
Milestone 3: Final Investigation (1 Day)
Conduct targeted investigations to address any remaining uncertainties.
Deliver a final report summarizing findings and actionable solutions.
Reporting Requirements
The reports should be concise and easy to understand, even for non-developers. Each daily report must include the following:
Areas Investigated: What parts of the stack were analyzed.
Confirmed Working: Components verified to be functioning correctly.
Hypotheses: Any ideas or observations about the potential root cause.
Attempted Fixes: Details of any changes made to resolve issues.
Live Testing Requests: If you wish to test something live, provide a clear explanation and wait for written approval before proceeding.
Next Steps: Planned activities for the following day.
Key Notes:
The reports should be delivered by the end of each working day.
All proposed live changes must be explicitly approved before implementation.
Documentation must prioritize clarity and accuracy to ensure ease of understanding.
We have a production stack on AWS that functions perfectly in our Native app, available on both the Play Store and App Store. However, our development environment cannot access it correctly. We need to identify the cause and implement a solution.
About the Stack
Our backend is hosted on AWS, leveraging a serverless architecture with the following services:
S3: Image storage
DynamoDB: Database
Cognito: User authentication
Amazon SNS: SMS verifications
IAM: Authorization
API Gateway & Lambda: API endpoints
Amplify Framework: API integration
CloudWatch: Logging and monitoring
The backend code is written in Node.js, and we use CloudFormation, YAML, and Serverless to manage the infrastructure.
Objective
Since this environment is live and serves our customers, all changes must be meticulously documented and carefully executed to ensure uninterrupted service. Under no circumstances should customer-facing functionality be disrupted.
During this investigation, it might become apparent that copying the stack (while retaining the Cognito database) or rebuilding it entirely could be a better solution. Therefore, the project is divided into three milestones to allow flexibility in the approach and scope of work.
Milestones
Milestone 1: Initial Investigation (2 Days)
Conduct an in-depth investigation of the current issue.
Provide detailed reports at the end of Day 1 and Day 2.
Milestone 2: Extended Investigation (2 Additional Days)
Continue the investigation based on findings from Milestone 1.
Propose refined solutions or confirm the root cause.
Provide detailed reports at the end of Day 3 and Day 4.
Milestone 3: Final Investigation (1 Day)
Conduct targeted investigations to address any remaining uncertainties.
Deliver a final report summarizing findings and actionable solutions.
Reporting Requirements
The reports should be concise and easy to understand, even for non-developers. Each daily report must include the following:
Areas Investigated: What parts of the stack were analyzed.
Confirmed Working: Components verified to be functioning correctly.
Hypotheses: Any ideas or observations about the potential root cause.
Attempted Fixes: Details of any changes made to resolve issues.
Live Testing Requests: If you wish to test something live, provide a clear explanation and wait for written approval before proceeding.
Next Steps: Planned activities for the following day.
Key Notes:
The reports should be delivered by the end of each working day.
All proposed live changes must be explicitly approved before implementation.
Documentation must prioritize clarity and accuracy to ensure ease of understanding.