SRE Guidance: Monitoring and Performance

Job ID: 40599814

Budget: ₹600 – ₹700 INR

I am handling the day-to-day reliability of a growing production stack and need a seasoned Site Reliability Engineer to coach me through two key domains: monitoring/alerting and performance optimization. I already manage the basic upkeep, but I want to elevate my skills so that incidents are caught sooner and services run leaner.

Here is what I’m hoping for:

• Regular screen-sharing or video sessions where we review my existing dashboards and alert rules, refine the signal-to-noise ratio, and discuss industry best practices.
• Deep-dive walkthroughs on performance tuning—profiling services, interpreting latency metrics, and translating findings into configuration or code changes.
• Actionable take-home steps after each session so I can apply what we discuss, then bring results back for feedback.

I work primarily in a Linux/containerized environment with common open-source tooling, but I’m open to adopting whatever stack you recommend—Prometheus, Datadog, Grafana, or other fit-for-purpose solutions.

To make sure we are a match, please tell me about:
• Similar mentoring or advisory roles you’ve done.
• Your approach to setting up meaningful alerts without alert fatigue.
• A brief example of how you diagnosed and fixed a tricky performance bottleneck.

We can start with a short engagement to align on goals; if the collaboration clicks, I’m happy to extend for ongoing guidance.