
Closed
Posted
Paid on delivery
I am handling the day-to-day reliability of a growing production stack and need a seasoned Site Reliability Engineer to coach me through two key domains: monitoring/alerting and performance optimization. I already manage the basic upkeep, but I want to elevate my skills so that incidents are caught sooner and services run leaner. Here is what I’m hoping for: • Regular screen-sharing or video sessions where we review my existing dashboards and alert rules, refine the signal-to-noise ratio, and discuss industry best practices. • Deep-dive walkthroughs on performance tuning—profiling services, interpreting latency metrics, and translating findings into configuration or code changes. • Actionable take-home steps after each session so I can apply what we discuss, then bring results back for feedback. I work primarily in a Linux/containerized environment with common open-source tooling, but I’m open to adopting whatever stack you recommend—Prometheus, Datadog, Grafana, or other fit-for-purpose solutions. To make sure we are a match, please tell me about: • Similar mentoring or advisory roles you’ve done. • Your approach to setting up meaningful alerts without alert fatigue. • A brief example of how you diagnosed and fixed a tricky performance bottleneck. We can start with a short engagement to align on goals; if the collaboration clicks, I’m happy to extend for ongoing guidance.
Project ID: 40599814
8 proposals
Remote project
Active 4 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
8 freelancers are bidding on average ₹969 INR for this job

As someone who has dedicated over a decade to mastering Linux and system administration, I have witnessed first-hand the transformative results that structured monitoring and performance optimization can produce. My experience aligns perfectly with your project's needs, having fine-tuned and managed numerous Linux-driven stacks and leveraged prominent open-source tools such as Grafana and Prometheus to deliver efficient, data-driven solutions. One of my key strengths is my ability to translate complex technical information into practical, actionable steps, a quality that would be invaluable in our potential collaboration. I'm deeply committed to fostering symbiotic relationships built on consistent dialogue and feedback sharing, precisely the environment you have cultivated for this project. In terms of your concerns regarding alert fatigue and performance bottlenecks, let me briefly illustrate how I addressed these challenges in the past. By thoroughly profiling services, evaluating latency metrics with precision, and implementing necessary changes, I successfully rectified a persistent performance issue for a client in the past. In doing so, I reduced extraneous alerts without compromising on critical warnings - an approach that marries well with your objectives.
₹600 INR in 1 day
2.4
2.4

Hello, I have hands-on experience with Linux, Docker, monitoring, and DevOps/SRE practices. I can help improve your monitoring dashboards, optimize alert rules, reduce alert fatigue, and enhance system performance. I work with Linux-based and containerized environments and have experience troubleshooting performance issues. My approach is to review your current setup, identify gaps, and provide practical recommendations using tools like Prometheus and Grafana. I am available for screen-sharing sessions, explain concepts clearly, and provide actionable steps after each session so you can apply them confidently. I am committed to helping you strengthen your SRE skills and improve production reliability. Looking forward to working with you. Best regards, Dhiraj Kumar
₹650 INR in 1 day
0.0
0.0

SRE mentor: refine alerts, optimize performance, actionable steps for reliable systems. Monitoring & performance expert—cut noise, boost reliability, mentor with proven results. Seasoned SRE coach: dashboards, alert tuning, latency fixes, clear guidance every session Looking forward to helping you strengthen your reliability practices. Best regards, Irfan ---
₹800 INR in 7 days
0.0
0.0

Hi, I’m a DevOps/SRE professional with 3+ years of hands-on experience in Linux, containers, AWS, Kubernetes, Prometheus, Grafana, and monitoring/alerting systems. I can mentor you through practical, real-world sessions focused on improving alert quality, reducing alert fatigue, troubleshooting incidents, and optimizing application and infrastructure performance. My approach is to focus on actionable signals tied to service health, SLOs, latency, errors, saturation, and user impact rather than simply increasing the number of alerts. We can review your existing dashboards and alert rules together, deep-dive into performance bottlenecks, and define clear take-home actions after every session. I’ll also help you understand the reasoning behind each recommendation so you can independently apply these practices going forward. I’d be happy to start with a short engagement and continue long-term if we’re a good fit.
₹650 INR in 7 days
0.0
0.0

Hello, You need a seasoned SRE to sharpen your monitoring and performance practices, and I can help with that directly. I'm Carlos Porter, a Senior Site Reliability Engineer with 16 years across Linux and cloud environments. I've built self-hosted Elastic Stack deployments for centralized log ingestion, configured dashboards and alert rules, and diagnosed production bottlenecks using tools like sar, tcpdump, and sosreport. I also led Linux migrations and coordinated incident response with platform teams, which gave me a practical eye for signal-to-noise ratio in alerting. My approach to meaningful alerts starts with understanding what matters to your services: latency, error rates, and saturation. I focus on actionable thresholds, not static numbers, and layer in runbooks so every alert has a clear response path. For performance bottlenecks, I recently traced a slow database query through log analysis and query profiling, then worked with the team to adjust indexing and connection pooling, cutting query latency by over 60%. I can review your current dashboards and alert rules in a screen-sharing session, walk through tuning techniques, and leave you with concrete steps to apply before our next call. If that sounds useful, let's set up a short call to align on goals. Carlos Porter Senior Site Reliability Engineer
₹2,000 INR in 7 days
0.0
0.0

Experienced AIML and MLOPS engineer, I am aws and microsoft azure certified engineer work on various client and infrastructure and also resolved very complex production issue and launch production infra in timely manner. i am work various it tools like linux windows git cicd terraform cloudflare ansible prometheus grafana docker k8s terraform trivy sonarqube argocd owasp observability open-telemetry and so on and still working on learning on cloud ai certification.
₹600 INR in 7 days
0.1
0.1

Hi there, As an SRE & Platform Engineering Leader, I can help you transform your monitoring from noisy alerts into a lean, proactive observability system and walk you through deep performance tuning. 1. Mentoring Experience: I regularly coach engineers 1-on-1 on observability, Linux performance, and containerized stacks—conducting live dashboard audits, pair-profiling, and establishing clear post-incident playbooks. 2. Eliminating Alert Fatigue: My approach relies on SRE Golden Signals (Latency, Traffic, Errors, Saturation). We alert on customer-impacting symptoms (like p95/p99 latency spikes), not benign causes. If an alert doesn't require immediate human action, it becomes a dashboard metric, not a page. 3. Fixing Performance Bottlenecks: I recently diagnosed a p99 latency spike (3000ms+) in a containerized stack where CPU/RAM appeared normal. Using Prometheus/Grafana metrics and perf/tracing tools, I isolated connection pool exhaustion and thread blocking. After tuning socket parameters and DB queries, p99 dropped to <150ms. Our Plan: Live Sessions: Interactive screen-sharing to audit dashboards, tune alert noise, and profile performance. Take-Home Steps: Clear action items after every session to apply changes, followed by feedback. I'm ready for a short initial session to align on your stack (Prometheus/Grafana/NewRelic) and goals. Best, Lokesh Gondane
₹1,250 INR in 1 day
0.0
0.0

Kolkata, India
Member since Oct 15, 2024
₹400-750 INR / hour
₹600-1500 INR
₹600-700 INR
₹600-700 INR
₹600-700 INR
₹600-700 INR
₹600-700 INR
₹600-700 INR
₹600-700 INR
₹600-700 INR