Summary
Highly accomplished Monitoring Engineer with 6 years of experience specializing in building robust, scalable observability platforms and driving site reliability. Proven expertise in architecting comprehensive monitoring solutions, automating incident response, and optimizing system performance across diverse cloud and on-premise environments. Adept at leveraging tools like Prometheus, Grafana, ELK Stack, and cloud-native services to enhance system uptime and streamline operational workflows.
Experience
Senior Monitoring EngineerZenithTech
- Architected and deployed a centralized observability platform using Prometheus, Grafana, and Alertmanager, reducing incident detection time by 30%.
- Developed automated remediation scripts in Python for critical service failures, decreasing Mean Time To Resolution (MTTR) by 25%.
- Led the migration of legacy monitoring systems to AWS CloudWatch and Splunk, improving data ingestion rates by 40% and reducing infrastructure costs by 15%.
- Implemented advanced anomaly detection algorithms to proactively identify performance degradation, preventing over 10 critical outages annually.
Monitoring EngineerQuantify Solutions
- Designed and implemented monitoring dashboards for 20+ critical microservices using Grafana, enhancing visibility for engineering and operations teams.
- Configured and managed alerting rules in PagerDuty and OpsGenie, reducing alert fatigue by 20% through smart aggregation and suppression strategies.
- Automated health checks and performance metric collection for containerized applications using Docker and Kubernetes, ensuring 99.9% uptime for key services.
- Collaborated with development teams to integrate application-level metrics (APM) into central monitoring, improving root cause analysis efficiency by 15%.
Projects
Cloud Cost Anomaly Detector
- Developed a Python-based serverless application (AWS Lambda) to monitor cloud spending patterns and detect anomalous spikes.
- Integrated with AWS Cost Explorer API and Slack notifications, reducing unexpected cost overruns by automatically alerting stakeholders.
Kubernetes Health Dashboard
- Built a custom web dashboard using React and a Go backend to visualize Kubernetes cluster health metrics from Prometheus.
- Provided real-time insights into pod status, resource utilization, and error rates, improving developer's ability to debug issues by 20%.
Log Aggregation & Search Tool
- Created a lightweight log aggregation system using Fluentd, Elasticsearch, and Kibana for personal projects.
- Enabled efficient centralized logging and full-text search capabilities, significantly speeding up debugging processes.
Education
Georgia Institute of TechnologyBachelor of Science in Computer Science
- Graduated with High Honors (Magna Cum Laude) for academic excellence.
- Maintained a 3.8 GPA, specializing in Distributed Systems and Cloud Computing.
- Awarded Dean's List for 6 consecutive semesters based on superior academic performance.
Skills
Monitoring & Observability
PrometheusGrafanaAlertmanagerELK Stack (Elasticsearch, Logstash, Kibana)SplunkCloudWatchDatadogPagerDuty
Cloud Platforms
AWS (EC2, S3, RDS, Lambda, EKS)AzureGCPTerraform
Scripting & Automation
PythonBashGoAnsibleChefPuppet
Containerization & Orchestration
DockerKubernetesHelmECS
CI/CD & DevOps
JenkinsGitLab CI/CDGitHub ActionsGitJiraConfluence
