Lead Site Reliability Engineer
RemoteUnited StatesseniorFull Time
- Posted
- today
- Source
- Himalayas
- Field
- Engineering
Skills
KubernetesTerraformPythonExcelLinuxAWSGo
Description
Category: Technology
Location:
-
Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
-
Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
-
Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.
-
Automate incident detection and response using automated runbooks or predefined workflows.
-
Write software as needed to support reliability or efficiency needs.
-
Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
-
Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.
-
Collaborate with development and operations teams on building reliable, scalable, and high-performance services.
-
Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.
-
Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.
-
Get involved in chaos engineering initiatives.
-
Participate on our on-call rotation.
-
Drive advanced alerting and anomaly detection applied to metrics
Requirements
-
3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.
-
Deep understanding of Linux systems, networking, and systems administration.
-
Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
-
Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
-
Strong skills in at least one programming language (Python, Go) to write production level code.
-
Strong skills in shell scripting using bash or similar. • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
-
Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.
Benefits
-
Competitive Salary Package:Receive a pay package that matches your skills and experience.
-
Vacation and Sick Leave credits:Enjoy vacation and sick leave credits to maintain work-life balance.
-
Health Coverage:Get medical, dental, and vision insurance for you and your dependents.
-
Government-Mandated Benefits:Full coverage of all statutory benefits like SSS, PhilHealth, and Pag-IBIG.
-
Learning Opportunities:Access training, certifications, and mentorship to grow your career.
-
Team Engagement:Join team-building activities and wellness programs.
-
Modern Tools:Use the latest technology to excel in your role.
-
Career Growth:Clear paths for promotion and professional development.
-
Inclusive Culture:Be part of a diverse, supportive, and collaborative global team.
-
Referral Rewards:Earn bonuses for bringing great talent to the team.
Details
Originally posted on Himalayas
JobMatch aggregates public listings. Always apply through the original posting.