JobMatch
← Back to jobs

Lead Site Reliability Engineer

iScale Solutions

RemoteUnited StatesseniorFull Time
Posted
today
Source
Himalayas
Field
Engineering

Skills

KubernetesTerraformPythonExcelLinuxAWSGo

Description

Category: Technology Location: - Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts. - Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services. - Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems. - Automate incident detection and response using automated runbooks or predefined workflows. - Write software as needed to support reliability or efficiency needs. - Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging. - Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow. - Collaborate with development and operations teams on building reliable, scalable, and high-performance services. - Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc. - Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar. - Get involved in chaos engineering initiatives. - Participate on our on-call rotation. - Drive advanced alerting and anomaly detection applied to metrics Requirements - 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments. - Deep understanding of Linux systems, networking, and systems administration. - Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools. - Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki. - Strong skills in at least one programming language (Python, Go) to write production level code. - Strong skills in shell scripting using bash or similar. • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration. - Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc. Benefits - Competitive Salary Package:Receive a pay package that matches your skills and experience. - Vacation and Sick Leave credits:Enjoy vacation and sick leave credits to maintain work-life balance. - Health Coverage:Get medical, dental, and vision insurance for you and your dependents. - Government-Mandated Benefits:Full coverage of all statutory benefits like SSS, PhilHealth, and Pag-IBIG. - Learning Opportunities:Access training, certifications, and mentorship to grow your career. - Team Engagement:Join team-building activities and wellness programs. - Modern Tools:Use the latest technology to excel in your role. - Career Growth:Clear paths for promotion and professional development. - Inclusive Culture:Be part of a diverse, supportive, and collaborative global team. - Referral Rewards:Earn bonuses for bringing great talent to the team. Details Originally posted on Himalayas

JobMatch aggregates public listings. Always apply through the original posting.