Site Reliability Engineer
RemoteWrocław, Dolnośląskie, PolandseniorFull-time
- Posted
- today
- Source
- LinkedIn (remote, Europe)
- Field
- Engineering
Skills
CommunicationGCPKubernetesLeadershipTerraformSecurityAnsibleDockerDevOpsAzureCI/CDAgileScrumAWS
Description
We are seeking a highly skilled Site Reliability Engineer to drive the design, development, testing, and operational excellence of enterprise technology solutions.
The ideal candidate will possess strong software engineering fundamentals, deep Software Development Lifecycle (SDLC) expertise, extensive test automation experience, and a proven track record in building and operating highly reliable, scalable, and resilient platforms.
Key Responsibilities:
- Lead the design, development, and deployment of high-quality software solutions following modern engineering best practices.
- Drive end-to-end SDLC processes, including requirements analysis, design, development, testing, deployment, and production support.
- Develop and implement automated testing frameworks and strategies to improve software quality, reliability, and release velocity.
- Design and maintain CI/CD pipelines to support continuous integration, automated testing, and continuous delivery.
- Apply SRE principles to improve system reliability, scalability, availability, and operational efficiency.
- Establish and monitor SLIs, SLOs, and error budgets to ensure service performance and reliability objectives are met.
- Collaborate with development, infrastructure, security, and operations teams to streamline software delivery and production support.
- Lead root cause analysis, incident management, and post-incident reviews to drive continuous improvement.
- Champion observability practices through monitoring, logging, alerting, and performance analysis.
- Mentor engineering teams on software engineering best practices, test automation, and reliability engineering.
Required Qualifications:
- 5+ years of experience in software engineering and enterprise application development.
- Strong expertise in Software Development Lifecycle (SDLC) methodologies and engineering best practices.
- Hands-on experience designing and implementing test automation frameworks and strategies.
- Proven experience with Site Reliability Engineering (SRE), production operations, and platform reliability.
- Experience with CI/CD pipelines, DevOps practices, and release automation.
- Strong understanding of system performance, scalability, availability, and resiliency principles.
- Experience with monitoring, observability, incident management, and operational excellence.
- Excellent analytical, troubleshooting, and problem-solving skills.
- Strong communication and leadership abilities.
Preferred Skills
- Cloud platforms (AWS, Azure, or Google Cloud Platform).
- Containerization and orchestration technologies such as Docker and Kubernetes.
- Infrastructure as Code (Terraform, Ansible, or similar).
- Monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, or Dynatrace.
- Agile, Scrum, and DevSecOps experience.
Technical Skills
- Software Engineering
- SDLC Management
- Test Automation
- Site Reliability Engineering (SRE)
- CI/CD and DevOps
- System Reliability & Performance Engineering
- Observability & Monitoring
- Incident & Problem Management
- Cloud and Container Technologies
- Automation & Infrastructure as Code (IaC)
For more information – please apply for this job.
Cavendish (Recruitment) Professionals Ltd are proud to be an equal opportunity employer and we believe that inclusivity begins with the candidate experience. All qualified applicants will receive consideration for employment regardless of, gender, race, age, sexual orientation, religion, or belief.
JobMatch aggregates public listings. Always apply through the original posting.