Platform Engineer
RemotePortugalseniorFull-time
- Posted
- today
- Source
- LinkedIn (remote, Europe)
- Field
- Engineering
Skills
PostgreSQLKubernetesTerraformSecurityAnsibleDevOpsCI/CDScrumLinuxNode.jsJiraAWS
Description
Requirements
- Minimum of 5 years of experience in DevOps, SRE or Platform Engineering, including at least 3 years of hands-on experience with Infrastructure as Code (IaC) in AWS production environments.
- Strong experience with Terraform (AWS Provider) and/or OpenTofu, including reusable modules, semantic versioning and provider upgrades.
- Hands-on experience with Ansible.
- Production experience with Amazon EKS, including cluster version upgrades, add-on upgrades and node group migrations.
- Solid knowledge of AWS services, particularly EC2, S3 and IAM.
- Strong Kubernetes skills, including Helm, ingress controllers, probes, Horizontal Pod Autoscaler (HPA) and PodDisruptionBudgets (PDBs).
- Experience administering PostgreSQL databases, including backup and restore procedures, high availability and failover testing.
- Experience upgrading self-hosted platforms, particularly HashiCorp Vault, Sonatype Nexus and GitLab running on virtual machines.
- Experience with CI/CD using GitLab, including pipelines, merge requests, branch protection and approval workflows, as well as Jenkins.
- Experience with monitoring and alerting using Prometheus/Grafana or equivalent tools.
- Experience implementing Kubernetes backup and restore procedures using Velero.
- Knowledge of security fundamentals, including IAM/RBAC, least privilege, secrets management with Vault and TLS.
- Strong Linux administration and Bash scripting skills.
- Familiarity with Jira, Confluence and Visual Studio Code.
- Ability to work autonomously, take ownership of technical activities and manage changes in production environments.
- Fluency in Portuguese and English.
Responsabilities:
- Ensure the availability, security, stability and efficiency of an AWS-based enterprise data platform.
- Plan and execute upgrades of self-hosted HashiCorp Vault, Sonatype Nexus and GitLab, ensuring backups, rollback plans and post-upgrade validation are in place.
- Upgrade Amazon EKS clusters and add-ons, and migrate node groups from Amazon Linux 2 to Amazon Linux 2023 without service downtime.
- Migrate the end-of-life ingress-nginx controller to a replacement solution while maintaining uninterrupted access to exposed services.
- Upgrade Terraform/OpenTofu and the AWS Provider, validating that execution plans are non-destructive before applying changes.
- Deliver critical version and dependency updates for components such as Kafka, Jenkins and Uptime Kuma.
- Design and implement PostgreSQL high-availability solutions, including failover testing and recovery validation.
- Strengthen the resilience of core platform components and regularly test Kubernetes backup and restore procedures using Velero, measuring recovery times
- Define and implement availability improvements for critical services, providing evidence that changes do not degrade service performance.
- Improve monitoring and alerting, and implement self-healing mechanisms, including probes, automatic restarts, autoscaling and automated remediation.
- Optimise EKS costs and capacity across pre-production and production environments through rightsizing, scheduled start/stop and tools such as Kubecost.
- Develop and maintain reusable Terraform and Ansible modules, ensuring semantic versioning, tagging, pipeline templates and automated testing.
- Manage infrastructure changes through GitLab merge requests and approval workflows, enforcing branch protection and change control processes.
- Secure secrets using Vault and enforce least privilege and segregation of duties.
- Investigate and remediate security findings identified by Checkmarx, Fortify and other security scanning tools in IaC modules.
- Produce technical designs and ITIL-aligned operational procedures covering upgrades, backup and restore, and observability.
- Maintain technical documentation in Confluence and prepare monthly progress reports.
- Take ownership of ongoing activities, conduct knowledge transfer sessions and collaborate with platform owners and third-party providers.
- Coordinate access requirements, maintenance windows and production change approvals.
- Work within a Scrum team, contributing to two-week sprints and using Jira and Confluence for task management, collaboration and documentation.
JobMatch aggregates public listings. Always apply through the original posting.