AI Infrastructure & Cluster Manager
RemoteLisbon Metropolitan AreaseniorFull-time
- Posted
- today
- Source
- LinkedIn (remote, Europe)
Skills
Stakeholder ManagementProject ManagementCybersecurityKubernetesLeadershipTerraformAnsibleDockerLinuxMachine LearningAI
Description
We are seeking an experienced AI Infrastructure & Cluster Management Leader (Manager/Senior Manager level) to drive high-impact transformation projects across cloud-native infrastructure, cluster management, and AI workload orchestration.
In this role, you will lead multidisciplinary engineering teams, architect large-scale distributed platforms, and act as a strategic technical advisor for global enterprise clients. You will play a pivotal role in optimizing high-performance computing resources (GPUs, memory, storage) and designing resilient, scalable systems for cutting-edge AI environments.
Key Responsibilities
- To lead large-scale transformation programs in infrastructure platforms, cluster management, and workload orchestration.
- To define architectures, standards, and best practices for high-scale platforms, specifically specifying and allocating computing resources (GPUs, memory, and high-speed storage).
- To manage and direct multidisciplinary engineering teams across complex, multi-project environments.
- To partner with senior client stakeholders to define operational models, capacity strategies, performance tuning, and scalability roadmaps.
- To drive initiatives around cluster management, scheduling, virtualization, containerization, and Infrastructure as Code (IaC).
- To evaluate and integrate emerging technologies in AI infrastructure to keep solutions at the forefront of industry standards.
What We Are Looking For
- Education: Bachelor’s or Master’s degree in Computer Engineering, Computer Science, Cybersecurity, or a related technical field.
- Experience: Proven track record in infrastructure platforms, AI infrastructure, or large-scale distributed systems, alongside demonstrated team leadership and project management experience.
- Technical Mastery: Deep hands-on knowledge of Kubernetes, containerization, and cloud-native ecosystems.
- Core Systems: Strong background in Linux systems, virtualization, and infrastructure automation.
- Value-Add Tech Stack:
Container & Cluster Management: OpenShift, Rancher, VMware Tanzu, Docker.
Infrastructure as Code (IaC): Terraform, Ansible, or equivalent tools.
Experience with AI/ML platforms and hardware scheduling (GPUs/HPC) is highly valued.
- Languages: Business fluency in English and Portuguese (written and spoken).
- Strong analytical problem-solving, critical thinking, strategic planning, and stakeholder management skills.
- Willingness to travel as project needs require.
What We Offer
- Personalised career progression plan tailored to your technical and functional goals.
- Direct exposure to large-scale, high-visibility national and international technology transformation projects.
- Unlimited access to world-class continuous learning platforms and sponsored certifications in Cloud and emerging technologies.
- Collaborative, inclusive, and innovation-driven work culture.
JobMatch aggregates public listings. Always apply through the original posting.