JobMatch
← Back to jobs

AI Infrastructure & Cluster Manager

Clan

RemoteLisbon Metropolitan AreaseniorFull-time
Posted
today
Source
LinkedIn (remote, Europe)

Skills

Stakeholder ManagementProject ManagementCybersecurityKubernetesLeadershipTerraformAnsibleDockerLinuxMachine LearningAI

Description

We are seeking an experienced AI Infrastructure & Cluster Management Leader (Manager/Senior Manager level) to drive high-impact transformation projects across cloud-native infrastructure, cluster management, and AI workload orchestration. In this role, you will lead multidisciplinary engineering teams, architect large-scale distributed platforms, and act as a strategic technical advisor for global enterprise clients. You will play a pivotal role in optimizing high-performance computing resources (GPUs, memory, storage) and designing resilient, scalable systems for cutting-edge AI environments. Key Responsibilities - To lead large-scale transformation programs in infrastructure platforms, cluster management, and workload orchestration. - To define architectures, standards, and best practices for high-scale platforms, specifically specifying and allocating computing resources (GPUs, memory, and high-speed storage). - To manage and direct multidisciplinary engineering teams across complex, multi-project environments. - To partner with senior client stakeholders to define operational models, capacity strategies, performance tuning, and scalability roadmaps. - To drive initiatives around cluster management, scheduling, virtualization, containerization, and Infrastructure as Code (IaC). - To evaluate and integrate emerging technologies in AI infrastructure to keep solutions at the forefront of industry standards. What We Are Looking For - Education: Bachelor’s or Master’s degree in Computer Engineering, Computer Science, Cybersecurity, or a related technical field. - Experience: Proven track record in infrastructure platforms, AI infrastructure, or large-scale distributed systems, alongside demonstrated team leadership and project management experience. - Technical Mastery: Deep hands-on knowledge of Kubernetes, containerization, and cloud-native ecosystems. - Core Systems: Strong background in Linux systems, virtualization, and infrastructure automation. - Value-Add Tech Stack: Container & Cluster Management: OpenShift, Rancher, VMware Tanzu, Docker. Infrastructure as Code (IaC): Terraform, Ansible, or equivalent tools. Experience with AI/ML platforms and hardware scheduling (GPUs/HPC) is highly valued. - Languages: Business fluency in English and Portuguese (written and spoken). - Strong analytical problem-solving, critical thinking, strategic planning, and stakeholder management skills. - Willingness to travel as project needs require. What We Offer - Personalised career progression plan tailored to your technical and functional goals. - Direct exposure to large-scale, high-visibility national and international technology transformation projects. - Unlimited access to world-class continuous learning platforms and sponsored certifications in Cloud and emerging technologies. - Collaborative, inclusive, and innovation-driven work culture.

JobMatch aggregates public listings. Always apply through the original posting.