Senior Platform Engineer - Azure Data & AI Platform
RemoteLondon, England, United KingdomseniorFull-time
- Posted
- today
- Source
- LinkedIn (remote, Europe)
- Field
- Engineering
Skills
LeadershipTerraformAnalyticsSecurityPythonDevOpsAzureCI/CDSparkUnitySQLGitMachine LearningAI
Description
Senior Platform Engineer - Azure Data & AI Platform
The Role
We are looking for senior platform engineers to build and operate our Azure-based data and AI platform. This is hands-on infrastructure and platform work - you'll be writing Terraform, designing network architecture, implementing MLOps (AI model) pipelines, and establishing DevSecOps patterns that engineering teams can use.
This isn't a "thought leadership" or "strategy" role. You'll be in the code, in the CLI, and in the infrastructure daily.
What You'll Actually Do
Azure Platform Engineering (40%)
-
Design and implement Azure landing zones, management groups, and subscription architecture
-
Build and maintain hub-spoke network topologies with proper segmentation and security controls
-
Implement Azure Policy, RBAC, and governance frameworks that balance security with developer productivity
-
Manage identity and access using Entra ID, service principals, managed identities
-
Establish monitoring, logging, and alerting with Azure Monitor, Log Analytics, and Application Insights
-
Cost management and FinOps practices - keeping cloud spend under control without hamstringing teams
Databricks & Data Platform (30%)
-
Deploy and configure Azure Databricks workspaces with Unity Catalog for data governance
-
Assisting both data and AI teams with pipeline development
-
Establish Databricks best practices: cluster policies, job scheduling, notebook standards, workspace organization
-
Integrate Databricks with ADLS Gen2, Azure SQL, Synapse, and other data services
-
Set up and maintain CI/CD for Databricks notebooks, jobs, and infrastructure
-
Work with data engineers on performance optimization, cost control, and platform capabilities
MLOps & AI Platform (20%)
-
Build ML model deployment pipelines using Azure ML, Databricks MLflow, or both
-
Implement model versioning, experiment tracking, and model registry patterns
-
Establish inference endpoints (batch and real-time) with proper monitoring and governance
-
Create reusable ML pipeline templates and infrastructure-as-code modules
-
Integrate AI services (Azure OpenAI, Cognitive Services) into platform offerings
-
Implement responsible AI guardrails: model monitoring, bias detection, explainability
DevSecOps & Platform Enablement (10%)
-
Creation of tools, features and dashboards for the developer platform
-
Build CI/CD pipelines in Azure DevOps or GitHub Actions with proper security scanning
-
Implement shift-left security: SAST/DAST, dependency scanning, infrastructure scanning, secrets management
-
Establish infrastructure-as-code standards with Terraform (or Bicep), including modules and policy enforcement
-
Create self-service tooling and automation for common platform tasks
-
Write documentation that engineers will read and use
-
Participate in on-call rotation for platform incidents
What We Need From You
Required
-
5+ years platform/infrastructure engineering - you've built production platforms, not just prototypes
-
Deep Azure knowledge - networking, IAM, storage, compute, PaaS services. You know the difference between service endpoints and private endpoints and when to use each
-
Databricks experience - you've deployed workspaces, configured Unity Catalog, optimized Spark jobs, managed costs
-
Infrastructure as Code - Terraform (preferred) or Bicep. You write modules, understand state management, know how to structure large IaC projects
-
CI/CD pipelines - Azure DevOps or GitHub Actions. You've built multi-stage pipelines with gates, approvals, and security scanning
-
Containerization experience - you know best practices when building and working with both application-based containers and containers holding ML/AI models
-
Security-first mindset - you understand defense in depth, least privilege, network segmentation, and don't treat security as an afterthought
-
MLOps fundamentals - model training vs inference, experiment tracking, model versioning, deployment patterns
-
Python and/or PowerShell - for automation, tooling, and platform utilities
-
Observability stack beyond basic metrics (distributed tracing, log aggregation patterns)
Strongly Preferred
-
Experience with Azure landing zones and CAF (Cloud Adoption Framework)
-
Microsoft Purview for data governance and cataloging
-
Experience with Delta Lake, Spark optimization, data quality frameworks
-
Azure networking certifications or equivalent deep knowledge
-
Container orchestration (AKS)
-
API design and management (API Management, App Gateway, Front Door)
What Actually Matters
-
Pragmatism over purity - you choose the right tool for the job, not the coolest one
-
Documentation discipline - you document as you build because you know future-you will thank today-you
-
Automation mindset - if you do it twice, you automate it
-
Everything-as-code - if it's not in git, it doesn't exist to you
-
Collaboration skills - you can translate between data scientists, engineers, and business stakeholders
-
Ownership mentality - you build it, you run it, you support it
-
Intellectual honesty - you say "I don't know" when you don't, and then you figure it out
What We Offer
Actual flexibility: Remote-first with occasional in-office travel for workshops/planning. We care about outcomes, not seat time.
Real learning budget: for conferences, training, certifications. We expect you to use it.
Tooling: You'll get the equipment and licenses you need to do the job properly.
Grown-up engineering culture:
-
PRs are required, branching is mandatory, tests matter
-
Blameless post-mortems when things break
-
Technical decisions driven by evidence and context, not politics or trends
-
We write RFCs for significant changes
Add the usual stuff here: How to apply, interview process.
JobMatch aggregates public listings. Always apply through the original posting.