Runtime and Performance Engineer
RemoteNetherlandsentryFull-time
- Posted
- today
- Source
- LinkedIn (remote, Europe)
- Field
- Engineering
Skills
Machine LearningJavaScriptKotlinUnreal EngineSwiftLinuxNode.jsRustC++AI
Description
Make our models run faster on Apple platforms, Android, and Windows, in our shared Swift core and its platform bindings.
About Desert Ant Labs
Desert Ant Labs is an on-device AI lab in Europe. We're building efficient frontier intelligence, bottom up. We make the fastest model for each task, built to run on the billions of phones, laptops, and browsers people already own. Every model is lightning fast and private, and has no per-call cost. Little brains in every product.
With our native SDKs for Swift, Kotlin, and JavaScript, you add a model to an app in a few lines of code. Data stays with your customer, and the feature works without a cloud provider. Every role here pushes the limits of what a phone or a laptop can do.
About the role
You speed up our models on Apple platforms, Android, Windows, and the web. You bring deep experience on at least one of those platforms. Maybe you built a rendering engine, a game or map renderer, a browser engine, or a video pipeline with a frame budget. You don't need machine learning experience.
Our SDKs share one core, written in Swift. Each model's pipeline runs as Swift on Apple platforms, Linux, and Windows. On Android, Kotlin calls the Swift core through JNI. In the browser the core runs as WebAssembly, and Node calls it through a C ABI. The core runs each model on Core ML, Core AI, MLX, LiteRT, or ONNX Runtime. We pick the runtime for each model and platform, and we switch when another one runs faster. You find out how fast each model can run on each chip, and make it run that fast.
You work from GPU kernels up to the SDK API. You also design the SDK API, so the default call runs a model the fastest way the device allows.
What you will do
- Profile and speed up inference on Apple platforms, Android, Windows, and the web: memory layout, SIMD, GPU scheduling, and cache behavior.
- Run as much of each model on the Neural Engine, NPU, or GPU as the hardware allows.
- Build the shared Swift core and its bindings to Kotlin, WebAssembly, and native code: memory budgets, routing, and fallbacks.
- Write Metal, WebGPU, or native kernels when a runtime runs a model too slowly. We have none yet, so you write the first.
- Work with researchers on model architectures that run fast on real devices.
You might be a fit if you
- Have deep experience on one or more platforms: Apple platforms in Swift and Metal, Android in Kotlin and the NDK, or Windows in C++ and DirectX.
- Have built performance-critical, low-level code, such as a rendering engine, a game or map renderer, or a browser engine.
- Have made software measurably faster on real hardware, and can explain the profile and the fix.
- Have shipped an SDK or runtime that other developers built on, and have changed an API after seeing developers use the API wrong.
- Read compiler output and profiler traces.
- Write Swift, including C interop and bindings to other languages.
- Work in C, C++, Kotlin, or Rust where a platform needs it.
- Write clean, tested code, and spot a weak change in review, whether a person or an agent wrote it.
- Plan and run your work through coding agents such as Claude Code, with a low tolerance for slop.
Bonus points
- Worked on map rendering such as Mapbox GL, on Chromium or WebKit rendering, or on the Unreal or Godot engines.
- Contributed to llama.cpp, MLX, or ONNX Runtime.
- Wrote GPU compute code on mobile.
How we work
We're a small, flat team, and we build the models, the SDKs, and the apps that use them. Detail and Subwave, our own apps, run our models in production.
Start what needs starting without waiting to be asked, and finish what you start. Take on work outside your role when a project needs you.
We ship quickly, so we cut scope until only the part users notice is left. Anyone can comment on your work or redo your draft, and we say early when work isn't ready. We read that feedback as help.
We judge the work by what shipped and what changed because of it. Nobody counts hours.
What we offer
- Salary and equity, based on level.
- In our Amsterdam office, or remote anywhere between Eastern Time in North America (UTC-5) and Central European Time (UTC+1), so everyone shares part of the working day.
Recruiters and agencies
We hire directly. We don't reply to emails from recruiters or agencies, and we don't want their outreach about this role or any other.
Location: Amsterdam or Remote (UTC-5 to UTC+1)
Apply on our site: https://desertant.com/jobs/runtime-and-performance-engineer/
JobMatch aggregates public listings. Always apply through the original posting.