Role Overview
Join a fully remote team within the EU to support and maintain highly available machine learning services in production. This 9-month contract, with extension possibilities, focuses on working closely with Applied Scientists to productionise ML models. Your primary objective is to ensure the underlying infrastructure remains scalable, reliable, and well monitored.
Key Responsibilities
- Support and troubleshoot real-time ML inference services in production.
- Build and maintain AWS infrastructure, CI/CD pipelines, and Infrastructure as Code.
- Manage environments using Kubernetes, Docker, autoscaling, and monitoring.
- Support GPU infrastructure, including NVIDIA/CUDA upgrades and troubleshooting.
- Manage real-time and batch data pipelines using Airflow, Databricks Workflows, or similar tools.
- Collaborate with Applied Scientists to move ML prototypes into production.
- Write and review production-quality Python.
- Participate in 24x7 on-call support and resolve live production incidents.
Key Skills & Experience
- Strong experience with Linux, Kubernetes, and Docker.
- Good knowledge of Python, Git, and CI/CD.
- Hands-on experience with Infrastructure as Code.
- Experience with Airflow, Databricks Workflows, or equivalent.
- Experience supporting high-load production environments and resolving live incidents.