Mid-Level ML Infrastructure Engineer - Remote
We are seeking a talented Mid-Level ML Infrastructure Engineer to support and maintain highly available machine learning services in production. This role involves close collaboration with Applied Scientists to productionize ML models and ensure that the underlying infrastructure is scalable, reliable, and well-monitored. This is a fully remote opportunity within the EU, available as a 9-month contract with possibilities for extension.
Key Responsibilities
- Support and troubleshoot real-time ML inference services in production.
- Build and maintain AWS infrastructure, CI/CD pipelines, and Infrastructure as Code.
- Manage environments using Kubernetes, Docker, autoscaling, and monitoring.
- Support GPU infrastructure, including NVIDIA/CUDA upgrades and troubleshooting.
- Manage real-time and batch data pipelines using Airflow, Databricks Workflows, or similar tools.
- Collaborate with Applied Scientists to transition ML prototypes into production.
- Write and review production-quality Python code.
- Participate in 24x7 on-call support and resolve live production incidents.
Key Skills & Experience
- Strong experience with Linux, Kubernetes, and Docker.
- Good knowledge of Python, Git, and CI/CD processes.
- Hands-on experience with Infrastructure as Code.
- Experience with Airflow, Databricks Workflows, or equivalent.
- Proven track record supporting high-load production environments and resolving live incidents.
Apply online using the form below. Only applications matching the job profile will be considered.
Work locationEastern Europe, Romania, Switzerland