Infrastructure Engineer / Infrastructure Engineeress

ETH Zurich - July 18, 2026

Apertus Engineer: Infrastructure

100%, Zurich, fixed-term

We are seeking a skilled infrastructure engineer to join the Apertus team. The ideal candidate will take ownership of the container image stack behind our pre-training, post-training, and serving workloads, collaborating with CSCS engineers to ensure that large-scale training on Alps remains stable and efficient. This role demands strong Linux and container skills, as well as experience in HPC environments, with the ability to work collaboratively across research, engineering, and operations teams.

Project Background

The Apertus project is a collaborative effort between EPFL, ETH Zürich, and CSCS, aiming to propel the next version of Apertus. The successful candidate will manage the container image stack, closely working with CSCS to maintain the stability and efficiency of large-scale training.

Our team trains open foundation models with hundreds of billions of parameters on thousands of GPUs, utilizing one of the largest AI-ready supercomputers in Europe. With over a dozen full-time engineers and collaborations with leading researchers from EPFL and ETH Zürich, we have released the Apertus 1 and Apertus 1.5 models and are committed to delivering fully open (open source), responsibly trained, multilingual, multimodal AI models for research and industry.

Apertus is developed on Alps, the supercomputing infrastructure of the Swiss National Supercomputing Centre (CSCS). The role requires comfort in an HPC environment and collaboration with researchers and infrastructure engineers.

Job Description

The engineer will ensure stability and throughput for the Apertus pre-training and post-training pipelines by maintaining the ML system images and collaborating with CSCS on the underlying infrastructure.

ML System Image Maintenance

  • Build, maintain, and upgrade container images for all core ML development phases: pre-training, post-training/alignment, and model serving/deployment.
  • Target the ARM-based (aarch64, Grace-Hopper) node architecture of Alps, managing the full dependency stack (CUDA, NCCL, PyTorch, training and serving frameworks).
  • Ensure image builds are reproducible, versioned, and documented, including continuous integration for builds and upgrades.
  • Validate images against reference pre-training and post-training workloads in collaboration with Apertus engineers, maintaining working launch examples.

Compute Partnership and Efficiency

  • Serve as the primary technical point of contact with CSCS engineers and researchers concerning compute, reliability, and efficiency.
  • Collaborate with CSCS staff to identify and implement improvements in efficiency and performance relevant to large-scale LLM training.
  • Contribute to systemic enhancements in CSCS-based resources (network, storage, scheduling) pertinent to large-scale LLM training.
  • Document and disseminate institutional knowledge about CSCS infrastructure and best practices for leveraging high-performance systems.

Infrastructure Stress Testing

  • Conduct stress tests of the infrastructure using representative pre-training and post-training workloads, validating stability and throughput after image upgrades, system maintenance, and configuration changes.
  • Collaborate closely with Apertus pre-training and post-training engineers to troubleshoot cluster-level issues affecting stability and throughput, including node failures, networking issues, storage performance, checkpointing, and scheduling.
  • Support the Apertus serving stack, which is built on the same images (operation of the serving stack is managed by a separate engineer).

Profile

Essential

  • MSc or PhD in Computer Science, Data Science, Artificial Intelligence, Machine Learning, or a related field.
  • Exceptional BSc candidates with strong engineering experience will also be considered.
  • Hands-on experience with HPC environments: job schedulers such as Slurm, shared filesystems, and multi-node GPU systems.
  • Strong Linux systems and container skills (Docker/Podman and HPC runtimes such as enroot or Apptainer).
  • Excellent collaboration and communication skills with the capability to work across research, engineering, and operations teams.
  • Prior hands-on experience in the core domains of this role is essential; this may be project or study-based experience, although formal work experience is preferred.
  • A high degree of flexibility: priorities, tools, and day-to-day tasks shift with training schedules, releases, and the fast-moving nature of the field.

Strongly Preferred

  • Familiarity with LLM training and serving frameworks such as Megatron-LM, PyTorch distributed, vLLM, or SGLang.
  • Experience building or adapting containers for ARM64/aarch64 platforms.
  • Familiarity with HPC networking and communication stacks: Slingshot, libfabric, NCCL, and its debugging.
  • Experience with parallel filesystems (e.g., Lustre) and storage performance tuning.
  • Experience setting up CI/CD pipelines for container image builds.

Nice to Have

  • Published research in domains relevant to this role or familiarity with recently published research on these topics.
  • Experience profiling distributed GPU workloads (Nsight, DCGM, communication benchmarks) and translating findings into infrastructure improvements.
  • Experience with Grace-Hopper (GH200) or other tightly coupled CPU-GPU architectures.
  • Contributions to open-source infrastructure or ML tooling.

Workplace

We offer a stimulating academic environment at one of the world’s leading technical universities. Collaborate with top researchers and engineers from EPFL, ETH Zürich, CSCS, and other Swiss institutions. Enjoy attractive employment conditions and comprehensive benefits, including pension plans from ETH Zürich and EPFL. Flexible working arrangements, including options for remote work, are available.

We Value Diversity and Sustainability

In line with our values, ETH Zurich encourages an inclusive culture. We promote equality of opportunity, value diversity, and nurture a working and learning environment that respects the rights and dignity of all our staff and students. Sustainability is a core value for us, and we consistently work towards a climate-neutral future.

Curious? So Are We.

We look forward to receiving your online application using the form below. Ensure your submission includes:

  • CV/Resume
  • Cover letter explaining your interest and qualifications
  • Academic transcripts
  • Contact information for 2-3 references
  • Links to GitHub repositories or other examples of your programming work (if available)

Further information about the ETH AI Center and the Swiss AI Initiative can be found on our website. For inquiries regarding the position, please contact Dr. Imanol Schlag at ischlag@ethz.ch. Please note that only applications meeting the job profile will be considered.

About ETH Zürich

ETH Zurich is among the world’s leading universities specializing in science and technology. Renowned for excellent education, cutting-edge research, and the direct transfer of new knowledge into society, we host over 30,000 individuals from more than 120 countries. Located in the heart of Europe, we aim to develop solutions for today and tomorrow's global challenges.

Location : Lausanne
Country : Switzerland

Application Form

Please enter your information in the following form and attach your resume (CV)

Only pdf, Word, or OpenOffice file. Maximum file size: 3 MB.