Apertus Engineer / Apertus Engineeress

ETH Zurich - July 18, 2026

Apertus Engineer: Evaluations

Location: 100%, Zurich, fixed-term

We are seeking a skilled engineer to join the Apertus evaluation effort. The ideal candidate will build and operate the evaluation codebase and pipelines that inform our training and release decisions, ensuring consistent results between training and serving. This role requires strong Python engineering, hands-on LLM evaluation experience, and the ability to work collaboratively in a research-focused environment.

Project Background

We train open foundation models with hundreds of billions of parameters on thousands of GPUs on one of the largest AI-ready supercomputers in Europe. The team consists of more than a dozen full-time engineers working alongside leading researchers from EPFL and ETH Zürich. Our achievements include the release of the Apertus 1 and Apertus 1.5 models, and we actively collaborate with over thirty academic partners to deliver fully open, responsibly trained, multilingual, multimodal AI models for both research and industry.

Apertus is trained and developed on Alps, the Swiss National Supercomputing Centre's (CSCS) supercomputing infrastructure. The selected candidate will be comfortable working in an HPC environment and collaborating with researchers and infrastructure engineers.

Job Description

The engineer will own the evaluation codebase and pipelines that inform training decisions and releases.

Evaluation Infrastructure

  • Build and maintain the evaluation codebase and pipelines for Apertus models, from checkpoints during training to released models.
  • Make evaluations run quickly and at scale: parallel execution on Alps, efficient use of inference backends, caching, and result tracking.
  • Reduce mismatch between evaluations during training and serving: ensure consistent tokenisation, chat templates, prompting, and sampling across evaluation harnesses and inference engines.
  • Debug evaluation failures, regressions, and inconsistencies across backends.

Benchmark Coverage

  • Integrate and execute the evaluations important to the project. Collaborating researchers and engineers will design new evaluations; this role will ensure they run reliably and at scale.
  • Cover image and audio evaluations alongside text within the same pipeline.
  • Integrate new benchmarks as the field evolves, collaborating with academic partners to onboard the benchmarks they create, and validate that metrics and harness implementations are trustworthy.

Comparative and Third-party Evaluation

  • Evaluate third-party services and other open and closed models against the same benchmark suite, producing directly comparable and reproducible results.
  • Provide evaluation results, reports, and dashboards that support training decisions and release decisions.
  • Work closely with engineers focused on safety, deployment, and community needs, integrating the evaluations they create into the shared pipeline.

Profile

Essential

  • MSc or PhD in Computer Science, Data Science, Artificial Intelligence, Machine Learning, or a related field. Exceptional BSc candidates with strong engineering experience will also be considered.
  • Strong Python and software engineering skills, including experience building robust data or evaluation pipelines.
  • Experience with LLM evaluation: established harnesses (e.g., lm-evaluation-harness) or custom benchmark tooling.
  • Excellent collaboration and communication skills, with the ability to work across research and engineering teams.
  • Prior hands-on experience in the core domains of this role is required—this can be project or study-based experience; formal work experience is preferred.
  • A high degree of flexibility: priorities, tools, and day-to-day tasks may shift with training schedules, releases, and the fast-moving landscape of the field.
  • Experience running evaluations at scale on GPU clusters (Slurm or similar) and with inference engines such as vLLM or SGLang.
  • Familiarity with agentic evaluation and agentic harnesses: tool use, sandboxed execution environments, benchmarks such as SWE-bench or similar.
  • Experience with multimodal (image or audio) model evaluation.

Strongly Preferred

  • An eye for statistical rigor: variance across runs, prompt sensitivity, significance of differences between models.

Nice to Have

  • Published research in domains relevant to this role, or familiarity with recently published research on these topics.
  • Experience with LLM-as-judge pipelines and their calibration.
  • Familiarity with benchmark contamination detection and decontamination practices.
  • Experience visualizing and communicating evaluation results to research teams.

Workplace

We offer a stimulating academic environment at one of the world's leading technical universities, access to Alps—the largest AI-ready supercomputer in Europe—and the opportunity to collaborate with top researchers and engineers from EPFL, ETH Zürich, CSCS, and other Swiss institutions.

We Offer

  • Attractive employment conditions and comprehensive benefits, including the ETH Zürich/EPFL pension plans.
  • Flexible working arrangements, including options for remote work.
  • Professional development opportunities, including conference attendance and specialized training.
  • The chance to contribute to open-source projects with global impact.
  • Being part of Switzerland's sovereign AI development, working on technology with national significance.
  • The role can be based either in Lausanne at EPFL or in Zürich at ETH Zürich.

We Value Diversity and Sustainability

In line with our values, ETH Zurich encourages an inclusive culture. We promote equality of opportunity, value diversity, and nurture a working and learning environment in which the rights and dignity of all our staff and students are respected. Sustainability is a core value for us, and we consistently work towards a climate-neutral future.

Curious? So Are We.

If you are interested in this opportunity, please apply online using the form below. Only applications matching the job profile will be considered.

For further information about the ETH AI Center and the Swiss AI Initiative, please visit our website. Questions regarding the position can be directed to Dr. Imanol Schlag at ischlag@ethz.ch (no applications).

About ETH Zürich

ETH Zurich is one of the world’s leading universities specializing in science and technology. We are renowned for our excellent education, cutting-edge fundamental research, and direct transfer of new knowledge into society. Over 30,000 people from more than 120 countries find our university to be a place that promotes independent thinking and an environment that inspires excellence. Located in the heart of Europe, we forge connections all over the world, working together to develop solutions for today's and tomorrow’s global challenges.

Location : Lausanne
Country : Switzerland

Application Form

Please enter your information in the following form and attach your resume (CV)

Only pdf, Word, or OpenOffice file. Maximum file size: 3 MB.