Apertus Engineer: Deployment
Location: 100%, Zurich, fixed-term
We are seeking a skilled engineer to oversee the technical release path of Apertus models. The ideal candidate will integrate Apertus into the open-source inference ecosystem, produce quantized variants, and ensure that each release works seamlessly for the community. This role requires strong Python and software engineering skills, experience with LLM inference stacks, and a proven track record of open-source contributions.
Project Background
We train open foundation models with hundreds of billions of parameters on thousands of GPUs using one of the largest AI-ready supercomputers in Europe. Our team includes over a dozen full-time engineers working alongside leading researchers from EPFL and ETH Zürich. We have successfully released the Apertus 1 and Apertus 1.5 models and collaborate with over thirty academic partners to deliver fully open, responsibly trained, multilingual, and multimodal AI models for research and industry.
Apertus is developed on Alps, the Swiss National Supercomputing Centre's (CSCS) supercomputing infrastructure. This position lies at the intersection of the training team and the open-source community, with the community manager managing social engagements and this role focusing on the technical release path.
Job Description
The engineer will ensure that Apertus releases work out of the box across the open-source LLM ecosystem, from server-grade inference to personal deployment.
Upstream Integration and Release Engineering
- Manage the technical release path of trained Apertus models, including checkpoint conversion and preparation of release artifacts (weights, configurations, tokenizers, model cards) in collaboration with the training team.
- Implement and support Apertus model architectures in community libraries such as Hugging Face Transformers, vLLM, SGLang, and llama.cpp, guiding these contributions through review to ensure community support.
- Verify day-0 compatibility of new releases with the major inference engines and model formats.
- Coordinate release timing and technical materials with the community manager.
Quantisation
- Produce quantized variants of released models (e.g., FP8, INT4/AWQ/GPTQ, GGUF) suitable for both server and personal deployment.
- Validate quantized variants against evaluation benchmarks to ensure that quality is maintained.
Documentation and Examples
- Provide example scripts and reference configurations to demonstrate how to serve and use Apertus models with vLLM, SGLang, and Transformers.
- Support personal and local deployment ecosystems such as LM Studio, Ollama, and llama.cpp.
- Maintain deployment documentation and troubleshooting guides for the community.
Profile
Essential
- MSc or PhD in Computer Science, Data Science, Artificial Intelligence, Machine Learning, or a related field; exceptional BSc candidates with strong engineering experience will also be considered.
- Strong Python and software engineering skills, including experience with open-source contribution workflows (pull requests, code review, CI).
- Experience with LLM inference stacks such as Hugging Face Transformers, vLLM, or SGLang.
- Strong collaboration and communication skills with the ability to work across research, engineering, and community-facing teams.
- Prior hands-on experience in the core domains of this role is required; project or study-based experience is acceptable, while formal work experience is preferred.
- A high degree of flexibility to adapt to shifting priorities, tools, and day-to-day tasks driven by training schedules, releases, and the fast-moving field.
- A track record of successful contributions to ML or inference libraries (e.g., Transformers, vLLM, SGLang, llama.cpp).
Strongly Preferred
- Experience converting models between formats and frameworks (e.g., Megatron-LM checkpoints, safetensors, GGUF).
- Familiarity with personal and local deployment tools such as LM Studio, Ollama, or llama.cpp.
- Experience writing developer-facing documentation and example code.
Nice to Have
- Published research in relevant domains or familiarity with recently published research on these topics.
- Experience quantizing models without performance degradation (FP8, INT4, AWQ, GPTQ) and evaluating quantized models.
- Experience with LLM evaluation harnesses and benchmark pipelines.
- Familiarity with GPU inference performance tuning and serving at scale.
- Experience with Apple Silicon / MLX or other consumer-hardware inference targets.
Workplace
The successful candidate can be based either in Lausanne at EPFL or in Zürich at ETH Zürich.
We Offer
- A stimulating academic environment at one of the world's leading technical universities.
- Access to Alps, one of the largest AI-ready supercomputers in Europe.
- The opportunity to work alongside and intersect with leading researchers in the field.
- Collaboration with top researchers and engineers from EPFL, ETH Zürich, CSCS, and other Swiss institutions.
- Attractive employment conditions and comprehensive benefits, including the ETH Zürich/EPFL pension plans.
- Flexible working arrangements, including options for remote work.
- Professional development opportunities, including conference attendance and specialized training.
- The chance to contribute to impactful open-source projects.
- Being part of Switzerland's sovereign AI development, working on technology with national significance.
Diversity and Sustainability
In line with our values, ETH Zurich promotes an inclusive culture. We advocate for equality of opportunity, value diversity, and nurture a working and learning environment that respects the rights and dignity of all our staff and students. Visit our Equal Opportunities and Diversity website to learn how we maintain a fair and open environment that allows everyone to grow and thrive. Sustainability is a core value for us, and we are committed to working towards a climate-neutral future.
Curious? So Are We.
Apply online using the form below. Only applications matching the job profile will be considered.
About ETH Zürich
ETH Zurich is one of the world’s leading universities specializing in science and technology. Renowned for excellent education, cutting-edge fundamental research, and the direct transfer of new knowledge into society, we host over 30,000 individuals from more than 120 countries, promoting independent thinking and an inspiring environment for excellence. Situated in the heart of Europe, we work collaboratively to develop solutions for today’s and tomorrow’s global challenges.