Overview
We are seeking a Data Research Engineer to join our AI team, where we are building the next generation of foundation models across various modalities: text, image, video, and audio. If you are passionate about exploring, designing, and constructing high-quality datasets to drive frontier AI models, this role is tailored for you.
At Microsoft AI, data is at the core of our innovation. In this role, you will collaborate closely with colleagues to curate, analyze, and evaluate diverse data sources that are critical for model development. You will lead initiatives in developing novel data collection strategies, enhancing dataset quality, understanding data-driven model behaviors, and ensuring alignment of datasets with ethical and societal values.
This is a cross-disciplinary, high-impact role ideal for engineers eager to push the boundaries of what AI can learn from data.
Microsoft’s mission is to empower every person and organization on the planet to achieve more. We unite with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day, we build on our values of respect, integrity, and accountability to foster a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
- Create high-quality datasets for training and evaluation; conduct experiments on new datasets to assess their impact and identify the most effective data.
- Develop and maintain scalable data pipelines for data ingestion, pre-processing, filtering, and annotation.
- Analyze real-world multimodal datasets to assess quality, diversity, and relevance, while identifying areas for improvement.
- Build tools and workflows for dataset auditing, visualization, and versioning.
- Collaborate with Safety, Ethics, and Governance teams to ensure datasets meet standards for quality, privacy, and responsible AI practices.
Qualifications
Required Qualifications:
- Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or a related technical field.
- Proficiency in Python.
- Experience with distributed data-processing frameworks such as Spark, Ray, and workflow-orchestration tools like Airflow.
- Experience working with datasets on a petabyte scale and managing trillions of rows.
- Experience building datasets for training foundation models, including language or multimodal models.
- Strong expertise in data analysis, data engineering, or both.
- Ability to communicate technical findings clearly and effectively to research, engineering, and product teams.
- Experience in evaluating dataset quality and measuring the impact of data through controlled model-training experiments.
Preferred Qualifications:
- Master’s degree in Computer Science or a related technical field, or equivalent experience.
- Experience working with large-scale, real-world datasets that are unstructured or semi-structured.
- Experience in evaluating dataset quality and measuring the impact of data through controlled model-training experiments.
- Experience with multimodal data, such as image, video, or audio data.
Apply online using the form below. Only applications matching the job profile will be considered.
Work locationZürich, Switzerland