Hey, I’m Elias Baumann, a data and validation machine learning engineer at Kaiko. I build data pipelines for frontier LLM training, covering deduplication, quality filtering, synthetic data generation and schema validation across terabyte-scale datasets. I also develop LLM-based tools for clinical text pseudonymization, orchestrating distributed data processing with Dagster, Ray and Spark.

Before joining Kaiko, I completed my PhD in computational pathology at the University of Bern, where I developed deep learning pipelines for gigapixel pathology images and published on biomarker discovery in colorectal cancer, including several first-author studies. I try to develop innovative solutions for problems, where the fundamental understanding of the respective domain is crucial, and therefore also love to learn everything about the problem domain.

I am currently based in Bern, Switzerland but am working in the Zürich office of Kaiko.

Experience

Kaiko.ai

Jul 2025 - Present

Data and Validation Machine Learning Engineer

University of Bern

Apr 2021 - May 2025

PhD Researcher, Computational Pathology

RadboudUMC Nijmegen

Mar 2023 - Aug 2023

Visiting researcher, Computational Pathology

Humboldt University of Berlin

2017 - 2020

Research Assistant, Chair of Information Systems

B.telligent

Nov 2015 - May 2016

Intern, Visual data discovery

Telefónica Germany & Co. OHG

Aug 2014 - Oct 2015

Intern, Mobile and web apps

Education

Ph. D. Computational Pathology

University of Bern

M. Sc. Information Systems

Humboldt University of Berlin

B. Sc. Computer Science

Technical University Munich

Competitions & recognition

SGPath

2023

Best Presentation

CoNIC Challenge

2022

2nd Place

AI For Climate Hackathon 2021

2021

Winner

Data Science Game

2018

Finalist (Winner of Fair Play Award)

Tools & technologies

  • Python, R, SQL

  • Dagster, Ray, Spark / Databricks

  • PyTorch, scikit-learn, XGBoost

  • Docker, Singularity

  • Azure, AWS

  • Pants, ruff, pyright

  • Git, Linux

Skills

  • LLM training data pipelines (deduplication, quality filtering, domain classification)
  • Synthetic data generation & prompt engineering (DSPy, litellm)
  • Clinical text pseudonymization & PII detection
  • Distributed data processing (Ray, Spark / Databricks, PyArrow, Daft)
  • Pipeline orchestration (Dagster)
  • Machine learning & deep learning (PyTorch, TensorFlow, Keras, scikit-learn, XGBoost)
  • Survival analysis & clinical statistics
  • Research from idea to published paper
  • Project Management

About this site

This site has been adapted from the original webiste by Matt Grey on Github which was made available under the GNU General Public License v3.