
Hey, I’m Elias Baumann, a data and validation machine learning engineer at Kaiko. I build data pipelines for frontier LLM training, covering deduplication, quality filtering, synthetic data generation and schema validation across terabyte-scale datasets. I also develop LLM-based tools for clinical text pseudonymization, orchestrating distributed data processing with Dagster, Ray and Spark.
Before joining Kaiko, I completed my PhD in computational pathology at the University of Bern, where I developed deep learning pipelines for gigapixel pathology images and published on biomarker discovery in colorectal cancer, including several first-author studies. I try to develop innovative solutions for problems, where the fundamental understanding of the respective domain is crucial, and therefore also love to learn everything about the problem domain.
I am currently based in Bern, Switzerland but am working in the Zürich office of Kaiko.
Experience
Kaiko.ai
Jul 2025 - Present
Data and Validation Machine Learning Engineer
University of Bern
Apr 2021 - May 2025
PhD Researcher, Computational Pathology
RadboudUMC Nijmegen
Mar 2023 - Aug 2023
Visiting researcher, Computational Pathology
Humboldt University of Berlin
2017 - 2020
Research Assistant, Chair of Information Systems
B.telligent
Nov 2015 - May 2016
Intern, Visual data discovery
Telefónica Germany & Co. OHG
Aug 2014 - Oct 2015
Intern, Mobile and web apps
Education
Ph. D. Computational Pathology
University of Bern
M. Sc. Information Systems
Humboldt University of Berlin
B. Sc. Computer Science
Technical University Munich
Competitions & recognition
SGPath
2023
Best Presentation
CoNIC Challenge
2022
2nd Place
AI For Climate Hackathon 2021
2021
Winner
Data Science Game
2018
Finalist (Winner of Fair Play Award)
Tools & technologies
Python, R, SQL
Dagster, Ray, Spark / Databricks
PyTorch, scikit-learn, XGBoost
Docker, Singularity
Azure, AWS
Pants, ruff, pyright
Git, Linux
Skills
- LLM training data pipelines (deduplication, quality filtering, domain classification)
- Synthetic data generation & prompt engineering (DSPy, litellm)
- Clinical text pseudonymization & PII detection
- Distributed data processing (Ray, Spark / Databricks, PyArrow, Daft)
- Pipeline orchestration (Dagster)
- Machine learning & deep learning (PyTorch, TensorFlow, Keras, scikit-learn, XGBoost)
- Survival analysis & clinical statistics
- Research from idea to published paper
- Project Management
About this site
This site has been adapted from the original webiste by Matt Grey on Github which was made available under the GNU General Public License v3.