Berlin, Germany Max Delbrück Center
Katarína Grešová
Machine learning for biology, built to be reused
My work is the loop from experiment to model to biological insight: benchmarks and data preparation at one end, model interpretation at the other, models of RNA regulation in between.
Training a good model is hard, and it is the part that gets all the attention. But on flawed data, or with no way to see what it learned, it still isn’t usable — and if it isn’t usable, why did we build it?
Seven years of software engineering taught me to build things other people can run, and I’ve learned a new field every few years since. Now I teach that, and I check it for people who need to know whether their own results hold.
Available for teaching, consulting and advisory work. Berlin, or remote across Europe. What that looks like
What you can ask me for Start a conversation
Independent work, alongside the research. Every one of these is something I have already done for somebody else — the note at the foot of each card says where.
-
Teach your team
Deep learning for people who need to use it, not publish about it.
Three days that take biologists from never having written a training loop to having one that works — or a half-day on interpretability for a team that already has models and cannot say what they learned. The notebooks go home with you, and they still run a year later.
Run as a three-day course at the University of Malta: seven notebooks, from k-mers through to interpreting a trained network, finishing with a hackathon on real miRNA data scored on a held-out set nobody had seen. See it
-
Audit what you have built
Find out whether the number is real before you spend on it.
Class imbalance a model can exploit without learning anything, composition bias, per-position give-aways, duplicate sequences, near-duplicate leakage between train and test. I score the dataset for the shortcuts available to a classifier and hand back a report you can read, a CSV you can put in CI, and a straight answer about which of your results survive.
Packaged as Genomic Benchmarks QC, the general form of a complaint I kept making. The benchmark suite behind it has been cited 175 times. See it
-
Build it with you
Sequence models, and the pipeline that lets somebody else rerun them.
RNA and genomic sequence models end to end — data preparation, training, interpretation — delivered as a package or a workflow rather than a folder of scripts. Seven years of production software engineering came before the research, which is the reason the handover actually works.
AlphaFind searches 200 million protein structures and is public. RBP-Tar and miRBench are too. Everything I build is used by somebody who isn't me. See it
- 11
- courses and workshops taught, students supervised
- 316
- citations across 16 papers, 5 as first or co-first author
- 182
- stars on the benchmark suite, cited 175 times
- 7
- years shipping production software in industry
Citation and star counts checked September 2026.
Teaching Everything I've taught
I teach in both directions — computation to biologists, biology to computer scientists — in university practicals, conference tutorials, a three-day deep learning course, and material anyone can work through on their own. Most of it is public, so you can see exactly what you would be booking.
-
Deep Learning for Genomics
Convolutional and recurrent models for biological sequence data, taught from first principles to a room of biologists who had not written a training loop before, and who had one working by the end of...
-
The Missing Skills
Project-based tutorials on the computational skills a science degree leaves out — version control, the shell, and making your work runnable by someone who isn't you. Written because the bottleneck in computational biology is...
-
Biology Crash Course
Molecular biology from the ground up, written for computer scientists who need enough of it to be dangerous. The mirror image of The Missing Skills, and between them the two directions I spend most...
Things I build Everything I've built
Everything here was used by somebody who isn't me — that is the only rule for getting on the list. Of the 14, 3 began as a complaint about an evaluation and 5 were shipped in industry, where somebody else's day was ruined if they broke.
-
Genomic Benchmarks QC
Automated quality control for genomic machine-learning datasets. It scores the biases, duplicate sequences and train/test leakage a classifier could exploit before you train on it — length differences, GC content, per-position give-aways, near-duplicate overlap...
-
miRBench
Benchmark datasets and a Python package for miRNA target-site prediction, built to remove the frequency-class bias that quietly inflates published scores. It ships the data, the splits and wrappers around the existing predictors, so...
-
Genomic Benchmarks
Eight curated datasets for genomic sequence classification across human, mouse and roundworm, each with a sensible split and a baseline model, so a new architecture has something honest to beat. Installable as a package...
-
Attribution sequence alignment
A way to make a model say what it noticed. Attribution gives you one importance score per nucleotide, which is not yet a motif; aligning those profiles across many sequences makes the patterns a...
-
AlphaFind and AlphaFind2
Structure-similarity search across the whole of AlphaFold DB. Pretrained networks compress each structure into an embedding, which turns an all-against-all comparison over more than 200 million structures into a query you can wait for....
-
RBP-Tar
A searchable database of experimentally determined RNA-binding protein sites, pulled out of published CLIP experiments and put behind a query interface — so you can ask what binds a transcript without reprocessing somebody else's...
Selected publications All publications
-
2026
SenCat: cataloging human cell senescence through multi-omic profiling of multiple senescent primary cell types
Molecular Cell86(13), 2605–2616.e8Shared second author
-
2025
miRBench: novel benchmark datasets for microRNA binding site prediction that mitigate against prevalent microRNA frequency class bias
Bioinformatics41(Supplement_1), i542–i551Co-first author
-
2024
RBP-Tar: a searchable database for experimental RBP binding sites
F1000Research12, 755First author
-
2023
Genomic benchmarks: a collection of datasets for genomic sequence classification
BMC Genomic Data24, 25First author
-
2023
Using attribution sequence alignment to interpret deep learning models for miRNA binding site prediction
Biology12(3), 369First author
Let's talk
A course you need run, a dataset you are not sure about, a model whose numbers look a little too good — those are the three I am quickest to answer. You don't need a worked-out proposal; a paragraph is plenty. I am also open to the right full-time role, and to conversations that are none of the above yet.
contact@katarinagresova.comBerlin, or remote across Europe.