Berlin, Germany Max Delbrück Center

Katarína Grešová

Portrait of Katarína Grešová

Machine learning for biology, built to be reused

My work is the loop from experiment to model to biological insight: benchmarks and data preparation at one end, model interpretation at the other, models of RNA regulation in between.

Training a good model is hard, and it is the part that gets all the attention. But on flawed data, or with no way to see what it learned, it still isn’t usable — and if it isn’t usable, why did we build it?

Seven years of software engineering taught me to build things other people can run, and I’ve learned a new field every few years since. Now I teach that, and I check it for people who need to know whether their own results hold.

Available for teaching, consulting and advisory work. Berlin, or remote across Europe. What that looks like

Start a conversation

Independent work, alongside the research. Every one of these is something I have already done for somebody else — the note at the foot of each card says where.

  • Teach your team

    Deep learning for people who need to use it, not publish about it.

    Three days that take biologists from never having written a training loop to having one that works — or a half-day on interpretability for a team that already has models and cannot say what they learned. The notebooks go home with you, and they still run a year later.

    Run as a three-day course at the University of Malta: seven notebooks, from k-mers through to interpreting a trained network, finishing with a hackathon on real miRNA data scored on a held-out set nobody had seen. See it

  • Audit what you have built

    Find out whether the number is real before you spend on it.

    Class imbalance a model can exploit without learning anything, composition bias, per-position give-aways, duplicate sequences, near-duplicate leakage between train and test. I score the dataset for the shortcuts available to a classifier and hand back a report you can read, a CSV you can put in CI, and a straight answer about which of your results survive.

    Packaged as Genomic Benchmarks QC, the general form of a complaint I kept making. The benchmark suite behind it has been cited 175 times. See it

  • Build it with you

    Sequence models, and the pipeline that lets somebody else rerun them.

    RNA and genomic sequence models end to end — data preparation, training, interpretation — delivered as a package or a workflow rather than a folder of scripts. Seven years of production software engineering came before the research, which is the reason the handover actually works.

    AlphaFind searches 200 million protein structures and is public. RBP-Tar and miRBench are too. Everything I build is used by somebody who isn't me. See it

11
courses and workshops taught, students supervised
316
citations across 16 papers, 5 as first or co-first author
182
stars on the benchmark suite, cited 175 times
7
years shipping production software in industry

Citation and star counts checked September 2026.

Everything I've built

Everything here was used by somebody who isn't me — that is the only rule for getting on the list. Of the 14, 3 began as a complaint about an evaluation and 5 were shipped in industry, where somebody else's day was ruined if they broke.

All publications

  1. 2026

    SenCat: cataloging human cell senescence through multi-omic profiling of multiple senescent primary cell types

    Carlos Anerillas, Gisela Altés, Katarína Grešová, Dimitrios Tsitsipatis, Krystyna Mazan-Mamczarz, et al., Manolis Maragkakis, Nathan Basisty, Myriam Gorospe

    Molecular Cell86(13), 2605–2616.e8Shared second author

  2. 2025

    miRBench: novel benchmark datasets for microRNA binding site prediction that mitigate against prevalent microRNA frequency class bias

    Stephanie Sammut, Katarína Grešová, Dimosthenis Tzimotoudis, Eva Maršálková, David Čechák, Panagiotis Alexiou

    Bioinformatics41(Supplement_1), i542–i551Co-first author

  3. 2024

    RBP-Tar: a searchable database for experimental RBP binding sites

    Katarína Grešová, Tomáš Racek, Vlastimil Martinek, David Čechák, Radka Svobodová, Panagiotis Alexiou

    F1000Research12, 755First author

  4. 2023

    Genomic benchmarks: a collection of datasets for genomic sequence classification

    Katarína Grešová, Vlastimil Martinek, David Čechák, Petr Šimeček, Panagiotis Alexiou

    BMC Genomic Data24, 25First author

  5. 2023

    Using attribution sequence alignment to interpret deep learning models for miRNA binding site prediction

    Katarína Grešová, Ondřej Vaculík, Panagiotis Alexiou

    Biology12(3), 369First author

Let's talk

A course you need run, a dataset you are not sure about, a model whose numbers look a little too good — those are the three I am quickest to answer. You don't need a worked-out proposal; a paragraph is plenty. I am also open to the right full-time role, and to conversations that are none of the above yet.

contact@katarinagresova.com

Berlin, or remote across Europe.