• Alba Refoyo Martinez
  • Work
  • CV
  • Publications

What I do

I’m part of a team that develops online training modules and offers support to researchers working with omics and clinical data. My work spans four connected areas: building tools, consulting, advising research groups, and teaching — all in service of making data science more reproducible, transparent, and FAIR.

All materials I create — from workshops to web-based exercises — are freely available on the project’s website.

  • App & Tool Development
  • Consulting
  • Research Advising
  • Teaching & Training

As part of the national Health Data Science Sandbox Project, I collaborate on building and maintaining containerized apps, Shiny tools, and training infrastructure used across Danish HPC systems — building training modules that combine notebooks, coding exercises, and interactive tutorials, all hosted on GitHub Pages using Quarto and fully version-controlled. A look at some of the projects I’ve been part of:

Health Data Science

Genomics

Transcriptomics

Browse all our apps on Docker Hub

I provide 1:1 and team consultations to researchers who need hands-on help getting their analyses running reliably — choosing the right tools, troubleshooting HPC and containerized environments, and setting up pipelines that will keep working long after the initial setup. The goal is always the same: leave the group more self-sufficient than I found them.

Beyond day-to-day troubleshooting, I advise research groups on the bigger picture: structuring projects for Research Data Management (RDM) and FAIR compliance, planning reproducible workflows before a project starts, and thinking through how data and tools will be shared, reused, and maintained over time.

I regularly teach workshops and contribute to creating online training materials — covering both practical tools like Git & GitHub, Snakemake & Nextflow, Docker & Conda, Cookiecutter, and Shiny Apps, and specific topics in genomics, transcriptomics, and clinical data — increasingly with a focus on AI tools throughout.

Computational RDM

Helping omics researchers improve how they manage and organize their data and projects on HPCs: RDM concepts, best practices in data organization, and tools to support effective data management.

HPC Best Practices

A go-to resource for researchers working on HPC, building workflows and pipelines, or running data science projects — all with reproducibility in mind: Snakemake, Nextflow, community pipelines like nf-core, and software management.

AI for Biomedical Researchers Active

A module I’m currently building to help researchers understand and apply AI tools — including LLMs — responsibly in their day-to-day biomedical research: practical use of LLMs, capabilities and limitations, and reproducible use of AI in scientific analysis.

Beyond the Sandbox modules, I teach PhD students and researchers practical aspects of HPC onboarding, pipeline development, software management, and Research Data Management (RDM).

Past Teaching Engagements
  • Analysis of large-scale datasets: PhD course - 2021
  • Interaction design (KOMIT) on the interaction and usability of user interfaces: Bachelor course - 2019

My teaching philosophy: I emphasize hands-on, project-based learning, where researchers actively engage with the tools they will use in their day-to-day work. I believe in breaking down complex concepts into simple, manageable parts and gradually building on them to promote deeper understanding and confidence.