Data Engineer

 

Description:

We're looking for a Data Engineer with a strong background in bioinformatics and genetics to help build the data pipelines supporting Our Future Health's growing Clinical Research Recruitment Service.

This is an opportunity to combine genetics, bioinformatics and modern data engineering, working with data at significant scale while contributing directly to research that could improve how diseases are prevented, detected and treated.

 

Our Future Health is an ambitious collaboration between the public, charity and private sectors, designed to help people live healthier lives for longer through better prevention, earlier detection and improved treatment of diseases. We will speed up the discovery of new methods of early disease detection, and the evaluation of new diagnostic tools, to help identify and treat diseases early, when outcomes are usually better. With over 2.7M volunteers across the UK, we're now the world's biggest health research programme of its kind, and our volunteer group is also more diverse than other, similar health research programmes.  

Technology and data are central to our mission. Our systems power web sites, clinics across the UK, secure analytics and research systems, pipelines that process highly sensitive health and genetic data, and we are continuing to grow our engineering capability to support this ambition. 

 

Our Clinical Research Recruitment Service helps life sciences and academic organisations identify potential participants for clinical studies. As the service grows, we need to turn scientific workflows and manual processes into robust, reusable and scalable production pipelines.

 

We're looking for a Data Engineer with a solid understanding and experience of bioinformatics, in particular tools and methods associated with genomic data. You can design, build and test pipelines using a range of different technologies. You know how to create repeatable and reusable products and can communicate to and between technical and non-technical stakeholders, with the ability to facilitate discussions and manage different perspectives within a multidisciplinary team including scientists, software engineers, product managers and other data engineers.  

Essential Duties and Responsibilities 

  • Support the build of re-usable data pipelines used to identify prospective clinical trial participants
  • Produce logic for data transformation steps as code, which meets the requirements for our end users and builds well curated, accessible and quality controlled data for analysis. 
  • Developing prototypes for pipelines for complex transformations drawing on existing workflows developed in industry and academia. 
  • Keep abreast of best practice in data engineering across industry, research and Government and facilitating the adoption of standards. 
  • Providing technical input into the upstream parts of the data pipeline, including the specification and transfer of data from data providers. 
  • Routine ad-hoc data curation activities requiring hands on development of bespoke ETL cleaning scripts using languages such as Python. 
  • Working with researchers to understand the data requirements and work with them to deliver the data needed for their projects. 
     

Requirements
 

We welcome applications from all who may not feel they match the full criteria, so if you have most of the below, we'd like to hear from you:  

  • Experience building and maintaining robust, scalable and efficient data pipelines. Capable of processing very large amounts of data based on feeds from multiple systems using a range of different technologies. 
  • Can listen to the needs of technical and business stakeholders and interpret them, and effectively manage stakeholder expectations.   
  • Detailed knowledge and understanding of genomic data (experience in genotyping and imputation is advantageous). 
  • Experience using bioinformatics file standards (VCF, BGEN etc) and tools (PLINK, bcftools, QCtools etc) 
  • Highly proficient in Python. 
  • Highly proficient in version control and Git/GitHub. 
  • Experience of workflow management tools, e.g. Nextflow, WDL/Cromwell, Airflow, Prefect, Dagster 
  • Understanding of containerisation (e.g. Docker) and deployment (e.g. Kubernetes). 
  • Good understanding of cloud environments (ideally Azure), distributed computing and scaling workflows and pipelines 
  • Understanding of common data transformation and storage formats, e.g. Apache Parquet. 
  • Awareness of data standards such as GA4GH ( https://www.ga4gh.org/) and FAIR (https://www.go-fair.org/fair-principles/). 
  • Experience with Spark, Databricks, data lakes. 
  • Follow best practices like code reviews, clean code and unit tests.  
  • Experience working in an agile development team. 

 

Organization Our Future Health UK
Industry Engineering Jobs
Occupational Category Data Engineer
Job Location London,UK
Shift Type Morning
Job Type Full Time
Gender No Preference
Career Level Intermediate
Experience 2 Years
Posted at 2026-08-24 10:40 pm
Expires on 2026-10-08