Description:
This is a hands-on role where you'll apply strong data science skills and an engineering mindset to real operational problems: safe outputs review, re-identification risk, privacy risks in trained models, synthetic data, and data-driven security controls. You will work closely with teams across Product, Security, Engineering, Science, Data Protection, and Researcher Operations as well as external experts to develop practical, scalable and evidence-based solutions.
You will be expected to work closely with users of these systems including researchers, Airlock reviewers, access governance reviewers and security specialists, to understand their workflows, pain points and risk decisions. A key part of the role will be turning complex privacy and security problems into trusted algorithms, evidence and automated support.
What you'll do
Develop data-driven approaches to disclosure control and safe outputs review, supporting the scaling of our TRE (Trusted Research Environment) Airlock. This may include building algorithms and automated review tools to classify outputs, detect potentially disclosive content, identify patterns of risk, and provide explainable decision support for human reviewers.
Work directly with Airlock reviewers and operational users to understand where automation can help, where human judgement is essential, and how tools should be designed to support consistent, auditable and proportionate decisions.
Contribute to our approach to de-identification and re-identification risk assessment, helping us assess how privacy risk changes across datasets, access models and analytical outputs.
Help us develop our approaches to safe AI using health data, including how we assess and manage privacy risks associated with trained models.
Explore and develop approaches to synthetic data generation, assessing how synthetic data can be used safely and usefully, and how to evaluate the privacy, fidelity and utility trade-offs of different approaches.
Collaborate with our Information Security team on data-driven security control assessment, threat detection and monitoring approaches, including identifying signals of risky behaviour, anomalous activity or misuse of data access environments, and cyber risk quantification.
Build prototypes and production-quality code, working with engineers to turn promising approaches into robust, maintainable and auditable tools.
Work closely with governance, legal, ethics and operational colleagues to ensure technical controls fulfil their requirements.
Keep up with emerging methods in statistical disclosure control, privacy-enhancing technologies, AI security, synthetic data, de-identification and privacy-preserving computation.
This role will be fully hybrid with the expectation we get together in our Holborn, London office at least once per month.
Requirements
We welcome applications from all who may not feel they match the full criteria, so if you have most of the below, we'd like to hear from you:
Significant experience applying data science, machine learning, statistical modelling or advanced analytics to complex real-world datasets.
Strong applied statistical expertise, including the ability to quantify risk and uncertainty, evaluate assumptions, design validation approaches, interpret imperfect or incomplete evidence, and communicate the limitations of statistical or machine learning models.
Strong Python skills and experience writing maintainable, production-quality code.
Experience working in cross-functional teams with software engineers, data engineers or platform teams to design and deliver data products, pipelines, analytical services or decision-support tools.
A user-focused approach to technical delivery: you are comfortable working with people who operate, review, govern or depend on data systems, and can translate their needs into technical requirements.
Applied machine learning experience and understanding of common privacy attacks against data and models, such as memorisation, membership inference, attribute inference, model inversion or leakage through model outputs.
Exposure to privacy-preserving machine learning, privacy-enhancing technologies or adjacent research areas (e.g. federated learning, secure aggregation, differential privacy, or confidential computing)
Experience working with sensitive, confidential or regulated data and a strong understanding of privacy, confidentiality or information security risks.
Ability to translate ambiguous operational, governance or security problems into clear data science questions and practical technical requirements.
Good communication skills, with the ability to explain complex technical concepts to non-specialist stakeholders.
A pragmatic, delivery-focused mindset
It would be a bonus if you have any of the following experience -
Working with health data or biomedical research data, electronic health records, and ideally genomic data
Developing models, algorithms or rule-based systems that support human decision-making, ideally where explainability, auditability and risk management are important
Anomaly detection, behavioural analytics, security monitoring or detection engineering.
Risk quantification methods from fields such as actuarial science, epidemiology, operational research, or cyber risk
Familiarity with UK data protection, research governance or health data access expectations.
Experience with PETs or privacy-preserving ML frameworks such as Flower, Opacus, TensorFlow Federated or similar
Cloud platforms, containerisation, CI/CD, MLOps or production ML systems.
Experience evaluating synthetic data using both utility and privacy metrics.
| Organization | Our Future Health UK |
| Industry | IT / Telecom / Software Jobs |
| Occupational Category | Senior Data Scientist |
| Job Location | London,UK |
| Shift Type | Morning |
| Job Type | Full Time |
| Gender | No Preference |
| Career Level | Intermediate |
| Experience | 2 Years |
| Posted at | 2026-08-24 10:45 pm |
| Expires on | 2026-10-08 |