UnaiteResearch proposals

Fellowship research programme

Research projects in machine learning for biology

Unaite Fellowship | Project proposals and methodological notes

Research projects in machine learning for biology

Three proposals: application-oriented evaluation, open biological knowledge graphs, and multimodal integration.

01

Foundation-model evaluation

Do pretrained models improve clinically and experimentally relevant predictions?

Contribution: datasets, evaluation protocols, and comparative results.
02

Biological knowledge infrastructure

How can heterogeneous biological evidence be made reproducible and directly usable?

Contribution: an open graph pipeline, baseline GNNs, and agent access.
03

Integration across biological scales

When does combining genetics, cell states, perturbations, and structures improve prediction?

Contribution: an integration method and evidence of its incremental value.

Proposed research directions · Project scope and methods remain to be defined with the fellows.

Scope of this presentation

These are research proposals, not completed studies. The first concerns evaluation methodology, the second reusable open-source infrastructure, and the third an open modelling question. They can be pursued independently; the knowledge graph could also provide structured prior information for the integration project. The evaluation designs below are starting points for discussion, not commitments to specific datasets or experimental collaborations.

Fellows

Christian Langridge

PhD student · University College London and Queen Mary University of London

Emeritus fellow

Raphaël Rubrice

Master's student · Fabian Theis lab, Helmholtz Munich
MVA, ENS Paris-Saclay & Université Paris-Cité

Emeritus fellow

Jacob Gonzale Isa

PhD student · University of Cambridge

Amine Ould

Master's student · ENS Paris-Saclay

Faith Ogundimu

PhD student · RCSI University of Medicine and Health Sciences
Furney Lab / Genomic Oncology Research Group

Students may propose their own projects

I believe students should be able to develop their own research questions and bring their own project proposals.

The three directions presented here are options. Project selection should reflect the student's interests, experience, and scientific objectives.

Scientific question
What is unknown, and what result would change our understanding?
Feasibility
Which data, supervision, computational resources, and experimental support are available?
Evaluation
Which comparison would establish a contribution, including an informative negative result?

The research question, initial scope, and expected outputs will be defined jointly.

Benchmarks from model capabilities to biomedical applications

Evaluate both intermediate capabilities and downstream utility: better representations do not automatically imply better biomedical predictions.

Candidate task Evaluation target
Intermediate model capabilities
Embedding quality Biological structure and transfer to unseen data
Denoising Reconstruction accuracy and preservation of biological signal
Perturbation prediction Expression changes under held-out interventions
Clinical, therapeutic, and experimental applications
Survival, diagnosis, traits Discrimination, prediction error, and calibration in external cohorts
Therapies, combinations, protocols Efficacy, adverse effects, and measured experimental outcomes
Illustrative donor-level split: all cells from donors A and B are training data; all cells from donor C are held out. The same held-out cohort evaluates a task-specific baseline and a pretrained model.
Example: hold out patients, not individual cells from the same patient.

Comparison: task-specific baseline → frozen pretrained representation → fine-tuned model; matched labels, splits, and tuning budgets.

Single-cell evaluations motivate explicit baseline and transfer tests. [1] [2] Therapeutics Data Commons provides relevant prior art. [3]

Possible Scienta Labs collaboration and NeurIPS submission: exploratory, not confirmed.

Proposed study design

The benchmark covers two complementary levels. Intermediate evaluations test embedding quality, denoising, and perturbation-response prediction. Application-level evaluations test survival analysis, disease diagnosis, trait prediction, therapy selection, combination safety, and protocol design. The aim is to measure whether gains at the first level translate into gains at the second; an intermediate score is not a substitute for a downstream outcome.

Embedding evaluations should distinguish biological structure from technical batch information and measure transfer to unseen datasets. Denoising requires a defined reference, such as technical replicates or controlled count thinning, and checks that rare states and biological variation are preserved rather than smoothed away. Perturbation prediction should evaluate changes relative to matched controls and hold out entire interventions where the claim is intervention generalization. Such predictions may themselves be useful in experimental work while remaining upstream of patient outcomes.

The benchmark should begin with a defined user and decision: for example, prognosis from patient-level molecular profiles, prediction of an adverse event for a drug pair, or selection of an experimental protocol. Different tasks require different modalities and foundation-model families. A transcriptomic model should not be assumed to address every endpoint in this list.

For survival analysis, report a censoring-aware discrimination measure, time-dependent Brier scores, and calibration at clinically meaningful horizons. For diagnosis, report AUROC and AUPRC alongside calibration and performance at a prespecified operating point. Continuous traits require regression metrics and external validation. These are proposed evaluation choices; no dataset or endpoint has yet been selected.

The split must match the generalization claim. Patient-level tasks require patient separation; site or time holdouts test further distribution shifts. Drug-pair holdouts test new combinations, while held-out compounds test a stronger extrapolation setting. Unlabelled drug-event pairs must not automatically be interpreted as verified negatives. Combination efficacy or synergy is not a substitute for combination safety.

Compare a well-tuned conventional method, frozen foundation-model features with a downstream predictor, and fine-tuning where feasible. Include ablations of pretraining when computationally affordable. Report tuning budgets, labelled sample sizes, runtime, and uncertainty with resampling at the independent experimental unit. Audit overlap with pretraining corpora where documented; otherwise disclose that contamination cannot be ruled out.

Protocol design requires observable experimental outcomes, constraints, or a validated expert assessment. Without these, evaluating generated text cannot establish improved experimental performance. The first release should cover only tasks for which data rights, reliable endpoints, and appropriate baselines can be established.

Relation to existing evaluations

Boiarsky et al. compare scBERT with logistic regression on cell-type annotation and perform pretraining ablations. Kedzierska et al. evaluate Geneformer and scGPT in zero-shot settings. These findings concern specific models and tasks; they do not establish that foundation models are generally ineffective. Therapeutics Data Commons already organizes therapeutics datasets and evaluation tasks. The proposed contribution must be defined by the additional endpoints, evaluation settings, or comparative evidence it supplies.

A reproducible biological knowledge graph for GNNs and agents

Publish the complete path from source records to versioned graph data, baseline training, and evidence-linked retrieval.

Source records become a typed graph of compounds, proteins, genes, and diseases. Each relation retains provenance and context. Outputs are GNN training and evidence-linked agent retrieval.Sources with provenance map to a typed biological graph, then to GNN training and agent retrieval with supporting records.
Illustrative schema. Relation labels distinguish molecular interactions, associations, and clinical evidence.

Reproducible construction

Versioned sources and licenses; explicit identifier mappings and transformations; evidence and context retained per relation.

Usable reference implementations

Graph exports, documented GNN baselines and task splits; agent queries that return source records with their answers.

PrimeKG provides an existing graph and construction pipeline; TxGNN demonstrates graph-based drug repurposing. [4] [5]

Starting point: an existing codebase. Contribution to establish: maintenance, provenance, and a reproducible training-and-retrieval package.

Research and engineering scope

The working hypothesis is that a maintained and documented integration pipeline can reduce the effort required to use biological graphs in machine learning and agent workflows. This is not a claim that no usable biological knowledge graph exists. PrimeKG already provides integrated data and reproducible construction instructions; TxGNN already supplies a substantial GNN application. Comparison with such resources is necessary before making a novelty claim.

Each source adapter should record the source release, retrieval date, license, input hash, identifier namespace, filtering rules, and transformation version. Preserve evidence-level records even when several sources describe the same relation. Distinguish absent information, explicit negative evidence, contradictory evidence, and model-derived predictions. Retain species, tissue or cell context, dose, and time where provided; missing fields should remain explicitly missing.

Deliverables and evaluation

The intended release includes source code, a data manifest, a typed graph schema, machine-learning exports, at least one baseline GNN configuration, and examples of agent retrieval. Some source licenses may allow construction code without allowing data redistribution. The release manifest should state those restrictions at source level.

Evaluate identifier-mapping coverage, retained provenance, reproducibility from pinned inputs, update latency, and training reproducibility. For link prediction, remove held-out edges and their inverse or duplicated equivalents before message passing; select time-, entity-, or relation-based splits according to the claim. Random-edge validation alone does not demonstrate prediction for unseen drugs or diseases. Agent evaluation should separately measure retrieval relevance, source support, and answer correctness. A plausible graph path is not itself evidence of a causal mechanism.

Joint modelling of genetics, perturbations, and molecular structure

Can explicit biological relationships improve prediction beyond concatenating representations from separate foundation models?

Sequence and GWAS, RNA-seq and perturbations, and co-folding plus binding-affinity predictions supply different evidence types. Compare concatenated embeddings with a joint relational model on the same held-out intervention effects.Genetics, cellular perturbations, and co-folding plus binding evidence feed two strategies: concatenated embeddings and a joint relational model. Both predict the same held-out perturbations.
Arrows denote proposed information flow. Associations, interventions, structures, and binding affinities retain distinct evidential status.

Candidate test

Predict the direction and magnitude of expression changes for held-out gene perturbations in a defined cell context.

Required comparisons

Single modalities, concatenated embeddings, and joint modelling; evidence-source ablations and uncertainty estimates.

Prior work: genetic effects + Perturb-seq (Ota et al.). [6] Complex structures (AlphaFold 3). [7] Joint structure and binding-affinity prediction (Boltz-2, preprint). [8]

Open issue: matched biological contexts and a task on which structure and binding evidence add measurable information.

What must be integrated?

Sequence models describe sequence-dependent properties. GWAS relates variation to phenotypes but does not by itself resolve the causal variant, target gene, or mechanism. Fine-mapping and regulatory evidence can help define uncertain variant-to-gene links. RNA-seq measures cellular states; Perturb-seq measures responses to interventions in particular contexts. Trajectories inferred from cross-sectional data require additional assumptions to be interpreted temporally. Co-folding supplies hypotheses about biomolecular complexes; binding models and assays add evidence about interaction affinity. Predicted affinity must be distinguished from experimentally measured affinity, with assay conditions, units, uncertainty, and model coverage recorded. Neither a plausible complex nor high predicted affinity establishes a cellular treatment effect.

The proposal is to investigate shared entities and structured relationships, not just feature concatenation. One possible architecture uses a gene/protein-centred graph, context-dependent regulatory relationships, and separate observation models for each modality. This is a candidate design, not an architecture selected in advance. Spatial transcriptomics could contribute tissue context where compatible data exist; it is not a prerequisite for an initial study.

Demonstrating incremental value

A first comparison could hold out entire gene perturbations and predict expression effects in a fixed cell context. Context transfer would require an additional holdout of cell types or experimental settings. Single-modality and concatenation baselines should use the same labels and splits; model-capacity and compute differences should be reported. Remove each evidence source, including predicted structures, to establish whether it contributes information. Evaluate changes relative to control as well as absolute expression so that baseline cell identity cannot dominate the score.

A positive result would establish predictive usefulness in the selected setting. It would not, by itself, identify a complete causal mechanism or establish clinical treatment benefit. If structural evidence provides no additional value, that negative result should constrain the proposed integration rather than be obscured by the aggregate performance of other modalities.

Relevant precedent and limitations

Ota et al. combine gene-level loss-of-function burden effects with Perturb-seq regulatory information to study blood traits. This is a specific precedent for linking genetic and perturbational evidence, not a demonstration of the full integration proposed here. AlphaFold 3 addresses biomolecular structure prediction; Boltz-2 jointly models structure and protein-small-molecule binding affinity. The proposed use of structural and binding evidence in cellular-response modelling remains to be tested.

Sources and scope

Fellow names and affiliations, the existing knowledge-graph codebase, and the possible Scienta Labs collaboration were supplied by Jérémie Kalfon. Publication and collaboration plans are tentative. Methods and evaluation criteria here are proposed refinements, not completed experiments. All diagrams are original conceptual schematics; they show no measured results.

The bibliography supports prior-work claims, not validation of the proposed projects. References in presentation mode are derived from citations on all six slides.

Editable Markdown · Build metadata

References

  1. Boiarsky et al. (2024). Deeper evaluation of a single-cell foundation model. Nature Machine Intelligence 6, 1443-1446.
  2. Kedzierska et al. (2025). Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biology 26, 101.
  3. Huang et al. (2021). Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development. NeurIPS, Datasets and Benchmarks.
  4. Chandak, Huang & Zitnik (2023). Building a knowledge graph to enable precision medicine. Scientific Data 10, 67.
  5. Huang et al. (2024). A foundation model for clinician-centered drug repurposing. Nature Medicine 30, 3601-3613.
  6. Ota et al. (2026; online 2025). Causal modelling of gene effects from regulators to programs to traits. Nature 650, 399-408.
  7. Abramson et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493-500.
  8. Passaro et al. (2025). Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. bioRxiv preprint.