Christian Langridge
PhD student · University College London and Queen Mary University of London
Emeritus fellowFellowship research programme
Unaite Fellowship | Project proposals and methodological notes
UNAITE FELLOWSHIP · JÉRÉMIE KALFON
Three proposals: application-oriented evaluation, open biological knowledge graphs, and multimodal integration.
Do pretrained models improve clinically and experimentally relevant predictions?
Contribution: datasets, evaluation protocols, and comparative results.How can heterogeneous biological evidence be made reproducible and directly usable?
Contribution: an open graph pipeline, baseline GNNs, and agent access.When does combining genetics, cell states, perturbations, and structures improve prediction?
Contribution: an integration method and evidence of its incremental value.These are research proposals, not completed studies. The first concerns evaluation methodology, the second reusable open-source infrastructure, and the third an open modelling question. They can be pursued independently; the knowledge graph could also provide structured prior information for the integration project. The evaluation designs below are starting points for discussion, not commitments to specific datasets or experimental collaborations.
RESEARCH GROUP
PhD student · University College London and Queen Mary University of London
Emeritus fellowMaster's student · Fabian Theis lab, Helmholtz Munich
MVA, ENS Paris-Saclay & Université Paris-Cité
PhD student · University of Cambridge
Master's student · ENS Paris-Saclay
PhD student · RCSI University of Medicine and Health Sciences
Furney Lab / Genomic Oncology Research Group
PROJECT SELECTION
I believe students should be able to develop their own research questions and bring their own project proposals.
The three directions presented here are options. Project selection should reflect the student's interests, experience, and scientific objectives.
PROPOSAL 01 · BENCHMARKS
Evaluate both intermediate capabilities and downstream utility: better representations do not automatically imply better biomedical predictions.
| Candidate task | Evaluation target |
|---|---|
| Intermediate model capabilities | |
| Embedding quality | Biological structure and transfer to unseen data |
| Denoising | Reconstruction accuracy and preservation of biological signal |
| Perturbation prediction | Expression changes under held-out interventions |
| Clinical, therapeutic, and experimental applications | |
| Survival, diagnosis, traits | Discrimination, prediction error, and calibration in external cohorts |
| Therapies, combinations, protocols | Efficacy, adverse effects, and measured experimental outcomes |
Comparison: task-specific baseline → frozen pretrained representation → fine-tuned model; matched labels, splits, and tuning budgets.
Single-cell evaluations motivate explicit baseline and transfer tests. [1] [2] Therapeutics Data Commons provides relevant prior art. [3]
The benchmark covers two complementary levels. Intermediate evaluations test embedding quality, denoising, and perturbation-response prediction. Application-level evaluations test survival analysis, disease diagnosis, trait prediction, therapy selection, combination safety, and protocol design. The aim is to measure whether gains at the first level translate into gains at the second; an intermediate score is not a substitute for a downstream outcome.
Embedding evaluations should distinguish biological structure from technical batch information and measure transfer to unseen datasets. Denoising requires a defined reference, such as technical replicates or controlled count thinning, and checks that rare states and biological variation are preserved rather than smoothed away. Perturbation prediction should evaluate changes relative to matched controls and hold out entire interventions where the claim is intervention generalization. Such predictions may themselves be useful in experimental work while remaining upstream of patient outcomes.
The benchmark should begin with a defined user and decision: for example, prognosis from patient-level molecular profiles, prediction of an adverse event for a drug pair, or selection of an experimental protocol. Different tasks require different modalities and foundation-model families. A transcriptomic model should not be assumed to address every endpoint in this list.
For survival analysis, report a censoring-aware discrimination measure, time-dependent Brier scores, and calibration at clinically meaningful horizons. For diagnosis, report AUROC and AUPRC alongside calibration and performance at a prespecified operating point. Continuous traits require regression metrics and external validation. These are proposed evaluation choices; no dataset or endpoint has yet been selected.
The split must match the generalization claim. Patient-level tasks require patient separation; site or time holdouts test further distribution shifts. Drug-pair holdouts test new combinations, while held-out compounds test a stronger extrapolation setting. Unlabelled drug-event pairs must not automatically be interpreted as verified negatives. Combination efficacy or synergy is not a substitute for combination safety.
Compare a well-tuned conventional method, frozen foundation-model features with a downstream predictor, and fine-tuning where feasible. Include ablations of pretraining when computationally affordable. Report tuning budgets, labelled sample sizes, runtime, and uncertainty with resampling at the independent experimental unit. Audit overlap with pretraining corpora where documented; otherwise disclose that contamination cannot be ruled out.
Protocol design requires observable experimental outcomes, constraints, or a validated expert assessment. Without these, evaluating generated text cannot establish improved experimental performance. The first release should cover only tasks for which data rights, reliable endpoints, and appropriate baselines can be established.
Boiarsky et al. compare scBERT with logistic regression on cell-type annotation and perform pretraining ablations. Kedzierska et al. evaluate Geneformer and scGPT in zero-shot settings. These findings concern specific models and tasks; they do not establish that foundation models are generally ineffective. Therapeutics Data Commons already organizes therapeutics datasets and evaluation tasks. The proposed contribution must be defined by the additional endpoints, evaluation settings, or comparative evidence it supplies.
PROPOSAL 02 · OPEN-SOURCE INFRASTRUCTURE
Publish the complete path from source records to versioned graph data, baseline training, and evidence-linked retrieval.
Versioned sources and licenses; explicit identifier mappings and transformations; evidence and context retained per relation.
Graph exports, documented GNN baselines and task splits; agent queries that return source records with their answers.
PrimeKG provides an existing graph and construction pipeline; TxGNN demonstrates graph-based drug repurposing. [4] [5]
The working hypothesis is that a maintained and documented integration pipeline can reduce the effort required to use biological graphs in machine learning and agent workflows. This is not a claim that no usable biological knowledge graph exists. PrimeKG already provides integrated data and reproducible construction instructions; TxGNN already supplies a substantial GNN application. Comparison with such resources is necessary before making a novelty claim.
Each source adapter should record the source release, retrieval date, license, input hash, identifier namespace, filtering rules, and transformation version. Preserve evidence-level records even when several sources describe the same relation. Distinguish absent information, explicit negative evidence, contradictory evidence, and model-derived predictions. Retain species, tissue or cell context, dose, and time where provided; missing fields should remain explicitly missing.
The intended release includes source code, a data manifest, a typed graph schema, machine-learning exports, at least one baseline GNN configuration, and examples of agent retrieval. Some source licenses may allow construction code without allowing data redistribution. The release manifest should state those restrictions at source level.
Evaluate identifier-mapping coverage, retained provenance, reproducibility from pinned inputs, update latency, and training reproducibility. For link prediction, remove held-out edges and their inverse or duplicated equivalents before message passing; select time-, entity-, or relation-based splits according to the claim. Random-edge validation alone does not demonstrate prediction for unseen drugs or diseases. Agent evaluation should separately measure retrieval relevance, source support, and answer correctness. A plausible graph path is not itself evidence of a causal mechanism.
PROPOSAL 03 · OPEN MODELLING PROBLEM
Can explicit biological relationships improve prediction beyond concatenating representations from separate foundation models?
Predict the direction and magnitude of expression changes for held-out gene perturbations in a defined cell context.
Single modalities, concatenated embeddings, and joint modelling; evidence-source ablations and uncertainty estimates.
Prior work: genetic effects + Perturb-seq (Ota et al.). [6] Complex structures (AlphaFold 3). [7] Joint structure and binding-affinity prediction (Boltz-2, preprint). [8]
Sequence models describe sequence-dependent properties. GWAS relates variation to phenotypes but does not by itself resolve the causal variant, target gene, or mechanism. Fine-mapping and regulatory evidence can help define uncertain variant-to-gene links. RNA-seq measures cellular states; Perturb-seq measures responses to interventions in particular contexts. Trajectories inferred from cross-sectional data require additional assumptions to be interpreted temporally. Co-folding supplies hypotheses about biomolecular complexes; binding models and assays add evidence about interaction affinity. Predicted affinity must be distinguished from experimentally measured affinity, with assay conditions, units, uncertainty, and model coverage recorded. Neither a plausible complex nor high predicted affinity establishes a cellular treatment effect.
The proposal is to investigate shared entities and structured relationships, not just feature concatenation. One possible architecture uses a gene/protein-centred graph, context-dependent regulatory relationships, and separate observation models for each modality. This is a candidate design, not an architecture selected in advance. Spatial transcriptomics could contribute tissue context where compatible data exist; it is not a prerequisite for an initial study.
A first comparison could hold out entire gene perturbations and predict expression effects in a fixed cell context. Context transfer would require an additional holdout of cell types or experimental settings. Single-modality and concatenation baselines should use the same labels and splits; model-capacity and compute differences should be reported. Remove each evidence source, including predicted structures, to establish whether it contributes information. Evaluate changes relative to control as well as absolute expression so that baseline cell identity cannot dominate the score.
A positive result would establish predictive usefulness in the selected setting. It would not, by itself, identify a complete causal mechanism or establish clinical treatment benefit. If structural evidence provides no additional value, that negative result should constrain the proposed integration rather than be obscured by the aggregate performance of other modalities.
Ota et al. combine gene-level loss-of-function burden effects with Perturb-seq regulatory information to study blood traits. This is a specific precedent for linking genetic and perturbational evidence, not a demonstration of the full integration proposed here. AlphaFold 3 addresses biomolecular structure prediction; Boltz-2 jointly models structure and protein-small-molecule binding affinity. The proposed use of structural and binding evidence in cellular-response modelling remains to be tested.
Fellow names and affiliations, the existing knowledge-graph codebase, and the possible Scienta Labs collaboration were supplied by Jérémie Kalfon. Publication and collaboration plans are tentative. Methods and evaluation criteria here are proposed refinements, not completed experiments. All diagrams are original conceptual schematics; they show no measured results.
The bibliography supports prior-work claims, not validation of the proposed projects. References in presentation mode are derived from citations on all six slides.