<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://www.jkobject.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://www.jkobject.com/" rel="alternate" type="text/html" /><updated>2026-07-13T12:44:00+00:00</updated><id>https://www.jkobject.com/feed.xml</id><title type="html">Jérémie</title><subtitle>My website and blog.</subtitle><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><entry><title type="html">The Gene Regulation Landscape</title><link href="https://www.jkobject.com/blog/gene-regulation-landscape/" rel="alternate" type="text/html" title="The Gene Regulation Landscape" /><published>2026-06-10T00:00:00+00:00</published><updated>2026-06-10T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/gene-regulation-landscape</id><content type="html" xml:base="https://www.jkobject.com/blog/gene-regulation-landscape/"><![CDATA[<p>I have always wanted to understand how a cell works.</p>

<p>Most of the time, even when textbooks go into molecular detail, the story is
still organized around dogmas. One chapter adds chromatin. Another adds
transcription factors. Later, RNA processing appears. Then translation, protein
degradation, signaling, condensates, non-coding RNAs, and so on. Each new
mechanism is real, but they often arrive as separate layers of complexity, not
as one formal picture of what we collectively know.</p>

<p>That is what I tried to build here: a summary of our current formal knowledge of
gene regulation, placed into one landscape.</p>

<p>While doing it, I noticed something that surprised me. Biology has many
measurements and many local names, but sometimes no clean conceptual object for
things that are probably views of the same underlying cellular structure. For
example, super-enhancers in ChIP-seq and transcriptional condensates in
microscopy are not strictly identical, but they are clearly not unrelated
either. In other places, molecular biology has long descriptive sentences for
mechanisms, but no short handle that makes the mechanism easy to reason about.</p>

<p>So I gave each main mechanism a short code name. The names are not meant to
replace the biology; they are handles that point back to precise glossary
entries. I hope the map is useful, and maybe sparks discussions.</p>

<p>The map is very large, so the embedded version below is mostly a preview. You can
<a href="/assets/images/gene-reg-v7.png">open the full-resolution zoomable map here</a>.
The companion <a href="/assets/documents/gene-regulation-landscape-details.md">technical notes</a>
contain the 1-to-1 glossary for every box name, the full mechanism catalogue,
link rationale, and legend details. The <a href="/assets/documents/gene-reg-v7.dot">Graphviz DOT source</a>
is also available.
A node-by-node <a href="/assets/documents/gene-regulation-node-audit/SYNTHESIS.md">literature audit</a>
now reviews the biological support, caveats, and proposed graph revisions for
all 37 mechanism boxes. Several arrows were added or softened after that audit:
for example, the cap now links to mRNA stability, growth and translation state
now regulate decoding tempo, codon optimality and translation surveillance now
link to mRNA decay, chaperone triage links back to ubiquitin routing, signaling
links to CRL licensing, and nuclear lncRNA guides link to repressive histone writers.
Other arrows are deliberately dashed because the evidence is contextual rather
than universal.</p>

<p>In the figure, each box title is a code name that maps 1-to-1 to the catalogue
entry in the technical notes. Solid arrows are the main mechanistic relations;
dashed arrows indicate contextual, feedback, or association-style links;
tee-headed arrows indicate repression; bold arrows mark especially important
coupling edges. Border styles mark meta-principles such as <code class="language-plaintext highlighter-rouge">LLPS</code> and <code class="language-plaintext highlighter-rouge">DECAY</code>:
these are recurring physical or regulatory motifs that appear across several
mechanisms, not separate boxes in the pathway.</p>

<p>The map is also intentionally reductionist. I did not try to draw every
important named event as its own box. Some biological phenomena are real and
important, but they are better understood as outputs of several lower-level
mechanisms already shown in the graph. For example, promoter/TSS choice depends
on transcription-factor grammar, chromatin accessibility, enhancer contacts,
and Pol II initiation/pausing. Transcription termination depends on Pol II
state, cleavage/polyadenylation, chromatin context, and RNA decay machinery.
Whether an RNA leaves the nucleus depends on capping, splicing, 3’ processing,
RNA-binding proteins, RNA marks, and quality control. In those cases, the goal
is not to add a new named event whenever biology has a label for one. The goal
is to ask whether the underlying decision mechanisms and their links are already
represented.</p>

<p>This is also why some familiar topics, such as R-loops, DNA torsional stress,
transcription-replication conflicts, nuclear bodies, broad epitranscriptomic
marks beyond m6A, repeats/transposons, motif grammar, and Pol I/Pol III
rRNA/tRNA biology are handled as caveats, extensions, or submechanisms in the
technical notes rather than as new boxes in the main landscape. That choice is
not a claim that they are unimportant. It is a claim about map resolution: this
figure prioritizes reusable regulatory mechanisms over named composite events.</p>

<p>Post-translational control is shown with representative feedback routes rather
than every substrate-specific event. <code class="language-plaintext highlighter-rouge">SWITCH</code>, <code class="language-plaintext highlighter-rouge">ROUTER</code>, and <code class="language-plaintext highlighter-rouge">DESTROY</code> can touch
many upstream programs; the map draws the main recurring routes to TF activity,
Pol II, stress translation, proteostasis, and mTOR/NF-κB-style feedback.
<code class="language-plaintext highlighter-rouge">LICENSE</code> is intentionally narrower: NEDDylation mainly activates cullin-RING E3
ligases, so its main graph role is to feed <code class="language-plaintext highlighter-rouge">ROUTER</code>.</p>

<p><a href="/assets/images/gene-reg-v7.png"><img src="/assets/images/gene-reg-v7.png" alt="Gene Regulation Landscape" /></a></p>

<h2 id="box-glossary">Box Glossary</h2>

<p>Each box in the image has exactly one entry here. Box titles are the short code
names; italic text inside the image marks genes or proteins when the label needs
examples.</p>

<table>
  <thead>
    <tr>
      <th>Code</th>
      <th>Layer</th>
      <th>Meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">ZONES</code></td>
      <td>3D genome</td>
      <td>A/B chromatin compartments.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">FENCES</code></td>
      <td>3D genome</td>
      <td>TAD boundaries and insulation by CTCF/cohesin.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">BRIDGES</code></td>
      <td>3D genome</td>
      <td>Enhancer-promoter loops.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">HUBS</code></td>
      <td>3D genome</td>
      <td>Super-enhancers / enhancer hubs, with Pol II/coactivator condensates treated as a related but not identical physical model.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">SILENCER</code></td>
      <td>Epigenetics</td>
      <td>DNA methylation and repressive chromatin memory.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">OPENER</code></td>
      <td>Epigenetics</td>
      <td>Histone acetylation that opens chromatin.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">WRITER-A</code></td>
      <td>Epigenetics</td>
      <td>Activating histone methylation marks.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">WRITER-R</code></td>
      <td>Epigenetics</td>
      <td>Repressive histone methylation marks.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">SHUFFLER</code></td>
      <td>Epigenetics</td>
      <td>ATP-dependent nucleosome remodeling.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">GUIDES</code></td>
      <td>Epigenetics</td>
      <td>ncRNAs that recruit chromatin regulators.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">KEYS</code></td>
      <td>Transcription</td>
      <td>Transcription factors, including pioneer factors.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">SCRIBE</code></td>
      <td>Transcription</td>
      <td>Pol II pausing, release, and CTD phosphorylation.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">SHIELD</code></td>
      <td>Co-transcriptional</td>
      <td>5’ capping and cap-dependent protection/export.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">SPLICER</code></td>
      <td>Co-transcriptional</td>
      <td>Alternative splicing coupled to Pol II kinetics.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">TRIMMER</code></td>
      <td>Co-transcriptional</td>
      <td>Alternative polyadenylation and 3’UTR choice.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">RECODER</code></td>
      <td>Co-transcriptional</td>
      <td>A-to-I RNA editing by ADARs.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">STAMP</code></td>
      <td>Post-transcriptional</td>
      <td>m6A RNA marking and reader-dependent fate choices.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">READERS</code></td>
      <td>Post-transcriptional</td>
      <td>RNA-binding proteins that tune RNA processing, stability, localization, and translation.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">DARTS</code></td>
      <td>Post-transcriptional</td>
      <td>miRNA-RISC targeting.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">SPONGE</code></td>
      <td>Post-transcriptional</td>
      <td>Cytoplasmic lncRNA/circRNA competition with miRNAs, kept as a strongly stoichiometry- and localization-dependent mechanism.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">CENSOR</code></td>
      <td>Post-transcriptional</td>
      <td>Nonsense-mediated mRNA decay.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">TIMER</code></td>
      <td>Post-transcriptional</td>
      <td>mRNA half-life, deadenylation, decapping, and decay.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">CLIPS</code></td>
      <td>Post-transcriptional</td>
      <td>RNA G-quadruplex structures that affect scanning and translation.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">VAULT</code></td>
      <td>Post-transcriptional</td>
      <td>Stress granules and P-bodies for RNA storage or decay.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">FORGE</code></td>
      <td>Translation</td>
      <td>Starting cap-dependent translation through mTOR/eIF4F.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">BRAKE</code></td>
      <td>Translation</td>
      <td>Slowing global translation during stress through ISR/eIF2α-P.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">DECOY</code></td>
      <td>Translation</td>
      <td>uORFs that divert scanning ribosomes and gate main ORF translation.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">BYPASS</code></td>
      <td>Translation</td>
      <td>Non-canonical initiation routes; viral IRESs are robust, while many cellular IRES-like claims need strict controls.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">TEMPO</code></td>
      <td>Translation</td>
      <td>Decoding kinetics, tRNA/codon effects, ribosome state, and their regulation by growth, initiation load, and stress context. Also covers decoding <em>fidelity</em>: alternate decoding can install non-genomic amino-acid substitutions that yield stable, abundant proteoforms.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">INSPECTOR</code></td>
      <td>Translation</td>
      <td>RQC/NGD/NSD surveillance of stalled or broken translation.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">SWITCH</code></td>
      <td>Post-translational</td>
      <td>Phosphorylation/O-GlcNAc switches for protein activity and interactions.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">ROUTER</code></td>
      <td>Post-translational</td>
      <td>Ubiquitin-chain logic that routes proteins to signaling, proteasome, or autophagy.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">TETHER</code></td>
      <td>Post-translational</td>
      <td>SUMOylation that tethers proteins into nuclear complexes, repression modules, or repair assemblies.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">LICENSE</code></td>
      <td>Post-translational</td>
      <td>Neddylation that licenses cullin-RING E3 ubiquitin ligases, regulated by CRL assembly and signaling context.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">DESTROY</code></td>
      <td>Post-translational</td>
      <td>Protein clearance through two fused outputs: proteasome and selective autophagy.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">MATURE</code></td>
      <td>Post-translational</td>
      <td>Protein folding, refolding, triage, and ER-stress UPR.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">PAR</code></td>
      <td>Post-translational</td>
      <td>PARP/PAR signaling at DNA damage and repair condensates.</td>
    </tr>
  </tbody>
</table>

<h2 id="term-glossary">Term Glossary</h2>

<p>These are the main non-gene, non-protein terms used in the figure and glossary.</p>

<table>
  <thead>
    <tr>
      <th>Term</th>
      <th>Meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>A/B compartments</td>
      <td>Large Hi-C chromatin domains; A is generally active/euchromatic, B is generally inactive/heterochromatic.</td>
    </tr>
    <tr>
      <td>TAD</td>
      <td>Topologically associating domain; a chromatin neighborhood insulated by boundaries such as CTCF/cohesin sites.</td>
    </tr>
    <tr>
      <td>Enhancer-promoter loop</td>
      <td>A 3D contact that brings a distal enhancer near a target promoter.</td>
    </tr>
    <tr>
      <td>Super-enhancer</td>
      <td>A dense enhancer cluster with high transcription-factor, Mediator, BRD4, and Pol II occupancy; an operational enhancer annotation, not automatically proof of a condensate.</td>
    </tr>
    <tr>
      <td>Condensate / LLPS</td>
      <td>Liquid-liquid phase separation; concentration of molecules into a dense phase without a membrane.</td>
    </tr>
    <tr>
      <td>DNA methylation</td>
      <td>Addition of methyl groups to cytosines, often linked to transcriptional repression at promoters.</td>
    </tr>
    <tr>
      <td>Histone mark</td>
      <td>A chemical modification on histones, such as acetylation or methylation, read by chromatin proteins.</td>
    </tr>
    <tr>
      <td>Chromatin remodeling</td>
      <td>ATP-driven repositioning, eviction, or exchange of nucleosomes.</td>
    </tr>
    <tr>
      <td>ncRNA</td>
      <td>Non-coding RNA; RNA that functions without being translated into protein.</td>
    </tr>
    <tr>
      <td>Pol II pausing</td>
      <td>Promoter-proximal RNA polymerase II stalling before productive elongation.</td>
    </tr>
    <tr>
      <td>CTD phosphorylation</td>
      <td>Phosphorylation of the Pol II C-terminal domain, coordinating transcription with RNA processing.</td>
    </tr>
    <tr>
      <td>5’ capping</td>
      <td>Addition of an m7G cap to nascent RNA, protecting it and enabling export/translation.</td>
    </tr>
    <tr>
      <td>Alternative splicing</td>
      <td>Regulated exon choice that produces multiple transcript isoforms from one gene.</td>
    </tr>
    <tr>
      <td>Alternative polyadenylation</td>
      <td>Choice of different cleavage/polyA sites, often changing 3’UTR length.</td>
    </tr>
    <tr>
      <td>A-to-I editing</td>
      <td>Adenosine-to-inosine RNA editing; inosine is read like guanosine by many machines.</td>
    </tr>
    <tr>
      <td>m6A</td>
      <td>N6-methyladenosine, a reversible RNA modification interpreted by reader proteins.</td>
    </tr>
    <tr>
      <td>RBP</td>
      <td>RNA-binding protein.</td>
    </tr>
    <tr>
      <td>miRNA-RISC</td>
      <td>MicroRNA loaded into the RISC complex to repress or destabilize target RNAs.</td>
    </tr>
    <tr>
      <td>ceRNA</td>
      <td>Competing endogenous RNA; an RNA that can buffer miRNAs by sharing target sites, when abundance, affinity, and colocalization are sufficient.</td>
    </tr>
    <tr>
      <td>NMD</td>
      <td>Nonsense-mediated decay; surveillance and degradation of transcripts with premature stop codons.</td>
    </tr>
    <tr>
      <td>Deadenylation / decapping</td>
      <td>Removal of the polyA tail and 5’ cap, usually committing an mRNA to decay.</td>
    </tr>
    <tr>
      <td>RNA G-quadruplex</td>
      <td>A guanine-rich RNA structure that can block scanning or alter RNA fate.</td>
    </tr>
    <tr>
      <td>Stress granule / P-body</td>
      <td>Cytoplasmic RNA-protein condensates involved in RNA storage, repression, or decay.</td>
    </tr>
    <tr>
      <td>Cap-dependent translation</td>
      <td>Canonical translation initiation through cap recognition and 5’UTR scanning.</td>
    </tr>
    <tr>
      <td>ISR / eIF2α-P</td>
      <td>Integrated stress response; phosphorylation of eIF2α lowers global initiation but favors selected mRNAs.</td>
    </tr>
    <tr>
      <td>uORF</td>
      <td>Upstream open reading frame in a 5’UTR that can divert scanning ribosomes.</td>
    </tr>
    <tr>
      <td>IRES / ITAF</td>
      <td>Internal ribosome entry site and its helper factors. Viral IRESs are strong examples; many cellular IRES claims need strict controls for cryptic promoters, splicing, readthrough, and RNA abundance.</td>
    </tr>
    <tr>
      <td>RQC / NGD / NSD</td>
      <td>Ribosome quality control, no-go decay, and non-stop decay; surveillance of stalled or abnormal translation.</td>
    </tr>
    <tr>
      <td>O-GlcNAc</td>
      <td>Reversible sugar modification on Ser/Thr residues that can crosstalk with phosphorylation.</td>
    </tr>
    <tr>
      <td>Ubiquitin chain</td>
      <td>A polymeric ubiquitin mark whose linkage type, such as K48 or K63, helps determine protein fate.</td>
    </tr>
    <tr>
      <td>SUMOylation</td>
      <td>Conjugation of SUMO proteins, often changing nuclear interactions or complex assembly.</td>
    </tr>
    <tr>
      <td>Neddylation</td>
      <td>Conjugation of NEDD8, especially to cullins, activating cullin-RING E3 ligases.</td>
    </tr>
    <tr>
      <td>Proteasome</td>
      <td>Protease complex that degrades many short-lived or damaged ubiquitinated proteins.</td>
    </tr>
    <tr>
      <td>Selective autophagy</td>
      <td>Lysosomal clearance of selected cargo such as aggregates, organelles, or ubiquitinated complexes.</td>
    </tr>
    <tr>
      <td>UPR</td>
      <td>Unfolded protein response; ER-stress response that expands folding capacity or slows translation.</td>
    </tr>
    <tr>
      <td>PAR / ADP-ribosylation</td>
      <td>Poly-ADP-ribose signaling, often used around DNA damage and condensate formation.</td>
    </tr>
  </tbody>
</table>

<h2 id="the-seven-layers">The Seven Layers</h2>

<p>The map follows the flow from DNA to RNA to protein:</p>

<ol>
  <li><strong>3D genome</strong>: <code class="language-plaintext highlighter-rouge">ZONES</code>, <code class="language-plaintext highlighter-rouge">FENCES</code>, <code class="language-plaintext highlighter-rouge">BRIDGES</code>, and <code class="language-plaintext highlighter-rouge">HUBS</code> represent A/B
compartments, TADs, enhancer-promoter loops, and super-enhancers that can
overlap with transcriptional condensates without being identical to them.</li>
  <li><strong>Epigenetics</strong>: <code class="language-plaintext highlighter-rouge">SILENCER</code>, <code class="language-plaintext highlighter-rouge">OPENER</code>, <code class="language-plaintext highlighter-rouge">WRITER-A</code>, <code class="language-plaintext highlighter-rouge">WRITER-R</code>, <code class="language-plaintext highlighter-rouge">SHUFFLER</code>,
and <code class="language-plaintext highlighter-rouge">GUIDES</code> cover DNA methylation, histone marks, chromatin remodeling, and
non-coding RNAs that guide chromatin complexes.</li>
  <li><strong>Transcription</strong>: <code class="language-plaintext highlighter-rouge">KEYS</code> are transcription factors; <code class="language-plaintext highlighter-rouge">SCRIBE</code> is Pol II,
promoter-proximal pausing, and the phosphorylation code of its CTD.</li>
  <li><strong>Co-transcriptional processing</strong>: <code class="language-plaintext highlighter-rouge">SHIELD</code>, <code class="language-plaintext highlighter-rouge">SPLICER</code>, <code class="language-plaintext highlighter-rouge">TRIMMER</code>, and
<code class="language-plaintext highlighter-rouge">RECODER</code> cover capping, alternative splicing, alternative polyadenylation,
and A-to-I RNA editing.</li>
  <li><strong>Post-transcriptional control</strong>: <code class="language-plaintext highlighter-rouge">STAMP</code>, <code class="language-plaintext highlighter-rouge">READERS</code>, <code class="language-plaintext highlighter-rouge">DARTS</code>, <code class="language-plaintext highlighter-rouge">SPONGE</code>,
<code class="language-plaintext highlighter-rouge">CENSOR</code>, <code class="language-plaintext highlighter-rouge">TIMER</code>, <code class="language-plaintext highlighter-rouge">CLIPS</code>, and <code class="language-plaintext highlighter-rouge">VAULT</code> cover m6A, RNA-binding proteins,
miRNAs, lncRNAs, nonsense-mediated decay, mRNA stability, RNA structures, and
cytoplasmic granules.</li>
  <li><strong>Translation</strong>: <code class="language-plaintext highlighter-rouge">FORGE</code>, <code class="language-plaintext highlighter-rouge">BRAKE</code>, <code class="language-plaintext highlighter-rouge">DECOY</code>, <code class="language-plaintext highlighter-rouge">BYPASS</code>, <code class="language-plaintext highlighter-rouge">TEMPO</code>, and
<code class="language-plaintext highlighter-rouge">INSPECTOR</code> describe cap-dependent initiation, the integrated stress
response, uORFs, non-canonical initiation, decoding kinetics, and ribosome
quality control. <code class="language-plaintext highlighter-rouge">TEMPO</code> now also folds in decoding <em>fidelity</em>: recent
proteogenomics across &gt;1,000 human samples found thousands of non-genomic
amino-acid substitutions from alternate ribosomal decoding — not explained by
DNA variants or A-to-I editing — producing proteoforms that are stable,
abundant, and tissue/cancer-specific
(<a href="https://www.nature.com/articles/s41586-026-10678-2">Tsour et al., <em>Nature</em> 2026</a>).
Because these proteins partly escape <code class="language-plaintext highlighter-rouge">INSPECTOR</code> surveillance and persist into
<code class="language-plaintext highlighter-rouge">MATURE</code>, this belongs on the speed-vs-fidelity axis of <code class="language-plaintext highlighter-rouge">TEMPO</code> rather than in
<code class="language-plaintext highlighter-rouge">RECODER</code> (which is transcript-level A-to-I editing).</li>
  <li><strong>Post-translational regulation</strong>: <code class="language-plaintext highlighter-rouge">SWITCH</code>, <code class="language-plaintext highlighter-rouge">ROUTER</code>, <code class="language-plaintext highlighter-rouge">TETHER</code>, <code class="language-plaintext highlighter-rouge">LICENSE</code>,
<code class="language-plaintext highlighter-rouge">DESTROY</code>, <code class="language-plaintext highlighter-rouge">MATURE</code>, and <code class="language-plaintext highlighter-rouge">PAR</code> cover phosphorylation/O-GlcNAc, ubiquitin,
SUMOylation, neddylation, proteasome/autophagy clearance, maturation/UPR,
and PARP/PAR signaling.</li>
</ol>

<p>Across the whole diagram, border styles flag recurring meta-principles. <code class="language-plaintext highlighter-rouge">LLPS</code>
marks liquid-liquid phase separation, and <code class="language-plaintext highlighter-rouge">DECAY</code> marks turnover or clearance.
They are not extra regulatory layers. They are reused in several places:
transcriptional condensates, stress granules and P-bodies, mRNA decay,
proteolytic condensates, autophagy, and DNA damage repair assemblies.</p>

<h2 id="the-useful-reduction">The Useful Reduction</h2>

<p>The full map contains 37 mechanism boxes, 2 meta-principles, and dozens of
interactions. But conceptually, most of gene regulation reduces to three
strategies.</p>

<p><strong>1. Control accessibility.</strong><br />
Make a substrate accessible or inaccessible to its molecular machinery.
Chromatin opening lets transcription factors bind. TADs constrain which
enhancers can contact which promoters. Stress granules temporarily remove mRNAs
from translation. miRNAs and lncRNAs tune whether an mRNA is available to the
ribosome.</p>

<p><strong>2. Write a reversible mark, then interpret it.</strong><br />
Histone methylation, DNA methylation, m6A, phosphorylation, ubiquitination,
SUMOylation: the mark alone is never the full story. The reader and the context
determine the output. m6A can promote translation or accelerate decay. A K48
ubiquitin chain points toward the proteasome; K63 often acts in signaling or
selective autophagy. Phosphorylation can activate a transcription factor or
create a degron.</p>

<p><strong>3. Couple two processes through kinetics.</strong><br />
Some regulation is not a static state but a timing problem. Pol II elongation
speed influences exon choice. SETD2 deposits H3K36me3 during elongation, linking
transcription to splicing. eIF2α phosphorylation globally slows translation but
selectively favors ATF4 through uORF logic. Codon usage changes ribosome speed
and can influence co-translational folding.</p>

<p>That is the central idea of the landscape: gene regulation is not just a list of
mechanisms. It is a multi-layer control architecture built from recurring design
patterns.</p>

<h2 id="why-this-matters-for-ai-biology">Why This Matters For AI Biology</h2>

<p>For AI biology, this kind of map is not only educational. It shows why
predicting “gene expression” cannot be reduced to reading a promoter sequence.</p>

<p>The output of a gene depends on chromatin state, 3D contacts, Pol II kinetics,
splicing, RNA modifications, RNA-binding proteins, translational control, and
protein lifetime. A model that wants to predict perturbation response, cell
state, or disease mechanism needs to represent at least part of this stack.</p>

<p>The lesson of the Gene Regulation Landscape is simple: gene expression is not a
scalar. It is the endpoint of a control system.</p>

<h2 id="sources-to-anchor-the-map">Sources To Anchor The Map</h2>

<ul>
  <li>Core &amp; Adelman, 2019, promoter-proximal Pol II pausing:
https://pubmed.ncbi.nlm.nih.gov/31123063/</li>
  <li>Naftelberg et al., 2015, transcription/chromatin/splicing coupling:
https://pubmed.ncbi.nlm.nih.gov/26034889/</li>
  <li>Wang &amp; He, 2014, dynamic RNA modifications:
https://pubmed.ncbi.nlm.nih.gov/25263552/</li>
  <li>Wang et al., 2015, m6A and translation efficiency:
https://www.cell.com/cell/fulltext/S0092-8674(15)00562-0</li>
  <li>Shi et al., 2017, YTHDF3 translation/decay:
https://pmc.ncbi.nlm.nih.gov/articles/PMC5339834/</li>
  <li>Sabari et al., 2018, coactivator condensation at super-enhancers:
https://pmc.ncbi.nlm.nih.gov/articles/PMC6092193/</li>
  <li>Robson et al., 2019, chromatin topology:
https://pubmed.ncbi.nlm.nih.gov/31324893/</li>
</ul>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="Cell Bio" /><category term="Comp-Bio" /><summary type="html"><![CDATA[A compact map of how cells decide which genes become proteins]]></summary></entry><entry><title type="html">Finishing the PhD</title><link href="https://www.jkobject.com/blog/finishing-the-phd/" rel="alternate" type="text/html" title="Finishing the PhD" /><published>2026-05-01T00:00:00+00:00</published><updated>2026-05-01T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/finishing-the-phd</id><content type="html" xml:base="https://www.jkobject.com/blog/finishing-the-phd/"><![CDATA[<p>It’s done. March 25, 2026, 1:30pm, Duclaux amphitheater at Institut Pasteur. 🎓</p>

<p>I started in October 2023. Took 2.5 years instead of two. Here’s what happened.</p>

<h2 id="what-i-set-out-to-do">What I set out to do</h2>

<p>The project was building foundation models for single-cell transcriptomics —
taking the pretraining ideas from NLP and applying them to gene expression at
scale. The goal: a model that could infer gene regulatory networks, classify
cell types zero-shot, denoise expression, and actually be useful to biologists.
Not just publishable, usable.</p>

<p>%% WIP$$$</p>

<p>I was split between two labs:
<a href="https://research.pasteur.fr/fr/team/machine-learning-for-integrative-genomics/">Laura Cantini</a>
at Institut Pasteur (computational biology, gene networks) and
<a href="http://www.gpeyre.com/">Gabriel Peyré</a> at ENS (applied mathematics, optimal
transport). Two very different worlds, two sets of expectations, two weekly
meetings.</p>

<p>Outside the lab, I wanted to keep training — endurance sports had become
important to me during this period, and I wanted to see more of the world while
I could.</p>

<h2 id="what-i-did">What I did</h2>

<h3 id="the-papers">The papers</h3>

<p>The thesis produced three papers:</p>

<p><strong>scPRINT</strong>
(<a href="https://www.nature.com/articles/s41467-025-58126-9">Nature Communications, 2025</a>)
— a large cell model trained on 50M+ cells for gene network inference. Novel
gene tokens from ESM2 protein embeddings, learned expression tokenization,
genomic positional encoding. State-of-the-art on GRN inference benchmarks,
zero-shot cell type classification, denoising. Code at
<a href="https://github.com/cantinilab/scPRINT">github.com/cantinilab/scPRINT</a>.</p>

<p><strong>Xpressor</strong> — a cross-attention architecture for learning across biological
scales. The idea: compress gene-level representations into cell-state vectors,
and fine-tune protein language models using cellular tasks. Improved cell-type
prediction (+28%) and embedding quality (+8%) over standard architectures.</p>

<p><strong>scPRINT-2</strong> (in revision, Nature Methods) — trained on 350M cells from 16
organisms, 25 TB of data. 42-model ablation study to figure out what actually
matters in scFM design. 75% zero-shot cell type classification on OpenProblems
(up from 47% with scPRINT-1). State-of-the-art denoising, batch correction,
cross-species generalization, counterfactual generation.</p>

<h3 id="the-tools">The tools</h3>

<p>Seven Python packages released:
<a href="https://github.com/cantinilab/scPRINT">scPRINT</a>,
<a href="https://github.com/jkobject/BenGRN">BenGRN</a> (GRN benchmarking),
<a href="https://github.com/cantinilab/GRnnData">GRnnData</a> (gene networks in AnnData),
and more for data processing, evaluation, and model serving. Everything
open-source under GPL-v3.</p>

<!-- TODO: add the full list of packages with links -->

<h3 id="the-conferences">The conferences</h3>

<!-- TODO: fill in the full list — Jérémie should add the specific conferences -->

<p>Over a dozen conferences and workshops across Europe and beyond. Oral
presentations, poster sessions, invited talks. Each one forced me to sharpen the
story and meet people working on adjacent problems.</p>

<h3 id="the-outreach">The outreach</h3>

<p>I mentored a student during the PhD. Wrote blog posts (you’re reading one), made
videos, gave talks to non-specialist audiences. I think this kind of work
matters — it’s how ideas spread beyond the 50 people who read your paper.</p>

<h3 id="the-people">The people</h3>

<p>The labs at Pasteur and ENS. Jules Samaran, Remi Trimbour, Geert Huizing at ENS.
Alex Wolf, Sergei Ribakov, Brice Rafestin for software help. The Nucleate
community, Whitelab Genomics, Blossom, dot Omics, Biographica. A lot of people
made this work possible.</p>

<h3 id="outside-the-lab">Outside the lab</h3>

<p>I trained for and ran a medium triathlon, a half-marathon, the Mont-Blanc trail,
and the Paris Marathon. Traveled to the UK, Portugal, Spain, Guadeloupe,
Vancouver, Thailand. I needed this. Long runs are good for thinking, and getting
away from the screen makes you come back sharper.</p>

<h2 id="the-defense">The defense</h2>

<p>March 25. Pasteur, Duclaux amphitheater.</p>

<p>I was stressed. I’d prepared well, but defending years of work in one afternoon
is a strange exercise. You try to make it look like everything was planned from
the start, when in reality half of the good ideas came from accidents.</p>

<p>The jury was rigorous but fair. Hard questions in the Q&amp;A — the kind that make
you think, not the kind designed to trip you up. That’s what makes a defense
feel worth it.</p>

<p>Having my mentoree there meant a lot. And the next week, I finished a marathon.
Different kind of endurance, same feeling at the end.</p>

<h2 id="what-i-learned">What I learned</h2>

<p>How to write papers. How to think about research impact beyond citations. How to
set goals when nobody is setting them for you.</p>

<p>The research community matters more than I expected. Not just for collaboration
— for sanity. The conferences, the DMs, the random conversations at poster
sessions.
<a href="https://www.jkobject.com/blog/the-phd-decision-grn-foundation-model/">Open source</a>
is how you actually have an impact: if scPRINT is useful, it’s because people
can use it.</p>

<p>Technically: a lot about transformers, diffusion models, optimal transport, the
math of self-supervised learning. But also about the gaps — the missing data,
the incomplete benchmarks, the things we still don’t know how to measure
properly.</p>

<p>Company research and a PhD are different. The PhD gave me time to go deep, be
wrong for months, follow a thread until it broke or became a paper. I wouldn’t
trade that.</p>

<h2 id="whats-next">What’s next</h2>

<p>I’m not going back into academia. Never planned to. 🚀</p>

<p>I’m thinking of building <strong>Jouvence</strong> — a company using AI and biology to extend
healthy human lifespan. The thesis work on foundation models, gene networks, and
cellular representation is directly relevant: if you can model how cells work
and how they break, you can start to intervene.
<!-- TODO: add 2-3 sentences about Jouvence's approach, from Notion pages --></p>

<p>I’m also looking at companies working on similar goals: longevity, disease
modeling, AI-driven high-throughput data generation. Places like Lilas
Bioscience and Xaira Therapeutics are doing interesting work in this space.</p>

<p>Lots to do.</p>

<hr />

<p>The full arc:
<a href="https://www.jkobject.com/blog/the-phd-decision-grn-foundation-model/">the PhD decision</a>,
<a href="https://www.jkobject.com/blog/a-year-in-the-phd/">a year in</a>, and now this.</p>

<p>Thanks to Laura, Gabriel, Juliette, my family, and everyone who was part of the
last 2.5 years.</p>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="PhD" /><category term="Comp-Bio" /><category term="Foundation Models" /><summary type="html"><![CDATA[2.5 years, two labs, one foundation model, and a defense. Here is what I learned.]]></summary></entry><entry><title type="html">How I managed thousands of datasets to build the scPRINT family of scRNA-seq foundation models</title><link href="https://www.jkobject.com/blog/lamindb-these/" rel="alternate" type="text/html" title="How I managed thousands of datasets to build the scPRINT family of scRNA-seq foundation models" /><published>2026-04-20T00:00:00+00:00</published><updated>2026-04-20T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/lamindb-these</id><content type="html" xml:base="https://www.jkobject.com/blog/lamindb-these/"><![CDATA[<p>At the start of my PhD, I was faced with what seemed like a mountain to climb:
build, largely alone, a foundation model for single-cell RNA-seq data. As anyone
in the field knows, building the model is not the hard part. Getting the data
is.</p>

<p>To train a cell foundation model that actually generalizes, you need thousands
of datasets. You need to find them, download them, harmonize gene names across
species, align cell type labels to controlled ontologies, preprocess everything
consistently, store it in a way that doesn’t collapse under its own weight, and
feed it to a model at scale. Managing a dozen datasets is already painful for
most computational biologists. I needed to handle thousands.</p>

<p>Thirty months ago, three things came at exactly the right moment. The Chan
Zuckerberg Initiative had made around 700 datasets easily accessible through
CellxGene. The LaminDB project gave me a way to manage large, heterogeneous
collections of biological data with metadata that actually meant something. And
Sergei Rybakov was building a loader for streaming single-cell data at scale.</p>

<h2 id="managing-scale-with-lamindb">Managing scale with LaminDB</h2>

<p>The core problem with large-scale single-cell data is not raw file size. It’s
the metadata. A dataset from a 2019 mouse lung study uses different gene IDs,
different tissue labels, and different cell type annotations than a 2023 human
heart study. Reconciling these across hundreds of datasets by hand is a losing
battle that compounds as the corpus grows.</p>

<p>LaminDB treats biological ontologies as first-class citizens. Every dataset I
ingested was linked to standardized terms: Cell Ontology for cell types, Uberon
for tissues, NCBI Taxonomy for species. That ontological consistency is what
made it possible to build scPRINT-2’s hierarchical classification loss, which
penalizes predictions based on their distance in the ontology graph rather than
just correct/wrong. The loss knows that “T cell” and “CD8-positive T cell” are
related in a way that “T cell” and “hepatocyte” are not. That knowledge came
from having the data structured correctly from the start.</p>

<p>Without LaminDB I would have spent months on this. With it, it took weeks.</p>

<h2 id="streaming-350-million-cells">Streaming 350 million cells</h2>

<p>Loading 350 million cells into memory is not an option. You need streaming,
shuffling across datasets, and batching that mixes cell types, species, and
sequencing technologies, without the dataloader becoming the bottleneck.</p>

<p><code class="language-plaintext highlighter-rouge">scDataLoader</code> handles this. It’s built on top of LaminDB’s <code class="language-plaintext highlighter-rouge">MappedCollection</code>
interface, which lets you treat hundreds of separate datasets as a single object
you can sample, filter, and iterate over. It streams directly from the artifact
store and integrates cleanly with PyTorch’s DataLoader. I was able to train
scPRINT-2<sup id="fnref:2" role="doc-noteref"><a href="#fn:2" class="footnote" rel="footnote">1</a></sup> on 350 million cells and 25 TB of data on a single cluster without
writing custom data infrastructure. That felt like a minor miracle at the time.</p>

<h2 id="beyond-training">Beyond training</h2>

<p>Once scPRINT<sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">2</a></sup> was published and colleagues and interns started using the
infrastructure, having a LaminDB instance meant they could reproduce my work
exactly: same artifacts, same lineage, same ontology mappings. Data lineage made
it easy to answer “which datasets went into this version of the model?” or “was
this processed before or after the normalization change?” without digging
through scripts.</p>

<p>It also meant I could serve processed data to the team with enough context
attached that they didn’t need me to explain what they were looking at.</p>

<h2 id="the-reflex-it-created">The reflex it created</h2>

<p>I used LaminDB throughout my PhD. It let me do a lot alone, in a reasonable
time, in a reproducible way. That’s a rare combination in this field.</p>

<p>These days, when I start a computational biology project, I set up a git repo
and a LaminDB instance. In that order, roughly.</p>

<h2 id="background">Background</h2>

<p>In fall 2023, Jeremie &amp; Alex met in CZI’s CellXGene Slack channel both trying to
figure out how to best manage metadata of thousands of scRNA-seq datasets.
Jeremie for his work on scRNA-seq foundation models, and Alex for his work on
LaminDB.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:2" role="doc-endnote">

      <p>Kalfon, J., Peyre, G., &amp; Cantini, L. (2026). scPRINT-2: Towards the
next-generation of cell foundation models and benchmarks. <em>bioRxiv</em>.
https://doi.org/10.64898/2025.12.11.693702 <a href="#fnref:2" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:1" role="doc-endnote">

      <p>Kalfon, J., Samaran, J., Peyre, G., &amp; Cantini, L. (2025). scPRINT:
pre-training on 50 million cells allows robust gene network predictions.
<em>Nature Communications</em>, 16, 3607.
https://doi.org/10.1038/s41467-025-58699-1 <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="Startups" /><category term="Comp-Bio" /><summary type="html"><![CDATA[Short stories about my professional experiences.]]></summary></entry><entry><title type="html">VCC starter pack</title><link href="https://www.jkobject.com/blog/vcc-starter-pack/" rel="alternate" type="text/html" title="VCC starter pack" /><published>2025-10-10T00:00:00+00:00</published><updated>2025-10-10T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/vcc-starter-pack</id><content type="html" xml:base="https://www.jkobject.com/blog/vcc-starter-pack/"><![CDATA[<p>To start the VCC, you need to know a few things. This is your starter pack:</p>

<p><img src="/assets/images/starterpack.png" alt="image of a starter pack" /></p>

<p>The main packages for single cell data analysis are <strong>scanpy</strong> - <strong>anndata</strong> -
<strong>pertpy:</strong> with many other easy to use packages being part of the
<strong><a href="https://scverse.org/">scverse</a></strong> ecosystem. In addition, both RNA-seq, csv,
and image datasets management across thousands of datasets is enabled with
<strong>lamindb</strong>.</p>

<p>To get started there is a couple nice tutorials in the packages referenced above
as well as some tutos
<a href="https://open.substack.com/pub/surajparmar/p/from-molecules-to-matrices-part-1">here</a>,
and <a href="https://huggingface.co/blog/virtual-cell-challenge">here</a> that are must
reads!</p>

<p><strong>For further readings:</strong></p>

<p>I will shamelessly plug my paper and package
<a href="https://www.nature.com/articles/s41467-025-58699-1">scPRINT</a> to understand
where we are at with single cell foundation model. I will also mention the
<a href="https://arcinstitute.org/news/virtual-cell-model-state">state</a> model that is
the first to delve a lot into perturbations and how to do multiple cells at a
time contrary to most other methods.</p>

<p>A key perturbation method that is not foundation model based is
<a href="https://academic.oup.com/bioinformatics/article/41/Supplement_1/i599/8199372">GPO-VAE</a>,
which encompasses very similar idea to <strong>scVI, cradle-VAE</strong>, <strong>GEARS</strong>,
<strong>chemCPA</strong>, <strong>Biolord</strong>, <strong><a href="https://arxiv.org/pdf/2505.14919">txpert</a></strong>, <strong><a href="https://www.nature.com/articles/s43588-025-00870-1">LPM</a></strong></p>

<p>Many of these tools and others have been benchmarked in a few papers:</p>

<ol>
  <li><a href="https://www.biorxiv.org/content/10.1101/2024.12.23.630036v2.full">A Systematic Comparison of Single-Cell Perturbation Response Prediction Models</a></li>
  <li><a href="https://www.biorxiv.org/content/10.1101/2024.12.20.629581v2.full">Benchmarking AI Models for In Silico Gene Perturbation of Cells</a></li>
  <li><a href="https://www.nature.com/articles/s41592-025-02980-0">Benchmarking algorithms for generalizable single-cell perturbation response prediction</a></li>
</ol>

<p>But we know already that many of the previous losses used to train models and
assess them on perturbation data suffered from issues as presented in this
paper: <a href="https://arxiv.org/abs/2506.22641">mode collapse in current models</a> and
this <a href="https://www.nature.com/articles/s41587-025-02777-8">paper</a>. This is also
because current perturbation dataset are very noisy.</p>

<p><strong>Things that many people don’t say and that should be our theory when designing
the models:</strong></p>

<ol>
  <li>Many perturbations do nothing.</li>
  <li>Many other perturbations kill the cell (but we don’t know it) and the readout
is just really poor quality stuff because of it, or cells that survived
somehow..</li>
  <li>Many perturbations have very similar phenotypic effect. Triggering something
like what happens when the cell has an issue in its genome.</li>
  <li>Moreover the RNA-guides used to decide which region of the genome is
perturbed have very different effect due to something called off-tagerting
which was extensively studied and corrected for in the DepMap project.</li>
  <li>Most datasets are super sparse (only some genes perturbed), or only a couple
cell lines, and often low depth.</li>
</ol>

<p>Fortunately, we have depmap and LINCs/L1000. which are genome-wide perturbations
in thousands of different cell models with guide-effect analysed However, while
some of them have 1000 genes assessed post perturbation most of the data is only
wether the cells died or not post perturbations. but we NEED to use this dataset
somehow. This in addition to what we learned above will be differentiating
factors. Another very similar
<a href="https://www.nature.com/articles/s41467-019-13805-y">dataset</a> done by Sanger on
the same set of cell lines and using very similar guides also exist and has many
similarities and some differences with depmap.</p>

<p>Additionally, an image-based dataset also exist, made by Recursion called
<a href="https://www.rxrx.ai/">RxRx</a>. But it is using similar guides / molecules and
cell lines that depmap and Sanger are using!</p>

<p>Arc hasn’t released the guide sequences yet but we might infer it. we know also
the genetic sequence of the cell line. We also know about the novel ideas for
creating better losses. Now how would we do it well?</p>

<p>Finally to know more about ideas for the future of this field, please read the
2024 markov bio
<a href="https://www.markov.bio/research/mech-interp-path-to-e2e-biology">blog post</a>,
which I agree a lot with.</p>

<p>Many perturbation datasets are easily available on lamindb’s own database
<a href="https://lamin.ai/laminlabs/pertdata">here</a>, plus
<a href="https://www.depmap.org/">depmap</a> (where you will need both the genedependency,
expression, guide efficacy, and model files) and <a href="https://clue.io/">LINCS</a> (see
the download instructions below).</p>

<h2 id="outline-of-a-project">Outline of a project</h2>

<ol>
  <li>get the lamindb and add all other datasets we need (depmap, lincs, xaira,
…) (3 weeks)</li>
  <li>generate a good set of train/val/test datasets to test some hypothesis (more
data, variability, adding chemical perturbation, multiple species, guide
information, phenotypic only …) (3 weeks - 2 months)</li>
  <li>reimplement all the scoring metrics used in previous papers</li>
  <li>Define a MIN and MAX (MAX = subset of perturbed cells, same results from
another dataset) (MIN=unperturbed cells, a linear model, a simple random
forest, a gene embedding KNN model, ChatGPT, …)</li>
  <li>answer some questions on identifiability, quality, and information content:
    <ol>
      <li>can we find perturbations that have no impact on cell state (might need to
compare to average of perturbations) (1 week)</li>
      <li>can we find perturbations that have the same effect? (cell death, growth
arrest, …) (1 week)</li>
      <li>what does correcting for off-targeting do in perturb-seq?</li>
      <li>how much more predictive accuracy do we get from using a depmap / rxrx /
10 genes / 1000 / … for different modalities, what information is only
available in one of them? (3 months)</li>
    </ol>
  </li>
  <li>looking at the benchmarking, get the best performing and/or simplest
VAE-based architecture (1 week)</li>
  <li>retrain it and test if adding more data, changing the losses, or other
parameters improve the model, check how much using LINCS &amp; Depmap help (3
months)</li>
  <li>simple fine-tune of scPRINT with this data (1 month)</li>
  <li>use a simple random forest model on crafted features to predict the
perturbation and see how it does (1 month)</li>
  <li>use a groundtruth dataset (literature and correlation (GENIE3)) to predict
the perturbation and see how it does (1 month)</li>
  <li>going further by mixing ideas of GPO-VAE to scPRINT in pseudo-bulk mode (3
months)</li>
  <li>going further by mixing ideas from STATE to scPRINT (3 months)</li>
</ol>

<h2 id="other-useful-dataset-ressources">other useful dataset ressources</h2>

<ul>
  <li><a href="https://virtualcellchallenge.org/datasets">VCC’s datasets</a></li>
  <li><a href="https://github.com/Peekxel/PRISM">https://github.com/Peekxel/PRISM</a></li>
  <li><a href="https://virtualcellmodels.cziscience.com/dataset/genome-scale-tcell-perturb-seq">CZI’s genome scale perturb seq</a></li>
  <li><a href="https://thevirtualcell.com/">Ginkgo’s dataset</a></li>
  <li><a href="https://www.kaggle.com/competitions/echoes-of-silenced-genes/leaderboard">Millya’s current challenge</a></li>
</ul>

<p>Another large dataset of perturbation, including some “unintended” double KO are
available in iPSCs thanks to XAIRA, called
<a href="https://plus.figshare.com/articles/dataset/Processed_data_for_X-Atlas_Orion_Genome-wide_Perturb-seq_Datasets_via_a_Scalable_Fix-Cryopreserve_Platform_for_Training_Dose-Dependent_Biological_Foundation_Models/29190726">orion</a>,
with a
<a href="https://www.biorxiv.org/content/10.1101/2024.11.28.625833v1">similar dataset</a>
published by the sanger institute.</p>

<p>Some dual guide KO was also done in colorectal cancer (CRC)
<a href="https://www.nature.com/articles/s41467-025-67256-9#Abs1">here</a>.</p>

<h3 id="lincs-download">LINCS download:</h3>

<p><code class="language-plaintext highlighter-rouge">wget https://s3.amazonaws.com/macchiato.clue.io/builds/LINCS2020/level5/level5_beta_all_n1201944x12328.gctx</code></p>

<p><code class="language-plaintext highlighter-rouge">wget https://s3.amazonaws.com/macchiato.clue.io/builds/LINCS2020/siginfo_beta.txt</code></p>

<p><code class="language-plaintext highlighter-rouge">wget -r ftp://ftp.ncbi.nlm.nih.gov/geo/series/GSE92nnn/GSE92742/suppl/GSE92742_Broad_LINCS_gene_info.txt.gz</code></p>

<p><code class="language-plaintext highlighter-rouge">wget -r ftp://ftp.ncbi.nlm.nih.gov/geo/series/GSE92nnn/GSE92742/suppl/GSE92742_Broad_LINCS_sig_info.txt.gz</code></p>

<p><code class="language-plaintext highlighter-rouge">wget -r ftp://ftp.ncbi.nlm.nih.gov/geo/series/GSE92nnn/GSE92742/suppl/GSE92742_Broad_LINCS_sig_metrics.txt.gz</code></p>

<h3 id="companies-to-whatchout-for-their-perturbation-datasets">Companies to whatchout for their perturbation datasets</h3>

<ul>
  <li>recursion</li>
  <li>xaira</li>
  <li>tahoe</li>
  <li>arc</li>
  <li>illumina
(<a href="https://www.illumina.com/company/news-center/press-releases/2026/fda84c92-b4b3-4691-a402-35555abe8605.html">yes indeed</a>)</li>
  <li>cellular intelligence</li>
  <li>czi</li>
  <li>retro bio</li>
  <li>ginkgo</li>
</ul>

<h2 id="edit-post-vccs-feedbacks">EDIT: post VCC’s feedbacks</h2>

<p>interestingly, post VCC, it seems that many people that did not knew much about the data (icluding the organizers?) learnt a lot about things that had been published months, even years ago… This is great, because it still means that many people know learned collectively, we are at a better place now than before. We could have gotten there more efficiently of course.. I adviced to read some of the blog posts published by contenders on the topic:</p>

<ul>
  <li><a href="https://blog.turbine.ai/p/how-did-we-get-a-regression-model?hide_intro_popup=true">regression model ftw</a></li>
  <li><a href="https://gmdbioinformatics.substack.com/p/arc-virtual-cell-challenge-my-diary">the VCC diary</a></li>
  <li><a href="https://giovannipalla.substack.com/p/virtual-cell-perturbation-metrics">the VCC metrics issues</a></li>
  <li><a href="https://newsletter.kiin.bio/p/arc-institutes-virtual-cell-challenge">are we just modelling cell’s fingerprints</a></li>
</ul>

<p>One thing that I find regretful is that no famous published models have been implemented openly to check for their abilities on this benchmark.</p>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="AI" /><category term="Comp-Bio" /><category term="Research" /><summary type="html"><![CDATA[If you do data science and machine learning, here is everything you need to know to get started on the Arc virtual cell challenge.]]></summary></entry><entry><title type="html">A year in the PhD</title><link href="https://www.jkobject.com/blog/a-year-in-the-phd/" rel="alternate" type="text/html" title="A year in the PhD" /><published>2025-03-05T00:00:00+00:00</published><updated>2025-03-05T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/a-year-in-the-phd</id><content type="html" xml:base="https://www.jkobject.com/blog/a-year-in-the-phd/"><![CDATA[<p>A year and a half in. When I wrote <a href="https://www.jkobject.com/blog/the-phd-decision-grn-foundation-model/">the PhD decision post</a>, I set four goals:</p>

<ol>
  <li>do it in 2 years and be prepared</li>
  <li>make as many connections as I can</li>
  <li>maximize impact on the community: make something useful</li>
  <li>enjoy it as much as possible</li>
</ol>

<p>Time to check in. 🧬</p>

<h2 id="the-environment">The environment</h2>

<p>My PhD is split between Institut Pasteur and ENS, which are very different places.</p>

<p>At Pasteur, I work with Laura Cantini. Pasteur is serious science, long corridors, old campus. Laura runs a small group at the intersection of ML and genomics. She pushes hard on rigor — every claim grounded, every figure telling a clear story. I’ve become a better scientist working with her.</p>

<p>At ENS, Gabriel Peyré does applied math — optimal transport, signal processing, things I’d admired from afar for a while. Having two PIs means two sets of expectations, two lab cultures, sometimes two conflicting intuitions about what matters. It took months to find a rhythm. What works: being very explicit about what I’m doing and where I’m stuck. Over-communicating rather than assuming.</p>

<p>The admin side is… well, French public research bureaucracy across two institutions. I probably spent two full working weeks last year on forms and approvals alone.</p>

<p>Having two advisors from such different backgrounds has been great, though. Academia runs on different incentives than industry — ideas and papers instead of delivery milestones and quarterly revenue. That took some adjusting.</p>

<h2 id="the-successes">The successes</h2>

<p>I’m proud of what got done this year. 🙌</p>

<p>The main thing is <a href="https://github.com/cantinilab/scPRINT">scPRINT</a> — a large foundation model for single-cell RNA sequencing. Getting it to work, getting it published, seeing people actually use it. I also shipped three more open-source tools around scPRINT for benchmarking, data processing, and evaluation.</p>

<p>I went to over 10 conferences and events across Europe. 4 oral presentations (still nerve-wracking), 3 poster sessions. I actually like poster sessions — the conversations are real, people stop because they want to, and you end up understanding your own work better by explaining it 30 times.</p>

<p>I’ve also met researchers I’ll be working with for years. Building that network has been one of the best parts of the PhD so far.</p>

<h2 id="the-difficulties">The difficulties</h2>

<p>Some things were harder than expected.</p>

<p>Switching projects completely when I started — I came from industry work on cell atlases and drug discovery, and going back to a more theoretical research mode meant resetting my idea of what “progress” looks like.</p>

<p>scPRINT opened more questions than it answered. Every experiment surfaced something new to investigate, and learning to say “this is out of scope for now” without feeling like I was cutting corners took discipline.</p>

<p>The isolation is real. You’re responsible for your own direction in a way that’s different from any job I’ve had. Some weeks nobody is telling you whether you’re on the right track. You have to figure that out yourself.</p>

<h2 id="the-good-and-the-bad-of-academia">The good and the bad of academia</h2>

<p>I see a lot of people either romanticizing academia or being cynical about it. I’ll try to be straight.</p>

<p>The freedom is real. No product manager asking if your idea is “strategically aligned”. I can spend two weeks on a mathematical question just because it might matter. I can collaborate with anyone across institutions and countries. The people I’ve met are great — generous with time and ideas.</p>

<p>But the hard stuff is also real. The bureaucracy I already mentioned. The lack of structure (great when you’re focused, bad when you’re stuck). Being a solo contributor on one main project for three years is a very different rhythm than industry. When you’re stuck, there’s no teammate to take over.</p>

<h2 id="what-is-next">What is next?</h2>

<p>Still aiming to finish in about 2 years from start, so the clock is ticking. 🚀</p>

<p>Looking at those 4 original goals: I think I’m on track. I shipped something useful, I made connections, I showed up at conferences, and I’ve mostly enjoyed it.</p>

<p>It’s getting harder though. The deeper you go into a research problem, the more you see what’s left. The project keeps growing in scope while the deadline doesn’t move.</p>

<p>More to come. 👀</p>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="PhD" /><category term="Comp-Bio" /><summary type="html"><![CDATA[What has happened in the last year? How does it feel like to be a PhD student?]]></summary></entry><entry><title type="html">What are large cell models?</title><link href="https://www.jkobject.com/blog/what-are-large-cell-models/" rel="alternate" type="text/html" title="What are large cell models?" /><published>2024-09-07T00:00:00+00:00</published><updated>2024-09-07T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/what-are-large-cell-models</id><content type="html" xml:base="https://www.jkobject.com/blog/what-are-large-cell-models/"><![CDATA[<p>In my recent paper <a href="https://www.biorxiv.org/content/10.1101/2024.07.29.605556v1">scPRINT</a> I define the term Large Cell Model but what are they?</p>

<p>AI is everywhere nowadays, and in biology it is helping us enter a new era of what <a href="https://x.com/jacobkimmel/status/1829550511768674401">jacob kimmel</a> describes as predictive biology.</p>

<p>In this new field, we use complex computational approaches together with high throughput data gathering to generate predictive models of physical phenoma like protein folding, binding, regulation, response, etc. 🧬</p>

<p>scPRINT, like scGPT, geneformer, scFoundation, scLM and more is part of a bread of AI tools 🤖 that try to better model the cell regulation and response using a large amount of parameters and data. In this context they are -large- -cell models-. 🦠</p>

<h2 id="why-this-name">Why This Name?</h2>

<p>The analogy to Large Language Models (LLMs) is intentional and I think it’s a good one. LLMs are trained on massive corpora of text and learn to model the “language” of human expression — grammar, semantics, context, all of it emergently, without being hand-coded. The hope with LCMs is exactly the same: train on massive datasets of single-cell transcriptomics and let the model learn the “language” of the cell. 🔤</p>

<p>A cell expresses thousands of genes at once. Which genes are active, how much, in what combination — this is the cell’s way of encoding its state, its identity, its response to the environment. If we squint, gene expression profiles look a lot like sentences in a high-dimensional language. So it makes sense to apply similar techniques. Transformers, masked token prediction, large-scale pretraining — the same toolkit that gave us GPT-4 and BERT, now applied to biology.</p>

<p>The “Large” part matters too. These models are trained on tens of millions of single-cell profiles. Geneformer (~40M parameters) used 30M cells for pretraining. scGPT (~50M parameters) used 33M. scFoundation pushed further with ~100M parameters and 50M cells. And the field keeps scaling. 📈</p>

<h2 id="what-can-they-actually-do">What Can They Actually Do?</h2>

<p>This is where things get interesting — and also where a lot of hype meets hard reality.</p>

<p>In principle, LCMs can do several useful things:</p>

<ul>
  <li><strong>Cell embeddings</strong>: Compress a high-dimensional gene expression profile into a compact vector that captures biological meaning. Useful for cell type annotation, clustering, comparison across datasets.</li>
  <li><strong>Perturbation prediction</strong>: Given a cell state, predict how gene expression would change if you knock out a gene, add a drug, or apply a cytokine. This is the holy grail for drug discovery. 💊</li>
  <li><strong>GRN inference</strong>: Extract which genes regulate which other genes — reconstructing the regulatory logic of a cell from the patterns the model learned.</li>
  <li><strong>Batch integration</strong>: Harmonize data from different labs, technologies, or conditions into a shared embedding space.</li>
</ul>

<p>In practice? It’s more nuanced. Recent benchmarking shows that in zero-shot settings (using these models without fine-tuning), both scGPT and Geneformer can be outperformed by simpler methods like selecting highly variable genes and running scVI or Harmony. That’s a bit sobering. 😅</p>

<p>But the fine-tuned performance is real. And unlike classic approaches, LCMs don’t need you to specify the regulatory logic upfront — they learn it from data. That’s the paradigm shift.</p>

<h2 id="how-are-they-different-from-classical-approaches">How Are They Different from Classical Approaches?</h2>

<p>Before LCMs, the computational biology toolkit included things like pseudotime analysis, diffusion maps, factor analysis, and simple regression models for GRN inference. These are great tools! But they require a lot of prior knowledge and don’t generalize well. You’d fit a model per dataset, per question.</p>

<p>LCMs change the frame: you pretrain once on the universe of available data, then adapt to specific questions. It’s the transfer learning paradigm that revolutionized NLP and computer vision — finally landing in single-cell biology. The bet is that the representations learned at scale will capture something more fundamental about how cells work.</p>

<h2 id="the-hard-challenges">The Hard Challenges</h2>

<p>I won’t sugarcoat it. LCMs face serious obstacles:</p>

<p><strong>Data quality and scale.</strong> Tens of millions of cells sounds like a lot, but the human body has ~37 trillion cells across hundreds of types and states. Most training data is heavily biased toward easy-to-measure cell types (blood, PBMCs, cell lines). Rare cell types, spatial context, and dynamic states are vastly underrepresented. 🔍</p>

<p><strong>Evaluation is hard.</strong> What does it mean for a cell model to be “good”? We don’t have ground-truth labels for most biological processes. Perturbation benchmarks help, but they’re limited in scale and often use the same few datasets everyone trains on. The field badly needs better, truly held-out evaluation.</p>

<p><strong>The measurement bottleneck.</strong> As I’ve <a href="/PhD/about-the-aivc-paper/">written before</a>, the fundamental constraint isn’t really the algorithm or even the compute — it’s our ability to measure cellular state accurately. Current scRNA-seq gives us a snapshot of messenger RNA, but misses proteins, non-coding RNAs, spatial organization, dynamic timescales… A model can only be as good as the data it learns from.</p>

<p><strong>What’s actually learned?</strong> Recent interpretability work using sparse autoencoders on Geneformer and scGPT suggests these models have internalized rich biological structure — pathway membership, protein interactions, cell type identity. But they capture mostly <em>co-expression</em> patterns, not genuine <em>causal regulatory logic</em>. Whether scaling will fix this or whether we need different architectures is an open question. 🤔</p>

<h2 id="where-is-this-going">Where Is This Going?</h2>

<p>I’m genuinely excited about the trajectory of this field, even with all the caveats above.</p>

<p>The next few years will likely bring bigger models trained on more diverse, multimodal data — combining transcriptomics with proteomics, epigenomics, spatial information. Better measurement technologies (total RNA-seq, multiplexed protein imaging, long-read sequencing) will give us richer training signal. And the benchmarking community is getting sharper about what “good” actually means.</p>

<p>The analogy to LLMs is instructive here too: early language models also seemed to fail on out-of-distribution tasks, often beaten by simpler methods. What changed? More data, better training objectives, and scale. I think we’re at an early BERT moment for cell biology.</p>

<p>LCMs won’t model a whole cell anytime soon — the biology is just too complex and our measurements too incomplete. But as tools for learning transferable, generalizable representations of cell state? They’re already changing how we do science, and the best is yet to come. 🚀</p>

<hr />

<p><em>If you want to dig deeper into how scPRINT specifically fits into this landscape, check out <a href="https://www.biorxiv.org/content/10.1101/2024.07.29.605556v1">the paper</a> or my other posts on <a href="/PhD/manage-grn-and-what-they-mean/">GRN inference</a> and <a href="/PhD/about-the-aivc-paper/">the vision for AI virtual cells</a>.</em></p>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="Comp-Bio" /><summary type="html"><![CDATA[A few thoughts on the term large cell models that I am using in my recent paper and my posts]]></summary></entry><entry><title type="html">Ancestry Bias in CRISPR Screens</title><link href="https://www.jkobject.com/blog/ancestry-bias-in-crispr/" rel="alternate" type="text/html" title="Ancestry Bias in CRISPR Screens" /><published>2024-06-21T00:00:00+00:00</published><updated>2024-06-21T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/ancestry-bias-in-crispr</id><content type="html" xml:base="https://www.jkobject.com/blog/ancestry-bias-in-crispr/"><![CDATA[<h1 id="ancestry-bias-in-crispr-screens">Ancestry Bias in CRISPR Screens</h1>

<p>🎉 Thrilled to share our paper on ancestry biases in CRISPR screens published in Nature Communications! This work, led by Sean Misek, reveals an important discovery about potential biases in CRISPR screening technology. You can read the full paper here: <a href="https://www.nature.com/articles/s41467-024-48957-z">https://www.nature.com/articles/s41467-024-48957-z</a></p>

<h2 id="how-it-started">How It Started</h2>

<p>🧐 While working in the Cancer Data Science group, I was exploring ways to extract interesting new features from our CCLE omics dataset that might reveal relationships between cancer vulnerabilities and genetics.</p>

<p>💡 One unexplored avenue was computing not just the cancer mutations in our cell lines, but also the germline mutations. The reasoning was twofold: some mutations might have been wrongly categorized as somatic, and we know there exist many germline mutations that can lead to cancer.</p>

<p>🔒 However, this idea faced some initial resistance as team members were concerned about releasing potentially identifiable patient-related data.</p>

<h2 id="the-discovery">The Discovery</h2>

<p>🚀 Despite the pushback, I decided to implement the germline analysis pipeline and generate the data while awaiting the final decision on its release. Fortuitously, Sean Misek, a postdoc at the Broad, approached us with the same idea, wanting to use germline mutations for GWAS analysis.</p>

<p>🕵️‍♂️ Sean’s analysis revealed surprisingly strong associations, but with genes completely unrelated to cancer. The breakthrough came when we examined how CRISPR guides were overlapping with these high association SNPs. We discovered that these SNPs were preventing the guides from binding to their target regions, creating false non-essentiality results. Even more significantly, when examining ancestry data we had recently computed, the affected cell lines were predominantly from patients of African descent.</p>

<h2 id="impact">Impact</h2>

<p>🤩 What started as a search for new cancer targets led to the discovery of a significant bias in our CRISPR guide library affecting 1-2% of the guides. This finding has now been incorporated into our CRISPR dataset and has important implications for ensuring more equitable and accurate research outcomes.</p>

<p>This project represents a true scientific journey - starting from an initially contested idea, leading to unexpected results, and culminating in an important discovery that helps address representation issues in cancer research. 🌍</p>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="Comp-Bio" /><category term="Research" /><category term="CRISPR" /><summary type="html"><![CDATA[How we discovered ancestry biases in CRISPR screens and published in Nature Communications]]></summary></entry><entry><title type="html">Managing GRNs and What They Mean</title><link href="https://www.jkobject.com/blog/manage-grn-and-what-they-mean/" rel="alternate" type="text/html" title="Managing GRNs and What They Mean" /><published>2024-06-17T00:00:00+00:00</published><updated>2024-06-17T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/manage-grn-and-what-they-mean</id><content type="html" xml:base="https://www.jkobject.com/blog/manage-grn-and-what-they-mean/"><![CDATA[<h1 id="managing-grns-and-what-they-mean">Managing GRNs and What They Mean</h1>

<p>In my recent paper <a href="https://www.biorxiv.org/content/10.1101/2024.07.29.605556v1">scPRINT</a>, I went into a deep dive on gene regulatory networks. Many findings are in the manuscript but I had some additional thoughts to share and wanted to present a tool I built to work with GRNs.</p>

<h2 id="what-are-gene-networks">What are Gene Networks?</h2>

<p>First we don’t use the term GRN most of the time but Gene Networks. This is because when GRNs are mentioned, they are almost always restricted to TF-gene connections. But gene expression / regulation is way more complex than this and driven by many more elements than just TFs. ♾️</p>

<p>Indeed, regulation happens with TF but also co-factors, cohesins and other proteins transporting, degrading, splicing the RNAs, shaping the DNA and working with TFs. Not only that but it is also well known that TFs themselves can be influenced and activated by pathways comprising many genes less related to transcription regulation. 🦠</p>

<p>Finally, a new realization in biology is that regulation also happens through RNAs interacting with each others, and with the DNA. So much happening but we are fixated on TF-genes… 🔍</p>

<h2 id="the-problem-with-simple-grns">The Problem with Simple GRNs</h2>

<p>The big issue in my opinion is when these simple computationally inferred GRN are used to perform simulation of gene expression or are conflated with a model of the cell! 😲</p>

<p>The cell is very complex and TF-gene binary connections won’t cut it, by a mile. ⛔ This is to me a big problem in the GRN and GRN inference field. We will need to better define what GRNs are and maybe use another term to talk about the TF-gene constrained networks: like TFGN instead of GRN.</p>

<p>Given all this information -and the still unknown amount of non coding RNAs in human cells- we have to accept that a true Gene Regulatory Network should implicate all RNAs, non coding RNAs, DNA regions, proteins, and likely other molecular products. 👩‍👩‍👧‍👦</p>

<p>Finally, connections are not binary, they are not weighted, molecular interactions are complex, can create complexes, partially interact and have conditionals. We will also need better representation of connections, likely using embedding representation of vertices. 🔀</p>

<h2 id="where-do-we-go-from-here">Where Do We Go From Here?</h2>

<h3 id="1-theoretical-foundations-">1. Theoretical Foundations 📋</h3>
<p>We need good theoretical definitions that most people agree on, however murky things still are to us.</p>

<h3 id="2-better-tools-and-infrastructure-">2. Better Tools and Infrastructure 🧰</h3>
<p>We need to build tools, visualisation, data-structures, and benchmarks that represent these definitions and concepts and help us interrogate the current and future datasets we might generate. This is what I tried to start with <a href="https://github.com/cantinilab/GRnnData">GRnnData</a> and <a href="https://github.com/jkobject/benGRN">benGRN</a>. But lot more work is needed.</p>

<h3 id="3-advanced-sequencing-methods-">3. Advanced Sequencing Methods 📏</h3>
<p>We need better sequencing methods than mRNA sequencing. Single cell is only a first step but we need to measure non coding and rare RNAs with much higher precision.</p>

<p>Who knows, it might be that a large-scale, high-depth, single-cell total-RNA-sequencing methodology with correct molecular amplification / reduction methods is all that we need to understand regulation… 🤔</p>

<h1 id="grnndata">GRnnData</h1>

<p>in scRNAseq data we have been fortunate to get many great python toolkits like anndata and scanpy, part of the scverse suite of tools. Allowing us to work together with a common set of standards. 📜</p>

<p>However, working with gene network methods, I have seen various ways to store them throughout the different papers and benchmarks. Often as some kind of tsv/csv/… file with some kind of a gene-gene list. This lack of standard made it quite hard for me and other members of the lab to work with gene networks 🙉</p>

<p><img src="/assets/images/grn1.png" alt="" /></p>

<p>Interestingly, there is a possible standard for it! 🎉AnnData contains the .varp field which is made to store var to var (e.g. genes to gene) relationships. However not many people use it…</p>

<p><img src="/assets/images/anndata.png" alt="" /></p>

<p>I thus decided to formalize this usage by creating <a href="https://github.com/cantinilab/GRnnData">GRnnData</a>: it is basically a way to important many different gene-gene format to an AnnData file. 💁 But it also contains more bells and whistles to work with gene networks like subsetting the network to some genes, extracting targets, and plotting the networks 💹.</p>

<p><img src="/assets/images/grn2.png" alt="" /></p>

<p>But <a href="https://github.com/cantinilab/GRnnData">GRnnData</a> can do more and integrates some utils functions doing things like clustering, centrality measures, enrichment and more. Go check <a href="https://github.com/cantinilab/GRnnData">GRnnData</a> if you work with gene networks!</p>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="Comp-Bio" /><category term="Research" /><summary type="html"><![CDATA[Some thoughts on Gene Regulatory Networks and their meaning in biology]]></summary></entry><entry><title type="html">AUPRC vs AP: Evaluating Binary Classification</title><link href="https://www.jkobject.com/blog/auprc-vs-ap/" rel="alternate" type="text/html" title="AUPRC vs AP: Evaluating Binary Classification" /><published>2024-06-09T00:00:00+00:00</published><updated>2024-06-09T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/auprc-vs-ap</id><content type="html" xml:base="https://www.jkobject.com/blog/auprc-vs-ap/"><![CDATA[<p>When analysing a classification task for my recent paper <a href="https://www.biorxiv.org/content/10.1101/2024.07.29.605556v1">scPRINT</a>, I fell into a fascinating “data-sciency” rabbit-hole about AUPRC, AUC and AP.</p>

<h2 id="understanding-classification-metrics">Understanding Classification Metrics</h2>

<p>When evaluating a binary classification model that outputs probabilities or has internal regularization, we need ways to assess its performance across different decision thresholds.</p>

<p><img src="/assets/images/auprc1.png" alt="Classification Example" /></p>

<h3 id="roc-curves">ROC Curves</h3>
<p>One common approach is plotting the ROC curve, which shows the tradeoff between true positive and false positive rates. The area under this curve (AUC/ROC-AUC/ROC/AUROC) is a popular metric:</p>

<p><img src="/assets/images/auprc2.png" alt="ROC Curve" /></p>

<h3 id="the-need-for-precision">The Need for Precision</h3>
<p>However, in some cases like predicting graph edges in sparse gene networks, precision becomes crucial. AUROC can miss important information about optimal cutoff points. 🚫</p>

<p><img src="/assets/images/auprc3.png" alt="Precision Example" /></p>

<h3 id="auprc-and-its-challenges">AUPRC and Its Challenges</h3>
<p>This is where AUPRC (PR-AUC) becomes valuable, focusing on precision and recall:</p>

<p><img src="/assets/images/auprc4.png" alt="AUPRC Curve" /></p>

<p>However, AUPRC comes with its own challenges: ⚠️</p>

<ol>
  <li>Random precision baselines differ between tasks, making direct comparisons difficult</li>
  <li>Jagged results require careful sampling of PR curves 😟</li>
  <li>In gene network inference, many links lack prediction values, requiring special handling 🔥</li>
</ol>

<h2 id="introducing-rauprc">Introducing rAUPRC</h2>

<p>To address these limitations, I developed <strong>rAUPRC</strong> (random-baseline-corrected AUPRC), implemented in <a href="https://github.com/jkobject/benGRN">benGRN</a>. It differs from vanilla AUPRC in three key ways:</p>

<p><strong>1. Baseline correction</strong> — Standard AUPRC is heavily influenced by the positive class ratio. A dataset with 1% positives will have a baseline AUPRC of ~0.01, making comparisons across tasks meaningless. rAUPRC normalizes for this: a random classifier always scores 0, a perfect classifier scores 1, regardless of class imbalance. 🎯</p>

<p><strong>2. Handling missing predictions</strong> — In gene network inference, most tools only predict a subset of possible links; the rest are simply absent (not predicted as negative). Vanilla AUPRC breaks here. rAUPRC handles this by treating missing predictions explicitly when drawing the recall axis, giving a fair picture even when coverage is partial. 🔍</p>

<p><strong>3. Curve interpolation</strong> — PR curves are notoriously jagged, especially at low recall. rAUPRC uses a modified interpolation that avoids the artificial inflation that naive linear interpolation can introduce between operating points.</p>

<p>In practice on GRN benchmarks, rAUPRC produced more stable rankings across different network densities and prediction tool outputs than either vanilla AUPRC or AP. 📊</p>

<h2 id="average-precision-ap">Average Precision (AP)</h2>

<p>AP is computed as:</p>

<p><img src="/assets/images/auprc5.png" alt="AP Formula" /></p>

<p>AP is a discrete approximation of AUPRC — rather than integrating the continuous curve, it sums precision at each threshold where a positive is retrieved, weighted by the change in recall:</p>

\[AP = \sum_n (R_n - R_{n-1}) \cdot P_n\]

<p>In practice, AP and AUPRC give very similar values when the PR curve is sampled densely. The key difference is that AP is purely a summary statistic (one number per run), while AUPRC can be computed from the full curve and lends itself better to baseline correction.</p>

<p><strong>AP is often recommended</strong> for object detection and ranking tasks (it’s the default in PASCAL VOC and COCO benchmarks). For gene network inference and other sparse biological tasks, rAUPRC is more appropriate. ✅</p>

<h2 id="when-to-use-what">When to use what</h2>

<table>
  <thead>
    <tr>
      <th>Metric</th>
      <th>Use when</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>AUROC</strong></td>
      <td>Classes are roughly balanced, you care about ranking overall</td>
    </tr>
    <tr>
      <td><strong>AUPRC / AP</strong></td>
      <td>Strong class imbalance, precision matters</td>
    </tr>
    <tr>
      <td><strong>rAUPRC</strong></td>
      <td>Gene networks, partial predictions, cross-task comparison</td>
    </tr>
  </tbody>
</table>

<p>For a deeper dive into the implementation, see the <a href="https://github.com/jkobject/benGRN">benGRN repo</a> and the benchmarking section of the <a href="https://www.biorxiv.org/content/10.1101/2024.07.29.605556v1">scPRINT paper</a>.</p>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="Data-Science" /><category term="Research" /><summary type="html"><![CDATA[A deep dive into classification metrics and their nuances]]></summary></entry><entry><title type="html">Enrichr, Prerank, GSEA or ssGSEA?</title><link href="https://www.jkobject.com/blog/enrichr-prerank-gsea/" rel="alternate" type="text/html" title="Enrichr, Prerank, GSEA or ssGSEA?" /><published>2024-02-19T00:00:00+00:00</published><updated>2024-02-19T00:00:00+00:00</updated><id>https://www.jkobject.com/blog/enrichr-prerank-gsea</id><content type="html" xml:base="https://www.jkobject.com/blog/enrichr-prerank-gsea/"><![CDATA[<p>Bioinformatician’s main tool for discovery has often been differential expression analysis. But between Enrichr, Prerank, GSEA, and ssGSEA, which tool should you use? Here is the quick reminder X-plainer. 🧬</p>

<h2 id="the-decision-tree">The Decision Tree</h2>

<p>The key question is: <strong>what shape is your data?</strong></p>

<hr />

<h3 id="enrichr--you-have-a-gene-list-nothing-else">Enrichr — you have a gene list, nothing else</h3>

<p><em>Enrichr</em> is when you just have a list of genes (can be small). No values, no conditions — just names. 📋</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>GENEA, GENEB, GENEC
</code></pre></div></div>

<p>Under the hood, Enrichr tests each gene set in its databases using a <strong>Fisher’s exact test</strong> (hypergeometric). It asks: is my list enriched for genes in pathway X more than expected by chance?</p>

<p>Best for: DE gene lists, hit lists from CRISPR screens, manually curated sets.</p>

<table>
  <tbody>
    <tr>
      <td>👉 <a href="https://maayanlab.cloud/Enrichr/">Enrichr</a></td>
      <td>Python: <a href="https://gseapy.readthedocs.io/en/latest/">gseapy.enrichr</a></td>
    </tr>
  </tbody>
</table>

<hr />

<h3 id="prerank--you-have-a-ranked-gene-list">Prerank — you have a ranked gene list</h3>

<p><em>Prerank</em> is when you have a continuous value per gene that you can rank — a fold change, a correlation, a t-statistic, anything. 📊</p>

<table>
  <thead>
    <tr>
      <th>Gene</th>
      <th>Value</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>GENEA</td>
      <td>12</td>
    </tr>
    <tr>
      <td>GENEB</td>
      <td>8</td>
    </tr>
    <tr>
      <td>GENEC</td>
      <td>4</td>
    </tr>
  </tbody>
</table>

<p>Prerank runs GSEA logic (the enrichment score / walking statistic) on your pre-ranked list, without needing raw expression data or phenotype labels. Useful when you already have a score but not the underlying samples.</p>

<p>Best for: correlation with a phenotype, output from another model, single-sample pseudo-bulk scores.</p>

<p>👉 Python: <a href="https://gseapy.readthedocs.io/en/latest/">gseapy.prerank</a></p>

<hr />

<h3 id="gsea--you-have-expression-data-with-two-conditions">GSEA — you have expression data with two conditions</h3>

<p><em>GSEA</em> works best when you have a matrix of gene expression values across multiple samples with a clear phenotype label (treated vs control, disease vs healthy, etc.). It computes its own gene ranking internally. 🔬</p>

<table>
  <thead>
    <tr>
      <th>Gene</th>
      <th>C1</th>
      <th>C2</th>
      <th>C3</th>
      <th>D1</th>
      <th>D2</th>
      <th>D3</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>GENEA</td>
      <td>12</td>
      <td>7</td>
      <td>3</td>
      <td>1</td>
      <td>1</td>
      <td>0</td>
    </tr>
    <tr>
      <td>GENEB</td>
      <td>8</td>
      <td>0</td>
      <td>6</td>
      <td>8</td>
      <td>1</td>
      <td>1</td>
    </tr>
    <tr>
      <td>GENEC</td>
      <td>4</td>
      <td>4</td>
      <td>3</td>
      <td>2</td>
      <td>3</td>
      <td>4</td>
    </tr>
  </tbody>
</table>

<p>The key advantage over Enrichr: GSEA doesn’t require you to define a hard cutoff (“top 200 DE genes”). It uses the full ranked list and identifies pathways enriched at the top or bottom. This makes it more sensitive and less arbitrary. ✅</p>

<p>Best for: bulk RNA-seq, any two-condition comparison with replicates (n ≥ 3 per group recommended).</p>

<table>
  <tbody>
    <tr>
      <td>👉 <a href="https://www.gsea-msigdb.org/gsea/index.jsp">GSEA software</a></td>
      <td>Python: <a href="https://gseapy.readthedocs.io/en/latest/">gseapy.gsea</a></td>
    </tr>
  </tbody>
</table>

<hr />

<h3 id="ssgsea--you-want-a-per-sample-enrichment-score">ssGSEA — you want a per-sample enrichment score</h3>

<p><em>ssGSEA</em> is for when you have many samples with no clear two-group contrast — or when you want a continuous enrichment score <em>per sample</em> rather than a comparison between groups. 🗂️</p>

<table>
  <thead>
    <tr>
      <th>Gene</th>
      <th>A</th>
      <th>B</th>
      <th>C</th>
      <th>D</th>
      <th>E</th>
      <th>F</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>GENEA</td>
      <td>12</td>
      <td>7</td>
      <td>3</td>
      <td>1</td>
      <td>1</td>
      <td>0</td>
    </tr>
    <tr>
      <td>GENEB</td>
      <td>8</td>
      <td>0</td>
      <td>6</td>
      <td>8</td>
      <td>1</td>
      <td>1</td>
    </tr>
    <tr>
      <td>GENEC</td>
      <td>4</td>
      <td>4</td>
      <td>3</td>
      <td>2</td>
      <td>3</td>
      <td>4</td>
    </tr>
  </tbody>
</table>

<p>Each sample gets its own enrichment score for each pathway, independently. The output is a sample × pathway matrix. Great for downstream analysis — clustering, survival analysis, correlating pathway activity with other variables.</p>

<p>Best for: large cohorts (TCGA, GTEx), single-cell pseudo-bulk, any analysis where you want pathway activity as a continuous feature.</p>

<p>👉 Python: <a href="https://gseapy.readthedocs.io/en/latest/">gseapy.ssgsea</a></p>

<hr />

<h2 id="quick-summary-table">Quick summary table</h2>

<table>
  <thead>
    <tr>
      <th>Tool</th>
      <th>Input</th>
      <th>Statistics</th>
      <th>Best for</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Enrichr</strong></td>
      <td>Gene list only</td>
      <td>Fisher / hypergeometric</td>
      <td>Small lists, no values</td>
    </tr>
    <tr>
      <td><strong>Prerank</strong></td>
      <td>Genes + score</td>
      <td>GSEA walking statistic</td>
      <td>Pre-computed rankings</td>
    </tr>
    <tr>
      <td><strong>GSEA</strong></td>
      <td>Expression matrix + 2 conditions</td>
      <td>GSEA walking statistic</td>
      <td>Bulk RNA-seq DE</td>
    </tr>
    <tr>
      <td><strong>ssGSEA</strong></td>
      <td>Expression matrix, no labels</td>
      <td>Per-sample enrichment</td>
      <td>Large cohorts, per-sample scores</td>
    </tr>
  </tbody>
</table>

<p>For most single-cell work: compute pseudo-bulk, run GSEA or Prerank per cell type. For single-cell pathway scoring directly, <a href="https://decoupler-py.readthedocs.io/">decoupleR</a> is worth a look. 🔍</p>]]></content><author><name>Jérémie Kalfon</name><email>jkobject@gmail.com</email></author><category term="PhD" /><category term="Comp-Bio" /><category term="Research" /><summary type="html"><![CDATA[A quick reminder of when to use which enrichment analysis tool]]></summary></entry></feed>