Book VII
Bioinformatics
single-cell data, gene regulatory networks, perturbation screens
Show how machine learning meets biological data, and where a perturbation turns a correlation into a causal question.
10 chapters0 propositions written
Read first
neural networks · causal inference
Mathematics assumed — follow these when a step stops making sense.
Distributions, discrete and continuous 0.PR.01 · Multiple comparisons 0.ST.04 · Rank, eigenvalues and the singular value decomposition 0.LA.04 · Kullback–Leibler divergence 0.IT.03
- Chapter 1VII.1Biological InformationBiological information is discrete, layered and measured indirectly, so every dataset here is a measurement rather than the thing itself.DNA · RNA · Proteins · Genes · Genomes · Cells
- Chapter 2VII.2Biological DataBiological counts are overdispersed and sparse, and a method that assumes otherwise fails in a direction it will not report.Sequencing · Expression · Single-cell data · Imaging · Spatial transcriptomics · Perturbation data
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 3VII.3Gene RegulationRegulation is cooperative and non-linear, and a correlation between two genes is not a regulatory edge.Transcription · Regulatory networks · Transcription factors · Enhancers · Gene regulatory networks
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 4VII.4Single-Cell BiologyA single-cell pipeline is a sequence of decisions, each of which changes the geometry a later clustering will read as biology.scRNA-seq · Cells as observations · Gene-expression matrices · Dimensionality reduction · Clustering · Cell types
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 5VII.5Representation Learning for BiologyThe machinery of a latent-variable model is domain-neutral; what makes it biological is the likelihood and what the nuisance terms absorb.Embeddings · Autoencoders · Variational autoencoders · Transformers · Foundation models for biology
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 6VII.6PerturbationA perturbation screen is an experiment in the exact sense of Book VI, and it identifies edges that observation leaves undecided.CRISPR · Perturb-seq · Interventions · Perturbation response · Causal interpretation
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 7VII.7Biological Sequences as GraphsSequence alignment is a shortest path through a grid, and genome assembly is an Eulerian path through a de Bruijn graph.
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 8VII.8Phylogenetics and Evolutionary GraphsA phylogeny is a tree inferred from distances or characters, and the processes that break tree-likeness are exactly the ones that matter most.
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 9VII.9Biological NetworksProtein, regulatory, metabolic and disease networks are four different graphs, and a centrality claim means something different in each.Graphs · Protein interaction networks · Gene regulatory networks · Graph neural networks
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 10VII.10Molecular Graphs and Drug DiscoveryA molecule is a graph over atoms, and how the data is split decides whether any reported score generalises.Molecular representation · Protein structure · Drug discovery · Disease modelling · Biomarker discovery
0/3 problems0/3 variants0/6 exercisesowes 9 more
Practical connection
Single-cell preprocessing, then clustering, then a de Bruijn assembler on a toy genome, then a GCN over a small interaction network.
The assembler must recover the sequence you traced by hand through the Eulerian path, base for base.
Verified againstVII.4.B01 · VII.7.B02 · VII.9.B02