A demonstration of bespoke visualization and analysis tools; not intended to be peer-reviewed — no novel scientific result or discovery is claimed. Independent of, and not endorsed by, the Royal Botanic Gardens, Kew. What this is, and how the numbers were computed →
Peter Repetti · July 16, 2026 · CC BY 4.0
Angiospermae · an interactive demonstration built on Kew's Tree of Life · release 4.0
Kew's Tree of Life places 353 nuclear genes across 20,475 flowering-plant samples spanning every order and family, with the gene trees and alignments all published. That is target capture, not whole genomes — about 260 kbp of sequence, under 0.1% of a genome. Those same loci are a coordinate system the whole tree already shares — and this demonstration reads them three ways. Pick a clade and see, across all 353 loci, which are recovered, which carry signal at that depth, which hold the clade together against the rest of the tree, and how divergent they are. Pin them to a sequenced genome and they land dispersed across its chromosomes, each inside a gene — broad genomic dispersion, consistent with the reduced linkage the coalescent tree assumes. Thread them through several genomes in the tree's own order and they become synteny anchors, showing where chromosomes keep their shape across a clade, and where a lineage has dealt them into a different hand.
Three orders carry real computed results: , and — click one here, or a green-highlighted tip in the tree below. For them, Recovery, Informativeness, Clade cohesion and Divergence are all measured from Kew's data; the rest of the tree is drawn for context, not clickable — the other orders carry no clade-specific data to show. (Node concordance is the one exception among the measures: a proposed placeholder, even for these three.) The panel opens on Poales with Clade cohesion showing, so the first thing you see is a real result.
Click a highlighted order, then choose a measure — Recovery, Informativeness, Clade cohesion, Node concordance or Divergence. The 353-bar panel recomputes for that clade — every bar is one Angiosperms353 locus, in Kew's gene order. For Poales and Fagales you can also drop a level to the family nested inside — Poaceae or Fagaceae — with the grain toggle above the measures, which reads the same 353 loci one level deeper. The note under the buttons says what the measure means and what to look for. And any single locus is a thread you can pull through the whole page: click one — a bar here, a table row, or a chromosome tick or ribbon in the genome panels below — and it stays lit in all of them at once, and stays lit as you switch clades. To reach a locus you can't see among the 353 bars — say PSAG (6498) — type its id or gene name into the find box above the chart. Shift-click bars or rows to pin several at once; each pinned locus shows as a chip you can release.
| Locus | Exemplar | Species | Genera | Cohesion · QCF |
|---|
What comes from Kew, and what this demonstration computes. The tree, its support values, the 353 locus identifiers, their exemplar Arabidopsis gene names, and per-locus species, genus and recovery figures are all from Kew's data. For Fagales, Chloranthales and Poales — and, one level down, the families Fagaceae and Poaceae — this demonstration computes three measures: Clade cohesion is a quartet concordance factor on each of the 353 gene trees — the fraction of quartets, drawn as two tips inside the clade against two from anywhere outside it, in which the clade's pair stays together; Informativeness is the observed fraction of parsimony-informative sites within the clade; and Divergence is the mean pairwise sequence divergence within the clade. Cohesion answers the clade's own premise: Poales is non-monophyletic on almost every gene tree, yet its cohesion stays above 0.99 — the size artefact that disqualifies monophyly counts, shown directly. Because "outside" spans the whole tree, it measures the clade's boundary against everything else, not resolution among its internal branches. Divergence ships the observed value; the Jukes-Cantor-corrected value is shown labelled beside it, never in its place. Node concordance remains simulated for every clade. The measured values are observed counts — no rate model, no assumed timescale.
The same loci, one level down — and the two families disagree about how much that matters. Drop Fagales to Fagaceae and every locus carries about a third as much signal: median parsimony-informative fraction falls from 0.36 to 0.12, median divergence from 0.084 to 0.025 — within the family the sampled tips are shallower and closer, so less varies. Drop Poales to Poaceae and the budget barely moves: informativeness 0.82 to 0.76, divergence 0.17 to 0.12, the ranking almost unchanged (rank correlation 0.93). Poaceae is a deep radiation that also makes up much of the sampled Poales, so the order and family views overlap heavily and dropping a level costs little — a contrast that reflects sampling depth and composition as much as absolute clade age. In both, the ranking of loci is broadly preserved, and Clade cohesion differs only in the tail: loci that fight the deep order tree resolve cleanly within the family — PSAG (6498) does it in both, rising from QCF 0.61 to 0.93 in the oaks and 0.75 to 1.00 in the grasses, because its conflict is with deep relationships, not the family. How much depth matters depends on which clade you are standing in.
Four questions a working systematist asks, and how to answer each by clicking. Every named locus is a measured result — select the clade in the tree, then the measure, and read it off.
Select your clade, open Informativeness to see which loci carry signal at that depth,
then Clade cohesion to see which of those hold the clade together against the rest of the
tree. Loci tall on both are
the workhorses. For Poales, genes 5335, 7021 and
5599 reach QCF 1.00 with ~84–90% informative sites — start a dataset there,
not from a locus you happened to have a primer for.
On Clade cohesion, hunt the rust bars: loci whose gene tree
places your taxa outside the clade. Cross-check Informativeness — if the same locus is
also tall there, the disagreement is real signal, not noise, and worth investigating.
Gene 6376 (INDL) is the deepest dip in Fagales
(QCF 0.62), Poales (0.72) and Chloranthales (0.64) alike — a locus that fights the
species tree across unrelated clades is a paralogy suspect you would want flagged before it
poisons a concatenated analysis.
Open Divergence: warmer bars are more divergent inside your clade. High divergence can
erode historical signal, but on its own it does not diagnose saturation — repeated substitutions,
paralogy, alignment error and compositional bias all look alike in a raw p-distance. What is worth
a second look is a locus that is both highly divergent and rust on Clade cohesion.
Gene 6376 (INDL) is the single most divergent locus in
all three clades — about 23–27% — and also the most discordant.
That combination makes it a quality-control priority, not a settled verdict: check for saturation,
paralogy, alignment error or compositional bias before a concatenated tree averages it in.
Recovery shows how much of each locus comes back, averaged tree-wide across mixed source
types (target capture, genome- and transcriptome-mined) — a coarse guide, not a per-clade or
per-collection prediction.
Genes 6507 and 6406 are the two lowest-recovery
loci tree-wide, at 43.8% and 46.4%; a locus that recovers poorly across the tree is a weak default choice, though it
may still work in a particular lineage. Read it beside Informativeness before dropping anything
from a panel.
The habit this builds. Notice that Poales looks absent under a monophyly check — non-monophyletic on all 353 gene trees — yet Clade cohesion shows it is 99.7% concordant. Reading cohesion instead of counting monophyly is the difference between "this clade falls apart" and "this clade is solid; three loci disagree, and here they are."
Kew's tree resolves every flowering-plant lineage from the same 353 nuclear markers — about 260 kbp, under 0.1% of a genome. Attach one whole genome to a single tip — Quercus robur — and this is what appears: every one of the Angiosperms353 loci at its real position on the twelve chromosomes, in genomic context. Hover a locus in the panel above, or a tick here, and it lights up in both. Do this wherever the tree meets a sequenced genome and the 353 markers stop being a list and become a map. Methods
What a genome layer adds. In the panel above a locus is an
identifier and a number; anchored to a genome it gains an address — locus
5599 (NERD1) sits on chromosome 6, inside gene LOC126733144, beside
its neighbours. And the pattern across all 353 is the assumption the tree rests on: they occupy just
· of the genome yet scatter across all twelve chromosomes, and they land
where the genes are — the gene-poorest quarter
(about · Mb, mostly pericentromeric) holds ·
of them, the gene-richest holds ·, and · of
353 fall inside an annotated gene. Loci this dispersed and this genic are consistent with the reduced
linkage the coalescent species tree assumes of them — dispersion lowers linkage without proving the
loci independent. That is one
branch's worth of a genome layer. Run the same
alignment on the other sequenced Fagaceae genomes and the shared loci become synteny
anchors, tying one genome to the next — which is the next panel.
Now thread the same 353 loci through five Fagaceae genomes, stacked in the order Kew's own tree gives them. All five assemblies report twelve chromosomes. Across chestnut, stone oak and the two oaks, those twelve are the same twelve — each locus stays on the matching chromosome, while the genome runs from 716 Mb to 964 Mb. The anchors keep their chromosome assignment as the genome grows; how much finer structure shifts between them is beyond what 353 markers resolve. Now beech, sister to the other four: still twelve chromosomes, but the anchors have been dealt into different hands. Methods & what's unresolved
How to read it. Each row is one genome, drawn to scale, so the bars lengthen as the genome grows. Each tick is one of the 353 loci at its real position; each ribbon follows one locus from the genome above to the genome below. Click any anchor to pin it — the locus stays lit here, on the oak chromosomes and in the 353-locus chart until you click it again. Colour marks the linkage group it occupies in Q. robur — six hues, hatched bars and dashed ribbons for groups 7–12, and every bar is labelled, so no group is told apart by colour alone. Between the lower four genomes the ribbons run in twelve clean bundles. Between beech and chestnut they cross.
What was measured, and what was not. Chromosome numbering is arbitrary in every assembly, so homology was computed from the shared loci themselves by maximum-weight matching — a one-to-one pairing by construction, so its credibility rests on the match purity below, not on the pairing itself. Measured against Q. robur, the other three of the lower four put · of shared loci on the matching chromosome, and each of the four resolves into exactly twelve syntenic blocks. Beech scores · and resolves into · blocks. That is not a broken assembly: the same mosaic appears in three independently assembled beech genomes — Sanger, Senckenberg and Genoscope — which agree with each other 98.0–99.4% and all disagree with the oaks. Which lineage did the rearranging, we cannot say. Beech is the sister lineage, so the change could sit on either stem branch; six Fagales outgroups all lean toward the oak arrangement, but only the two Juglandaceae clear a bootstrap interval and they are a single family, not independent replicates. Reported unresolved. These 353 anchors resolve which chromosome a locus sits on — not fine-scale order, inversions, or block boundaries. Chromosomes are drawn in whichever orientation agrees with the reference, since assembly orientation carries no meaning. Loci with a strong second hit are drawn at their best hit and counted in the table.
What these two panels could become across the tree. Both run on one fact: the same 353 loci are recovered at every tip, so they are a coordinate system every angiosperm genome already shares. Through them a newly sequenced genome pins to the tree with no whole-genome alignment. Stack any clade's genomes in Kew's topology and the beech mosaic stops being a single finding and becomes a question you can ask anywhere — where do the linkage groups hold, and where has a lineage dealt them into a different hand? Count how many times each anchor appears in a genome and repeated hits flag a locus for duplication follow-up — though 353 deliberately low-copy markers can only raise the question, not diagnose whole-genome duplication, which needs genome-wide syntenic evidence. The empty branches carry information too: only about 1,222 of some 13,600 flowering-plant genera have any genome at all, so the map shows which lineages, if sequenced, would tie the most of the tree together. Each of these is a direction these tools could open across the full tree — an illustration of what Kew's data could power, not a result claimed here.
Where each figure's numbers come from, how they are computed, and what to distrust. Every measured value is observed, not modelled, and nothing is computed in the browser — the page embeds a precomputed table and fetches nothing. Each entry names the script that reproduces it. Informativeness, Clade cohesion and Divergence are measured from Kew's published alignments and gene trees for three orders — Fagales, Chloranthales, Poales — and two families within them, Fagaceae and Poaceae. For every other clade, and for Node concordance throughout, the bars are a placeholder.
treeoflife.kew.org/api;
bulk files at
sftp.kew.org/pub/paftol/current_release.-nan for some values. They are local
posterior probabilities (0–1).fasta/alignments/{gene}.dna.aln.fasta (~8.1 GB, streamed, never stored),
restricted to the taxa in the selected clade.MIN_SCORED = 4).MIN_SCORED = 4). In a
densely sampled clade a sparse or short locus can therefore post a high fraction on few taxa;
read it beside occupancy and total aligned length, and treat it as a within-clade ranking, not
an absolute yield.analysis/clade_locus.py (orders),
analysis/family_clade.py (families), via informativeness.py.tree/gene/{gene}.tree (Newick, internal
labels are support percentages 0–100).analysis/qcf.py (self-tested against brute force).analysis/clade_locus.py (orders),
analysis/family_clade.py (families), via informativeness.py —
the mean pairwise divergence is computed alongside the informative-site count.family field of
Kew's specimen metadata instead of order; the computation is otherwise identical.family_clade.py --family Poaceae → merge_family.py
→ embed.py.GCF_932294415.1
(dhQueRobu3.1), 12 chromosomes.minimap2
(spliced, best hit per locus), plus gene density along each chromosome.GCA_964211995.1,
Castanea sativa
GCF_040712315.1,
Lithocarpus litseifolius
GCA_040182985.1,
Q. lobata
GCF_001633185.2,
Q. robur
GCF_932294415.1.minimap2;
native chromosome → reference linkage group computed from the shared loci by maximum-weight
bipartite matching (one-to-one by construction; its credibility is the match purity below,
not the pairing). Stacked in Kew's
own topology, (Fagus,(Castanea,(Lithocarpus,(Q. lobata, Q. robur)))).analysis/synteny.py, checks in
analysis/synteny_checks.py.You do not have to program to check whether the analysis is sound. What matters in a script is not the machinery but a handful of decisions — thresholds, what counts as data, what gets measured — and those read in plain terms. Two examples, one you can read and one you can't.
A script you can read. This is the core of
informativeness.py, behind the Informativeness bar — lightly simplified, the array
bookkeeping trimmed, the decisions left as they run:
MIN_SCORED = 4 # a column needs ≥4 of your clade's species present to count at all # For each alignment column, count how many A, C, G, T appear across your clade's # sequences. Only real bases count — gaps (-), N and ambiguity codes are treated as # missing. `present` = number of species with a real base in that column. scored = present >= MIN_SCORED n_states = bases_present(column) # how many of A,C,G,T appear at all: 1 = constant n_states_ge2 = bases_in_two_or_more(column) # how many appear in ≥2 species variable = count_columns( scored & (n_states >= 2) ) # species differ here pis = count_columns( scored & (n_states_ge2 >= 2) ) # ...differ usefully for a tree pis_frac = pis / cols_scored # ← the Informativeness bar
Four decisions, all legible without reading the machinery: the ≥4 floor (score a locus only where at least four of your clade's species have real sequence — raise it for a sparse clade and rerun); what counts as data (only A, C, G, T, so a gappy column can't masquerade as variable); the variable-vs-informative distinction (a lone odd base varies, but a tree can't use it); and the denominator (scored columns, not full length, so a locus isn't charged twice for missing data). A systematist can accept, reject, or change any of them without touching the array code.
A script you can't — in plain words. Clade cohesion is
computed by qcf.py, which is genuinely dense: it counts quartets, and there are far too
many to list (Poales against the rest of the tree is about 260 trillion), so it totals them by a
bookkeeping trick in a single pass over the tree. You do not need to follow that. What it does — and
what you can question — is this. It takes one gene tree and one clade, considers every way of picking
two species from your clade and two from outside it, and asks whether the gene tree keeps your two
together; cohesion is the fraction where it does, over the quartets the tree actually resolves.
The choices inside are legible even though the code isn't: "outside" means every other tip on the
tree, gymnosperms included; weakly supported branches (below bootstrap 30) are collapsed
first, so a shaky branch gets no vote; and unresolved quartets are excluded from both the
count and the total. The fast counting is checked against a slow, obviously-correct version that
really does enumerate every quartet, on small trees, and the two must agree. Those choices — what
"outside" is, the collapse threshold, dropping ties — are exactly what a systematist would weigh, and
none of them requires reading the counting code.