Sampling
One small area of a working farm, sampled once: about hectares of former orchard, historically disturbed - it was used for informal dumping before restoration began - and now being brought back. Sheep have been put through it times a year for around years. None of the plants sampled was sown. Soil was taken from beneath five kinds of plant growing in patches within that area - chicory, fescue, ragweed, red clover and white clover - giving samples in all.
The five plants were the comparison, and the fact that they shared one area was what made the comparison work. Soil type, slope, drainage, weather and the broad history of this ground were the same for all cores. So when the communities differed, the usual excuses - different soils, different fields, different days - were ruled out by how the sampling was done rather than by argument afterwards. What is not ruled out is the ground within the area: with cores at least three meters apart, a plant cannot be separated from the spot it grew in.
The reverse of that strength is the limit: nothing here can be attributed to how the farm is run, or to restoration. There is no undisturbed reference ground, no earlier time point, and no second field. An earlier version of this site described the study as a survey of a regenerative pasture; that was wrong, and it was corrected on 1 August 2026. The sheep matter for a specific reason: animals drop nitrogen in patches, those patches were not mapped, and nitrogen is what the archaeal result later gets read against.
The protocol was deliberately plain. Each patch was found and marked on the day of collection, chosen where one plant had the most stems within about thirty centimeters and little else competed, and kept at least three meters from the next so no two samples sat on the same ground. Sampling ran over two consecutive days in July 2025. At each spot a single vertical plug was cut with a forty-five-centimeter spade, so every horizon in the profile was represented; the whole plug - around four liters - went into one unused plastic bag. The spade was wiped clean of visible soil between spots but not chemically disinfected. The bags were frozen at roughly −18 °C and stayed frozen throughout. About two weeks on, each was thawed, mixed through and subsampled into sterile tubes with a metal spoon cleaned with ethanol between samples, sealed, and shipped overnight on dry ice to the sequencing provider.
One of the white-clover samples came from a compacted vehicle track. It was kept and reported as an incidental observation, not treated as a designed contrast - compaction changes soil profoundly, and folding it in silently would be smuggling a second experiment into the first.
In the lab
DNA was extracted from each sample as a mixture - everything present, plant and animal and microbe together. Two short marker genes were then copied (amplified) out of that mixture and sequenced. A marker gene is a stretch of DNA that every organism in a group carries and that differs just enough between them to tell species apart, so copying one is how you make a census of one group and ignore everything else.
| Marker | What it reads | Region | Forward primer | Reverse primer | Reference database |
|---|
Extraction, library preparation and sequencing were done by a commercial provider (Novogene), using a soil-specific extraction kit and paired-end Illumina sequencing; the provider merged the read pairs and returned them. All samples were sequenced for both markers. For the eukaryotic marker the white clover sample and one chicory sample yielded too little amplified DNA to sequence, leaving , and one ragweed sample later fell below the depth the others were leveled to, leaving for the diversity analyses. That is why the food-web results covered fewer patches than the bacterial ones. The raw reads are public in the NCBI Sequence Read Archive under accession .
The letters in those primer sequences that are not A, C, G or T are deliberate:
M means "A or C", H means "anything but G". A primer
is made degenerate on purpose so that one ordered molecule matches the slightly
different versions of the same gene carried by organisms whose last common
ancestor lived billions of years ago. It is also why a primer cannot be found
in a read with a plain text search, which matters for
verifying public data further down.
No fungal ITS marker was ever sequenced. A folder delivered under that name contains the same sequences as the 18S run, matched against a fungus-only database - which forces every read, including protists and animals, onto a fungal name. Those outputs are artifacts and are not used. Fungi are described here only at the coarse level 18S actually supports.
The pipeline
Raw sequencing output is millions of short, error-prone reads. Turning them into a table of organisms and counts was done in , the standard toolkit for this kind of survey, in three steps:
- – reads per core, after quality filtering denoised by
- exact sequences, each treated as one organism named by a classifier
- genera abundant enough to test analyzed in
- Denoising. modeled the machine's error pattern and corrected reads back to the exact sequences that were really there, rather than lumping similar reads into approximate groups. The result is called an ASV - one distinct sequence, treated as one kind of organism. The older convention was to cluster reads at 97% similarity and call each cluster a species, which both merges organisms that differ and invents ones that do not exist; it also produces clusters that cannot be compared between studies, where an exact sequence can.
- Classifying. Each ASV was compared against a curated reference library with a classifier, which reported the most specific name the evidence supported - often a genus, sometimes only a phylum, sometimes nothing. Two consequences run through everything below: a fragment this short frequently cannot resolve a species, so genus is the working unit of this study, and the unnameable fraction is not noise to be dropped but soil we cannot identify.
- Analysis. Everything after that - diversity, ordination, the guild abundances, every figure and every number on this site - ran in .
Every tool, database and statistical test named on this page is listed with its reference on the sources page.
What came out
distinct bacterial and archaeal sequences across samples, at between and reads per sample.
A table like that is compositional, and everything downstream is shaped by it. A sequencing run returns a roughly fixed budget of reads, so what comes out are shares of a total that has nothing to do with how much life was in the ground. If one group doubles, every other group must appear to shrink even though nothing happened to it. That is why the clover result on the home page needed a section of stress tests rather than a correlation coefficient, and the reason one of those tests re-ran the whole thing on log-ratios instead of percentages.
Equalizing sequencing depth
The deepest sample was read times as deeply as the shallowest. A sample read more deeply shows more kinds of organism for no biological reason at all, so comparing diversity across samples of unequal depth partly measures the sequencing run rather than the soil.
Diversity was therefore computed on tables subsampled down to a common depth of reads for the bacterial marker (and for the eukaryotic one, which dropped one ragweed sample) - the shallowest sample's own count, which is the deepest level at which no sample has to be discarded. The subsampling used the multivariate hypergeometric distribution, which is exact sampling without replacement rather than the multinomial approximation usually used for it, and the whole draw was repeated times and averaged so that no reported value depends on one lucky shuffle.
The check that did not come out the way it usually does
Rarefaction is meant to break the link between how deeply a sample was read and how rich it looks. Here it barely moved it: richness tracked the original read depth at before subsampling and still after. The reading is that the deeply sequenced samples were genuinely richer rather than merely better sampled - depth and richness are correlated in this soil, not only in the machine. The confound has not been removed.
Is there as much life under one plant as another?
Three metrics, because "how much life" has three defensible meanings and a result that holds under one but not the others is a result about the metric. Each was tested across the five plants with a Kruskal–Wallis test, the rank-based alternative to ANOVA - chosen because with two or three cores per plant there is no way to establish that anything is normally distributed.
| Metric | What it counts | H | df | p |
|---|
All three are reported for a reason. A community can hold a great many kinds of organism that are all close relatives; richness and Shannon would call that diverse, and Faith's phylogenetic diversity - which measures the total length of evolutionary branch present in a sample - would not. Finding every plant equal under all three is a stronger statement than finding them equal under any one of them.
A flat result is not proof of equality. With cores across five groups these tests could only have detected a large difference. "No difference was detected" is what they say; "the plants host equal diversity" is more than this design can support.
Is it the same life?
A different question, and the one carrying the home page's result. It was answered by measuring how different every pair of samples was, then asking whether cores under the same plant were more alike than chance would make them - a test called PERMANOVA. The manuscript's version, run in R over reshufflings of the plant labels, gave R² = , pseudo-F = , p = , and that is the figure the pages quote. The pipeline in this repository re-runs the same test over reshufflings with the random seed fixed at , so that a reviewer reproduces its exact p-value rather than one near it; its result is in the table below and agrees to the second decimal.
Averaged over every pair, two cores under the same plant differed by and two cores under different plants by (Bray–Curtis, where 0 is identical and 1 shares nothing at all). That is same-plant pairs against different-plant pairs - a real gap, and a narrow enough one to need a test.
| Distance measure | What it counts | pseudo-F | R² | null | excess | p |
|---|
Why the raw R² is not the number
PERMANOVA's R² is routinely read as "the fraction of variation the grouping explains", which invites reading as most of it. That is wrong. R² here is a ratio of sums of squares whose expected value under a meaningless grouping is not zero but roughly (k − 1) / (n − 1) - with five plants and cores, - because within-group sums run over far fewer pairs than the total does.
So it was compared against itself. Shuffling the plant labels times put the empirical null mean at , within a thousandth of that analytic floor, and its 95th percentile at . The observed value sat above that percentile, and what gets reported is the of it chance does not account for (p = by that effect-size test, agreeing closely with the pseudo-F permutation p-value of ). The gap between the raw figure and the excess is the single most common way a study of this size is oversold.
A significant PERMANOVA has one notorious alternative explanation: groups can come out "different" merely because one is more internally variable than another, and the test cannot tell that from a genuine difference in location. A separate test - PERMDISP - compares each sample's distance to its own group's center. In the manuscript's analysis the dispersions did differ (p ), so the PERMANOVA result may reflect a difference in composition, a difference in spread, or both, and it is reported as suggestive rather than conclusive. Two sensitivity runs said the same: without the single white clover core the plant effect sat at R² = , p = ; without the two chicory cores as well it was R² = , p = .
The two implementations disagree here, and the manuscript's is
canonical. This repository's own dispersion test, over
permutations
(F = ,
p = ), did not reach
significance where R's betadisper did. Both are the same idea
with different centering and permutation schemes; the discrepancy has not
been resolved, and until it is, the pages take the more cautious reading.
The phylogeny, and why two of those rows exist
Bray–Curtis treats every ASV as equally unrelated to every other. Two samples that differ only in which strain of Nitrososphaera they carry look exactly as different, under that measure, as two samples that differ by an entire phylum. UniFrac repairs this by weighting each difference by the length of evolutionary branch separating the taxa, and Faith's phylogenetic diversity does the same for the single-sample metrics. Agreement between the two families is evidence that a result is about community biology rather than about a choice of arithmetic.
That needs a tree, and the 16S run shipped one - which no analysis in
this project had ever opened, for a mundane reason. The exported count table
had been renamed to ASV00001-style identifiers while the tree kept
the pipeline's original sequence hashes, so the two files shared not one label
and any straightforward join produced an empty table with no error raised. A
stored mapping between the two naming schemes is the bridge; the loader applied
it, refused to continue if the join came back empty rather than reporting
zeros, rarefied to reads, then trimmed the
tree down to exactly the sequences that survived.
Rarefying first is not optional here. Unweighted UniFrac and Faith's PD both count rare branches, and a more deeply read sample is simply more likely to have stumbled onto one, so an uncorrected comparison across a × depth range would partly measure the sequencing run. The two UniFrac variants answer different questions, which is why both are in the table above: the weighted one is dominated by which organisms are abundant, the unweighted one by which rare lineages are present at all. The weighted variant gave the largest excess R² of the three measures and the unweighted the smallest - the shape of a result driven by shifting abundances among common taxa rather than by wholesale replacement of rare ones.
Putting names to jobs
Sequencing says who is present, not what they are doing. To get from names to nitrogen, genera known to perform a given job were listed in advance and their combined share of the reads was used as a proxy for that job's capacity: genera of ammonia-oxidizing archaea, of ammonia-oxidizing bacteria, nitrite-oxidizers and nitrogen-fixing genera.
Matching was done case-insensitively against the genus field of the reference lineage, as a substring, so one list entry absorbs species-level suffixes and the reference database's own naming variants. The nitrogen-fixer list was deliberately the strict one: canonical fixers only. An earlier version included the broad facultative candidates - Methylobacterium, cyanobacteria, Herbaspirillum and similar - which inflated the fixer estimate several-fold. Those were excluded, and the figure on the home page is the conservative one.
Two limits come with the approach and are stated wherever the result appears: a gene being present is not the same as a gene being switched on, so this measures potential; and read abundance is not cell count, which gets a check of its own below.
One list entry was wrong for a year. The reference database spells one genuine ammonia-oxidizing genus Nitrocosmicus; the list looked for Nitrosocosmicus, with an extra syllable. The two never matched, and a real ammonia oxidizer worth 0.27% of all reads was dropped from every archaeal count - a list that silently returns nothing looks exactly like a list that correctly returns nothing. Correcting it moved the home page's headline ratio from 143 to 148. The independent database described just below agrees the correction was right: our count moved from 8.53% to 8.79% of reads, and that database, which knows nothing of our lists, says 8.97%.
The same table, annotated by strangers
As an independent check the table was re-annotated with FAPROTAX , a published database built by other people, from other literature, that knows nothing of our lists. It put ammonia oxidation at % of reads against our hand-curated %, and nitrogen fixation at % against our %.
Coverage is the limit of that agreement, so it is stated too: the database can say something about % of reads on average, ranging from % to % across samples, and recovers of its functional groups in this soil. Two unrelated methods agreeing about two-thirds of the reads is corroboration; it is not corroboration about the other third. Separately, of functional groups compared between microsites, survived correction for having asked that many questions at once.
Sorting the archaea by the acidity they can tolerate
Ammonia-oxidizing archaea are not one kind of organism with one preference. One lineage, Nitrosotalea and its relatives, is an obligate acidophile: it grows between about pH 4.0 and 5.5 and not outside that. Another, Nitrososphaera and Nitrosocosmicus, prefers near-neutral soil and is the group usually found around plant roots. Since public soil maps place this site near pH 5.5, which lineage dominates is a test of whether acidity explains the headline - and it is a test the existing reads can run.
The split was counted at order rank, not genus, and that choice is doing real work. About 92% of the archaeal reads here resolved no further than the family Nitrososphaeraceae: at genus rank most of them were simply blank, and a genus-level count would discard them. The acidophile boundary happens to fall one rank higher - Nitrosotaleales against Nitrososphaerales - which the reads did resolve. So the question can be answered at all only because the line worth drawing sits above the rank where these reads give out.
Both lineages were counted the same way as each other, which is what a comparison requires. Their combined total was slightly larger than the genus-strict archaeal figure quoted elsewhere, because counting at order rank also picks up the family-only reads that have no genus for the guild list to match. The two numbers are not interchangeable and are not presented as though they were.
Two limits travel with the result wherever it is quoted. Which organisms are present is not how active they are, and a neutral-soil lineage can persist in acid soil while doing less work. And the pH itself is modeled from public maps, not measured in this soil.
The third ammonia oxidizer this study cannot see
The headline ratio counts archaea against bacteria, which was for a long time the whole story. It is not: some Nitrospira bacteria carry out the complete oxidation of ammonia to nitrate on their own, and they have a higher affinity for ammonia than either of the two groups in the ratio - meaning they compete best in exactly the low-ammonia soil this study infers.
They cannot be separated here. Telling a complete oxidizer from an ordinary nitrite-oxidizing Nitrospira requires reading the ammonia-processing gene itself, and this study read a different gene. So the total Nitrospira share is reported as an upper bound on the complete oxidizers rather than a measurement of them, and the headline is described as a comparison between the two canonical ammonia-oxidizing guilds rather than between everything that oxidizes ammonia.
Reads are not cells
The home page's headline is a ratio of read shares, and there is a dull reason a ratio of read shares can mislead: a genome carries anywhere from one to roughly fifteen copies of the 16S gene, and fast-growing bacteria tend to carry more of them. A group with more copies per cell contributes more reads at the same number of cells. If the archaea-to-bacteria ratio were a copy-number artifact, this is the shape it would take.
So each sample's community-weighted mean copy number was computed from , a published per-taxon copy-number database, covering % of reads. It ranged from to copies per genome across the thirteen cores, so there was real variation for the check to find.
It found nothing. Copy number did not track the archaeal share (ρ = ), did not differ between legume and non-legume cores (p = ), and showed no gradient along the community's dominant axis. The ratio is not a copy-number effect. Reads still are not cells, and the page says so wherever the number appears.
Testing whether the clover result is really about nitrogen
Ammonia-oxidizing archaea were scarcer under clover, and clover fixes nitrogen, so the tidy story is that leaked nitrogen leaves the archaea less work to do. The correlation was strong (ρ = , Spearman's rank correlation, as everywhere ρ appears on these pages: it ranks the cores rather than using their values, so one unusual core cannot manufacture a slope). Three cheap explanations can produce a number like that with no mechanism behind it, and each one got a test.
- In a fixed total, everything correlates with everything. The archaeal share spanned % to % of reads here, so it was a large and highly variable slice of that fixed total, and almost any group that does not rise with it is pushed the other way by arithmetic alone. The fixers' correlation was therefore ranked against the same correlation computed for every reasonably abundant genus: all that averaged at least % of reads and appeared in at least of the cores. The nitrogen-fixers ranked , and unrelated genera tracked the archaea at least as strongly as . This is the test the tidy story fails, and it is why the home page says we cannot yet say why.
- The fixed-total artifact specifically. The whole comparison was re-run on centered log-ratio abundances, which are free of the sum-to-a-constant constraint (zeros replaced by half the smallest observed non-zero value before taking logarithms). The association survived (ρ = , p = ), with genera still past the threshold. Compositional closure did not manufacture it, then: the association is real. What fails is its specificity, which is a separate problem.
- A two-group difference wearing a gradient's clothes. The cores split into legume and non-legume, and a correlation across all thirteen looks like a smooth slope even when it is really a single step between two clusters. So it was re-tested inside each group, where no step between the groups can contribute.
| Group | Cores | ρ | p |
|---|
The two-group contrast itself was significant (p = , Mann–Whitney, which like Kruskal–Wallis assumes nothing about the shape of the distribution). Wherever many things were tested at once - every genus against an axis, every functional group between microsites - p-values were adjusted by the false-discovery procedure, implemented in this repository rather than imported, and pinned by its own tests.
What the archaea are actually tracking
That negative result points at something. If many unrelated genera all follow the archaeal share, the likely explanation is that the whole community varies along one dominant direction and the archaeal share is a readout of position along it. That is directly checkable: correlate the archaeal share against each of the leading axes of the ordination - the same principal-coordinates projection the home page plots, which turns the table of pairwise dissimilarities into positions.
| Axis | Share of all variation | ρ | p |
|---|
The archaea rode the community's single dominant gradient, which by itself carried % of the variation between cores. They were not responding to one neighbor; they moved with everything at once. What sets that gradient is the open question this study ends on.
Other people's data
The cross-farm comparison on the home page reprocessed a public dataset from raw reads through this same pipeline - the same denoiser, the same reference database, the same guild definitions - rather than comparing published percentages. Numbers from two different pipelines are not comparable, and quoting them side by side is a mistake this project inherited and had to undo.
The two studies' tables were then aligned by collapsing both to genus, because reference lineages still differ above genus even after identical reprocessing. Phylum-level figures have to be summed from the uncollapsed lineages instead, since collapsing to genus discards the phylum prefix. The comparison's control was internal: organic versus conventional management within the published study, where the wet lab and the sequencing run are shared and no cross-study batch effect can contribute to the difference.
Checking what a public file actually contains
Public metadata records only "amplicon". Accessions were therefore verified before use by reading the primer still attached to the submitted reads: stream the head of several runs off the archive mirror - a bounded read, so whole files never download - and match the starts of reads against known primer sequences. Three complications each caused a false negative in an earlier version and are now handled explicitly. Primers were searched anywhere within the leading bases, because libraries carry barcodes and spacers ahead of the primer. Both members of a pair were searched in both read files, because single-end submissions are sometimes the reverse read. And where a submitter stripped the primers off, 16S V4 is still recognizable from the conserved motif that opens the region, reported as a weaker call.
Sampling one run is not enough, which is the lesson of the accession that forced all this. It was registered as samples of bacterial 16S and ranked as the best available public test of the archaea-to-bacteria ratio. Read by their primers, those runs held different markers:
| Marker on the reads | Runs | Share |
|---|
Only four runs were bacterial, and they were the wrong four: every unfertilized control was deposited under a different marker, so the nitrogen gradient that made the study attractive cannot be tested there at all. Sampling a single run would have called the whole accession fungal, which is no better than trusting the registry. The verifier exits non-zero on a mismatch, so it gates the pipeline rather than warning about it.
Where this soil is, according to public maps
None of this is a measurement of our thirteen cores. It is what published map products say about this location, which is a prior about the neighborhood. Our own laboratory pH is still pending. The distinction matters, because "pH was not measured" and "pH was not measured, and two independent surveys put it near " are different states of knowledge, and the second is the one we are in.
Soil pH drives the archaea-to-bacteria balance nearly as strongly as nitrogen does, which makes an unmeasured pH the largest single alternative explanation for this study's headline. Two independent sources were queried at the farm point and agreed with each other: a global machine-learned soil model put surface pH at (±), and the national soil survey's mapped components gave a depth-weighted - a consensus of , which is classed . Acidic soil favors archaeal ammonia oxidizers, so this is a live competing explanation for the ratio.
The same sources put the surface at % clay, % silt and % sand, with g/kg organic carbon, g/kg total nitrogen and a cation-exchange capacity of . The survey does not map this ground as one uniform soil either: it maps components with different properties, including different pH.
| Component | Share | Drainage | Surface pH | Clay % | Organic matter % |
|---|
Two things are deliberately not on this page. The farm's coordinates, and the survey's names for the map unit and its component soil series - a named series is mapped over a small area, so publishing one would narrow the location to roughly a county. Everything above that bears on the science is here.
Reproducibility
Every random draw in the analysis - the rarefaction subsampling, both PERMANOVA permutation tests, the R² null, PERMDISP - was seeded at . That is deliberate, and it has a cost worth naming: a permutation p-value genuinely does vary from run to run, and fixing the seed hides that variation in exchange for a number a reviewer can reproduce exactly. The unseeded spread was checked, and for the main test it was roughly 0.015–0.025 - comfortably inside the same conclusion.
Statistics that do not depend on a random draw - pseudo-F, every R², every correlation - are deterministic regardless of the seed, which is why those are the ones the test suite pins. The suite recomputes every headline number on this site from the raw tables on every commit; the badge in the top bar is that suite's live status, and a red badge means a number on these pages is in doubt.
| Package | Version |
|---|
What isn't documented
Four gaps. Three of them are records nobody kept; the fourth is a limit of the method itself.
Parts of the wet lab are not recorded. The extraction kit, library preparation, sequencing platform and read length were not passed on with the data, and are being requested. They do not change any result already reported, but a reader cannot fully reproduce the lab work without them.
The 18S taxonomy cannot be regenerated. It arrived as a finished object with no script, log, or reference file, so the exact database release cannot be pinned down. It is used for coarse description only, and no conclusion on this site rests on it alone.
This soil's own chemistry has not been measured. No pH, no nitrate, no ammonium, no organic matter from a laboratory - only the mapped priors above. Every mechanism this study proposes would be settled or refuted faster by one round of soil chemistry than by more sequencing. The acidity test above works around the gap by asking which kind of archaea are present rather than what the pH is, and it argues that acidity is not the explanation - but an argument from which organisms turned up is still not a measurement.
Nothing here measures activity. Every functional statement on both pages is about which organisms are present and in what proportion. A gene that is present may be switched off; an organism that is abundant may be dormant. Nitrogen cycling was never watched happening in this soil, and no number on this site is a rate.