Soil microbiome studies

The community moved along one gradient

The tidiest story this survey found is that ammonia-oxidizing archaea were scarce under clover because the clover leaks nitrogen and leaves them less to do. It is a good story, and it failed a test the data can run against it. The association was real; it was just not the archaea answering to the legumes. The whole community moved along a single dominant direction, and the archaeal share was a readout of position on it.

The tidy story

Across the thirteen cores, the archaeal share of the DNA fell where the nitrogen-fixing rhizobia rose. The correlation was strong - Spearman's ρ = , ranking the cores rather than trusting their raw values - and the two-group split between legume and non-legume patches was itself significant (p = ). Read forward, that is a mechanism: fixed nitrogen reaches the soil, the archaea have less ammonia to oxidize, and their numbers fall. It is the kind of result a paper is built around.

So ask the cheap question first. Would a number that strong turn up even if the archaea were not answering to the legumes at all? Here it would, and the rest of this page is why.

The result was real, but it was not specific

A correlation is only evidence for the story told about it if it is specific to the players in that story. So the fixers' correlation with the archaea was ranked against the same correlation computed for every other reasonably abundant genus - all that cleared a small abundance and prevalence floor. If the nitrogen mechanism were doing the work, the fixers should sit at the top of that ranking.

They did not. The nitrogen-fixers ranked , and unrelated genera - groups with no known part in nitrogen fixation - tracked the archaea at least as closely, past a threshold of ρ = . A story that singles out the legumes cannot explain why so many groups that have nothing to do with them follow the archaea just as faithfully.

Two duller explanations were ruled out before drawing that conclusion. The first is compositional closure: in a fixed total of reads, a slice as large and variable as the archaeal share (% to % here) pushes almost everything else the other way by arithmetic alone. Re-run on centered log-ratio abundances, which are free of that constraint, the association survived (ρ = , p = ), with genera still past the line. Closure did not manufacture it, so the association was genuine. What failed was its specificity.

What the archaea were actually tracking

That negative result pointed at something. If dozens of unrelated genera all follow the archaeal share, the simplest reason is that the whole community varies along one dominant direction and the archaeal share is simply a position along it. That is directly checkable. The leading axis of the ordination - the same projection the home page plots - accounted for % of the variation between cores, and the archaeal share tracked it tightly: ρ = , p = .

So the clover correlation and the fescue correlation and a dozen others were not a dozen separate facts about a dozen plants. They were all the same fact: the community sorted along a single axis, the plants sat at different places on it, and the archaea rose and fell with the axis. The legume signal was real because the legumes really did sit at one end - but the mechanism the number seemed to name is not established, because the same number fell out of the community's overall structure without any nitrogen story at all.

And it was not a simple growth-strategy gradient either

The obvious next question is what that dominant axis is. One standing candidate is a life-history gradient: soils are often described as running from slow, resource-poor communities of oligotrophs to fast, resource-rich communities of copiotrophs, and fast growers tend to carry more copies of the 16S gene per genome. If the axis were that gradient, community-weighted copy number should climb along it.

Each core's community-weighted mean copy number was computed from , a published per-taxon database, covering % of reads. It ranged from to copies per genome across the thirteen cores, so there was real variation for the check to find.

It found very little. Copy number did not differ between legume and non-legume cores (p = ), showed no clean gradient along the dominant axis, and barely tracked the archaeal share (ρ = ). The same test ruled out the dull artifact it was first run to catch - the archaea-to-bacteria ratio was not an effect of copy number - but it does not name the axis. Whatever organizes this community, a plain copiotroph-to-oligotroph reading is not it.