How do you take a census of something you can't even see?
Step 1: Take soil samples from beneath five kinds of plants (cover crops) growing within a small area of a working regenerative farm: a former orchard, which was historically disturbed and now being restored, that sheep pass through a few times a year. Dig out each sample by hand, making sure the whole sample reaches down to about forty five centimeters below the surface.
Step 2: Freeze the samples immediately after extraction on the same day, in an industrial walk-in freezer.
Step 3: Aliquot soil from each of the samples, put them into sterile tubes, seal the tubes with Parafilm, and store in dry ice on the way to the sequencing lab.
From that spoonful we pulled all the DNA at once - the plants', the worms', the bacteria's, everything mixed together.
Then we photocopied one particular gene, called 16S, millions of times. Every bacterium carries it and no plant or animal does, so copying it was how we ignored everything that is not a microbe. A sequencing machine read those copies, and matching them against a reference library turned them into a list of names and counts.
Out of thirteen samples came distinguishable kinds of bacteria and archaea, from between and DNA reads per sample.
Every patch had a similar variety of life. Just not the same life.
Walking the orchard, I looked for spots where a single plant clearly owned an area on ground. Each spot also had to sit at least three meters from the next, so that no two samples came from the exact same area. That rule turned up two clovers, a grass, and two broad leaved plants.
Did some plants host more life than others? They did not. The number of kinds of life was statistically flat across all five plants (p = ), and so was diversity by two other measures (Shannon p = , Simpson p = ), and flat again under a measure that accounts for how closely related the organisms are (p = ).
Each dot is one soil core. The bar is that plant's average.
But that is only a count. When I looked at which microbes were present and not just by how many, the samples began to sort themselves by the plant growing above them.
Samples that contained similar communities sat close together. The two directions are the largest and second-largest ways these samples differ: they carry % and % of all the variation.
The plant explained more of that separation than chance would. Two cores under the same plant differed by on average, two under different plants by , and a permutation test on the whole set gave R² = (p = , shuffles). Because there were five plants and only thirteen cores, pure chance already produced an apparent R² of about , so the raw figure overstated the effect; the histogram below shows how far above chance it sits.
Two follow-up checks, and neither was clean. The spread of cores was not the same under every plant (dispersion p ), so the test may be reading a difference in composition, a difference in variability, or both. And the result leans on the smallest groups: drop the one white clover core and it sat at p = ; drop the two chicory cores as well and it was gone (p = ). It did survive a completely different way of measuring how different two communities are, one that counts how distantly related the organisms are rather than just how many they share (p = ).
So the amount of life was about the same under every plant, and the mix of it probably was not. The plant above seemed to pick the neighbors below - suggestive rather than settled, on this many cores.
How a plant does that is not a mystery, though it's not something these cores can show. Roots leak: sugars, acids and amino acids seep out into the millimeter of sorril around them, and different plants leak different things. A microbe that can live on what one plant puts out may be unable to use what the next one does, so the plant feeds some of its neighbors and not others. That is the accepted explanation, established elsewhere - nothing leaking from these roots was collected or measured, and the assay that would test it here has not been done.
The analysis shuffled the plant names across the cores times and recomputed the effect each time. Every gray bar is how many of those shuffles landed at that value. Press the button to pull one out.
What a press actually does. It draws one value from the shuffles the analysis already ran - sampling the real null, not simulating a new one in your browser.
Archaea first: to 1
Turning ammonia into nitrate is one of the essential jobs in any soil - it is most of what "nitrogen cycling" means. Two very different kinds of life can do it: ordinary bacteria, and archaea, an ancient and completely separate branch of life. Which of the two dominates says something about the conditions they live in.
Archaeal ammonia oxidizers tend to be favored where ammonium is scarce, and the bacterial ones respond more strongly when nitrogen is added - though pH and much else shape the balance too. Here the bacteria barely registered: % of the DNA read was archaeal against % bacterial. Most of the archaeal reads ( of those percentage points) could be named only to family, not to a genus. What the ratio says about this soil's nitrogen is a reading, not a measurement; the sections below take it apart.
How do we know the list of archaea is right?
That headline depends on a hand-built list of which organisms oxidize ammonia, so it is fair to ask who built the list. We re-ran the same samples through an independent published database that was assembled by other people for other reasons. It put ammonia oxidation at % where our list said %, and nitrogen fixation at % against our %. Two unrelated methods, the same answer.
But couldn't the soil just be sour?
There is an obvious objection to all of this. Archaea also take over from bacteria in acid soil - nothing to do with nitrogen. Public soil maps put this ground at around pH 5.5, which is firmly on the acid side. So perhaps the ratio is not a story about nitrogen at all. Perhaps it is just geology.
The archaea answered this themselves, because they are not all alike. One group of them can only live in acid - put it in ordinary soil and it dies. The other group prefers ordinary soil. If sour ground were the reason this soil is full of archaea, the acid specialists should be a large share of them.
They were % of the archaea here. The ordinary-soil group outnumbered them roughly to one (% of all DNA read, against %).
These were the wrong archaea for sour ground. Acidity may be part of the story; on its own it is not enough.
There is a second way at this, and it does not use our sequences at all. Nobody has measured the acidity of this field - but two research stations within km have been sending soil to a laboratory for a decade: , on open ground km away, and the , under forest at . Between them they have measurements. If the maps are right about this neighborhood, those measurements should sit around the modeled value.
The dashed line is what public soil maps estimate for this farm. The bars are measured samples from the two nearest research stations: the box is the middle half, the line inside it the median, the whiskers the full range. Measured in a .
They did not. The nearer station's median was and the further one's was - the modeled prior for this farm fell of a unit below the nearest of them, and outside the middle half of both. Less acid, in other words, than the maps had this ground.
These are not this field. They are the closest measured soil, and kilometers off, and one of the two is under woodland - soil pH turns over across a single hillside, and the range on the chart shows it doing exactly that. What this establishes is narrower than the pH of this field: the number the acid explanation leans on is a model's guess, and where anybody has checked nearby, that guess reads low.
The size of this argument. It weakens the simplest rival explanation. It does not prove the nitrogen one, and it is no substitute for putting a meter in this soil, which still has not been done - pH, ammonium, nitrate and moisture were all unmeasured. Reading which organisms are present also tells you who is there, not how hard they are working.
Under clover, fewer archaea and more rhizobia
Clover makes its own nitrogen fertilizer, using bacteria housed in nodules on its roots. If some of that nitrogen leaks into the surrounding soil, the ammonia-oxidizing archaea nearby should have less work to do - and should be less abundant. That is what the samples showed, and it is the clearest pattern in the study. Under the four legume cores (red and white clover) the archaeal ammonia-oxidizer proxies averaged % of the DNA read; under the nine non-legume cores, %, with no overlap between the two sets (Wilcoxon W = , p = , q = ). By plant, the archaea ran from % under ragweed and % under chicory down to % under red clover and % under the single white clover core.
The clovers' partners ran the other way. The bacterial groups that contain the known nitrogen-fixing rhizobia averaged % under the legumes and % under the rest (W = , p = , q = ): highest under white clover at % and red clover at %, against –% under the non-legumes. The bacterial ammonia oxidizers, meanwhile, were rare everywhere and did not differ between the two groups (% against %, q = ). Two nitrogen-cycling groups, opposite associations with the same plant type.
Each dot is one core. Orange cores were taken under a legume - red or white clover, the plants that make their own nitrogen.
What this is, and what it is not. These were taxonomic proxies: groups of organisms named from one gene, whose share of the DNA reads was being compared. The rhizobial genera include lineages that nodulate clover and lineages that do not, and the gene read here cannot tell them apart, so a higher share is a clover-associated rhizobial signal and not evidence of more nitrogen fixation. Likewise the archaea's share says who is present, not how much ammonia is being oxidized. Ammonium, nitrate, pH and moisture were never measured in this soil, so the reciprocal pattern is biologically plausible and consistent with what legumes do, and it is a hypothesis for the next study rather than a demonstrated effect of legumes.
A food web, not a soup
A second set of DNA reads, tuned to pick up organisms with more complex cells, showed this was not just a bag of bacteria. There were fungi, there were protists that hunt and eat bacteria, and there were microscopic animals - nematodes and mites - that eat the protists. Something was eating something else down there, which is more than a list of residents sharing space. Fungi were % of these reads, and % of the fungal reads were Ascomycota, a group full of the decomposers that break down plant litter. Land-plant DNA was another %, the protist group Cercozoa %, and microscopic animals %.
Share of the complex-celled organisms in each patch. Based on of the cores: the white clover core and one chicory core yielded too little of this second gene to sequence, and one ragweed core fell below the depth the others were leveled to.
That there were predators is clear; how many is not. These reads name protists only to broad groups, and the broad groups mix hunters with scavengers and parasites - what an organism eats cannot be read off this marker at the level it resolves. The most abundant protist group here, for instance, is about half bacteria-eaters and half other things in soils where anyone has checked. So the food web was real - the animals and the hunting lineages were genuinely present - but the exact share of it that lives by eating bacteria is past what these counts can settle.
The same question, asked of completely different organisms
Everything about which plant changes what lives beneath it, up to here, came from one gene read out of bacteria and archaea. That invites an obvious worry: what if the pattern is a quirk of that one gene, or of the reference library used to name it? The second marker could answer it, because it shares none of those parts - different gene, different library, and it sees an entirely different branch of life.
It did not find the same thing. Across the cores this second test covered once they were leveled to a common depth, the fungi, protists and animals did not separate by the plant overhead (R² = , p = ), and cores under the same plant were barely more alike than cores under different ones ( against ). The spread of cores differed between plants here too (p = ). Their diversity was flat as well (richness p = , Shannon p = ).
Read this as inconclusive, not as a contradiction. Ten cores across four plants was less than the bacterial test had, one plant was represented by a single core, and a p of was not far from the cutoff. It means the second gene neither confirms nor refutes the first; the plant effect rests on the bacteria and archaea alone.
Ragweed, the plant the farm asked about
Common ragweed is the volunteer here that anyone would pull, and whether it is doing anything to the soil beneath it was the practical question behind the sampling. On the broad measure, it was not: the three ragweed cores did not hold a distinct bacterial and archaeal community from the other ten (R² = , p = ; dispersion p = ). It did have the highest archaeal ammonia-oxidizer share of any plant, which is the pattern above seen from the other end.
One narrower thing did turn up, among the fungi. A lineage assigned to Plectosphaerella averaged % of the eukaryote reads under ragweed, against % under red clover and % under fescue, with no overlap among those three groups - and the single chicory core, left out of the test because one core is not a group, carried %. Ragweed and chicory are both in the daisy family, which raises the possibility of a plant-family association worth testing properly.
Two methods, one answer each. One differential-abundance method supported the ragweed difference after correction for multiple testing (q = ); a second did not (q = ). The gene used cannot say what this Plectosphaerella is doing, and the genus holds both harmless soil fungi and plant pathogens. It is a lead for a replicated study, not evidence that ragweed helps or harms the soil.
So which plant is best for the soil?
This is the question the whole thing is for, and the honest answer is that these thirteen teaspoons cannot settle it. That is worth explaining, because “we found no difference” and “we could not have found one” look identical in a results table and mean opposite things.
Two of the five plants cannot be ranked at all. There was core of soil under white clover and under chicory. A single core has nothing to be consistent with, so there is no honest way to place it above or below anything. Their numbers appear in the charts above because leaving them out would be its own distortion - but they are not evidence about their plant.
Asking which bacteria each plant favors is beyond this sample size, and that is arithmetic rather than bad luck. With three or four patches per plant there are only so many ways the samples can be shuffled, which puts a hard floor of under the smallest p-value the test can return - no matter how large the real difference is. Against kinds of bacteria, the usual correction for testing that many things at once means of them would have to separate the plants perfectly before the first one counted. So a blank result there says nothing about soil. It says thirteen.
Grouping those bacteria by what they do rather than who they are cuts the count to , and that is few enough to work with. It is where the nitrogen result above comes from, and it is the only per-plant comparison in this study with room to breathe.
One measurement here is a recognized test of soil health, and it found nothing - this time meaningfully. The small animals in the food web above - the nematodes and mites - are a long-used indicator of how a soil is doing. Across the plants, no tier of that web differed (p = at best). Unlike the bacteria, this test did have the room to find a difference, so finding none is a real answer: the shape of the food web under these plants was the same.
The averages are tempting - red clover did come out top for small animals and ragweed bottom. But two cores taken under the same plant differed more than the plants differed from each other ( as much signal as noise). Ranking the plants on those averages would be ranking them on chance.
Nor can the published literature close the gap. We looked up what is known about of the bacterial groups that most distinguish these plants. Only of them was studied in a way that transfers to this question. Nearly all “good for soil” findings come from adding a cultured strain to a pot and watching what happens, which shows what an organism can do - not what it means to find more of it in ground nobody touched. The clearest warning is a bacterium reliably more abundant in disease-resistant soils, which failed to protect anything when it was actually added.
Two of those look-ups also point the wrong way for a tidy story. The bacterium most characteristic of ragweed - the one plant here nobody would defend - is a known biological control agent. And the ammonia-eating archaea that were so abundant under the non-clovers give off roughly half the greenhouse gas their bacterial counterparts do, so a soil full of them is not, on that count, a damaged soil.
The plants did seem to grow different communities, and the clovers certainly kept different company among the nitrogen-cyclers - that is the finding above. What does not follow is that any of them is better. Nothing here measures the things people mean by healthy soil: its carbon, how it holds together in the rain, what it grows. Settling this needs more cores under every plant, a measurement of the soil itself, and ultimately the experiment nobody has run - removing a plant and watching what the ground does next.
What we cannot say yet
Thirteen cores from one small area, sampled once over two days, are a baseline and not a verdict. The numbers throughout are shares of DNA reads, not counts of organisms and not rates of anything: DNA cannot tell a live cell from a dead one, or a busy microbe from a sleeping one. No undisturbed reference ground and no earlier time point were sampled, so nothing here can say whether the soil is recovering, or how it has changed. Each core mixed the whole profile from the surface to about forty-five centimeters, so these were bulk soils and not the thin rhizosphere layer around the roots. And thirteen cores over five plants, one of them sampled once and one twice, cannot pull a plant apart from the spot it happened to grow in.
The plant effect leans on the smallest groups. The composition result came with unequal spread between plants (dispersion p ) and faded when the one white clover core and the two chicory cores were set aside (p = without both). It can be argued either way, and it is reported as suggestive rather than settled.
The soil's chemistry was never measured. pH, ammonium, nitrate, moisture and organic matter are all unknown for these cores, and every one of them is known to move the balance of archaeal and bacterial ammonia oxidizers. So the nitrogen pattern above cannot be separated from whatever the soil chemistry was doing under each plant.
The archaea moved with a lot of the community, not just the clovers' partners. A further check asked, of common bacterial groups, which tracked the archaea most closely; the nitrogen-fixers ranked , and unrelated groups tracked them about as well. The archaea also followed the community's single largest gradient (correlation ). Read the reciprocal pattern as one visible strand of a broader shift under clover, not as a private conversation between two groups.
Nobody recorded where in the patch the cores came from. The records carry which plant and which replicate, and nothing about position. The cores were kept at least three meters apart within one patch, which bounds this to a matter of meters rather than fields - but ground is not uniform even at that scale, so a core that happened to sit on an old dung deposit or a wet hollow carries that difference labeled as the plant above it, and there is no way to check.
Nor was it recorded how the five plants were laid out inside the patch - whether they grew mixed together throughout, or each kind held its own corner. That distinction decides how much the caveat above costs. Mixed together, an uneven patch of ground spreads its unevenness across all five plants, which makes a real difference harder to find rather than easier - the result would be understating itself. Sorted into corners, “which plant” and “which part of the patch” become the same question, and nothing in this data can pull them back apart. Somebody who stood in that patch knows the answer; the records do not, and nobody has been asked yet. It is the cheapest open question here by a wide margin.
Grazing animals fertilize in patches, and they do reach this one. Sheep have been put through this area times a year for about years. Stock drop nitrogen unevenly, in dung and urine, and none of those patches was mapped - so a core sitting on an old dung deposit carries a local flush of ammonium that has nothing to do with the plant above it, and nothing in the data can tell the two apart. Of everything on this list, this is the one that bears most directly on the archaeal result, because ammonium is exactly what that result is read against.
The white clover finding rests on one sample. A single white clover patch was sampled, which is a data point and not a group; read nothing here as a finding about that plant.
The ratio counts two competitors, and there is a third. Some bacteria do the whole ammonia-to-nitrate job by themselves rather than handing it over halfway, and they are especially good at it when there is very little ammonia around - which is what we think is the case here. The kind of DNA reading used for this study cannot tell them apart from a close relative that does an unrelated job, so we can only say their combined share is %. That is small, but it is not nothing, and it means the ratio compares the two best-known ammonia oxidizers rather than every organism doing the work.
The fungi were read through a general marker. No fungus-specific barcode was sequenced, so fungi were named only as far as the general eukaryote gene allows, and the root-partner fungi (mycorrhizae) that need their own primers were not properly surveyed at all; their small share here should not be read as scarcity. As with every amplicon survey, the numbers are shares of reads, shaped by which primers were used, and not counts of organisms.
An open question, and the numbers behind it
Published work has sampled soil the same way this was sampled. Put on the same axis, this soil read high - noticeably more archaea, and more of the group that tends to dominate ground left undisturbed. That is worth asking about, but it does not answer anything on its own.
Each dot is one soil sample. The bar is the group average.
What this figure cannot settle. Start with the mismatch in kind: the published work sampled managed pasture, and these cores came from a recovering orchard area. On top of that they are different ground, in a different place, run in a different laboratory - any one of which can move the numbers on its own, with no biology involved. And the sharpest check in that published work cuts against the easy reading: within it, organic and conventional management were not distinguishable on either measure (p = and ). So a gap this size is not management alone. This puts the number in context - it says the archaeal count here is worth explaining - and it does not score this ground against anyone else's.
Only measurements would close the gap, and the ones that would are listed under what would settle it below.
The wider landscape, from the air
Everything so far has leaned on soil that had to be dug up and put in a tube. There is one line of evidence that needs no laboratory at all: satellites have been photographing this field, and the fields around it, every few days for years. From those pictures you can measure how green a parcel is, and - more usefully - whether it stays green.
That second one used to be the point of this section. Ground kept under continuous living cover should not swing between lush and bare across a season; ground that is tilled, or cut and left, should. The section was built to test that prediction as a stand-in for the management claim the study could not make directly.
That reasoning no longer holds, and the section is kept only as landscape description. Two things broke it. The study makes no management claim at all any more, so there is nothing for a stand-in to stand in for. And these satellites measure a square roughly 260 m on a side, while the cores came from an area far smaller than that, in a spot nobody wrote down. Whatever this chart shows, it is the ground around where the soil was dug, not the ground it was dug from. Read the rest of this section as a description of the neighborhood, which is all it was ever entitled to be.
Each point is one parcel across cloud-free satellite passes. Right is greener. Down is steadier - less change between one pass and the next.
On greenness this field was unremarkable: it read greener than only % of the parcels on the ring around it. On steadiness it stood out - steadier than % of them, and one of that never once dropped below the bare-soil line across passes.
Both turn on what the other parcels are, and the answer is mostly: trees. Of the parcels greener than this field, are woodland - and woodland reads greener than grass, which is a fact about trees, not farming. Steadiness runs the same way in reverse: deciduous trees shed their leaves and grass does not, so every open parcel here looks steadier than every wooded one for reasons of leaf phenology alone.
Set beside the few parcels that are actually similar, the margin nearly vanishes. Only of the surrounding parcels are open pasture or hay at all. This field's greenness swung by %; the steadier of those two by %, and the steadiest patch of woodland by %. So the field was nominally the steadiest of three similar parcels - by a margin far too small, and a sample far too tiny, to carry any weight.
The claim predicts cover that persists rather than cover that peaks, and that is the shape these parcels have. Which one holds it best is beyond what nine boxes on a satellite image can settle.
Nothing is known about how those parcels are run. They are fixed boxes on a ring around this field, not property lines, and not one has been established as conventionally farmed - so none is being judged here. Some may not even be independent of this field: the operation leases and grazes ground it does not own, so a ring parcel could be under the same management as the one in the middle. Eight parcels is too few to test anything regardless, which is why no line is drawn through them and no p-value quoted. And the box in the middle is not where the soil came from, which is the objection that retired this comparison.
What this ground has been
One more public record, and the one piece of this section that still does a job. A national survey has classified every field in the country, year by year, since . In all of those years it has put this parcel in pasture or hay - never once in a tillage class, and in its worst year still % grass by area.
This rules out the single most disruptive thing that can happen to a soil community - the plow - for every year anyone has been keeping score, and it comes from a source with no stake in the answer. Two limits on it. The survey draws at thirty meters, coarser than the sampled area itself, so strictly this describes the ground around the cores rather than under them. And what it settles is the plow: how the ground has been grazed is a separate question this record does not reach.
Forty years of it
That survey says what class this ground was put in. A second says what was actually growing on it, every year since - longer than anyone now managing this farm has been. Landsat scenes tied to some seventy-five thousand field plots estimate, each year, how much of a parcel was perennial grass, annual grass, shrub, tree, litter or bare dirt.
Each parcel's average cover across years. The two bars are measured independently and are not shares of one whole.
It answers the question the satellite figure could only raise. This parcel averaged % perennial forbs and grasses over four decades and % trees. The parcels that read greener averaged % trees: they are woods, and have been the whole time. The greenness figure above is very largely a measurement of that.
On its own record the ground reads well. Bare dirt averaged % of the parcel, and in years it has crossed % exactly time - in , at %. Across the whole record bare ground drifted down by points a decade while perennial cover drifted up by .
Set beside the two other open parcels, this field was the barer ground. It averaged % bare, where the non-woodland parcels around it averaged %. It is the same shape as the greenness figure: the comparison flatters this ground when the neighbors are woods, and not when they are pasture. Two parcels cannot establish that it is worse, any more than they could that it is better - but a section that sets the woodland aside when woodland flatters the field has to set it aside when it does not.
There is a step in the record, and it is probably not about this farm. Split at , the parcel's perennial cover fell points - a finding, until you notice that of the neighboring parcels fell across the same year, and that the year was picked by eye from this field's own series, not by any test. The lost cover also went nowhere: bare ground, shrub, tree and litter all held steady across the step, and real land-cover change has to land somewhere. Between a boundary picked from the same series it is being tested on and cover that vanishes without turning up anywhere else, the likelier story is the model moving rather than the field. So the step is written down here instead of charted.
This is a model's estimate of what a parcel's surface looked like from orbit, not a measurement of the soil under it. Its worth is that it is long, independent of everyone involved, and agrees with the two shorter records above: this has been grass, continuously, for as far back as anybody can see.
What would settle it
Three things, none of them exotic. A measurement of this soil's acidity, which drives the archaea-to-bacteria balance nearly as strongly as nitrogen does. The archaea themselves already argue that acidity is not the whole story, but that is an inference, and a pH meter would be an answer.
A set of soils given known, deliberately different amounts of nitrogen - if the ratio moves with the nitrogen, that settles the mechanism. It means running an experiment, though, rather than looking harder at anyone's field.
And a test of a different idea entirely. Some plants release compounds from their roots that shut down ammonia oxidizers directly, and how strongly they do it varies between grasses and clovers - the very plants growing side by side in this patch. If that is what is happening here, the clover pattern is not about nitrogen at all: it is the plant switching the archaea off. Testing it means washing the roots and seeing whether what comes off them stops the reaction in a dish - bench work on plants already growing here.
And, before any of that, the ordinary things a next survey would do: the same number of cores under every plant, laid out so that plant and position are not the same variable, sampled more than once, with the soil's chemistry measured beside each core and, ideally, the nitrogen-cycling genes and rates themselves rather than the names of the organisms that carry them.