Skip to main content

Agriculture and statistics: How they shaped each other

By Gustavo A. Roa, Postdoctoral Researcher, Department of Agronomy, Kansas State University, Manhattan, KS
September 29, 2026
Photo by Gabriel Paiao. This image was modified to remove some text labels using Adobe Firefly AI tools.
Photo by Gabriel Paiao. This image was modified to remove some text labels using Adobe Firefly AI tools.
Long before the era of big data, agricultural researchers were grappling with uncertainty, variation, and the need for reliable evidence. Their solutions laid the foundation for many of the statistical methods now used across science.

From field experiments and crop yields to experimental design, mixed models, spatial statistics, genomics, remote sensing, and artificial intelligence, agriculture and statistics have continually shaped how we understand and improve agricultural systems.

Agriculture is often described as an applied science, while statistics is viewed as a collection of tools used after an experiment has been completed. Historically, however, the relationship has been much deeper.

Agriculture presented statisticians with a fundamental problem: nature is variable. Soils differ across a field, weather changes from year to year, crops respond differently to their environments, animals differ genetically and environmentally, and agricultural experiments often take months or years to complete on limited land.

Under these conditions, a difference between two treatment means does not automatically represent a treatment effect. It may reflect soil variation, weather, location, measurement error, or other sources of background variation.

This challenge made agriculture an important setting for the development of modern statistics, while statistical advances transformed how agricultural experiments were designed, analyzed, and interpreted. 

The relationship is two-way:

Agriculture has not simply applied statistics, and statistics has not simply served agriculture. Agricultural problems have helped create and refine statistical methods, while statistical advances have transformed how agricultural questions are studied.

Before Fisher: Agriculture reveals the problem of variation

The connection between agriculture and quantitative information predates modern statistics.

Governments needed information about land, crops, livestock, harvests, prices, and food supplies. Sir John Sinclair's Statistical Account of Scotland, published from 1791 to 1799, provides an early example of large-scale systematic collection of social and agricultural information. Gower (1988) places this period within the broader development of scientific agriculture and quantitative investigation.

During the 19th century, agricultural experimentation also became more systematic. Researchers compared crop varieties, rotations, manures, fertilizers, and cultivation practices.

The Broadbalk wheat experiment at Rothamsted Research, established in 1843. Source: Rothamsted Research. 

Rothamsted Experimental Station became one of the most important centers of this work. John Bennet Lawes and Joseph Henry Gilbert established long-term experiments on crop nutrition and fertilizer management. The Broadbalk wheat experiment, established in 1843, continues today. 

These experiments produced extraordinary scientific records, but they also exposed a fundamental difficulty: field plots are not identical.

Soil fertility, drainage, moisture, slope, and previous management can vary across a field. Neighboring plots may be more similar to one another than plots farther apart. Early uniformity experiments helped demonstrate the structured nature of field variation.

By the beginning of the 20th century, agricultural researchers needed ways to separate treatment effects from field variation.

Fisher and the transformation of agricultural experimentation

The decisive chapter began in 1919, when Ronald A. Fisher joined Rothamsted.

 Fisher and Wishart's 1930 work on the arrangement and statistical analysis of field experiments at Rothamsted. 

Fisher worked with the station's long-term agricultural datasets for approximately 14 years, using replicated field experiments with heterogeneous soils, multiple treatments, and long-term observations.

Fisher's contribution was larger than any individual statistical test. He helped establish the principle that statistical inference begins with experimental design. These statistical ideas, in turn, gave agricultural researchers better ways to design experiments, account for field variation, and draw reliable conclusions from their results

Randomization, replication, and blocking

Agricultural researchers had often arranged treatments systematically. Fisher promoted randomization, assigning treatments by chance so that treatment placement was not systematically associated with unknown field characteristics.

He also emphasized replication, which allows experimental variation to be estimated, and blocking, which allows known environmental gradients to be controlled through comparisons within relatively homogeneous groups.

These ideas transformed field experiments and later became fundamental throughout experimental science. Randomization, replication, and blocking are not procedures added after data collection. They determine what comparisons the experiment can support (Parolini, 2015; Street, 1990).

Randomization, replication, and blocking in an agricultural field experiment. Image created by Gustavo Roa with assistance from ChatGPT (OpenAI).

 

ANOVA and the partitioning of variation

Fisher also helped establish analysis of variance (ANOVA), which provided a framework for partitioning variation among the sources represented in an experimental design (including treatments, blocks, interactions, and residual error) rather than treating variation simply as an obstacle.

This shifted the question from simply comparing means to partitioning and quantifying sources of variation.

A framework developed around agricultural experiments subsequently became fundamental across biology, medicine, psychology, engineering, and many other sciences.

When agriculture became more complicated

Agricultural experiments quickly exposed limitations in simple experimental designs because crops do not experience fertilizer, variety, irrigation, and planting date independently; these factors can interact.

Factorial designs allowed researchers to study several factors simultaneously and estimate their interactions.

But physical constraints created another challenge. Some treatments cannot be applied to small plots.

Irrigation, tillage, fumigation, or machinery treatments may require relatively large areas, while varieties or fertilizer treatments can be applied to smaller subplots. This led to split-plot designs, in which one factor is assigned to large main plots and another to smaller subplots.

Example of a split-plot design. Factor A is assigned to whole plots, while factor B is assigned to subplots within each whole plot. Image edited by Gustavo Roa with assistance from ChatGPT (OpenAI).

The two factors therefore have different experimental units and sources of error. Fisher's agricultural work included experiments that were subsequently analyzed as split plots, and the design became standard in experimental methodology.

As experiments became larger, researchers also needed ways to handle many treatment combinations without creating enormous heterogeneous blocks. Confounding, incomplete-block designs, and lattice designs provided ways to preserve important comparisons while making experiments physically manageable.

Breeders may need to compare hundreds or thousands of varieties, making conventional complete blocks inefficient. Patterson and Williams developed alpha designs, a class of resolvable incomplete-block designs that allowed large variety trials to be conducted with smaller blocks and better control of environmental variation (Patterson & Williams, 1976).

The pattern is consistent:

An agricultural constraint became a statistical problem, and the statistical design evolved to solve it.

Agricultural problems have repeatedly motivated the development and use of statistical methods. Image created by Gustavo Roa with assistance from ChatGPT (OpenAI).

 

The bull problem and the rise of mixed models

Animal breeding presented statisticians with a particularly interesting problem. A dairy cow can be evaluated through her milk production. A bull does not produce milk, yet a bull may influence the milk production of hundreds or thousands of daughters.

Its genetic value therefore cannot be measured directly through the trait itself. It must be predicted from information about relatives, daughters, pedigree, herd, year, and environment. This problem became an important application of mixed models in quantitative genetics.

Mixed models allow systematic effects, such as herd and year, to be represented separately from random effects such as animals and genetic values.

Charles Roy Henderson's work was particularly influential in developing mixed-model methods for animal breeding. His mixed-model equations provided a framework for estimating fixed effects and predicting random effects, including genetic breeding values (Henderson, 1975). This led to Best Linear Unbiased Prediction, or BLUP, which became fundamental to genetic evaluation in livestock and later plant breeding.

The importance of BLUP is that it changed what statistical analysis could accomplish.

Statistics was no longer limited to estimating observed averages or treatment effects. It could predict an unobserved genetic value by combining information from multiple sources while accounting for environmental variation and genetic relationships.

REML and the problem of variance components

Mixed models introduced another important challenge: estimating variance components. How much variation is genetic, environmental, or residual?

Patterson and Thompson introduced restricted maximum likelihood (REML) in the early 1970s in connection with statistical problems arising from agricultural incomplete-block experiments (Patterson & Thompson, 1971). REML provided a way to estimate variance components while accounting for fixed effects.

It subsequently became a standard method for mixed-model analysis and quantitative genetics.

From field variation to spatial statistics

Agricultural fields also helped reveal that observations are not necessarily independent.

A soil property, drainage pattern, disease outbreak, or moisture gradient can affect neighboring plots. As a result, observations close together may be more similar than observations far apart.

Example of within-field spatial variability in corn grain yield. 

Early uniformity trials demonstrated this spatial structure, and Papadakis (1937) later developed a nearest-neighbor approach for field experiments. Subsequent work produced increasingly sophisticated spatial methods for agricultural trials (Edmondson, 2005; Verdooren, 2020).

The same problem appears in precision agriculture, where yield monitors generate thousands of observations across a field, satellite imagery provides millions of pixels, and soil sensors produce dense measurements. But more observations do not make spatial dependence disappear.

Modern agricultural statistics therefore uses spatial models, geostatistics, and spatiotemporal methods to account for the structure of these data, allowing researchers to better understand and manage variation across agricultural landscapes. 

Agriculture, breeding, and the problem of generalization

Agriculture also confronted statisticians with a problem that remains central today: results do not necessarily generalize across environments.

A variety that performs well in one location may perform poorly somewhere else. A drought-tolerant genotype may have an advantage under water stress but little advantage under favorable conditions.

This is the problem of genotype × environment interaction (G × E).

Multi-environment trials therefore require statistical approaches that can represent variation among environments and differences in genotype responses.

The question is not simply which variety has the highest average yield. It is also which varieties are stable and where recommendations can be generalized.

This problem connects classical agricultural experimentation to modern quantitative genetics, genomic prediction, and ultimately the broader challenge of generalizing predictions beyond the environments used to train them.

From precision agriculture to artificial intelligence

As agricultural experiments became larger and statistical models became more complex, computation became essential. Modern agricultural research may combine field experiments, soil measurements, weather records, yield monitors, satellite imagery, drone observations, genomic data, and economic information. Agriculture has become a data-rich science, creating new opportunities for statistical analysis while introducing new challenges in prediction, validation, and generalization.

Modern precision agriculture provides perhaps the clearest continuation of this historical story.

GPS systems, yield monitors, drones, sensors, and satellites allow researchers to measure agricultural processes at unprecedented spatial and temporal resolution.

Machine-learning models can predict yield, identify crop stress, classify land cover, detect disease, and support management decisions.

But statistical thinking remains essential. A model trained on one set of farms may fail on another, perform well on training data but poorly in future seasons, or produce overly optimistic validation results because of spatial dependence. Likewise, a strong prediction does not necessarily establish causation, and a highly precise estimate can still be wrong.

The fundamental questions are therefore familiar:

How were the data collected? What is the experimental or sampling unit? What sources of variation matter? How uncertain is the result? Does it generalize?

The technologies have changed. The statistical questions have not.

From yield to sustainability

Agricultural research is also asking broader questions than it did a century ago.

Yield remains important, but researchers increasingly evaluate productivity alongside profitability, nutrient-use efficiency, water use, soil health, greenhouse-gas emissions, and resilience to climate variability.

These outcomes are often related, and treatment effects may depend on environmental conditions, creating additional challenges for modeling and inference.

Long-term experiments help distinguish short-term fluctuations from persistent changes. Rothamsted's long-term experiments are an important example of how agricultural datasets can support questions that cannot be answered in a single growing season.

The statistical challenge has therefore expanded from estimating treatment effects to understanding agricultural systems under uncertainty.

The statistician and the agricultural scientist

The relationship between statisticians and agricultural scientists has always been collaborative. As agricultural questions have become more complex, this collaboration has become increasingly important.

Statistics is most valuable when it is involved from the beginning of a study, helping researchers translate biological questions into appropriate experimental designs, sampling strategies, and analytical approaches.

At the same time, statisticians benefit from understanding the biological processes, practical constraints, and management decisions that shape agricultural research. Agricultural scientists bring the questions and context, while statistical methods provide ways to design studies, quantify variation, evaluate uncertainty, and draw conclusions from complex data.

Effective collaboration can sometimes be challenging. Differences in terminology, priorities, and ways of approaching a scientific problem can make communication difficult. Kozak (2016) highlighted these challenges in communication between agricultural scientists and statisticians. But these challenges can be overcome through communication, shared understanding, and collaboration throughout the research process.

This partnership has evolved alongside agricultural science. Early field experiments helped motivate new approaches to experimental design. Later, agricultural problems contributed to the development of mixed models, spatial methods, quantitative genetics, and statistical computing. Today, the same collaboration is expanding into remote sensing, genomics, machine learning, and large-scale agricultural data.

Looking ahead, this partnership will become even more important as agricultural research continues to generate larger and more complex datasets. The future of agricultural research will depend not only on better data and more powerful models, but also on bringing biological and statistical perspectives together from the beginning.

The wheat sheaf

There is an appropriate symbol for this history. Founded in 1834, the Statistical Society of London was the world's first statistical society. Its early seal featured a sheaf of wheat, reflecting the close connection between agriculture and the emerging field of statistics. In 1887, the society received a Royal Charter and became the Royal Statistical Society, the name it carries today.

Early and later wheat sheaf seal of the Statistical Society. Source: Royal Statistical Society, via Wikimedia Commons. 

The symbol is an appropriate metaphor for statistics: agriculture produces observations, while statistics helps organize them, compare them, quantify uncertainty, and turn them into knowledge.

A relationship that continues

The history of agriculture and statistics is therefore much more than the familiar story of Fisher and ANOVA.

Agricultural research repeatedly created problems that required better ways to design experiments, separate sources of variation, predict outcomes, and quantify uncertainty. At the same time, these statistical advances gave agricultural researchers new ways to design experiments, understand variation, predict outcomes, and make decisions.

Today, an agricultural experiment may involve hundreds of plots, dozens of environments, thousands of genetic markers, satellite imagery, automated sensors, and machine-learning models.

Although the scale has changed dramatically, the fundamental scientific problem has not:

How can we separate meaningful biological patterns from the variation surrounding them, and make decisions that remain defensible beyond the data we have observed?

That question helped shape modern statistics. It remains central to agricultural research today.

Edmondson, R. N. (2005). Past developments and future opportunities in the design and analysis of crop experiments. The Journal of Agricultural Science, 143(1), 27–33. https://doi.org/10.1017/S0021859604004472

Gower, J. C. (1988). Statistics and agriculture. Journal of the Royal Statistical Society: Series A (Statistics in Society), 151(1), 179–200. https://doi.org/10.2307/2982191

Henderson, C. R. (1975). Best linear unbiased estimation and prediction under a selection model. Biometrics, 31(2), 423–447. https://doi.org/10.2307/2529430

Kozak, M. (2016). Communication between agricultural scientists and statisticians: A broken bridge? Scientia Agricola, 73(6), 505–511. https://doi.org/10.1590/0103-9016-2015-0399

Papadakis, J. S. (1937). Méthode statistique pour des expériences sur champ. Bulletin de l'Institut d'Amélioration des Plantes à Salonique, 23, 1–30.

Parolini, G. (2015). In pursuit of a science of agriculture: The role of statistics in field experiments. History and Philosophy of the Life Sciences, 37(3), 261–281. https://doi.org/10.1007/s40656-015-0075-9

Patterson, H. D., & Thompson, R. (1971). Recovery of inter-block information when block sizes are unequal. Biometrika, 58(3), 545–554. https://doi.org/10.1093/biomet/58.3.545

Patterson, H. D., & Williams, E. R. (1976). A new class of resolvable incomplete block designs. Biometrika, 63(1), 83–92. https://doi.org/10.1093/biomet/63.1.83

Street, D. J. (1990). Fisher's contributions to agricultural statistics. Biometrics, 46(4), 937–945. https://doi.org/10.2307/2532439

Verdooren, L. R. (2020). History of the statistical design of agricultural experiments. Journal of Agricultural, Biological and Environmental Statistics, 25(4), 457–486. https://doi.org/10.1007/s13253-020-00394-3

Connecting with us

The article is a contribution from the ASA, CSSA, and SSSA Graduate Student Committee. If you would like to give us feedback on our work or want to volunteer to join the committee to help plan any of our activities, please email Bala Subramanyam Sivarathri at [email protected] (send message), the 2026 chair of the committee! If you would like to stay up to date with our committee, learn more about our work, contribute to one of our CSA News articles or suggest activities you would like us to promote, watch your emails, connect with us on X (Twitter), or view the committee page.


Text © . The authors. CC BY-NC-ND 4.0. Except where otherwise noted, images are subject to copyright. Any reuse without express permission from the copyright owner is prohibited.