Skip to content
Research Atlas

Collecting Environmental Data · Lesson 1 of 5

Why Sampling Matters

~13 min

The Concept

The Bluewater Basin Research Station has fifteen groundwater monitoring wells spread across the watershed. Your supervisor tells you this with obvious pride, it took three years of fieldwork and two rounds of grant funding to install them all. Then she asks the question that stops you cold: 'If we'd put them all in the same corner of the basin, would that tell us anything useful about the whole watershed?'

The answer is no. But almost every bad environmental dataset in history, data that led to incorrect conclusions and misspent policy dollars, came from exactly this mistake: measuring in the easy places, not the right places.

Sampling is choosing which measurements to take from a larger system you can't measure entirely. You'll never install a sensor at every square metre of Bluewater Basin, so where you do install them determines what you can truthfully claim to know.

A good sample is representative (it captures the range of conditions that actually exist in the system). A bad sample is biased (it over-represents convenient or interesting locations and quietly ignores everything else). The dangerous thing about a biased sample is that the data it produces looks just as 'clean' as data from a good sample, the bias is invisible in the numbers themselves.

The Analogy

Imagine grading how much a class likes broccoli, but you only ask the three kids sitting closest to the salad bar. You will get a clean, confident number. It will also be completely wrong about the rest of the class. Where you choose to ask matters just as much as how many people you ask.

Why Real Researchers Care

Sampling design is one of the most consequential decisions in any environmental study. The celebrated 'reproducibility crisis' in science, where many published results couldn't be replicated, is partly attributable to convenience sampling: testing only in accessible locations, recruiting only the people who showed up, or measuring only on days when weather was cooperative. The statistical methods in later missions can only be as trustworthy as the sampling design behind them.

Quick Check

Q1. What makes a sample biased?

Q2. Why can't you detect sampling bias just by looking at the data spreadsheet?

Your Goal

Look at the map of Bluewater Basin's fifteen wells. Identify one part of the watershed that appears under-sampled, somewhere you'd add a sixteenth well if you could. Explain why that location matters.

Hint: Think about where the agricultural runoff would most likely reach groundwater, is there a well near that transition zone?

Teach It Back

Explain the difference between a 'representative' and a 'convenient' sample, using an example from outside Bluewater Basin.