Skip to content
,
Statistics for analysts · Part 1

Sampling and bias for analysts

13 min read
Editorial featured image for Sampling and bias for analysts. Title text reads Sampling and bias for analysts.

A sample is a slice of the people you care about, and it is only safe to talk about the slice you actually measured. Say a coworker in product posts a chart in your team chat with a confident caption: “Customers love the new checkout.” The conversion rate did jump, but the data is only last week’s mobile sessions in the US. Finance is already writing the win into the forecast. Then someone in support asks a quiet question: did we measure all customers, or only the ones who got far enough to load the new page?

That is sampling and bias in one meeting. You rarely get a clean count of everyone, because what you get is a slice. The slice can still be useful. It becomes dangerous when you pretend it is the whole world, or when the missing people are exactly the ones who would have changed the story.

This post opens a short series called Statistics for analysts, written for people who build dashboards, write SQL (the language used to pull rows out of a database), and get asked whether a number is “significant” before anyone has said who the number covers. We start with sampling and bias because every later topic, from group comparisons to confidence intervals to A/B tests, inherits whatever sample you chose or whatever sample your product handed you.

Population, sample, and the story you are allowed to tell

Start with four labels, and write them down once for every analysis that will leave your laptop. Each one answers a different question about who is and is not in your numbers.

  • Population: the full set of units you care about for the decision (all active subscribers in the third quarter, all warehouse products that ship weekly, or all support tickets opened last month).
  • Sampling frame: the list or system you can actually reach (users who have an email address, events stored in your data warehouse, or accounts that still exist after a system move).
  • Sample: the units you actually measure (the 2,400 survey responses, the 10% of sessions that had tracking switched on, or last Tuesday’s group of new users).
  • Unit of analysis: what one row means (user, account, session, order, day). Mixing units is a classic way to invent growth, because counting sessions in one place and users in another makes the total look bigger than it is.

If the population and the frame differ, say so, and do the same if the frame and the sample differ. Most “the data shows” arguments fail because someone quietly swapped one layer for another. The decision may still be fine, but the wording should match how far the evidence actually reaches.

Rule of thumb: Your conclusion may only be as wide as the population you can honestly name. Anything wider is a guess with a chart attached.

This habit pairs with the work of defining metrics, which the metrics series covers. You cannot sample “engagement” until you have defined what counts as engaging. Sampling sits on top of definitions, so a bad definition gives you a beautifully drawn sample that answers the wrong question.

How samples form in real analytics work

Textbooks love simple random samples. Product and business data usually give you one of these instead.

Every row in one system (still not “everyone”)

You pull every row from a table, and that feels complete. It is complete only for what that system recorded. Deleted users, events that failed to log, visitors whose browsers blocked tracking, and regions where the feature never shipped are all still missing. A full warehouse table covers the frame, which is not always the same as the whole population.

Convenience and operational samples

Think of last week’s logs, the accounts your customer success managers happen to know, or the tickets tagged “billing.” These samples are easy to get and often useful for putting out fires. They are weak evidence for claims about “all customers,” unless you can argue that the easy path has nothing to do with the outcome you care about. It usually does.

Probability samples (surveys, panels, some research pulls)

Here every person has a known chance of being picked. That lets you weight the answers, estimate your error, and speak more carefully about how far the result applies. Most day-to-day product analytics is not like this. When you do run a survey, write down the response rate and note who never opened the email, because people who do not respond act as their own filter on the sample.

Experiment assignment (designed sample)

An A/B test shows half your visitors one version and half another, which creates a test group and a comparison group on purpose. Random assignment helps you say what caused a change inside the people who took part. It does not fix a narrow eligibility rule, so if only power users join the experiment, your result is about power users. Later posts in this series and the experimentation culture series go deeper on this, and sampling is where the honesty starts.

Filled sampling layers: population, frame, sample, and who got excluded
Filled sampling layers: population, frame, sample, and who got excluded

Bias types analysts actually meet

Bias here means a steady tilt in one direction, not a rude comment in a spreadsheet. Random noise averages out as you collect more data, but bias does not, so more biased data can leave you more confident in the wrong story.

Selection bias

The process that gets units into your sample correlates with the outcome. Say you study feature adoption using only users who opened the product at least five times. Heavy users look like they love everything, while light users who bounced never show up in the table. Your adoption rate then describes the people who stuck around, not everyone who signed up.

Survivor bias

Survivor bias means you only see the units that lasted. The classic business version is averaging the revenue of current customers while ignoring the accounts that already left. The classic career version is studying “what successful startups did” without looking at the failed ones that did the same things. In analytics, a retention table that starts at “users still active in month 6” hides the early exits, and those exits are where the real product risk sits.

Response and self-selection bias

Surveys, reviews, and optional feedback forms attract more than their share of people with strong feelings or free time. A 12% response rate is not a small random sample, because those who answered chose themselves. If you publish “customers rate us 9.1,” add who was invited, who answered, and how the answers differ by plan tier if you can measure that.

Measurement bias

Here the measuring tool itself tilts the answer. Events recorded in the browser miss users who run ad blockers, while counts recorded on the server (the computer that runs the app) may miss frustration that only happens in the browser. A customer satisfaction (CSAT) question on a 1 to 5 scale, where “5” is labeled “ecstatic” and there is no 4.5 option, still shapes the scores. Timezone cutoffs, currency conversions, and events that arrive late all measure something close to the thing you named, but not quite the thing itself.

Confounding dressed up as sampling

Say you compare two markets, but one of them also got a promotion and a different sales approach. Your “sample of Germany versus France” is really a sample of two different treatments. The next post covers comparing groups in depth. For now, notice that your sample description should list the other conditions that travel with the group, not only the country label.

Bias flavorMonday morning tellWhat to write in the chart note
Selection“We only looked at users who…”Eligibility rule and who failed it
SurvivorAverages on current base onlyInclude churned / inactive definition
ResponseOptional survey or review dataInvite count, response rate, channel
MeasurementTracker gaps, late eventsKnown coverage holes by platform
Convenience“I used the export I already had”Why this slice, and what is out of scope

Coverage, missingness, and “null is a person”

Missing data is not only a data quality ticket, because missing values are often a sampling story too. If mobile web has more blank device fields, then any split by device type is also a split by how well the tracking works. If new markets have incomplete customer records, “complete profiles convert better” may only mean that complete profiles come from older markets with better operations, and not that filling in the form causes a sale.

When you drop blank values without a write-up, you quietly redefine the sample. When the drop is large, report the rate both ways: once with the unknowns counted as their own category, and once without them, along with how many you excluded. The data quality series pairs well here: quality problems become statistical problems the moment you summarize.

Worked example: the NPS slide that flattered the product

Imagine your company sells software to other businesses. The product team wants a headline Net Promoter Score (NPS) for the board, which is a number built from one survey question about how likely people are to recommend you. Operations pulls every response from the in-app survey for the last 90 days and gets +42, with the caption “Customers are promoters.” You are the analyst who has to decide whether that sentence is allowed.

You rebuild the path that led to that number, step by step.

  1. Population claimed: all customers.
  2. Frame: users who logged into the app, because the survey only appears inside the product.
  3. Trigger: the survey appears after three successful workflows in a week, which favors happy users.
  4. Response: 18% of those who saw the prompt answered.
  5. Unit: the scores are per user, but board decisions are per account, and power users in large accounts answer more often.

Next you build a transparent version of the table. The numbers below are made up for teaching and are not a real company result.

SliceInvited or eligibleRespondedNPS
All active accounts (CRM)12,400 accountsn/aunknown
Users who logged in (90 days)41,200 usersn/aunknown
Saw in-app prompt9,800 users1,764+42
Enterprise plan respondersn/a610+51
Free or trial respondersn/a220+9
Accounts with a churn risk flagn/a95-8

With these made-up teaching numbers, an honest version for the board reads: “Among users who completed three workflows and chose to answer an in-app survey (about 4% of logged-in users), NPS was +42, and higher on enterprise seats. We do not yet have a representative sample of all accounts, and accounts flagged as likely to leave score poorly.” That sentence is longer, but it is also true, and being able to defend it is worth the extra words.

Result card comparing claimed population all customers versus actual sample of in-app survey responders with NPS plus 42
Result card comparing claimed population all customers versus actual sample of in-app survey responders with NPS plus 42

A short SQL pattern for documenting the funnel into the sample

You do not need fancy statistics software to start. Count how many people passed through each step that created the sample, and keep those counts next to the metric.

WITH base AS (
  SELECT user_id, account_id, plan_tier
  FROM users
  WHERE last_login_at >= CURRENT_DATE - INTERVAL '90' DAY
),
prompted AS (
  SELECT DISTINCT user_id
  FROM survey_impressions
  WHERE survey_id = 'nps_inapp_v3'
    AND impressed_at >= CURRENT_DATE - INTERVAL '90' DAY
),
answered AS (
  SELECT user_id, nps_score, submitted_at
  FROM survey_responses
  WHERE survey_id = 'nps_inapp_v3'
    AND submitted_at >= CURRENT_DATE - INTERVAL '90' DAY
)
SELECT
  (SELECT COUNT(*) FROM base) AS logged_in_users,
  (SELECT COUNT(*) FROM prompted) AS saw_prompt,
  (SELECT COUNT(*) FROM answered) AS responded,
  ROUND(100.0 * (SELECT COUNT(*) FROM answered)
    / NULLIF((SELECT COUNT(*) FROM prompted), 0), 1) AS response_rate_pct
;

Put those four numbers in the slide footer, since people argue less with a footer that shows the arithmetic.

How much sample is “enough”?

Enough for what? A quick gut check on a product can use a small, biased sample if everyone agrees the goal is “what are angry power users saying” and not “what does the market believe.” A pricing decision that moves millions needs a sample designed to match the stakes. Sample size calculators matter later, when you plan experiments, but before you worry about size you should fix who is in the sample. A huge biased sample is still biased, and a modest representative sample can beat a giant pile of whatever was easiest to collect.

A practical move is to say what it would cost to be wrong, then say how far the sample reaches. If those two feel mismatched, stop polishing the chart and redesign the measurement plan.

Writing sample notes people will actually read

Keep a four-line note under any number that leaves your team.

  1. Who is in: eligibility in one sentence.
  2. Who is out: the largest excluded groups.
  3. When: date range and any odd seasons (launches, outages, holidays).
  4. How measured: event source, survey channel, or definition link.

If you cannot fill in those lines, you are not ready to defend the number. Related habits live in data stewardship and in clear metric contracts. A sampling note is that same care applied to a statistical claim.

Common mistakes

  • Calling a warehouse extract “unbiased” because it was large. A big sample can still leave the same people out.
  • Dropping blank values without reporting how many. You redefined the sample in silence.
  • Mixing units (sessions, users, and accounts) mid-story, so growth appears from double counting.
  • Generalizing from log-in users to all customers when many customers use the product offline, share seats, or only connect through an API (a way for software to talk to software).
  • Treating early adopters as the voice of the mass market after a beta that required an invitation.
  • Treating experiment traffic as product-wide truth when eligibility was narrow.
  • Publishing survey scores without the response rate, as if the people who stayed silent were picked at random.

How to practice

  1. Pick one dashboard (a screen of charts that updates as new data comes in) number you show every week. Write its population, frame, sample, and unit in four lines.
  2. Estimate, even roughly, who is left out. Say the biggest excluded group out loud to a teammate.
  3. Add a footer with coverage or response rates for any survey-like metric.
  4. Find one place where dropping blank values changes a rate by more than a few points, and write down both versions.
  5. Optional: rebuild the funnel of “who ends up in the metric” in SQL for one important score.

Series notes: This is Part 1 of Statistics for analysts. The next post covers comparing groups without fooling yourself, including baselines, mix shifts, and fair like-for-like comparisons. For a broader map of skills, use the Learn hub. When AI helps you draft SQL for these funnels, still verify its filters the way the post on checking AI-written SQL shows.

Quick recap

  • Every analysis has a population, a frame, a sample, and a unit. Name them.
  • Bias is a steady tilt in one direction, and more data does not automatically fix it.
  • Product data is often whatever was easy to collect, or every row in one system, and rarely a textbook random sample.
  • Missing values and eligibility rules are sampling decisions, not only engineering bugs.
  • Honest footers (who is in, who is out, when, how measured) prevent overclaiming.
  • Match sample reach to decision stakes before you polish the chart.

Sources

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: