Skip to main content
note
This page was translated from the German original, partly by machine. Some passages may read awkwardly or contain inaccuracies. When in doubt, please read the original.

19. Evaluating Empirical Probability

 

Summary

This lecture from Allan Di Donato's Critical Thinking course examines how to evaluate empirical probability, the first of several videos on certainty and probability in inductive reasoning.

What Empirical Probability Means

Empirical probability is the examination of empirical evidence to determine the probable conclusion in an inductive situation. Empirical means derived from or relying upon experience or observation. Probability is the likelihood that something is the case: a claim supported by evidence strong enough to establish a presumption, but not necessarily proof. To evaluate an inductive argument built on empirical data, Di Donato proposes four considerations: extension, the nature of the situation, the manner of procedure, and our background knowledge.

Extension and the Nature of the Situation

Extension asks how many cases are examined, that is, how broad the sample is. The more cases, the stronger the induction. A presidential approval rating based on 20,000 people is far more reliable than one based on 20.

The nature of the situation asks how representative the evidence is, that is, how closely the testing environment mirrors the real world. If the 20,000 polled were all Republicans, or if the urban/rural split did not match the actual population, the results would be skewed.

These two factors work together. To estimate daily red-meat consumption by the average American, polling every citizen would be the most accurate (a possible complete induction) but is impractical; a sample is less accurate in principle yet far more workable. Who you sample also matters: people leaving a butcher shop would lead to an overestimate, residents of San Francisco to an underestimate, while a random selection from across the country offers the best chance of a true cross-section.

Sample Size, Bias, and Margin of Error

Larger samples yield better inductions, but acquiring data gets harder. Of 200, 1,500, or 20,000 people, 1,500 is the practical "sweet spot." Whatever the size, the goal is to avoid a biased sample, one that fails to represent the target in some relevant way. The difficulty is that you do not always know in advance which features are relevant (educational level might matter for studying cat burglars; handedness or hair color probably would not). Bias is best avoided through a random sample, in which every member of the population has an equal chance of being selected, and a varied sample, which contains variety among its members.

Because selection is random, no sample is guaranteed to be perfectly representative. The range of random variation between samples is the margin of error: larger samples produce a smaller margin. The related idea of confidence level is the probability that the margin of error is accurate. As sample size grows, the margin of error shrinks, but with diminishing returns. The decline is steep up to about 1,200–1,500, after which the curve flattens and the margin is already below the 3% mark. Pushing much below 3% is generally not worth the added cost and effort.

Manner of Procedure and Background Knowledge

Manner of procedure asks how the evidence was gathered and examined: How many similarities and differences were studied? Were all explanations and all the data considered? Was there a control group? In a polling scenario, were all respondents asked the same question, was the wording neutral, and were all responses counted equally?

New data must also be weighed against our background knowledge. Does it contradict what we believe, or does it enable a better explanation than before? The historical example is the observed redshift of distant stars: it conflicted with the steady-state model of the universe and ultimately led scientists to adopt the expanding-universe model.

Applied to red meat and colon cancer, the four study designs each carry trade-offs. A random sample of the general population rarely captures enough cancer cases; a sample of meat-eaters lacks a non-meat-eating control; a sample of cancer patients can overestimate the link because we ignore the many meat-eaters who never get cancer; and a long-term observational study of meat-eaters captures future cases but omits non-eaters unless designed in.

Correlation Is Not Causation

Even when an observational study reveals an apparent pattern, correlation is not causation. A correlation has many possible explanations: a shared deficiency could drive both the craving and the risk, or confounding lifestyle factors (less fiber, less exercise, other unhealthy foods) could be the real cause. Establishing causation requires more than observation; it requires the experimental method (and Mill's methods), topics for later videos.

Di Donato closes with fallacies tied to empirical reasoning:

  • Biased generalization: a sample large enough yet still unrepresentative of the target.
  • Hasty generalization: an overconfident conclusion drawn from too small a sample ("these 10 movies are poorly written, so most movies are" is weaker than concluding "some are").
  • Anecdotal evidence: a hasty generalization built on one or two persuasive stories; effectively a sample of one or two.
  • Self-selection fallacy: overestimating a conclusion from a self-selected sample (e.g., a website or call-in poll) that over-represents the strongly motivated.
  • Slanted question: wording chosen to elicit a particular response (asking about "protection from dangerous second-hand smoke" versus "infringing on a citizen's right to smoke").
  • Weak analogy: an argument by analogy that breaks down on too many points.
  • Vague generalities: statements too vague to be meaningful, where the group or attribute is so loosely defined that membership cannot be judged; remedied by a sampling frame or precising definition.

The next video turns from inductive probability to certainty, comparing inductive certainty with other kinds.