Within Alternatives

Did Training Work, or Did the Task Change?

Performance gains are not strong evidence for training until rival explanations such as easier tasks or changing conditions are tested.

15 sources 3 graphics

On this page

  • Why before and after gains can mislead
  • Control groups and unchanged tasks as separators
  • What to compare before crediting an intervention
Preview for Did Training Work, or Did the Task Change?

Introduction

When people perform better after training, it is tempting to conclude that the training caused the improvement. However, a simple before-and-after comparison cannot distinguish between genuine learning and other explanations. Participants may have become familiar with the task, the second assessment may have been easier, external conditions may have changed, or merely taking part in the study may have influenced behaviour. Control groups exist to separate these competing explanations. By comparing people who received an intervention with similar people who did not, researchers can ask a more precise question: did the intervention produce an improvement beyond what would have happened anyway? This comparison is one of the strongest practical safeguards against mistaking an easier task or changing circumstances for evidence that training worked.[Wikipedia]WikipediaRandomized controlled trialRandomized controlled trial

Control Groups illustration 1

Why before-and-after gains can mislead

A rise in performance between two tests is real, but its cause is often uncertain. Several rival explanations can produce apparent improvement even when the intervention has little or no effect.

Common alternatives include:

  • Practice effects. People often perform better simply because they have seen the task before and understand its format.
  • Task differences. The second version of a test may inadvertently be easier, even if intended to measure the same skill.
  • Changing conditions. Better equipment, different instructions, reduced stress, or more favourable timing can all improve performance independently of training.
  • Regression to the mean. Individuals selected because of unusually poor or unusually good initial performance often move closer to their typical level on later testing, even without intervention.[American Journal of Clinical Nutrition]ajcn.nutrition.orgAmerican Journal of Clinical NutritionBest (but oft-forgotten) practices: identifying and accounting for…by DM Thomas · 2020 · Cited b…

Suppose a company introduces a new reasoning course and employees score higher on a follow-up assessment. If the second assessment contains simpler questions or participants already know what to expect, the improvement may exaggerate—or even completely explain—the apparent success of the course. Looking only at the trained group cannot distinguish these possibilities.

This illustrates an important principle in analytical thinking: evidence becomes stronger when it rules out plausible rival explanations, not merely when it is consistent with a preferred story.

Control groups and unchanged tasks as separators

A control group provides a benchmark for what would have happened without the intervention. Ideally, both groups experience the same testing schedule, environmental conditions and measurement methods, with the intervention itself being the principal difference between them. Random assignment strengthens this comparison by making the groups similar in both measured and unmeasured characteristics before the study begins.[Wikipedia]WikipediaRandomized controlled trialRandomized controlled trial

Consider two groups taking the same reasoning assessment twice:

  • The training group completes a six-week programme.
  • The control group continues normal activities.
  • Both groups sit equivalent assessments under the same conditions.

If both groups improve by roughly the same amount, practice, familiarity or easier testing become more convincing explanations than the training itself. If the trained group improves substantially more than the control group, the intervention becomes a more credible explanation because the shared influences have largely been accounted for.

The key comparison is therefore the difference in improvement, not simply whether improvement occurred.

Control Groups illustration 2

Why control groups sometimes improve as well

An important misconception is that control groups should remain unchanged. In reality, they often improve for entirely understandable reasons.

Research across behavioural and health interventions has repeatedly found meaningful gains among control participants. In some studies, simply measuring behaviour, asking repeated questions, or increasing participants’ awareness appears to influence subsequent performance. This phenomenon overlaps with measurement reactivity and the broader family of effects sometimes described under the Hawthorne effect.[PMC]pmc.ncbi.nlm.nih.govSystematic review of the Hawthorne effect: New concepts are…by J McCambridge · 2014 · Cited by 3787 — This study aims to (1) elucid…

For example, systematic reviews of physical activity interventions have found that many control groups increased their reported activity levels despite receiving no substantial exercise programme. Factors such as repeated measurement, self-monitoring and participant awareness contributed to these improvements. Without a comparison group, researchers could easily have attributed these gains to an intervention that participants never actually received.[ResearchGate]researchgate.netExplanations have been…Read more…

Rather than undermining the value of control groups, these findings demonstrate why they are indispensable. They reveal background improvements that would otherwise be mistaken for treatment effects.

What to compare before crediting an intervention

When evaluating whether training genuinely worked, several comparisons are more informative than the headline “scores went up”.

Prioritise questions such as:

  • Did both groups complete equivalent tasks under comparable conditions?
  • Were both groups tested at the same times?
  • Did the control group also improve? If so, by how much?
  • Was the intervention group’s improvement substantially larger than the control group’s?
  • Could any remaining differences be explained by unequal starting points, participant drop-out or changes in measurement?

These questions encourage attention to the quality of the comparison rather than the size of the apparent gain alone.

Where randomised control groups are impractical, well-designed comparison groups, matched cohorts or quasi-experimental designs can still provide stronger evidence than isolated before-and-after measurements because they retain the central logic of comparing change against a credible counterfactual.[Wikipedia]WikipediaOpen source on wikipedia.org.

Control Groups illustration 3

A practical habit for better analytical thinking

Outside formal experiments, few decisions involve perfect control groups. Nevertheless, the underlying reasoning remains valuable.

Whenever someone claims that an intervention worked, ask what happened to an equivalent group that did not receive it. If no comparison exists, consider what background changes could plausibly explain the result instead:

  • Was everyone improving over time?
  • Did the task become easier?
  • Were expectations different on the second attempt?
  • Did measurement itself influence behaviour?

Thinking in these terms shifts attention from impressive-looking gains to diagnostic evidence. The question is no longer, “Did performance improve?” but rather, “Did it improve more than we would reasonably expect without the intervention?” That distinction is central to evaluating rival explanations and avoiding the common mistake of crediting training for improvements that arose because the task, the measurement or the surrounding conditions changed instead.

Amazon book picks

Further Reading

Books and field guides related to Did Training Work, or Did the Task Change?. Use these as the next step if you want deeper reading beyond the article.

BookCover for Bad Science

Bad Science

By Ben Goldacre

Uses accessible examples to explain why control groups and fair comparisons matter when evaluating interventions.

BookCover for The Book of Why

The Book of Why

By Judea Pearl, Dana Mackenzie

Focuses on causation rather than simple association, matching the page's emphasis on evaluating whether training truly caused improvement.

eBay marketplace picks

Marketplace Samples

Example marketplace items related to this page. Use the search link to explore similar finds on eBay.

UsingUSA

Endnotes

1. Source: Wikipedia
Title: Randomized controlled trial
Link:https://en.wikipedia.org/wiki/Randomized_controlled_trial

2. Source: Wikipedia
Title: Impact evaluation
Link:https://en.wikipedia.org/wiki/Impact_evaluation

3. Source: ajcn.nutrition.org
Link:https://ajcn.nutrition.org/article/S0002-9165%2822%2901002-4/fulltext

Source snippet

American Journal of Clinical NutritionBest (but oft-forgotten) practices: identifying and accounting for...by DM Thomas · 2020 · Cited b...

4. Source: pmc.ncbi.nlm.nih.gov
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC3969247/

Source snippet

Systematic review of the Hawthorne effect: New concepts are...by J McCambridge · 2014 · Cited by 3787 — This study aims to (1) elucid...

5. Source: researchgate.net
Link:https://www.researchgate.net/publication/230685038_Control_Group_Improvements_in_Physical_Activity_Intervention_Trials_and_Possible_Explanatory_Factors_A_Systematic_Review

Source snippet

Explanations have been...Read more...

6. Source: Wikipedia
Link:https://en.wikipedia.org/wiki/Quasi-experiment

Additional References

7. Source: semanticscholar.org
Link:https://www.semanticscholar.org/paper/How-to-deal-with-regression-to-the-mean-in-studies-Yudkin-Stratton/aabeede14c2aae82753a089e5cffd04b76854037

Source snippet

How to deal with regression to the mean in intervention...Decision-makers should always consider RTM to be a viable explanation of the o...

8. Source: steve.psy.gla.ac.uk
Link:https://steve.psy.gla.ac.uk/hawth.html

Source snippet

Hawthorne, Pygmalion, Placebo and other effects of...by SW Draper · Cited by 67 — Three different reasons for a control group showing re...

9. Source: api.repository.cam.ac.uk
Link:https://api.repository.cam.ac.uk/server/api/core/bitstreams/01a33193-239f-46fe-8389-b0b0c52502a1/content

Source snippet

alization, triangulation, and control. Educational Assessment, Evaluation and...

10. Source: static1.squarespace.com
Link:https://static1.squarespace.com/static/5b75fc713917ee2070a01bfb/t/5c5494a6fa0d605ae549b60c/1549046951415/Alan%2BBoneva%2BErtac%2BGrit%2B2019.pdf

Source snippet

from a Randomized Educational Intervention on Gritby S Alan · 2019 · Cited by 693 — We show that a targeted educational intervention impl...

11. Source: youtube.com
Title: Analysis of Competing Hypotheses (ACH): Finding Plausible Answers
Link:https://www.youtube.com/watch?v=xt4EnzvGA4w

Source snippet

Analysis of Competing Hypotheses (ACH): A Structured Analytic Technique (SAT) for FinCrime...

12. Source: files.eric.ed.gov
Link:https://files.eric.ed.gov/fulltext/ED138639.pdf

Source snippet

Groups; Educational Action research often...by LH Cross · 1977 — Action research often necessitates the use of intact groups for the com...

13. Source: youtube.com
Title: Abduction (Inference to the Best Explanation)
Link:https://www.youtube.com/watch?v=7TeM7rBhRmw

Source snippet

Analysis of Competing Hypotheses (ACH): Finding Plausible Answers...

14. Source: youtube.com
Link:https://www.youtube.com/watch?v=Y-J0FYOQRMY

Source snippet

Evidence Interpretation Diagnostics...

15. Source: frontiersin.org
Link:https://www.frontiersin.org/journals/public-health/articles/10.3389/fpubh.2022.832523/full

Source snippet

Attention to Progression Principles and Variables of...by G Stassen · 2022 · Cited by 9 — The objective of this systematic review was to...

Topic Tree

Follow this branch

Parent topic

Alternatives What Else Could Explain This?

Related pages 5