HomeLearningLibraryEngineering
Back to Library
Thursday, July 9, 2026
Surface Scan

Endurance Training Response Is Not A Fixed Trait

A single training block is evidence about one context, not proof that you are a permanent responder or non-responder.

How to use this

Read the surface scan first. Switch to deep dive only if you want more mechanics and nuance.

Done state

Mark as read when you can explain the core model back in one or two sentences.

Next move

After finishing, either go deeper, ask questions below, or return home for the next recommendation.

What Is This?

A common training story says people have stable response types:

same training stimulus -> responder / non-responder trait -> predictable adaptation

A 2026 Journal of Applied Physiology study makes that model less safe.

Forty-two untrained middle-aged adults completed two similar 8-week endurance-training periods, separated by 8 weeks of detraining. The study measured several outcomes: hemoglobin mass, skeletal-muscle citrate synthase activity, capillaries per fibre, maximal oxygen uptake, and 15-minute maximal mean power output.

The important split was this:

baseline measures were highly reproducible
training adaptations were poorly reproducible within the same people

Baseline VO2max and 15-minute power were very stable between periods. But the size of each person's adaptation was not. In other words, the same person could look like a strong responder in one block and a weak responder in another.

The useful model is:

adaptation = person x stimulus x context x timing x measurement

Not:

adaptation = fixed responder type

Why Does It Matter?

Training data tempts you to over-read one block.

If FTP rises quickly after a threshold block, the easy story is: “I respond well to threshold.” If VO2max does not move after a block, the easy story is: “I am a low responder.” If Garmin shows a weak training-status signal, the easy story is: “that stimulus does not work for me.”

The paper argues against that reflex. A single block is evidence, but not identity.

For Jamie, the practical consequence is simple:

judge training by repeated patterns across blocks, not by one response snapshot

This is especially relevant for cycling because power meters, heart-rate data, Garmin labels, sleep scores, and perceived effort create a high-resolution illusion. More data can make a weak inference feel stronger than it is.

The Study Design

The headline paper was designed to test reproducibility of adaptation inside the same individuals.

Structure:

  • Participants: 42 middle-aged, untrained men and women.
  • Training: two similar 8-week endurance-training periods.
  • Washout: 8 weeks of detraining between blocks.
  • Compliance: participants completed about 24 sessions in each period, with minimal within-person variation.
  • Outcomes: hemoglobin mass, citrate synthase activity, capillaries per fibre, VO2max, and 15-minute maximal mean power output.
  • Question: if a person adapts strongly in block one, do they adapt strongly again in block two?

The answer was mostly no.

Baseline measures were reproducible. The reported intraclass correlation for baseline VO2max was 0.98, and for baseline 15-minute power it was 0.96. That matters because it means the people did not become random between blocks. Their starting levels were stable.

The adaptation responses were different. Reproducibility of within-individual adaptations was poor:

Outcome Reported adaptation reproducibility
Hemoglobin mass ICC 0.00 [0.00, 0.30]
Citrate synthase activity ICC 0.15 [0.00, 0.33]
Capillaries per fibre ICC 0.00 [0.00, 0.36]
VO2max ICC 0.19 [0.00, 0.37]
15-minute maximal mean power ICC 0.20 [0.00, 0.37]

The strongest conclusion is not “training does not work.” Training worked at group level. The conclusion is narrower and more useful:

the size and pathway of one person's adaptation did not behave like a stable trait under repeated exposure

The Responder Trap

The old mental model is attractive because it compresses messy biology into clean categories:

responder
non-responder
high VO2 responder
low strength responder
threshold person
volume person

Those labels may sometimes describe a pattern. The trap is treating them as a diagnosis after one block.

There are three reasons this goes wrong.

1. Measurement error can look like biology

A 2018 Journal of Applied Physiology paper on repeated testing argued that observed exercise responses mix true interindividual differences with random error and method choices. In its proof-of-concept training study, only 11 of 20 participants were consistently classified across analytical approaches.

That is the warning label on any responder classification:

classification can be partly a property of the testing method

2. Baseline testing itself can shift

A separate 2018 study on VO2 peak reproducibility found high reproducibility across repeated tests, but also a learning effect: the first pre-training test was lower than the third. That changed how people were classified after training.

For training interpretation, this matters because a single “before” test may not be a clean before. It can include unfamiliarity, pacing, equipment, motivation, and protocol learning.

3. Biology is dynamic, not a static knob

The 2026 repeated-block study found stable baselines but unstable adaptation magnitudes. That points to a dynamic system: sleep, fueling, accumulated fatigue, stress, readiness, muscle state, hematological state, season, and training history may alter what the same nominal block produces.

So the better question is not:

am I a responder to this stimulus?

It is:

under what conditions does this stimulus produce the adaptation I want?

How To Interpret A Training Block

A block should be treated as an experiment, not a verdict.

Better frame:

block -> response signal -> confidence update -> next block design

Not:

block -> permanent identity label

For a cycling block, track at least four classes of evidence.

1. Performance output

Examples:

  • power at threshold;
  • 5-minute power;
  • 15- or 20-minute maximal mean power;
  • power after accumulated kilojoules;
  • repeatability across intervals.

2. Cost of output

Examples:

  • heart rate for a given power;
  • RPE for a given power;
  • decoupling over long rides;
  • recovery time after sessions;
  • next-day legs.

3. Context

Examples:

  • sleep;
  • carbohydrate availability;
  • illness;
  • heat;
  • work stress;
  • life load;
  • accumulated fatigue;
  • consistency of execution.

4. Measurement quality

Examples:

  • same bike / power meter / protocol;
  • sufficient familiarization;
  • similar fatigue state;
  • enough repeated tests to avoid over-reading noise;
  • comparable environmental conditions.

The point is not to drown training in data. The point is to avoid making identity claims from unstable evidence.

Practical Takeaways For Jamie

Do not call yourself a non-responder after one block

One weak response means:

this block, under these conditions, produced this measured result

It does not mean:

my physiology cannot adapt to this stimulus

Repeat the training question before changing the whole model

If a block under-delivers, first check:

  • Was the stimulus actually completed?
  • Was recovery sufficient?
  • Was fueling sufficient?
  • Was the test comparable?
  • Was the block long enough?
  • Was the target adaptation plausible for the dose?

Then decide whether to repeat, modify, or abandon it.

Track adaptation pathways, not just headline FTP

The 2026 paper measured several pathways: blood, muscle oxidative enzyme activity, capillarization, VO2max, and power. They did not all behave as one neat package.

For Jamie's cycling, the equivalent is to separate:

fresh power
fatigued power
aerobic efficiency
repeatability
recovery cost
subjective resilience

One number cannot carry the whole interpretation.

Use Garmin as signal, not judgement

Garmin-style training labels can help detect trends. They should not become identity labels. If a block looks flat, use it as a prompt to inspect context and measurement, not as a verdict on trainability.

Why Smart People Get This Wrong

They confuse precision with certainty

A power meter can give exact watts. That does not mean the causal story is exact.

They confuse group effects with individual predictions

A training intervention can work on average while individual response sizes remain hard to predict.

They compress biology into personality

“I'm a volume person” or “I'm not a VO2 person” may be useful shorthand after years of repeated evidence. It is bad science after one block.

They ignore the repeatability problem

If the same person does not reproduce the same adaptation response in two similar blocks, then one block cannot support a strong trait claim.

What This Does Not Prove

This does not prove:

  • individual differences do not exist;
  • genetics do not matter;
  • all endurance programs are equivalent;
  • training should be random;
  • responder labels are always useless;
  • the exact findings generalize to trained cyclists, elite athletes, or every training mode.

The HERITAGE Family Study, for example, found substantial heterogeneity in VO2max response to standardized training and evidence of familial aggregation. That older result still matters. The 2026 paper does not erase individual variation. It sharpens the question:

which parts of response are stable traits, which are measurement artifacts, and which are dynamic context-dependent states?

Validation Surface

  • Primary validation: Odden et al., 2026, Journal of Applied Physiology, PMID 42241643, DOI 10.1152/japplphysiol.00154.2026.
  • Independent methodological support: Hecksteden et al., 2018, Journal of Applied Physiology, on repeated testing and instability of individual-response classification.
  • Measurement caution: Leifer et al., 2018, Clinical Physiology and Functional Imaging, on VO2 peak reproducibility, familiarization, and classification sensitivity.
  • Counterweight: Bouchard et al., 1999, HERITAGE Family Study, showing heterogeneous VO2max response and familial aggregation.
  • What remains uncertain: whether the same poor adaptation reproducibility appears in well-trained cyclists, longer blocks, different modalities, or individualized training prescriptions.

How To Use This

Use this decision rule after a training block:

if response is strong:
    preserve the stimulus, but do not assume permanent responder status
if response is weak:
    inspect execution, recovery, fueling, and measurement before abandoning the stimulus
if response is mixed:
    identify which pathway changed and which did not
if two or three repeated blocks converge:
    update the training model with higher confidence

A better training log question:

what conditions made adaptation more or less likely this time?

Key Terms

  • Responder / non-responder: a classification based on whether a measured outcome changes beyond a threshold after training.
  • Trainability: the capacity to improve in response to a training stimulus.
  • Intraclass correlation coefficient (ICC): a statistic used here to estimate how reproducible individual differences were across repeated blocks.
  • VO2max: maximal oxygen uptake; a common marker of cardiorespiratory fitness.
  • Citrate synthase activity: a marker related to skeletal-muscle oxidative capacity.
  • Capillaries per fibre: a marker of local oxygen-delivery infrastructure in muscle.
  • Detraining period: a washout period intended to return participants closer to baseline before retesting.

Recall Questions

  1. What was stable in the 2026 repeated-block study: baseline physiology or adaptation magnitude?
  2. Why is a single training block weak evidence for a permanent responder/non-responder label?
  3. How can measurement error or familiarization change responder classification?
  4. What is the difference between “training works on average” and “this individual's response is predictable”?
  5. What should Jamie inspect before abandoning a training stimulus that produced a weak response?

Best Resources To Learn More

  • Start with the 2026 Journal of Applied Physiology paper for the repeated-block result.
  • Read Hecksteden et al. for the methodological problem of classifying individual exercise response.
  • Read the HERITAGE Family Study paper as a counterweight: individual variation and familial aggregation are real, but not the whole story.

Sources

  • Odden IU, Hamarsland H, Odden TU, Hammarström D, Rustaden AM, Koll L, Hanestadhaugen M, Larsen S, Ellefsen S, Hansen J, Nygaard H, Lundby C. “Limited reproducibility of individual physiological adaptations to repeated endurance exercise training.” Journal of Applied Physiology. 2026;141(1):215-231. PMID: 42241643. DOI: 10.1152/japplphysiol.00154.2026. https://doi.org/10.1152/japplphysiol.00154.2026
  • Hecksteden A, Kraushaar J, Scharhag-Rosenberger F, Theisen D, Senn S, Meyer T. “Repeated testing for the assessment of individual response to exercise training.” Journal of Applied Physiology. 2018;124(6):1567-1579. PMID: 29357481. DOI: 10.1152/japplphysiol.00896.2017. https://doi.org/10.1152/japplphysiol.00896.2017
  • Leifer ES, et al. “Reproducibility of peak oxygen consumption and the impact of test variability on classification of individual training responses in young recreationally active adults.” Clinical Physiology and Functional Imaging. 2018;38(6):938-944. PMID: 28960784. DOI: 10.1111/cpf.12459. https://doi.org/10.1111/cpf.12459
  • Bouchard C, An P, Rice T, Skinner JS, Wilmore JH, Gagnon J, Pérusse L, Leon AS, Rao DC. “Familial aggregation of VO2max response to exercise training: results from the HERITAGE Family Study.” Journal of Applied Physiology. 1999;87(3):1003-1008. PMID: 10484570. DOI: 10.1152/jappl.1999.87.3.1003. https://doi.org/10.1152/jappl.1999.87.3.1003

Want more depth?

If the surface scan feels useful, request a deep dive and turn this into a heavier explanatory piece.

What next?

Back to Home

Get the next recommended module or article.

Open Learning

Switch from standalone reading into guided progression.

Questions & Answers

Back to Library