On alternating practice days, is my next-morning soreness at least 0.5 points lower after 10 minutes at 12 °C than after no immersion?
Soreness dropped 0.9 points, and I trust that number about half as far as I can throw the trough.
Dates
Sep 01 – Sep 26, 2025
Duration
20 practice days
Pre-registered primary
Next-morning soreness (0–10)
FN-03 · 01
The question
When preseason two-a-days start, the whole team climbs into the troughs, on the theory that cold equals recovery. The Library article on cold-water immersion says something more specific: it probably reduces next-day soreness, and it may simultaneously blunt long-term training adaptation when it is parked after every strength session. I tested the first clause on the only athlete I have full access to.
The design is the best one available to a person who owns exactly one body: alternation. Ten practice days followed by immersion, ten matched practice days without, interleaved by a calendar set in advance, so I could not quietly assign the brutal practices to whichever condition needed the help.
FN-03 · 02
Pre-registration
- Protocol
- Immersion days: within 15 minutes of practice ending, seated to the navel in the team trough at 12 °C (thermometer, adjusted with bagged ice), 10 minutes on a timer. Control days: normal cool-down, lukewarm shower, no cold exposure of any kind.
- Duration
- 20 practice days across four weeks: 10 immersion, 10 control, strict alternation fixed by calendar before day one.
- Primary outcome
- Next-morning soreness, 0 to 10 in half-point steps, rated before getting out of bed, averaged across quads, calves, and shoulders.
- Secondary outcomes
- Same-evening soreness, 0 to 10
- Subjective freshness at the next practice warm-up, 1 to 5
- Counts as null
- A mean difference under 0.5 points between conditions is a null. And for the record, written here before any data exists: I expect the ice to work. Remember that sentence when you read the result.
- Note to self
- The expectation clause above is deliberate. If the result agrees with what I predicted, the agreement is evidence about my prediction as much as about the ice.
Saved as a dated, unedited note before day one. The goalposts above are the goalposts the data was judged against; nothing in this record was touched after the first measurement.
FN-03 · 03
The protocol
Logistics were the hard science here. Twelve degrees is not what a hose produces in September; it is what a thermometer and repeated bags of ice produce. Every session logged between 12.0 and 12.5 °C, ten minutes on a phone timer, in to the navel because volleyball soreness lives in legs and low back.
The alternation held for all twenty days, because it was fixed by calendar before day one and I never let myself renegotiate it. That rule is the entire integrity of this design, such as it is: no condition ever got chosen after I knew how hard practice had been, and the one evening I badly wanted to swap (a brutal Wednesday scrimmage that landed on a control day) is exactly the evening the rule existed for.
Water temp
12.0–12.5°C
Duration
10:00min
Immersion days
10
Control days
10
FN-03 · 04
The log
The actual daily data, charted with the intervention window shaded, plus the raw log as a table under each chart. Same arrays feed both, so they cannot disagree.
Conditions alternate by calendar, so adjacent points share a training week. Axis shows 2.5 to 6 of a 0 to 10 instrument. Ratings move in half-point steps; the rater knew the condition every single morning.
›View the raw log · 20 rows
| practice day | Date | Phase | After immersion (0–10 scale) | Control (no cold) (0–10 scale) |
|---|---|---|---|---|
| 1 | Sep 01 | intervention | 4 | · |
| 2 | Sep 02 | intervention | · | 4.5 |
| 3 | Sep 03 | intervention | 3.5 | · |
| 4 | Sep 04 | intervention | · | 5 |
| 5 | Sep 05 | intervention | 4.5 | · |
| 6 | Sep 08 | intervention | · | 4 |
| 7 | Sep 09 | intervention | 3 | · |
| 8 | Sep 10 | intervention | · | 5.5 |
| 9 | Sep 11 | intervention | 4 | · |
| 10 | Sep 12 | intervention | · | 4.5 |
| 11 | Sep 15 | intervention | 3.5 | · |
| 12 | Sep 16 | intervention | · | 4 |
| 13 | Sep 17 | intervention | 3 | · |
| 14 | Sep 18 | intervention | · | 5 |
| 15 | Sep 19 | intervention | 4.5 | · |
| 16 | Sep 22 | intervention | · | 4.5 |
| 17 | Sep 23 | intervention | 3.5 | · |
| 18 | Sep 24 | intervention | · | 5.5 |
| 19 | Sep 25 | intervention | 4 | · |
| 20 | Sep 26 | intervention | · | 4 |
FN-03 · 05
What happened
Immersion mean
3.75/ 10
Control mean
4.65/ 10
Difference
−0.9pts
Pre-reg bar
0.5pts
Immersion mornings averaged 3.75; control mornings averaged 4.65. The difference is 0.9 points in the ice's favor, comfortably past the pre-registered 0.5 bar. On paper, that is an effect, and if I were selling troughs this write-up would end here.
Here is why it is filed as inconclusive instead. I rate soreness in half-point steps. I knew the condition every single morning, because you do not forget sitting in 12-degree water. I expected the ice to work and wrote that expectation into the sealed record before day one. A 0.9-point gap on an unblindable subjective scale, produced by a rater who predicted the gap, is exactly the size of gap that expectation can manufacture without any physiology involved. The result agrees with the literature, but agreement with the literature is also precisely what my expectation would generate on its own. I cannot distinguish those two worlds from inside this dataset, and honesty about that is the entry.
There is also the half of the question this design never touched. The Library's second clause, the one about cold blunting adaptation when it follows strength work, is invisible at this timescale. If the ice bought me 0.9 points of comfort at a hidden cost in training adaptation, this experiment could not see the invoice, and four weeks is not long enough for the bill to arrive.
I predicted the result I got, which is exactly why I do not get to celebrate it.
FN-03 · 06 · Mandatory in every entry
Why this is weak evidence
This section is the module’s actual thesis. The result above is the least important thing on this page; the list below is why.
01
Unblindable, structurally
You cannot placebo a 12-degree trough. Even proper trials struggle here; sham immersion with inert 'recovery additives' changes what subjects believe, not what the water does. My version had no sham at all: the subject, the rater, and the believer were the same person in the same cold water.
02
The expectation is on the record
The sealed pre-registration says I expected the ice to work. For a subjective primary, that converts a confirming result into weak evidence almost by definition: the experiment was run by someone whose rating hand knew the hypothesis.
03
Load was alternated, not equalized
Calendar alternation controls day-of-week structure, but Wednesday scrimmages were heavier than Tuesday technique days, and with n = 10 per condition the actual training load per arm does not average out cleanly. I did not quantify load, so I cannot adjust for it.
04
A coarse instrument wearing precise clothes
Means like 3.75 imply a precision the half-point scale does not possess. The entire between-condition difference is less than two scale steps of a self-report.
05
Comfort was the only endpoint
No morning jump test, no creatine-kinase, no performance measure of any kind. This experiment measured how I felt about my legs, which is a real thing, and also the single easiest thing in sports science to bias.
FN-03 · 07
What I would do differently
- 01Add one objective anchor: a standardized morning jump or grip test, rated before I could talk myself into anything.
- 02Randomize condition order within weeks instead of strict alternation, so heavy-practice patterns cannot line up with one arm.
- 03Log practice load (session RPE times minutes, or jumps counted) and analyze the difference with load on the table.
- 04Run the adaptation question properly: an off-season lifting block with ice after every session versus none, tracking strength gain rather than comfort. That is the experiment whose answer would actually change behavior, and it is a much harder one.
FN-03 · 08
The verdict
The verdict, for me
InconclusiveInconclusive, leaning 'probably some real comfort effect, inflated by an unknown amount of expectation.' Behavior changed anyway, in the direction the mechanism argues: the trough is now reserved for tournament weekends and back-to-back days, when tomorrow's comfort is worth more than marginal adaptation. On ordinary strength days I skip it, because the Library's arithmetic about adaptation costs is not something my little dataset gets to overrule.
The non-verdict, for you
A 0.9-point subjective improvement from a subject who wanted one is not a reason for you to buy a trough or an ice barrel.
It is a live demonstration of why soreness scales are where recovery marketing goes to look better than it is. Every cold-plunge testimonial you have ever read was generated by this exact mechanism, minus the confession.
FN-03 · Cross-references
Where this entry leads
The group evidence, the mechanism, and the instruments behind this experiment live in the other modules.
Education, not diagnosis. This is a student-authored science platform. Nothing here replaces a physician, a physical therapist, or an athletic trainer. Sudden severe pain, numbness, an inability to bear weight, or visible deformity means stop reading and get seen.