How many nights do I need? It is the first question anyone asks about a self-experiment, and the intuitive answer — more — assumes each night adds a fixed quantity of information. Nights from one person are not independent draws: a short night tends to follow a short night, a bad week drags a run of them down together. Once measurements are correlated in time, the number you have and the amount you know come apart, and what closes the gap is arrangement.
The same 400 measurements, seven ways
Wang and Schork worked this through directly [1]. They fixed the total at 400 measurements split between two interventions, set alpha at 0.05 and the effect to be detected at 0.3 standard deviations, and varied two things: how the 400 were carved into alternating periods, and the strength of AR(1) serial correlation ρ between consecutive measurements. Everything else held.
With ρ = 0, arrangement made no difference whatsoever. Every design from one long block per intervention to forty short ones returned power of 0.851. Independent observations are interchangeable, so only the count matters. Turn the correlation on and the designs come apart.
The ladder between the two is monotone: 0.126 with one period per intervention, 0.176 at four periods, 0.275 at ten, 0.433 at twenty, 0.681 at forty. At a milder ρ = 0.5 it runs from 0.330 to 0.674 across the same range. No extra data was collected at any point.
“the stronger the serial correlation between the observations, the greater the reduction in power to detect an intervention effect. However, this can be mitigated to some degree by breaking up the periods during which the patient is on each intervention”
A long block of correlated measurements is worth a much smaller number of independent ones, since each night is largely predicted by the night before. Switching often makes the comparison repeatedly, across spans too short for the drift to accumulate.
What trials actually do
A methodological review of 74 randomised n-of-1 trials published between 2011 and 2023 gives the shape of practice [2]. Median period length was 14 days, interquartile range 5 to 28. Thirty-two trials (43.2%) used washouts, median 7 days.
Six periods of fourteen days is twelve weeks, arranged close to the shape the simulation ranks near the bottom. For drug trials that is usually forced: titration and washout take time, and some outcomes need a week to move. For a nightly-measured outcome and an intervention that clears overnight, none of it binds, and the long-block shape gets copied anyway.
Cycles are the unit, not nights
Senn's tutorial on paired-cycle designs makes the counting explicit: pairs of periods, one on each treatment, "have been referred to as cycles", and the treatment effect is estimated from differences taken within them [3]. That makes the number of cycles, not the number of nights, the thing that sets how much the estimate can be pinned down.
“For n-of-1 trials, however, there are typically few degrees of freedom per patient”
His worked example runs three cycles per patient, which is why he advises pooling variances across patients rather than trusting any one person's. Counting in cycles makes the shortage visible: twelve weeks at fourteen-day periods is three cycles, whatever the night count says.
What this does not establish
The power figures are properties of a model, not measurements of anybody. Wang and Schork assume AR(1) dependence, decaying geometrically with lag; real sleep and HRV series may carry weekly periodicity, trends and shocks it cannot represent, and the paper never tests the assumption against wearable data.
Those tables also assume no carryover. Short alternating periods are precisely where a real carryover does most damage, so shortening them trades a variance problem for a bias one — the forty-period design wins in a world where the intervention stops when you stop taking it.
And the whole distance between 0.126 and 0.681 is governed by ρ — a property of your own outcome that you cannot know before collecting it. We found no published estimates of serial correlation in consumer wearable sleep metrics that we were willing to cite. Any required number of nights depends on a quantity nobody has measured for what you are tracking.
On the analysis side, Tang and Landes show the standard t-test is anti-conservative under positive serial correlation: it rejects the null more often than its stated error rate allows [4]. The test most people reach for is the one most likely to tell them the supplement worked. We give their direction and not their figures: the preprint and published versions we read describe different simulation grids.
Why this is an n-of-1 problem
In a parallel-group trial the errors come from different bodies and are, near enough, independent, so the participant count is a fair index of what has been learned. That is why population research spends its design effort on recruitment. In a trial run on one body the errors are a time series, and the observation count overstates the information by a factor that depends on how sticky the outcome is. The design effort has to go into the ordering instead.
What follows is not a target number of nights. It is that switching frequency is the lever, that the count worth quoting is cycles, and that a protocol stating a duration without an arrangement has not answered the question it appears to answer.
Sources
- 1.Wang Y, Schork NJ. Power and Design Issues in Crossover-Based N-Of-1 Clinical Trials with Fixed Data Collection Periods. Healthcare (Basel). 2019;7(3):84. doi:10.3390/healthcare7030084. PMID:31269712. Link ↗
- 2.Hawksworth O, Chatters R, Julious S, Cook A, Biggs K, Solaiman K, Quah MCH, Cheong SC. A methodological review of randomised n-of-1 trials. Trials. 2024;25(1):263. doi:10.1186/s13063-024-08100-1. PMID:38622638. Link ↗
- 3.Senn S. The analysis of continuous data from n-of-1 trials using paired cycles: a simple tutorial. Trials. 2024;25(1):128. doi:10.1186/s13063-024-07964-7. PMID:38365817. Link ↗
- 4.Tang J, Landes RD. Some t-tests for N-of-1 trials with serial correlation. PLOS ONE. 2020;15(2):e0228077. doi:10.1371/journal.pone.0228077. Link ↗