01Start with ordinary days
A baseline is a stretch of ordinary days, recorded before anything changes, that everything afterwards is compared with. Single-case research makes it the first step: until it shows a predictable pattern, there is nothing steady to judge a change against (Kratochwill et al., 2010).
How long it needs depends on how much the days vary. US guidance for drug trials starts with one to two weeks of screening, partly to train people in the daily ratings (US FDA, 2012). In a study where people rated their symptoms several times a day for a week, each person’s average for days 1 to 3 agreed well with that for days 5 to 7 (Bosman et al., 2025).
Record the baseline in exactly the same way as the test, at the same time of day. How a note is made can move the numbers by itself, as Write it down on the day shows.
02Days vary on their own
In a study where people kept a daily bowel diary for up to three months, most had both hard and loose stools, with about three swings between them a month on average (Palsson et al., 2012). Yet each person’s overall mix was much the same from month to month.
Single-case research describes a baseline by its level (the average), its trend (drifting up or down) and its variability (how far days stray from the average) (Kratochwill et al., 2010). A baseline that is still drifting, or swinging widely, needs more ordinary days before it can fairly be compared with what follows.
03Why a bad patch misleads
After an extreme reading, the next one usually sits nearer the usual level. This is called regression to the mean, and it can make ordinary ups and downs pass for a real change (Barnett et al., 2005). Noisy readings make it stronger, and so does looking only at readings picked for being extreme.
A test started in a bad patch is the one-person version: the days that follow tend to be better whatever was changed, and the change easily gets the credit. A paper on decisions in health care warns that anything aimed at people far from average will tend to look successful, though as a group tendency, not a certainty for any one case (Morton and Torgerson, 2003). A series of readings makes it easier to spot (Kratochwill et al., 2010). So a baseline should cover ordinary days, not only the worst ones, and a test is easier to read if it starts near the usual level.
04Decide what counts, before you start
Set the bar in advance: one chosen after seeing the numbers makes it too easy to find what you hoped for. The bar says what size of change would matter; it can’t show that a change that size wasn’t chance, which takes repeats. Drug trials are asked to define a “response” in their plan, before the results are in (US FDA, 2012).
In pooled data from ten pain trials, a fall of about 2 points on a scale of 0 to 10, or about 30%, matched people rating themselves much or very much improved (Farrar et al., 2001). A study of abdominal pain rated from 1 to 10 found much the same: 2.2 points, about 30%, in people who said they had improved a little (Spiegel et al., 2009).
Why a percentage? The same 2 points means more from a low start: a fall from 8 to 6 is a quarter, while a fall from 3 to 1 is two-thirds. In the pooled trials, people who started higher needed a bigger fall in points, but the percentage that mattered held steady (Farrar et al., 2001). So when starting levels differ, a percentage is the fairer yardstick.
These figures come from pain that was easing: a starting point, not a rule for every kind of rating, or for a change for the worse. The drug-trial guidance says its 30% bar rests mainly on research in other kinds of long-term pain (US FDA, 2012). It mostly applies the bar to weekly averages, and counts someone as a responder only if it is met in at least half the weeks or days.
05What the studies found
- In 2010, a panel set standards for judging single-case studies, for the US What Works Clearinghouse (Kratochwill et al., 2010). Reading a study starts with a predictable baseline, described by its level, trend and variability, and repeated readings make regression to the mean easy to spot.
- In its 2012 guidance for drug trials in irritable bowel syndrome, the US Food and Drug Administration recommends a screening period of one to two weeks and daily 0 to 10 ratings of the worst abdominal pain (US FDA, 2012). A weekly “responder”, defined in advance, has a weekly average at least 30% below baseline.
- In a 2025 study of 230 people with irritable bowel syndrome, people rated their symptoms from 0 to 10 at up to ten random moments a day for a week (Bosman et al., 2025). Each person’s averages for days 1 to 3 and days 5 to 7 agreed well. In the 162 who also kept an end-of-day diary, its scores ran 0.8 to 1.8 points above their momentary averages.
- In a 2012 study, 185 people with irritable bowel syndrome rated every bowel movement, and their symptoms each night, for up to 90 days (Palsson et al., 2012). Most had both loose and hard stools, yet each person’s mix stayed much the same from month to month. A questionnaire at enrolment overstated their symptoms compared with the diaries.
- A 2005 methods paper describes regression to the mean: after an extreme reading, the next one usually sits nearer the average, so ordinary ups and downs can pass for a real change (Barnett et al., 2005). Noisy measurements, and following up only people picked for an extreme first reading, make it stronger.
- A 2003 paper on decisions in health care warns that, because of regression to the mean, anything aimed at people far from average will tend to look successful (Morton and Torgerson, 2003). It calls this a group tendency, and suggests averaging repeated measurements.
- In a 2001 analysis of ten placebo-controlled trials of a pain medicine, covering 2,724 people with long-term pain, daily 0 to 10 pain ratings were compared with people’s own rating of their change (Farrar et al., 2001). A fall of about 2 points, or about 30%, matched people rating themselves much or very much improved, and the percentage held at every starting level. The authors’ listed affiliations include a drug company.
- In a 2009 study of 277 people with irritable bowel syndrome, abdominal pain was rated from 1 to 10 (Spiegel et al., 2009). Three months later, the 19 who said they had improved a little had dropped 2.2 points on average, about 30%; the authors note that this rests on a small group.
06Sources on this page
Each source is listed once, with links to its record. The reference list says what was read for each, and when it was checked.
- Barnett et al., 2005Barnett AG, van der Pols JC, Dobson AJ. 2005. Regression to the mean: what it is and how to deal with it. International Journal of Epidemiology 34(1):215–220.DOI 10.1093/ije/dyh299 · PubMed 15333621
- Bosman et al., 2025Bosman M, Vork L, Jonkers D, Snijkers J, Topan R, Aziz Q, et al.; ESM study group. 2025. Results From a Psychometric Validation Study: Patients With Irritable Bowel Syndrome Report Higher Symptom Burden Using End-of-Day Vs Real-Time Assessment. American Journal of Gastroenterology 120(5):1098–1107.DOI 10.14309/ajg.0000000000003091 · PubMed 39311432 · PMC12043265
- Farrar et al., 2001Farrar JT, Young JP Jr, LaMoreaux L, Werth JL, Poole MR. 2001. Clinical importance of changes in chronic pain intensity measured on an 11-point numerical pain rating scale. Pain 94(2):149–158.DOI 10.1016/S0304-3959(01)00349-9 · PubMed 11690728
- Kratochwill et al., 2010Kratochwill TR, Hitchcock J, Horner RH, Levin JR, Odom SL, Rindskopf DM, et al. 2010. Single-case designs technical documentation. What Works Clearinghouse, US Institute of Education Sciences. Version 1.0 (pilot), June 2010.ies.ed.gov
- Morton and Torgerson, 2003Morton V, Torgerson DJ. 2003. Effect of regression to the mean on decision making in health care. BMJ 326(7398):1083–1084.DOI 10.1136/bmj.326.7398.1083 · PubMed 12750214 · PMC1125994
- Palsson et al., 2012Palsson OS, Baggish JS, Turner MJ, Whitehead WE. 2012. IBS patients show frequent fluctuations between loose/watery and hard/lumpy stools: implications for treatment. American Journal of Gastroenterology 107(2):286–295.DOI 10.1038/ajg.2011.358 · PubMed 22068664 · PMC3855407
- Spiegel et al., 2009Spiegel B, Bolus R, Harris LA, Lucak S, Naliboff B, Esrailian E, et al. 2009. Measuring irritable bowel syndrome patient-reported outcomes with an abdominal pain numeric rating scale. Alimentary Pharmacology & Therapeutics 30(11–12):1159–1170.DOI 10.1111/j.1365-2036.2009.04144.x · PubMed 19751360 · PMC2793273
- US FDA, 2012US Food and Drug Administration, Center for Drug Evaluation and Research. 2012. Guidance for Industry. Irritable Bowel Syndrome — Clinical Evaluation of Drugs for Treatment. US Food and Drug Administration. Final guidance, May 2012 (docket FDA-2010-D-0146).fda.gov
n1lab is a notebook for one-person experiments. You keep a short record each day, and the app keeps it in order.
Check with your practitioner or doctor before you change what you eat or take.
Last checked 10 October 2026. This page is general information about research methods, not medical advice: see what n1lab is not, in the terms. Spotted a mistake? Email hello@n1lab.app.