01A few ratings from 0 to 10
Start with what matters to you: in clinical one-person trials, the person being tested often helps choose what is measured (Shamseer et al., 2015). Give each thing its own rating from 0 to 10. Number scales are quick, and in most studies that compared them, people completed them more reliably than line or word scales (Hjermstad et al., 2011).
One number can carry real information: on a daily 0 to 10 pain rating, a fall of about 2 points, or about 30%, matched people saying they felt much better (Farrar et al., 2001). Several things can be rated side by side, too. One widely used set of quick 0 to 10 symptom ratings has been tested for validity and translated into more than 20 languages (Hui and Bruera, 2017).
Separate ratings show more than one overall question can. The US Food and Drug Administration’s guidance for trials of new medicines for irritable bowel syndrome asks for a daily 0 to 10 pain rating (US FDA, 2012). It no longer recommends one overall question as the main measure, because that can’t show one symptom improving while another worsens (US FDA, 2012). The guidance is for trials comparing groups of people, not one-person tests, but the point carries over.
02Name the main measure before you start
Every extra rating is one more number that can move by chance and look like a result. So research names one main measure before a trial begins. The reporting standard for one-person trials asks reports to define their main and other outcomes in advance (Vohra et al., 2015). Yet in a review of 100 reports of such trials, 79% didn’t say which outcome was the main one (Vohra et al., 2015).
Choosing afterwards is a known trap. In group trials, changes to the planned outcomes after a trial has started have been linked to the direction of the results, tilting the evidence towards favourable findings (Shamseer et al., 2015). Guidance for group trials also asks for the size of change that counts to be set in advance (Irvine et al., 2016; US FDA, 2012).
So write down the main rating, and the smallest change that would matter to you, before the first day. Keep rating the others, to see whether anything else got worse: the FDA’s guidance asks the same of trials aimed at one symptom (US FDA, 2012).
03The cost of tracking too much
Tracking more can feel more thorough, but it has costs. In a study of 52 recorded talks by keen self-trackers, tracking too many things was a common mistake: some gave up altogether, worn out, and others had more data than they could analyse (Choe et al., 2014). The authors also suggest that tracking many things at first can help people work out which to keep (Choe et al., 2014). The aim is a short list that gets filled in every day.
04Keep the wording the same
A rating is only as steady as its question. The studies in one review labelled the ends of their pain scales in 24 different ways, and whether that changes the numbers people give hasn’t been tested (Hjermstad et al., 2011). But on one widely used symptom checklist, some people misread items and scored them backwards, and a revised version with clearer wording was easier to understand (Hui and Bruera, 2017). People also read the same scale differently: one person’s 6 out of 10 can be another’s ordinary day (Hui and Bruera, 2017). So compare your ratings only with your own.
The time a rating covers is part of its wording too: the FDA’s guidance asks about the worst of the past 24 hours, while the revised checklist asks about now (US FDA, 2012; Hui and Bruera, 2017).
So write each question down once, with what 0 and 10 mean and the time it covers, and use the same words every day. Rate at the same time each day, too, as Write it down on the day explains. Keep the list the same: a rating added halfway through has no earlier days to compare with. If a question has to change, note the day, and don’t compare numbers across it.
05What the studies found
- In 2012, the US Food and Drug Administration issued guidance for group trials of new medicines for irritable bowel syndrome (US FDA, 2012). It asks for a daily 0 to 10 rating of the worst abdominal pain in the past 24 hours, with response defined in advance. Its suggested line for pain, a weekly average at least 30% below the starting level, was borrowed from research on other long-term pain and hasn’t been validated for this use.
- In 2016, an expert committee set out how trials for functional gut disorders should be designed (Irvine et al., 2016). It said the main analysis should count who reaches a clinically meaningful change in a self-rated outcome, defined in advance, and that trials should be registered before they start.
- In a 2011 review of 54 studies comparing pain-intensity scales in adults, number scales were completed more reliably than line or word scales in 15 of the 19 studies that compared this (Hjermstad et al., 2011). The ends of the scales were labelled in 24 different ways, and the authors said whether this affects the scores still needs testing.
- In a 2001 analysis of 10 trials of a pain medicine, with 2,724 people with long-term pain, daily 0 to 10 pain ratings were compared with how much better people said they felt (Farrar et al., 2001). On average, a fall of about 2 points, or about 30%, matched feeling much or very much improved; the 30% figure held at any starting level.
- In a 2009 study of 277 people with irritable bowel syndrome, a single rating of the day’s abdominal pain, from 1 to 10, lined up with wider measures of severity and quality of life (Spiegel et al., 2009). Three months later, the 19 who said they had improved a little had dropped 2.2 points on average, about 30%, the smallest change that mattered to them.
- In a 2017 review, two researchers described the Edmonton Symptom Assessment System, a short set of 0 to 10 ratings of several symptoms used in palliative care and other clinics (Hui and Bruera, 2017). It has been validated and translated into more than 20 languages. Studies they reviewed found some items were misread, and that a revised wording was easier to understand.
- In 2015, an international group published CENT, a standard for reporting one-person trials, which asks for main and other outcomes to be defined in advance (Vohra et al., 2015). It cited a review of 100 such reports in which 79% didn’t say which outcome was the main one.
- In a companion paper, the same group explained each item, with examples (Shamseer et al., 2015). It notes that in group trials, changes to planned outcomes after the start have been linked to the direction of the results, and that in one-person trials the person often helps choose what is measured.
- In a 2014 study, researchers analysed 52 recorded talks in which self-trackers at Quantified Self meetups presented their own data (Choe et al., 2014). Speakers tracked about three things on average; tracking too many was a common pitfall, leading some to give up and others to more data than they could analyse. The speakers were volunteers, probably more skilled than most self-trackers.
06Sources on this page
Each source is listed once, with links to its record. The reference list says what was read for each, and when it was checked.
- Choe et al., 2014Choe EK, Lee NB, Lee B, Pratt W, Kientz JA. 2014. Understanding quantified-selfers' practices in collecting and exploring personal data. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, 1143–1152.DOI 10.1145/2556288.2557372
- Farrar et al., 2001Farrar JT, Young JP Jr, LaMoreaux L, Werth JL, Poole MR. 2001. Clinical importance of changes in chronic pain intensity measured on an 11-point numerical pain rating scale. Pain 94(2):149–158.DOI 10.1016/S0304-3959(01)00349-9 · PubMed 11690728
- Hjermstad et al., 2011Hjermstad MJ, Fayers PM, Haugen DF, Caraceni A, Hanks GW, Loge JH, et al.; European Palliative Care Research Collaborative (EPCRC). 2011. Studies comparing Numerical Rating Scales, Verbal Rating Scales, and Visual Analogue Scales for assessment of pain intensity in adults: a systematic literature review. Journal of Pain and Symptom Management 41(6):1073–1093.DOI 10.1016/j.jpainsymman.2010.08.016 · PubMed 21621130
- Hui and Bruera, 2017Hui D, Bruera E. 2017. The Edmonton Symptom Assessment System 25 Years Later: Past, Present, and Future Developments. Journal of Pain and Symptom Management 53(3):630–643.DOI 10.1016/j.jpainsymman.2016.10.370 · PubMed 28042071 · PMC5337174
- Irvine et al., 2016Irvine EJ, Tack J, Crowell MD, Gwee KA, Ke M, Schmulson MJ, et al. 2016. Design of Treatment Trials for Functional Gastrointestinal Disorders. Gastroenterology 150(6):1469–1480.e1.DOI 10.1053/j.gastro.2016.02.010 · PubMed 27147123
- Shamseer et al., 2015Shamseer L, Sampson M, Bukutu C, Schmid CH, Nikles J, Tate R, et al.; CENT Group. 2015. CONSORT extension for reporting N-of-1 trials (CENT) 2015: Explanation and elaboration. BMJ 350:h1793.DOI 10.1136/bmj.h1793 · PubMed 25976162
- Spiegel et al., 2009Spiegel B, Bolus R, Harris LA, Lucak S, Naliboff B, Esrailian E, et al. 2009. Measuring irritable bowel syndrome patient-reported outcomes with an abdominal pain numeric rating scale. Alimentary Pharmacology & Therapeutics 30(11–12):1159–1170.DOI 10.1111/j.1365-2036.2009.04144.x · PubMed 19751360 · PMC2793273
- US FDA, 2012US Food and Drug Administration, Center for Drug Evaluation and Research. 2012. Guidance for Industry. Irritable Bowel Syndrome — Clinical Evaluation of Drugs for Treatment. US Food and Drug Administration. Final guidance, May 2012 (docket FDA-2010-D-0146).fda.gov
- Vohra et al., 2015Vohra S, Shamseer L, Sampson M, Bukutu C, Schmid CH, Tate R, et al.; CENT Group. 2015. CONSORT extension for reporting N-of-1 trials (CENT) 2015 Statement. BMJ 350:h1738.DOI 10.1136/bmj.h1738 · PubMed 25976398
n1lab is a notebook for one-person experiments. You keep a short record each day, and the app keeps it in order.
Check with your practitioner or doctor before you change what you eat or take.
Last checked 10 October 2026. This page is general information about research methods, not medical advice: see what n1lab is not, in the terms. Spotted a mistake? Email hello@n1lab.app.