Personal experiments are attractive for a good reason: the question is often about you, not about the average person in a study.

Would blocking notifications during a work session help you finish more of the work you intended? Do you edit more accurately in silence or with familiar music? Does one recurring study schedule produce better recall than another?

Those are reasonable questions. They are also easy to answer badly.

Human beings are good at noticing a result after they already know which condition produced it, remembering the dramatic days, changing several things at once, stopping when the preferred answer looks good, and then constructing a convincing explanation after the fact.

A small experiment does not eliminate those problems. It can make them harder to hide.

Start with the decision

Before collecting data, finish this sentence:

If condition A is meaningfully better than condition B, I will...

If there is no plausible action at the end of the sentence, the experiment may be interesting but not useful.

Good personal experiments usually compare low-risk, reversible choices that can actually be repeated: two notification settings, two work routines, two study schedules, two ways of structuring a recurring task.

This is not a method for independently testing prescription changes, unsafe exposures, extreme diets, deliberate sleep deprivation, or other interventions where being wrong could cause harm. Those questions deserve appropriate clinical or professional guidance.

Write the question before you see the answer

"Does this help me?" is too loose.

A better question names the conditions, the outcome, and the comparison. For example:

During comparable 60-minute editing sessions, do notifications off or notifications on produce more correctly completed pages?

Now there is something to measure and something to compare.

Writing the question first matters because the data will contain many possible stories. If you measure duration, errors, pages completed, mood, interruptions, time of day, and perceived difficulty, some variable will probably look interesting by chance. The primary question should not be selected after looking at all of them.

Choose one primary outcome

You can record several things, but choose one outcome that will carry the decision.

A useful outcome is close to the capability you actually care about. If the goal is better editing, "minutes spent at the desk" may be easy to measure but weakly connected to success. Correctly completed pages, errors caught, or a comparable finished unit may be closer.

Secondary measures can help explain what happened. They should not become emergency replacement outcomes when the primary result is disappointing.

Formal N-of-1 guidance makes the same basic point in a more rigorous setting: outcomes, periods, measurement frequency, and analysis should be planned in advance so the result can be interpreted rather than reconstructed afterward. The AHRQ N-of-1 User's Guide is a useful reference for the underlying design principles.

Change one important thing at a time

Suppose condition A is "notifications on, normal desk, afternoon" and condition B is "notifications off, library, morning, coffee first." If B wins, you do not know which difference mattered.

A cleaner comparison keeps the surrounding routine as similar as practical while changing the factor under investigation.

Real life will never be a laboratory. The aim is not perfect control. It is to avoid building the conclusion into a bundle of simultaneous changes.

When an important difference cannot be held constant, record it. A session interrupted by a fire alarm should not quietly become evidence against the study method used that day.

Randomize the order when you reasonably can

Order creates hidden bias.

If every A session happens on Monday and every B session happens on Friday, the experiment also compares Mondays with Fridays. If A always comes first, improvement from practice may make B look better. If the harder tasks happen to land in one condition, the method may be blamed for the task.

In a simple personal experiment, randomization can be modest. Pair comparable sessions and use a coin flip or random-number generator to decide whether each pair runs A then B or B then A.

The point is not statistical theater. It is to prevent yourself from assigning conditions in a way that feels neutral but quietly favors the expected winner.

The CENT explanation for N-of-1 trials describes randomization, counterbalancing, period effects, and treatment order for exactly this reason: time itself can confound repeated comparisons within one person.

Repeat the comparison

One good day is an anecdote with numbers.

The advantage of an N-of-1 style design is repeated comparison. If A and B can be alternated several times, you can ask whether the difference appears again rather than whether one memorable session went well.

There is no universal number of repetitions. The useful amount depends on how noisy the outcome is, how often the activity occurs, how large a difference would matter, and how much burden the experiment creates. AHRQ's statistical design chapter notes that the number of measurements within one person effectively determines the study's sample size, while more periods and more measurements have different tradeoffs.

For a lightweight experiment, a preplanned set of several paired comparisons is more defensible than "keep going until I feel sure."

Ask whether the effect carries over

Some conditions end when the session ends. Others do not.

Turning notifications off for an hour has little reason to affect tomorrow's notification condition. A learning technique that changes what you have already learned is different. Once knowledge has been acquired, you cannot return to the original state and learn the same material again as though nothing happened.

This is a carryover problem in crossover research. The CENT guidance explains why formal N-of-1 trials sometimes use washout periods so the effects of one condition do not contaminate the next.

For an ordinary personal experiment, the practical lesson is simpler: do not use an A/B crossover design when A permanently changes the thing B is supposed to test. Use different but comparable tasks, or choose a different design.

Do not treat adjacent observations as independent worlds

A difficult week can make several consecutive sessions bad. Practice can make later sessions better. Deadlines, travel, workload, weather, sleep, and changing task difficulty can create time trends that have nothing to do with the condition being tested.

Repeated observations from one person also tend to be correlated with nearby observations. N-of-1 methodology calls this autocorrelation. Both the AHRQ statistical guide and CENT warn that treating repeated measurements as though they were unrelated observations can produce misleading inference.

You do not need to fit an autoregressive model to compare two desk routines. You do need to look at the sequence, not merely average every A row against every B row and forget that the rows happened over time.

Decide the stopping rule before the preferred answer appears

Stopping is another place where expectation can enter.

If you planned eight comparable sessions, do not stop at session four because A has won three times. Do not quietly extend to twelve because B is ahead and you expected A to win.

Changing the plan can be legitimate. Life happens. The important part is to record the change and its reason instead of pretending the new rule was always the rule.

A preplanned stopping point is especially useful because small personal experiments are rarely able to settle tiny differences. If the result is still ambiguous after the amount of effort you were willing to spend, "not enough evidence to change" is a valid conclusion.

Look at size and consistency, not just significance

The question is usually practical: is the difference large and consistent enough to justify changing what you do?

If condition A wins five of six comparable pairs but only by a trivial amount, the result may not matter. If B produces a large improvement but only on two unusually easy days, confidence should be lower.

The AHRQ guide makes an important point for individual trials: significance testing can be less pertinent than producing the information the person needs to make a future decision.

For a personal routine, a simple table of paired results, the size of each difference, and the sequence over time can often tell you more than a decorative p-value.

Keep the conclusion narrower than the experiment

A result from six editing sessions does not establish a law of cognition.

It may establish something useful and smaller: under the kinds of tasks and conditions represented in this experiment, one setup produced better outcomes often enough that it is worth preferring for now.

That conclusion can later be revised. Different work, a different season, more experience, or a changed environment may produce a different answer.

The point of a personal experiment is not to prove that you have discovered a permanent truth about yourself. It is to make one decision with less self-deception than intuition alone would allow.

A small template

  1. Decision: What will I do differently if A wins?
  2. Conditions: What exactly are A and B?
  3. Primary outcome: What one measure decides the comparison?
  4. Comparable units: Which sessions, tasks, or days are fair to pair?
  5. Order: Can I randomize or counterbalance A and B?
  6. Carryover: Does one condition change the next period?
  7. Duration: How many comparisons will I complete before looking for a verdict?
  8. Exceptions: What disruptions or protocol changes will I record?
  9. Decision rule: What size and consistency of difference would actually matter?
  10. Conclusion: What does this experiment support, and what does it not support?

A relevant Ulix tool

Track Analysis

Track Analysis can serve as a lightweight event log for a personal experiment. Each session can be recorded with a timestamp, condition, outcome, and brief context, then exported to CSV for comparison. It does not randomize conditions or turn an observational log into an experiment for you. The design still matters.

Sources