A productivity system is working if it improves the thing you built it to improve, at a maintenance cost worth paying.

Task managers count tasks. Calendars fill with blocks. Note systems accumulate notes. Those are signs that a system is being operated, not yet signs that it is useful.

Name the failure it should fix

Start with a concrete problem: commitments get forgotten, important work loses to urgent small tasks, projects stall because the next action stays vague, or useful information cannot be found when needed.

Then choose an outcome outside the system. If commitments are the problem, count missed commitments. If retrieval is the problem, notice failed searches and retrieval time. If important work is being crowded out, measure whether planned high-priority work reaches meaningful milestones.

Tasks captured, notes created, tags assigned, inboxes processed, and review streaks can be operating measures. They are not automatically outcomes.

Measure the cost too

Maintenance includes capture, filing, tagging, rescheduling, reviewing, cleaning stale items, and moving information between tools.

A useful rough equation is: value = useful failures prevented + useful work enabled - maintenance and attention cost.

If a weekly review prevents missed commitments and makes Monday obvious, it may be worth the time. If a dashboard takes the same time to maintain and nothing changes when you skip it, it probably is not.

Test it during an ordinary busy week

A system that works only during a quiet setup week is not robust. The core loop should survive deadlines and interruptions: important commitments still get captured, the next useful action is findable, and missing one review does not require a cleanup project.

Monitoring matters only when it changes action

Progress monitoring does have evidence behind it. Harkin and colleagues meta-analyzed 138 randomized studies involving 19,951 participants and found that interventions increasing monitoring improved goal attainment on average. The meta-analysis is here.

That supports feedback, not decorative dashboards. Seeing that you are behind should change something: the next action, the schedule, the priority, the request for help, or the plan itself.

Look for fewer repeated failures

The strongest evidence may be boring. You stop forgetting the same things. You spend less time reconstructing stalled projects. Important tasks disappear less often beneath incoming requests. You recover faster after a disrupted week.

Locke and Latham's review of decades of goal-setting research likewise keeps the focus on performance through direction, effort, persistence, and strategy rather than on the apparatus surrounding a goal. Their review is here.

Do not let the metric replace the goal

A zero inbox can become more compelling than difficult work. Completing many small tasks can feel better than advancing one uncertain project. A writing tracker can reward word count when the real need is revision.

When the visible metric starts making the underlying goal worse, it has stopped being a useful proxy.

Run a small test

Choose one recurring failure and one maintenance cost. Observe the baseline, then use the system through enough ordinary cycles to encounter the situations it was built for. If you want a more deliberate comparison, use the design in How to Run a Small Personal Experiment Without Fooling Yourself.

The simplest useful question is: What happens better, faster, more reliably, or with less effort because this system exists?

If the only answer is that the system itself is well maintained, you have measured the wrong output.

A relevant Ulix tool

Track Analysis

Track Analysis can keep a lightweight timestamped record of a few outcomes or recurring failures during a test period and export the history to CSV. Record only the measures that could change the verdict.