Skip to content
← All posts

Do budgeting apps work?

Do budgeting apps actually work?

Somebody ran a randomized controlled trial on 9,035 people to find out. Spending did not drop. Here is what the study tested, what it did not test, and why it is evidence against the app I am building as much as any other.

If budgeting apps have never stuck for you, you are not undisciplined. You are the typical case, and there is a decent experiment to show it.

The study

In 2020, researchers at Common Cents Lab and Irrational Labs ran a randomized controlled trial with Clarity Money, a personal finance app, on 9,035 people.

They split users into groups:

  • one group got a single spendable number, a simplified figure for what remained for the period
  • one group got full category budgeting, the traditional envelope approach
  • the rest got neither, as a control

Then they watched what people actually spent.

Spending did not drop. Not for the group given the single number. Not for the group given categories. The intervention that was supposed to change behavior did not change it.

That is a real randomized trial, at a sample size most consumer-finance research never reaches, and it says the thing most of the industry would rather not print.

Why this is not surprising once you look at it

A budget is a plan. Plans decay.

The specific way a household budget decays is worth naming, because it is not a character flaw:

It requires maintenance you have no reason to enjoy. Categorizing transactions is data entry. It is the same work as expense reports, which nobody does voluntarily either.

The inputs move without telling you. A direct debit changes amount. A subscription renews at a new price. Payday shifts because the last working day fell on a weekend. Each of those quietly invalidates the plan, and none of them announces itself.

The feedback arrives too late to act on. You learn in a monthly review that you overspent on eating out. The information arrives weeks after every decision it could have changed.

Falling behind is terminal. Miss a week of categorizing and you now face an hour of backlog to get a number you are not confident in. Most people quietly stop there, and then feel bad about it, which is its own tax.

So the app is not failing because the math is hard. It is failing because it hands you a job, and the job competes with everything else in your life every single day.

What the study does not say

I want to be careful here, because this study gets over-claimed in both directions.

It does not say budgeting is pointless. It says that showing people a budget inside an app did not reduce their spending in this trial. People who genuinely enjoy the method, and there are plenty, get real value from it. If YNAB works for you, it works, and the fact that it works partly because it demands attention is the point of it.

It does not say all apps are equivalent. It tested particular interventions in one product over one period. It is evidence about a category, not a verdict on every possible design.

It does not measure everything worth measuring. Not overdraft fees avoided, not anxiety, not whether people knew where they stood. It measured spending. Spending is a reasonable outcome to pick, and it is not the only reason someone opens a money app.

The uncomfortable part, for me

I am building a budgeting app. Its core feature is a single number telling you what you can spend today.

That is one of the arms that did not work.

I do not think it is honest to cite this study for the half that indicts everyone else and skip the half that points at me, so: the strongest available evidence says that showing you a spendable number, on its own, will not change what you spend.

What I think is actually different, and why I might be wrong

My read of the failure is that all three arms put the same job on the person: remember to look.

A budget you have to remember to check is one more thing to remember. The single-number arm made the checking cheaper. It did not remove the checking.

So the bet Ralphy makes is that the thing worth automating is not the arithmetic. It is the watching. The app should notice that a bill changed, that a tight stretch is coming, that a paycheck landed early, and speak up before the moment you would have wanted to know, rather than waiting to be opened and consulted.

On most days you should not need to open it at all.

That is a hypothesis. It is a reasonable one, it follows from the shape of the failure rather than from wishful thinking, and it has not been through a randomized trial. If somebody runs one and it comes out flat, I would rather have said this in advance.

If you are choosing an app

A few things I would actually weigh, given the above:

Does it require ongoing data entry to stay correct? If yes, be honest with yourself about whether you will do it in month four. Most people will not, and that is fine, it just means picking differently.

Does it tell you things without being opened? Notifications that fire on a real event beat a dashboard you have to remember to visit.

Does it show its working? A number you cannot check is a number you will stop believing the first time it looks wrong. And it will look wrong at some point, because your money is messy.

Does it fail safe? Working from a pending deposit that has not landed, or a card balance that excludes pending charges, will flatter you at exactly the wrong moment. I wrote about which balance to trust separately.

Is the price proportionate? Some are $100 a year and more. If a tool exists mainly to stop one bad week a year, it should be priced like it.

And if the answer is that none of them fit, the pen-and-paper version genuinely works. It takes about fifteen minutes a week.


If you want the version that does the watching for you, Ralphy is in beta, or you can try a morning of it in your browser with no download and no bank connection.

Ralphy does this part for you.

One number each morning, worked out around every bill and paycheck still coming. Two weeks free, then $3.99/month.

Get early accessTry the demo