A thirty-five-year-old forecasting method, two eager machine-learning rookies, a leak that almost fooled everyone, and the honest scoreboard that came after.
There’s a desk at every bank that nobody throws a party for.
The person who sits there doesn’t launch products. Doesn’t demo anything flashy in front of the board. Every quarter, they answer one question and go home: how many of these loans are going to go bad next quarter?
For years, that desk had one occupant. Everyone just called it ARIMA.
ARIMA was old. Not “outdated” old. Reliable old. The kind of old that makes younger models nervous, because it’s still right more often than they’d like.
ARIMA didn’t need much. Three numbers, actually. p, d, q. That’s it. That’s the whole personality.
“That’s it?” someone always asked, the first time they saw it.
“That’s it,” ARIMA said. “I look back a couple of quarters. I check if the trend needs smoothing out. And I write down what I see.”
ARIMA watched five accounts, quarter after quarter, for thirty-five years:
Every quarter, the drill was the same. Look at everything known so far. Guess the next four quarters. Wait. Get graded. Do it again, a little more history each time, sliding forward one quarter at a time. Fifty-nine times per account, for thirty-five years of quarterly data. That’s called a walk-forward backtest, and it’s a brutal, honest way to be graded: no peeking at the future, ever, at any point.
Nobody expected ARIMA to be a genius. It was graded against two lazy alternatives: just guess it’s the same as last quarter (simple-naive), and just guess it’s the same as this time last year (seasonal-naive). Beating a lazy guess isn’t supposed to be hard.
ARIMA didn’t just beat them. It embarrassed them.
| Account | ARIMA’s error, as a fraction of the lazy guess’s error |
|---|---|
| CRE | 0.135 (errors 87% smaller) |
| Business | 0.236 |
| All Loans | 0.269 |
| Credit Card | 0.325 |
| Mortgage | 0.348 |
A number below 1.0 means “better than the lazy guess.” Every single account, every single quarter for thirty-five years, ARIMA was comfortably below 1.0.
Nobody clapped. That’s just what the job looked like, every quarter, for years.
Then, one Tuesday, two new hires arrived.
Their names were Xander and Gigi. Everyone had heard of them. Xander went by his full name when he wanted to sound serious: XGBoost. Gigi’s was LightGBM, but nobody called her that unless something had gone wrong.
They didn’t work like ARIMA. ARIMA looked at a few numbers and wrote an equation. Xander and Gigi built trees: thousands of tiny yes/no questions, stacked on top of each other, voting on an answer.
“Was last quarter above 4%? If yes, was the quarter before that also rising? If yes…” and so on, hundreds of times, hundreds of trees.
They’d read about themselves winning competitions. Real ones. Big ones. They were, by reputation, the future.
Surya, who ran the desk now, decided to give them a fair shot: the exact same walk-forward drill ARIMA had been doing for years. No shortcuts, no head start. Same five accounts. Same fifty-nine graded quarters each.
There was one immediate problem: Xander and Gigi couldn’t read a raw number the way ARIMA could. So Surya gave them a study guide instead: what the account looked like 1, 2, 3, and 4 quarters ago, what it looked like 8 quarters ago (two years, worth checking for slower patterns too), and how choppy the last year had been.
“That’s all you get,” Surya said. “No looking anywhere else. Same information ARIMA has, just handed to you differently.”
Xander cracked his knuckles. Gigi didn’t say anything. She just started building trees.
The first results came back.
They were stunning. Numbers so good Surya actually sat back in the chair.
They were also, it would turn out, a lie.
The tell came from somewhere small: a routine experiment to see if Xander and Gigi could get even better with some tuning. Just twenty tries each at adjusting their own settings, evaluated fairly, on data they’d never touched.
One run came back with a number so low it didn’t look like a forecast error anymore. It looked like a typo.
Surya stared at it for a long moment.
Nobody predicts loan delinquency rates that well. Not from five numbers and a rolling average. Not in five tuning attempts. Not ever, really.
When a result looks too good to be true, it’s not a discovery. It’s a leak.
So Surya went looking. And found it.
Here’s what had happened. Xander and Gigi weren’t just handed “what happened 1, 2, 3, 4, 8 quarters ago.” They were handed a study guide with the answer key stapled in by accident, right at the edge, where nobody had checked.
Every training example was supposed to teach them: given the past, guess four quarters ahead. But near the very end of each training window, a few examples’ “four quarters ahead” landed inside the quarters they were about to be tested on that same round. A handful of times, for a handful of quarters, they weren’t predicting the test. They were copying it: their own upcoming answer, quietly sitting in their own homework.
Nobody had done this on purpose. It was one line of logic, quietly wrong, in exactly the kind of place these things always hide: a boundary condition, off by a few rows, at the edge of a sliding window.
Surya fixed it. Rebuilt the study guide so it stopped at exactly the right line, every time. Not one row of the future allowed to leak backward, ever. Then ran everything again, from scratch.
The stunning numbers disappeared. What was left was the truth.
| Account | ARIMA | Xander (XGBoost) | Gigi (LightGBM) |
|---|---|---|---|
| All Loans | 0.269 | 0.601 | 0.510 |
| Credit Card | 0.325 | 0.770 | 0.833 |
| Business | 0.236 | 0.660 | 0.643 |
| Mortgage | 0.348 | 1.663 | 1.490 |
| CRE | 0.135 | 0.537 | 0.515 |
ARIMA won. Every account. Not close.
Xander didn’t say anything for a while.
“Maybe it’s just how the coin landed,” he finally said. “Maybe on a different week, we’d have taken one.”
So Surya brought in a referee whose entire job is answering exactly that question: the Diebold-Mariano test, which looks at fifty-nine graded quarters and asks: is this gap real, or did ARIMA just get lucky?
The referee’s verdict, across every account, every forecast horizon, both rookies: 27 out of 40 times, not a coincidence. Real. Statistically real, at the strictest quarter-ahead call, every single time: ten out of ten. The verdict got a little less certain the further out they were asked to guess, which made sense; further out, everyone’s guesses get noisier, ARIMA’s included.
But there wasn’t a single one of the forty match-ups where the numbers leaned toward Xander or Gigi. Not once. Just some where the lead wasn’t big enough yet to call it, officially, beyond doubt.
Xander and Gigi didn’t quit. They tried two things.
First, they stopped competing separately and pooled what they knew. Instead of Xander learning Mortgage alone from its own 142 quarters, he learned from all five accounts at once, roughly 700 examples, telling them apart with a simple tag on each one: this one’s Mortgage, this one’s CRE. More to learn from. Shared patterns across accounts that move together.
It helped. Genuinely. On four of the five accounts, the pooled version beat the solo version by a wide margin.
It still didn’t beat ARIMA. Not once.
Second, they asked for a tutor. Twenty rounds of careful, honest tuning, searched for on the early three-quarters of their training history, tested only on the untouched final quarter of it, so the tutor couldn’t quietly cheat either.
The tutor helped Xander and Gigi in four of ten tries. In the other six, tutoring actually made things worse.
“That’s… not encouraging,” Gigi said.
“It’s honest, though,” Surya said. “You don’t have enough homework for a tutor to matter much. You’d need a lot more of it before fine-tuning starts paying off reliably. That’s not a flaw in you. It’s arithmetic.”
There was one more thing Surya wanted to know. Not how well Xander and Gigi were guessing: what they were even looking at when they guessed.
There’s a technique for this. You don’t ask the model. You watch it, very carefully, across every decision it makes, and add up which piece of information it leaned on the hardest. It’s called SHAP, and it doesn’t let a model lie about its own reasoning.
The answer came back the same way, ten times out of ten, every account, both Xander and Gigi:
Last quarter’s number. Just that. Somewhere between 47% and 59% of their entire decision, every single time, was just: what was it last quarter.
Xander looked almost embarrassed. “That’s… barely a model. That’s just looking at yesterday.”
“It’s not just that,” Surya said. “You also glanced at the quarter before, and the one before that, and a two-year-back check, and the recent choppiness. But yes. Mostly, you were looking at yesterday.”
Here’s the quiet part nobody had said out loud yet: that’s what ARIMA does too. ARIMA’s whole equation is mostly “yesterday, weighted correctly.” It just writes that down directly, in three numbers, instead of discovering it the hard way across a few hundred trees and eighty training rows.
Xander and Gigi hadn’t been wrong. They’d just spent a lot of machinery re-deriving something the veteran already had built into its bones on day one.
One last thing needed checking: not who guessed closer, but who was honest about how sure they were.
ARIMA said “I’m 95% confident the real number lands in this range,” quarter after quarter. Checked against fifty-nine real outcomes, it was right 98% to 100% of the time. If anything, a little too modest, a bit wider than it needed to be.
Xander and Gigi tried to build their own version of that confidence, from their own track record of past misses. Checked the same way, their “95% confident” was actually right only 75% to 83% of the time.
That’s not a small gap. That’s a confidence interval that lies about itself, and it lies hardest exactly when a bank can least afford it. The two-decade run of data made clear that both the mess and the model’s uncertainty about it get bigger right when a downturn is already happening. That’s the one moment a “95% confident” range needs to actually mean it. Xander and Gigi’s didn’t, not yet.
So here’s the honest version of the story, the one that doesn’t flatter anybody on purpose.
Classical ARIMA beat machine learning on this project, and it wasn’t close. Significantly so on most accounts. Not because Xander and Gigi were bad at their jobs. Because with only about 140 quarters of history per account, a three-number model can be estimated reliably, while a model built from hundreds of tree-splits needs more life experience than that to find anything ARIMA hadn’t already found. SHAP just made it official: both of them, veteran and rookies alike, were leaning on the same one thing.
The honest takeaway isn’t “machine learning is bad.” It’s “the tool has to match the size of the job.” Xander and Gigi didn’t need a better trick. They needed more to learn from, not more clever ways to slice the same 140 quarters. Real, new information. A regressor ARIMA never had access to either. Unemployment, maybe. A rate curve. Something the veteran couldn’t see coming from three numbers alone.
A full walk-forward backtest, complete with a caught-and-fixed data leakage bug, Diebold-Mariano significance testing, conformal prediction intervals, and SHAP interpretability. Full methodology, code, and results in the technical write-up and on GitHub.