How I Built an LSTM Deposit Forecasting Model at SVB UK Using PyTorch
When I joined SVB UK, the treasury team was forecasting deposit flows with moving averages and fixed 2026-10-6 20:6:20 Author: hackernoon.com(查看原文) 阅读量:2 收藏

When I joined SVB UK, the treasury team was forecasting deposit flows with moving averages and fixed percentage multipliers. I'm not saying that to be dismissive. For a lot of banks, that probably works fine. SVB is not a lot of banks.

First, Why SVB Deposits Are Nothing Like Normal Deposit Data

Before I get into anything technical, I want to spend some time on this. Not because it's an interesting background colour, but because if you don't understand it, none of the modelling decisions make sense.

SVB's clients aren't retail customers. There are no salary credits, no monthly direct debits, no consumer behaviour patterns to anchor on. The deposit base is almost entirely institutional: PE firms parking cash between deals, VC funds sitting on dry powder, portfolio companies burning runways. Deposits move in large, irregular chunks tied to funding cycles and market sentiment. The kind of thing that's very hard to quantify and very easy to be wrong about.

In the early weeks, I was genuinely optimistic about macro variables. Moody's interest rate data, GDP figures. It seemed reasonable that institutional investors would respond to the same macro signals. We tested it properly. The correlation just wasn't there. And even if we'd found something marginal, you then inherit this whole other problem: now you're retrieving, cleaning, and version-controlling external data every single month just to run the model. Tracking data vintages for backtesting. Chasing down why this month's Moody's pull looks slightly different from last month's. For the accuracy improvement we were actually seeing, it wasn't close to worth it.

So we stripped it back to one feature: the previous month's end-of-month deposit balance. When we landed there it felt almost embarrassingly simple. It took a few weeks to stop second-guessing it.

Models We Tried Before Landing On LSTM

I want to be honest about the path here because it wasn't clean, and I think the messy version is more useful than the "here's what we built" version.

We started with ARIMA. It made sense. It's interpretable, auditors understand it, the treasury team could follow the logic without an ML background. That matters more in a regulated environment than people realise until they're actually in one. But ARIMA kept falling over on the non-linear volatility in our data. Institutional deposit flows don't decompose neatly. It would get the trend direction broadly right and then completely miss the magnitude on a bad month. And the bad months are exactly when the forecast matters most.

From there we moved to a standard LSTM, then tested a bidirectional LSTM to see if looking at the sequence in both directions helped. The bidirectional version was genuinely interesting. I think there's a version of this problem where it's the right call. But for monthly end-of-balance prediction with our specific data it didn't justify the added complexity. The standard LSTM beat ARIMA on Mean Squared Error, which is what the treasury actually needed. That was the call.

Looking back I'd probably push to LSTM sooner. Though I also think the ARIMA work forced us to understand the data structure in ways that made the LSTM behave better. Those two things are probably both true and I'm not sure which one wins.

On PyTorch, Briefly

Honestly this section could be short because the answer isn't that interesting. I just find it easier to work with.

TensorFlow is fine. I've used it. But debugging in TensorFlow always felt like I was one abstraction layer away from where the actual problem was sitting. With PyTorch I can just step through it. When something breaks you can usually find it without spending forty minutes reading docs about graph execution modes.

The research thing also mattered practically. We were testing new architectures every few months and most papers now ship with PyTorch code. Not having to mentally translate from a different framework saves more time than you'd think, especially when you're doing it under deadline.

Speed was faster than what we compared against. Not dramatically. But it was never the slower option, which eventually just stopped being something I thought about.

How I Explain Gradient Descent To People Who Don't Do ML

The model risk team needed to understand what the model was actually doing. Not the maths, the intuition. I've explained gradient descent enough times now that I've settled on a version of this.

You're at the top of a mountain, blindfolded. Your job is to reach the lowest point in the valley. You can't see anything but you can feel the slope under your feet. So you take a step in whichever direction feels steepest downward. Check the slope again. Take another step. You keep going until the ground under you is flat. That's the bottom.

That bottom is where the difference between our forecast and the actual deposit balance is as small as it can get. The model starts with random weights, measures how wrong it is, nudges those weights slightly in the direction that reduces the error, and repeats until the improvement per step gets too small to matter.

The maths underneath based on a sample linear equation(for ease of understanding):

y_predicted = w * x + b

where w = weight and b = bias.

As per MSE formula,

loss = (1/n) * Σ(y_predicted - y)²

where y represents actual/true values

Expanding the loss function,

loss = (1/n) * Σ((w * x + b) - y)²

Calculating the partial derivatives for weight and bias terms,

∂loss/∂w = (2/n) * Σ((w*x + b - y) * x)

∂loss/∂b = (2/n) * Σ(w*x + b - y)

Updating the weight and bias values as follows:

w_new = w_old - learning_rate * ∂loss/∂w

b_new = b_old - learning_rate * ∂loss/∂b

PyTorch handles all of this through its computational graph. Forward pass tracks every operation, backward pass runs them in reverse to compute the gradients. You don't derive this by hand in practice. But understanding what it's doing matters when a model risk reviewer asks you to justify your optimisation approach at 4pm on a Thursday and you need to give them something they can actually write down.

The Real Challenge: Synchronizing Data From Multiple Business Systems

Everyone wants to talk about model architecture. The thing that actually ate most of this project was getting reliable data into the pipeline every single month.

SVB UK's deposit data came from multiple business lines, each sitting in a different system with a different publishing schedule. That meant manually synchronising retrieval across sources every month. When one feed was late, and it happened more than I'd like to admit, usually on a Friday afternoon when I had other things going on, the whole pipeline stalled. You're either waiting or you're deciding whether to run on incomplete data and documenting why, which carries its own audit implications either way.

Beyond timing, the data needed real work before it was usable. Missing values, stationarity checks, seasonality and trend handling. None of it is technically hard but it's slow and it has to be done consistently every time. One month early on I caught mid-backtesting that we'd handled a gap in one of the business line feeds differently than the previous run. Small thing. Took half a day to unpick and verify nothing downstream had been affected.

Data operations consumed at least as much time as the modelling work. Probably more. If I was doing this again that's the first thing I'd plan properly rather than assuming it would sort itself out.

Keeping The Model Honest After It Went Live

The accuracy bar was high and the validation process reflected that.

We ran one-month rolling backtests tracking actual versus forecast continuously, and nine-quarter historical backtests to check for accuracy threshold breaches over longer windows. A model that looks fine month-to-month can be quietly drifting in a direction you don't want, and nine quarters is usually where that shows up if it's there.

Every few months we also tested a new architecture against the same historical window. Not because the LSTM was performing badly, but because "the model is fine" and "the model is the best available option" aren't the same question, and in a production environment it's easy to stop asking the second one once the first is answered.

That ongoing work was heavier than I'd anticipated. People tend to treat model launch as the finish line. It really isn't.

The Governance Side, Which I'm Still Figuring Out

This is the part I find hardest to write about clearly, partly because I'm still working out what I actually think.

The basics: every run version-controlled, full logs of model versions and training data vintages, documented justifications for the architecture and optimisation choices. The model risk team needed to be able to follow the reasoning, not just see the results.

What I didn't fully appreciate until it bit me on a different project: a model that was performing well by every metric we tracked still failed governance review. The issue was a design decision made in week one that never got properly documented. By the time the review came around, the person who made the call had moved teams and nobody could reconstruct the reasoning clearly enough to satisfy the reviewer. The model got pulled. I still think it was probably the right model. It didn't matter.

So now I over-document. Not because I think it's the most efficient use of time, but because I've seen what happens when you don't and I'd rather not repeat it.

Did It Actually Work

Before this, SVB UK had no ML-based deposit forecasting. Static methods, that was it.

The LSTM beat them, and the improvement was most pronounced at longer horizons. Three months out and beyond. That's where static approaches break down hardest and also where better forecasting had the most real impact on liquidity planning. The one-month numbers improved too, but that wasn't where the treasury team noticed it.

What I'd Do Differently

Data infrastructure. I'd push harder on this in week one. We made modelling assumptions early that depended on clean, timely, consistently-formatted inputs and we didn't fully have that for a while. Unwinding assumptions mid-project is the kind of slow that's hard to describe until you've done it.

The single-feature thing I came around to later than I should have, if I'm honest. Kept feeling like we were leaving something on the table. Maybe we were, marginally. But I eventually realised that in a regulated environment you're not just optimising for MSE. You're optimising for MSE plus the ability to explain the whole thing to someone who has no particular interest in understanding it. That changes what "better" means.

Ongoing validation: just put it in the plan from day one and stop treating it as something to figure out later. Later always arrives faster than expected.

I keep thinking there's some tidy conclusion here about how the technical problem is never actually the hard part. Maybe there is. I've thought that at the end of the last few projects and it hasn't made me any better at predicting where the time goes.

Conclusion

Building this deposit forecasting model taught me that operational challenges often outweigh technical ones in regulated environments. The project succeeded because we prioritized simplicity and explainability over complexity. For anyone tackling similar ML initiatives in banking: invest in data infrastructure early, plan for governance overhead, and remember that the best model is the one that keeps working reliably in production.


文章来源: https://hackernoon.com/how-i-built-an-lstm-deposit-forecasting-model-at-svb-uk-using-pytorch?source=rss
如有侵权请联系:admin#unsafe.sh