STORM Financial AI: Learning Across Stocks and Through Time
STORM combines temporal and cross-asset representations for financial prediction. Understand its architecture, ablations and the limits of historical backtests.
6
Select a figure to open it at full size.
A financial forecasting model has two related jobs. It must recognize how an asset’s behavior develops over time and how that asset relates to the rest of the market. A method that handles only one can miss information carried by the other.
STORM, a factor-model architecture from Yilei Zhao and colleagues, learns these two views separately before combining them. Its research value lies in that design and in the experiments that test whether both views contribute. Its backtests are historical simulations, not a forecast of returns available to an investor today.
The Core Insight: Two Views of the Same Market
The model receives a 64-day window of price and technical-indicator features. One branch follows individual assets through time, working with four-day patches. The other examines assets together within a day, looking for relationships across the market.
Both branches use learned numerical representations. Each representation is mapped to a nearby entry in a finite codebook, a process called vector quantization. Instead of retaining an unconstrained vector for every input, the system learns a reusable vocabulary of patterns. The reported configuration uses 512 codebook entries.
The model encourages those entries to be diverse and reconstructs price information from the representations. These training objectives aim to preserve useful structure and avoid having all inputs collapse onto the same few codes. They do not prove that each code corresponds to a named economic factor such as value or momentum.
Original Figure 2 separates the time-series and cross-sectional paths before they are combined. The architecture learns numerical factors; it does not establish causal market drivers. Zhao et al., page 4.
Cross-attention then combines information from the two branches, while an alignment objective encourages related representations to agree. The combined information feeds a model of future returns.
There is an important distinction between training and use. During training, the architecture can use observed future returns to learn its target distribution. At prediction time, the corresponding inference path must work from past information. A faithful implementation has to preserve that separation; otherwise, a strong backtest may simply reveal that future information leaked into the forecast.
What the Research Actually Shows
The experiments use 28 Dow Jones stocks and 408 S&P stocks, with 152 features derived from prices and technical indicators. The stated training period runs from April 2008 through March 2021; testing runs from April 2021 through March 2024.
The authors first examine whether the model’s predictions rank subsequent stock returns usefully. Rank correlation asks whether stocks predicted to do better tend to appear higher in the realized-return ordering. It is not the fraction of trades that make money.
In the reported S&P experiment, STORM’s mean rank correlation is 0.062, compared with 0.057 for HireVAE. In the Dow experiment, the values are 0.065 and 0.058. These are modest correlations whose practical value depends on costs, portfolio construction and stability.
The ablation study asks a more specific question: what happens when a component is removed? Removing the temporal branch reduces the reported rank correlation to 0.053 for S&P and 0.055 for Dow. Removing the cross-sectional branch produces 0.054 and 0.053. Under this evaluation, both branches contribute to the complete model’s result.
Original Table 4 connects the architecture to a test: removing either view weakens the reported ranking result. Correlation is not trading accuracy or a return percentage. Source, page 8.
That is stronger evidence for the design than a high-level assertion that combining space and time should help. It remains evidence within the selected datasets and historical split, rather than a guarantee that the same components will help every market or horizon.
A Better Forecast Does Not Settle the Trading Decision
The paper also feeds predictions into portfolio management and individual-asset trading tests. The portfolio strategy selects five assets and permits three replacements at a rebalance. Its reported S&P annualized return is 18.8%, compared with 7.9% for FactorVAE under the paper’s setup. The backtest assumes a transaction-cost rate of one basis point.
Those details matter because the result combines the predictor with a trading rule and a cost assumption. Different turnover, execution costs or asset availability can change the outcome. The paper does not fully resolve how a reader should reconstruct a point-in-time stock universe and validation process from the reported description alone.
Risk results also resist a simple “better everywhere” interpretation. In the individual-asset tests, STORM’s reported maximum drawdown for Microsoft is 0.210, compared with 0.182 for the Transformer baseline. For Intel, it is 0.227, compared with 0.149 for LSTM. A larger drawdown means a deeper peak-to-trough loss during the test. Higher returns in other comparisons do not erase these counterexamples.
The architecture’s codebook is another boundary. A discrete code can make a representation easier to organize without making it an economically interpretable or causal factor. A technical team should distinguish inspectable model components from an explanation of why markets moved.
Real-World Applications: Evaluate the Representation Before the Portfolio
The most useful first question is whether the two-view representation improves an existing forecasting pipeline under a controlled comparison. Keep the same assets, dates, features and evaluation rules. Compare a simple baseline, a temporal-only model and the combined model before introducing a new portfolio strategy.
Use chronological validation and a final untouched period. Check feature timestamps, corporate-action treatment and whether the asset universe could have been known at the time. Any operation using future returns belongs on the training or evaluation side of the boundary, never in live inputs.
Then examine whether any ranking improvement survives realistic costs and turnover constraints. Report periods of deterioration and drawdowns, not just aggregate return. A useful result should persist across reasonable evaluation choices rather than depend on one favorable interval. The same discipline applies to AI financial-risk models, where prediction quality and investment usefulness are separate tests.
Implementation Frameworks
Qlib’s strategy and backtesting interfaces provide a way to connect prediction signals with trading decisions and evaluate a declared strategy. They are useful evaluation infrastructure; they do not reproduce STORM’s dual-codebook architecture automatically.
Start by implementing the same simple strategy for all competing signals, with explicit costs and a fixed chronological split. Add the research architecture only after the baseline pipeline passes timestamp and feasibility checks. Save predictions before backtesting so changes to portfolio rules cannot silently become changes to the forecasting experiment.
An adoption gate should require stable out-of-period improvement, acceptable turnover and risk, and a reproducible account of the data available at each decision. No model or trading strategy was executed for this article.
TechClarity’s View
STORM offers a useful architectural hypothesis: learning patterns across assets and through time can produce a better representation than relying on either view alone. Its ablations provide evidence for that hypothesis within the reported tests.
The appropriate next step is reproduction under a tightly controlled data pipeline. The research is a candidate for improving a forecasting component, not a reason to accept the headline backtest as an investment expectation.