How Much Supply-Chain History Is Enough for a Quant Backtest?

September 14, 2026

Altsets

Research by Altsets Research

Share

Calendar years alone do not determine statistical power because different network strategies consume different numbers of independent customer events, relationship changes, companies, clusters, and market regimes.

Data used:Altsets Supply Chain Intelligence: 90k+ entities, 400k+ relationships, 20+ years of history.

Key findings

  • Historical depth should be matched to the unit that generates the signal, such as company-month observations, customer earnings events, relationship formations, or relationship dissolutions, rather than judged from calendar years alone.
  • A long dataset still needs multiple regimes, independent network clusters, untouched holdout data, and honest accounting for repeated strategy trials because more snapshots do not eliminate overfitting or network dependence.

There is no fixed number of years that makes a supply-chain backtest adequate. Enough history means enough independent events, relationship changes, companies, and network clusters for the specific hypothesis, not merely a long calendar span.

Match the history requirement to the event that generates the signal

Altsets has monthly point-in-time relationship history spanning 2006 through 2026. That depth can support long-horizon research, but it should not be interpreted as one universal answer to "how much data is enough?" A static customer-concentration feature generates one cross-sectional observation per company per formation date. A customer-earnings propagation strategy generates observations around scheduled customer events. A relationship-formation feature generates observations only when a new edge appears under the historical observation rules. The effective sample sizes can differ dramatically even when all three strategies use the same twenty-year database.

A useful research plan begins by identifying the unit of evidence. If the hypothesis is that high supplier concentration predicts future downside, the unit may be company-month observations with appropriate controls for network dependence. If the hypothesis is that Nvidia earnings predict connected suppliers, the unit is closer to independent Nvidia earnings events crossed with supplier exposures. If the hypothesis is that relationship dissolution predicts future supplier weakness, the unit is the set of confirmed dissolution events. Counting years before defining the unit can give the researcher false comfort.

Regime coverage matters because economic mechanisms are conditional

A long dataset is valuable partly because it contains different environments. Customer and supplier dependencies can behave differently during recessions, shortages, rapid growth, trade restrictions, inflation shocks, and calm periods. A strategy tested only during one favorable era can mistake a regime-specific effect for a stable relationship. This is why r/algotrading discussions repeatedly push new researchers to think about walk-forward testing and multiple regimes rather than choosing an arbitrary number of years.

The right response is not to include every ancient period automatically. The business structure itself can change. A 2007 semiconductor network may not represent today's AI supply chain, and a feature based on modern relationship coverage may have lower-quality history in earlier years. The researcher should compare results across historical subperiods and ask whether the mechanism should plausibly survive the structural changes between them. Long history is most valuable when it expands the set of independent regimes without quietly changing what the feature means.

More rows cannot compensate for repeated testing on the same history

A researcher can have millions of company-month observations and still overfit if the same twenty-year period is used repeatedly to select features, thresholds, and holding rules. Every new strategy variant consumes information from the sample. Bailey and coauthors formalize the probability of backtest overfitting, while Harvey, Liu, and Zhu show why multiple testing raises the threshold for believing newly discovered return factors.

Supply-chain data increases this danger because the graph supports an enormous number of plausible features. The researcher can test customer concentration, supplier concentration, shared nodes, centrality, relationship age, edge churn, two-hop exposure, event propagation, and hundreds of combinations. A long historical dataset makes that exploration possible but does not make repeated mining free. Untouched holdout periods and pre-specified hypothesis families remain necessary no matter how many snapshots exist.

Independent events and network clusters should be reported alongside years

A strong supply-chain backtest can report the calendar span and the structure of the evidence. For example: twenty years of monthly snapshots, several thousand securities over time, a certain number of unique customer-supplier relationships, a certain number of relationship-change events, and a certain number of distinct customer-event clusters. The exact numbers depend on the dataset and research universe, but the principle is universal. The reader should be able to tell whether the result is supported by broad independent evidence or by many repeated observations around a small number of economically related nodes.

This is especially useful when comparing strategies. A network-centrality factor may have a huge cross-sectional panel but relatively slow feature evolution. A customer-event strategy may have fewer observations but cleaner event identification. A relationship-churn model may have fewer still, but each observation can represent a more substantial structural change. The amount of history required is therefore a property of the hypothesis, not a product specification.

The conclusion is to count the evidence your strategy actually consumes

There is no universal rule that a supply-chain strategy needs five, ten, or twenty years. The right historical depth is enough to observe many independent instances of the mechanism, enough market regimes to challenge it, and enough untouched data to evaluate it after the research choices are frozen. Twenty years can be inadequate for a rare-event hypothesis and excessive for a relationship whose economic meaning changed completely halfway through the sample. Count events, clusters, regimes, and trials alongside calendar time.

The survivorship-bias guide explains why a long history still needs the correct historical security universe. The network-dependence guide explains why a large row count can overstate the amount of independent statistical evidence.

For relationship definitions and evidence limits, read the Altsets methodology.

Sources

Methodology

Read the methodology for this research.