Research article

Get dataset information

Look-Ahead Bias in Equity Backtests Using Current Supply Chain Data

Cite Dataset

Look-Ahead Bias in Equity Backtests Using Current Supply Chain Data

Abstract

A historical quantitative signal can use perfectly clean price data and still contain future information when its network structure is reconstructed from relationships observed only at the end of the sample. The distortion is different from ordinary price look-ahead. The contaminated object is the adjacency matrix and, when economic edge weights are also backfilled, the cross-sectional exposure measure itself. This paper measures that distortion with eight current commercial relationships spanning Intel, Boeing, Nike, and PepsiCo and quarterly Altsets point-in-time observations from September 2021 through March 2026. Each edge receives a bilateral dependency weight built from Relationship Size, Supplier Revenue %, and Customer Cost %. A focal firm's weighted economic degree is then computed from the true historical graph and from the current graph projected backward. Across the 19 historical quarters, only 4.1 of the eight current relationships are present on average. The mean Spearman rank correlation between point-in-time and backfilled firm signals is 0.399, and the two highest-ranked firms agree only 63.2% of the time. Normalized absolute exposure error averages 71.4%. An exact decomposition attributes most of the signed distortion to edge existence rather than historical reweighting. A dollar-neutral diagnostic portfolio formed from the signal produces a mean absolute annual return-series divergence of 20.5 percentage points over 2022 through 2025, even though average returns under the two constructions are nearly identical. The empirical pattern is therefore less consistent with a simple claim that backfill makes returns look better than with a broader result: current-network backfill can manufacture historical stability, suppress signal turnover, change cross-sectional ranks, and materially alter the path of a network-conditioned backtest.

Keywords: Look-ahead bias; Temporal networks; Cross-sectional ranking; Portfolio turnover; Signal contamination

Problem

A network signal requires two historical objects: the state of the firms and the state of the links connecting them. Quantitative research often treats the second object more casually than the first. Prices are timestamped, corporate actions are adjusted, and financial variables may be lagged to publication, while a present-day customer-supplier map is attached to every past observation as though commercial adjacency were permanent. That operation introduces information unavailable at the historical decision date even when no future return enters the model. It also creates a relationship-level analogue of survivor conditioning: the researcher begins with edges that are visible now and asks how the firms connected by those edges behaved earlier. Brown, Goetzmann, Ibbotson, and Ross showed that sample truncation by survival can generate apparent predictability in performance studies, while Bailey, Borwein, Lopez de Prado, and Zhu formalized a separate class of backtest distortions arising from model selection and repeated experimentation. The network problem studied here occurs earlier in the research pipeline. The information set itself has been changed before signal estimation begins.

The finance literature supplies an economic reason to care about that contamination. Cohen and Frazzini documented return predictability across firms connected by principal-customer relationships, consistent with delayed incorporation of news about economically linked firms. Barrot and Sauvagnat showed that firm-specific shocks can propagate through production relationships, especially when inputs are difficult to substitute. Herskovic's production-network asset-pricing model makes network concentration and sparsity sources of systematic risk. These results differ in identification and scale, but all assign financial meaning to the topology or economic importance of commercial links. If the graph itself is dated incorrectly, the researcher is changing a state variable that prior research has treated as economically consequential.

Temporal-network theory sharpens the distinction. Holme and Saramaki model networks in which edges become active and inactive through time, so temporal structure enters the mathematical representation rather than being absorbed into a static graph. Time aggregation discards information about when a connection exists and can create network properties that the temporal system never possessed at a given date. A commercial network has the same basic problem even when the object of interest is weighted degree rather than a time-respecting path. A supplier added in 2025 cannot contribute to a firm's 2022 network centrality merely because the pair is visible in 2026. A relationship that existed in 2022 with a small economic weight cannot be represented by its mature 2026 weight without importing later information about the scale and dependence of the pair.

The primary hypothesis is therefore directional for exposure but two-sided for returns. Conditional on selecting relationships from the current graph, backfilling should generally overstate historical network exposure because current edges that were absent from earlier point-in-time snapshots enter with positive weights. Historical reweighting can offset that bias when a surviving relationship was stronger in the past than it is currently, so total error need not be positive for every firm-date observation. The predicted portfolio effect is different. A change in ranks, relative magnitudes, or turnover should alter a signal-conditioned return series, but economic theory gives no reason for the contaminated series to be systematically higher. A backfilled graph can improve, damage, smooth, or reverse apparent historical performance depending on which firms receive the artificial exposure.

Signal

The empirical panel uses eight current Altsets relationships, two for each of four focal firms. The selection is deliberately heterogeneous: Intel is paired with SK Hynix and Samsung Electronics on the supplier side; Boeing with Hyundai Glovis as a supplier and Korean Air as a customer; Nike with Teijin as a supplier and MAP Aktif Adiperkasa as a customer; and PepsiCo with Toyota Industries as a supplier and Seven & I Holdings as a customer. The sample is small enough that every edge can be inspected directly while still permitting a cross-sectional ranking. It also avoids making the result a disguised semiconductor-industry study. The point-in-time panel contains quarterly relationship states and the three Altsets dependency metrics used here: Relationship Size, Supplier Revenue %, and Customer Cost %.

Focal firmCounterparty roleCounterpartyFirst complete snapshotCurrent size, $mSupplier revenue %Customer cost %
IntelSupplierSK Hynix2021-09-281,159.61.352.70
IntelSupplierSamsung Electronics2023-06-28896.30.351.98
BoeingSupplierHyundai Glovis2021-09-28138.90.600.12
BoeingCustomerKorean Air2022-06-28833.80.7936.58
NikeSupplierTeijin2025-12-28129.22.070.39
NikeCustomerMAP Aktif Adiperkasa2021-12-2814.00.011.98
PepsiCoSupplierToyota Industries2023-06-284.90.010.08
PepsiCoCustomerSeven & I Holdings2021-09-281,987.91.943.33

Public company evidence is consistent with treating commercial topology as a changing state variable. Nike reported 146 strategic Tier 2 suppliers in fiscal 2023, 169 in fiscal 2024, 184 in fiscal 2025, and 205 in fiscal 2026. Those filings also show changes in the number and composition of finished-goods factories and contract manufacturers. Boeing provides an even more concentrated illustration of weight evolution. Korean Air announced a commitment for up to 50 Boeing widebody aircraft in 2024, followed in 2025 by additional orders and a record commitment for 103 Boeing jets. A current Boeing-Korean Air edge therefore embeds an economic relationship whose scale changed materially near the end of the sample. Projecting its mature state backward imposes later commercial intensity on earlier dates.

For relationship i,ji,j at quarter tt, let Vij,tV_{ij,t} denote Relationship Size measured in millions of dollars, sij,ts_{ij,t} Supplier Revenue % expressed as a decimal fraction, and cij,tc_{ij,t} Customer Cost % expressed as a decimal fraction. The edge dependency weight is

wij,t=ln⁡(1+Vij,t)sij,tcij,t.w_{ij,t} = \ln(1+V_{ij,t}) \sqrt{s_{ij,t}c_{ij,t}}.

The logarithm prevents a multi-billion-dollar relationship from dominating solely through absolute scale. The geometric mean of the two percentage measures gives high weight to relationships with meaningful dependence on both sides while penalizing a pair that is economically large for one firm and negligible for the other. Relationship Size still distinguishes two pairs with similar proportional dependence. The construction has no fitted coefficient, return target, or optimized parameter. Its purpose is to create a reproducible weighted-network object whose historical state genuinely requires all three dependency variables.

Let Aij,t=1A_{ij,t}=1 when the exact directed relationship is present with complete metrics in the point-in-time snapshot and zero otherwise. The historical economic degree of focal firm ii is

Ci,tPIT=∑jAij,twij,t.C^{PIT}_{i,t} = \sum_j A_{ij,t}w_{ij,t}.

The strict current-network backfill freezes the terminal adjacency and terminal economic weights:

Ci,tBF=∑jAij,Twij,T,C^{BF}_{i,t} = \sum_j A_{ij,T}w_{ij,T},

where TT is the current Altsets snapshot. The quantity is indexed by tt only because it is being assigned to historical dates; its value is constant through the backtest. The comparison is intentionally strict. It represents the common research shortcut in which a precomputed current commercial network is joined to historical market data.

The data permit an exact decomposition. Define a counterfactual that retains the historically correct edge set but replaces its relationship weights with current weights:

Ci,tW=∑jAij,twij,T.C^{W}_{i,t} = \sum_j A_{ij,t}w_{ij,T}.

Then

Ci,tBF−Ci,tPIT=(Ci,tBF−Ci,tW)+(Ci,tW−Ci,tPIT).C^{BF}_{i,t}-C^{PIT}_{i,t} = \left(C^{BF}_{i,t}-C^{W}_{i,t}\right) + \left(C^{W}_{i,t}-C^{PIT}_{i,t}\right).

The first term is edge-existence error. The second is relationship-weight error among edges that were actually present at tt. Because the centrality measure is linear in the weighted adjacency matrix, the decomposition requires no residual interaction term. For a current-edge-selected sample, the edge term is mechanically nonnegative before an edge appears historically. The weight term can have either sign.

Three diagnostics measure what this difference would do to an investment signal. Spearman correlation compares the point-in-time and backfilled cross-sectional ranks. Top-half overlap records whether the same two of four firms occupy the high-centrality portfolio. Exposure error is

Dt=∑i∣Ci,tBF−Ci,tPIT∣∑iCi,tBF.D_t = \frac{\sum_i\left|C^{BF}_{i,t}-C^{PIT}_{i,t}\right|} {\sum_i C^{BF}_{i,t}}.

Finally, continuous portfolio weights are obtained by centering each signal and normalizing gross exposure to one:

qi,t=Ci,t−Cˉt∑k∣Ck,t−Cˉt∣.q_{i,t} = \frac{C_{i,t}-\bar C_t} {\sum_k |C_{k,t}-\bar C_t|}.

The portfolio is dollar neutral, with positive positions in above-average centrality firms and negative positions in below-average firms. No transaction costs are subtracted in the primary return comparison because the object of interest is the return-series difference generated solely by network construction. Turnover is measured separately, so any linear cost assumption κ\kappa can be applied as κ\kappa times turnover without changing the network comparison.

Distortion

The current network is a poor representation of the typical historical cross-section in this sample. Across the 19 quarterly dates from September 2021 through March 2026, an average of 4.1 of the eight current relationships is present in the actual point-in-time graph. The median is four, with quarterly counts ranging from two to seven. A researcher backfilling the current network therefore inserts roughly half of the terminal edge set into the average earlier snapshot. The normalized absolute exposure error DtD_t averages 71.4% and has a median of 82.0%. Measured in signed terms relative to the total current centrality mass, the backfilled construction exceeds the point-in-time signal by 63.5% on average. Edge existence accounts for 58.7 percentage points of that average signed discrepancy, while substituting current weights for the historical weights on relationships already present contributes 4.8 percentage points. Historical reweighting is locally important in individual quarters, but relationship membership is the dominant signed distortion over the full period.

The ranking error is smaller than the exposure error and still large enough to change portfolio construction. Mean Spearman correlation between the current-network ranking and the point-in-time ranking is 0.399, with a median of 0.400. The quarterly correlation ranges from -0.800 to 0.949 before the terminal current state. The two highest-ranked firms agree only 63.2% of the time, and 13 of 19 quarters contain at least one top-half membership difference. The extreme quarter is June 2024. Point-in-time centralities rank Nike first, Intel second, Boeing third, and PepsiCo fourth. Backfilling the current graph produces Boeing first, PepsiCo second, Intel third, and Nike fourth, giving a Spearman correlation of -0.800 and zero overlap between the two high-centrality sets. At that historical date, Nike's MAP Aktif relationship carries a large contemporaneous dependency weight, while several relationships that later become important in the current graph are absent from the point-in-time snapshot.

The magnitude comparison also shows why rank correlation is an incomplete audit. In December 2023, the point-in-time and backfilled specifications select the same two highest-ranked firms and produce a Spearman correlation of 0.949. Yet normalized absolute exposure error is 95.9%. Boeing's point-in-time centrality is approximately 0.033 and PepsiCo's is approximately 0.0005, with Intel and Nike at zero in this eight-edge sample. The current graph assigns approximately 0.375 to Boeing, 0.194 to PepsiCo, 0.191 to Intel, and 0.048 to Nike. The ordering looks similar because two firms remain above two others, while the economic magnitude represented by the signal has been almost entirely rewritten. A backtest validation process based only on rank stability or constituent overlap could therefore approve a network whose portfolio weights are still severely contaminated.

Turnover changes for the same reason. The continuous point-in-time portfolio has average quarterly turnover of 0.364 over the historical window, where turnover is one-half of the absolute change in portfolio weights. The strict backfilled portfolio has zero network-signal turnover because the current weighted graph has been copied unchanged to every historical date. Some of that result follows directly from the strict counterfactual, but the economic interpretation is less trivial: a present-day graph removes the formation, disappearance, and reweighting events that a historical network strategy would actually have experienced. A backtest can consequently appear easier to trade, more stable, and less sensitive to commercial reconfiguration even before transaction costs or execution assumptions enter the calculation.

Returns

The return experiment is a diagnostic of path dependence rather than a test of expected-return alpha. For each December 28 snapshot from 2021 through 2024, the centered network signal determines dollar-neutral weights held against the next complete calendar year's stock return for Intel, Boeing, Nike, and PepsiCo. The market series are adjusted historical stock-price changes reported by Macrotrends. Its stock-history pages report the underlying series as adjusted for splits and dividends. Intel returned -46.64%, 94.56%, -59.57%, and 84.04% in 2022 through 2025; Boeing returned -5.38%, 36.84%, -32.10%, and 22.67%; Nike returned -29.04%, -6.01%, -29.10%, and -13.81%; PepsiCo returned 6.78%, -3.29%, -7.60%, and -1.84%.

Holding yearPoint-in-timeCurrent graph backfilledDifference
202221.4%11.5%9.9 pp
2023-23.7%18.3%-42.0 pp
20240.3%-1.1%1.3 pp
202543.8%15.0%28.9 pp
Mean10.4%10.9%-0.5 pp
Annual standard deviation28.9%8.5%20.4 pp

The average return is almost unchanged, which is informative. The backfilled version is not mechanically an optimistic-return machine in this sample. Its main empirical effect is to create a different return process. The mean absolute annual difference between the two series is 20.5 percentage points, while their four-year arithmetic means differ by only 0.5 percentage points. The backfilled portfolio also has an annual sample standard deviation of 8.5%, compared with 28.9% for the point-in-time version. That compression follows from a current graph that assigns historically stable weights to relationships whose actual network state changed sharply. A researcher examining only average performance could miss most of the contamination; a researcher examining volatility, drawdowns, annual attribution, or capacity could reach a materially different conclusion.

The 2023 observation makes the mechanism especially visible. December 2022 top-half membership agrees under the two network constructions, so a binary portfolio-overlap test reports no problem. The continuous point-in-time signal nevertheless produces a 2023 return of -23.7%, while the current-network signal produces +18.3%. The 42.0 percentage point difference comes from the relative magnitudes assigned inside the same four-stock universe. The result therefore cannot be reduced to a story in which backfill merely adds a few different portfolio constituents. Relationship weights affect how much capital the signal allocates even when the broad ranking looks acceptable.

Tests

A less severe counterfactual freezes only the current topology. Historical relationship weights are retained whenever an edge actually exists at date tt; when a current edge is absent from the historical graph, its current weight is inserted. This construction allows genuine historical reweighting among observed links and therefore removes the automatic zero-turnover property of the strict backfill. Distortion remains. Across the 19 historical quarters, mean Spearman correlation with the point-in-time signal is 0.417, top-half overlap is 71.1%, and normalized exposure error is 58.7%. Average quarterly signal turnover falls to 0.186 from 0.364 under the point-in-time graph. In the four annual return observations, mean absolute return-series divergence rises to 25.9 percentage points, and the topology-backfilled portfolio's annual standard deviation is 10.8% versus 28.9% under the historical graph. The central result therefore survives when historical weights are allowed to move.

A second test removes Relationship Size from the edge-weight formula and uses only the geometric mean of Supplier Revenue % and Customer Cost %. If the original result were driven primarily by the logarithmic size term or by one or two exceptionally large dollar relationships, rank and exposure errors should contract sharply. They do not. Mean historical rank correlation is 0.447, top-half overlap is 65.8%, and normalized absolute exposure error is 69.3%. The three-metric specification remains preferable for the main analysis because absolute economic scale is part of dependency, but the backfill result is not created by the specific transformation applied to Relationship Size.

These calculations also define conditions under which current-network backfill would be relatively harmless. The bias should be small when edge membership is persistent, current and historical dependency weights are similar, and the signal is insensitive to moderate changes in cross-sectional magnitude. A null result would therefore be plausible in a mature network with long-lived counterparties and nearly constant commercial shares. The present sample has the opposite structure. Several terminal relationships enter late, others disappear for long intervals, and the economic weight of observed links changes enough to alter ranks and portfolio sizing. The empirical result should be read as a measured distortion for this eight-relationship panel, rather than as an estimate of a universal percentage error for all commercial-network strategies.

The mechanism also separates this problem from conventional backtest overfitting. No alternative parameter set was searched to maximize the reported performance, and the edge-weight formula was fixed without reference to future returns. The contamination occurs because the backfilled model conditions historical exposures on a graph observed later. Cohen and Frazzini's economic-link return effect, Barrot and Sauvagnat's shock propagation evidence, and Herskovic's network asset-pricing results all give commercial structure financial content. Once the network is part of the state vector, dating the network incorrectly is economically equivalent to dating another predictive state variable incorrectly.

Conclusion

The empirical comparison identifies a specific failure mode in network-based quantitative research. A current commercial graph projected backward does more than alter a historical picture of who traded with whom. It changes economic centrality, relative ranks, portfolio membership, capital weights, turnover, and the realized path of a signal-conditioned return series. In the eight-edge panel studied here, the average historical snapshot contains only 4.1 of the eight current relationships. The resulting strict backfill produces 71.4% mean normalized absolute exposure error, a 0.399 mean rank correlation, and only 63.2% top-half overlap. The signed decomposition assigns most of the discrepancy to edge existence, while relationship-weight changes provide an additional source of error.

The return experiment gives the distortion an economically interpretable scale. Point-in-time and backfilled portfolios have nearly the same four-year average annual return, yet their annual paths differ by 20.5 percentage points on average in absolute terms and exhibit sharply different volatility. A backtest built from current relationships can therefore deliver a cleaner and more stable history without producing a conspicuously inflated mean return. That feature makes the error harder to detect from headline performance statistics. The graph can be wrong while the backtest remains plausible.

For quantitative work using commercial relationships, the relevant information set at date tt is the weighted adjacency matrix that existed at tt. Historical prices alone cannot make a network signal historical. The network itself is a dated financial variable.

References

Altsets supply chain dataset.

Bailey, D.H., Borwein, J.M., Lopez de Prado, M., and Zhu, Q.J. 2017. "The Probability of Backtest Overfitting." Journal of Computational Finance.

Barrot, J.N., and Sauvagnat, J. 2016. "Input Specificity and the Propagation of Idiosyncratic Shocks in Production Networks." Quarterly Journal of Economics 131(3), 1543-1592.

Brown, S.J., Goetzmann, W., Ibbotson, R.G., and Ross, S.A. 1992. "Survivorship Bias in Performance Studies." Review of Financial Studies 5(4), 553-580.

Cohen, L., and Frazzini, A. 2008. "Economic Links and Predictable Returns." Journal of Finance 63(4), 1977-2011.

Herskovic, B. 2018. "Networks in Production: Asset Pricing Implications." Journal of Finance 73(4), 1785-1818.

Holme, P., and Saramaki, J. 2012. "Temporal Networks." Physics Reports 519(3), 97-125.

Methodology

Read the methodology for this research.

Cite this research

Altsets Research. "Look-Ahead Bias in Equity Backtests Using Current Supply Chain Data." Published September 28, 2026. https://www.altsets.com/research/look-ahead-bias-equity-backtests-current-supply-chain-data