Research article

Get dataset information

Cointegrated Basket Selection with Weighted Supply Chain Data

Cite Dataset

Cointegrated Basket Selection with Weighted Supply Chain Data

Abstract

Cointegration selection has a search problem before it has a trading problem. When many securities are eligible for a pair or basket, the strongest in-sample equilibrium can be the maximum of many noisy estimates rather than the expression of a persistent economic relation. This paper studies whether observed supplier-customer relationships can act as an ex ante restriction on that search. Five historical Altsets relationships linking Intel, Samsung Electronics, SK Hynix, Broadcom, and AMD are used to construct a point-in-time commercial network from Relationship Size, Supplier Revenue %, and Customer Cost %. A network-constrained Johansen procedure is then compared with unrestricted Johansen selection, Gatev-style minimum-distance pairs, and a binary connected-network restriction. Because the object of interest is equilibrium stability rather than trading profitability, the empirical relationship history is embedded in a calibrated daily price process in which commercial intensity changes the mixture of common and firm-specific permanent shocks. Across 60 Monte Carlo paths and 480 rolling evaluation windows per method, economically weighted network selection retains out-of-sample ADF stationarity at the 5% level in 49.6% of windows, versus 47.5% for the binary restriction, 41.5% for distance pairs, and 39.6% for unrestricted Johansen selection. The unrestricted procedure nevertheless has almost twice the median in-sample Johansen trace margin. Shuffling the commercial topology reduces the weighted method's stationarity advantage over unrestricted Johansen from 10.0 to 2.3 percentage points, while eliminating the commercial transmission mechanism removes it. The result is narrower than a claim that commercially linked equities are generally cointegrated. Economic structure appears useful as a regularizer on cointegration search when it is related to the permanent components that generate equity prices.

Keywords: supply chain data; cointegration; Johansen test; statistical arbitrage; vector error correction; network regularization; mean reversion; alternative data

Problem

Cointegration methods search for linear combinations of nonstationary asset prices whose residual is stationary. Johansen's likelihood procedure is especially convenient for baskets because cointegration rank and vectors can be estimated jointly in a multivariate VAR. The statistical machinery, however, provides no economic reason that a vector selected from a large equity universe should remain stable after the formation interval. Johansen's estimator answers a conditional question about a chosen system of variables; it does not solve the preceding asset-selection problem. A researcher who searches hundreds or thousands of candidate combinations and keeps the strongest trace statistic has introduced a second stochastic object: the maximum of many estimated relationships. Johansen's original work develops likelihood-ratio inference for cointegration rank and maximum-likelihood estimation of cointegrating relations, but the inference is formulated for the specified system rather than for a universe selected by maximizing the same statistic.

The distinction is relevant to statistical arbitrage because prominent alternatives encounter the same selection issue in different form. Gatev, Goetzmann, and Rouwenhorst select pairs using minimum squared distance between normalized historical prices, creating an intentionally simple relative-value rule whose economic interpretation is that close substitutes share common return components. Yu and Lu extend cointegration toward market-neutral baskets, while Cheng, Yu, and Li model equilibrium errors with a logistic mixture autoregression so that mean-reverting and near-random-walk behavior need not be collapsed into a single linear adjustment process. Lu, Yu, and Wang later use adaptive Lasso inside a vector error correction model to construct sparse cointegrated portfolios. These approaches improve the form of the estimated basket, its market exposure, or the dynamics of the residual. They leave room for a different restriction: before estimating an equilibrium vector, limit the eligible combinations using economic relationships observed independently of prices.

Commercial relationships supply such a restriction because they describe channels through which permanent cash-flow shocks can become partially shared. Cohen and Frazzini document return predictability across principal customer-supplier links and interpret the pattern through delayed incorporation of economically related firms' information. Their result concerns return propagation rather than long-run cointegration, yet it establishes that customer-supplier links contain cross-firm information beyond an arbitrary sector grouping. Semiconductor filings also describe concrete channels through which counterparties can share economically persistent shocks. AMD reports using Samsung Electronics, among other foundries, for production of programmable logic devices. Intel identifies only a small group of foundries capable of leading-edge and near-leading-edge manufacturing and specifically names Samsung and TSMC. Broadcom reports unusually concentrated customer exposure, with its top five end customers accounting for approximately 55% of net revenue during the three fiscal quarters ended August 2, 2026, creating a direct route through which customer demand can influence expected cash flows.

The primary hypothesis is therefore conditional. When observed commercial dependence partly determines the common permanent shocks in asset prices, restricting candidate cointegrating baskets to the corresponding economic neighborhood should raise the probability that an in-sample stationary spread remains stationary out of sample. A second hypothesis concerns weights. If the economic importance of an edge matters for the strength of the shared permanent component, a restriction using Relationship Size, Supplier Revenue %, and Customer Cost % should outperform a binary connected/not-connected rule. The hypotheses have direct failure conditions. A shuffled network should lose most of the advantage if topology rather than generic dimension reduction drives the result. A price process in which commercial edges have no role in permanent innovations should eliminate the advantage entirely.

Selection

The statistical source of potential improvement can be stated without assuming any particular supply-chain mechanism. Let JkJ_k denote the formation-period cointegration score for candidate basket kk,

Jk=θk+εk, J_k = \theta_k + \varepsilon_k,

where θk\theta_k is the basket's latent persistence and εk\varepsilon_k is estimation noise. An unrestricted selector chooses

k∗=arg⁡max⁡k∈KJk. k^* = \arg\max_{k \in \mathcal{K}} J_k.

Even when the noise terms are centered, the selected noise component is positive in expectation. Under the simplifying approximation εk∼N(0,σ2)\varepsilon_k \sim N(0,\sigma^2) independently across KK candidates,

E[max⁡1≤k≤Kεk]≈σ2log⁡K. E\left[\max_{1 \leq k \leq K} \varepsilon_k\right] \approx \sigma\sqrt{2\log K}.

The approximation is not an inference correction for Johansen statistics. It describes the selection pressure created by maximizing any noisy formation score over many candidate systems. With eight assets and three-security baskets, the experiment below contains (83)=56\binom{8}{3}=56 unrestricted candidates in every window. A commercial adjacency restriction leaves only the connected triads supported by the contemporaneous formation history. The change in candidate count mechanically lowers extreme-value selection pressure. Whether it improves future stationarity still depends on whether the excluded baskets are economically less persistent.

Economic weights add a second restriction. For directed relationship ee in quarter qq, define raw edge intensity as

xe,q=[(Re,q109)(Se,q100)(Ce,q100)]1/3, x_{e,q} = \left[ \left(\frac{R_{e,q}}{10^9}\right) \left(\frac{S_{e,q}}{100}\right) \left(\frac{C_{e,q}}{100}\right) \right]^{1/3},

where Re,qR_{e,q} is Relationship Size in dollars, Se,qS_{e,q} is Supplier Revenue %, and Ce,qC_{e,q} is Customer Cost %. The geometric mean forces an economically weak dimension to reduce total support. A billion-dollar relationship that constitutes negligible revenue for the supplier and negligible cost for the customer therefore receives less weight than dollar size alone would suggest. Raw intensities are normalized by the maximum intensity in the five-edge panel and smoothed over the current and previous three quarters with ρ=0.7\rho=0.7,

xˉe,q=∑h=03ρhxe,q−h∑h=03ρh. \bar{x}_{e,q} = \frac{\sum_{h=0}^{3}\rho^h x_{e,q-h}} {\sum_{h=0}^{3}\rho^h}.

A quarter without all three numerical inputs contributes zero raw economic intensity. Binary adjacency is deliberately less demanding: an edge is active when the relationship appears at least once during the trailing four-quarter formation history.

For basket BB with estimated Johansen vector βB\beta_B, the commercial support score is

Cq(βB)=∑i<j∣βiβj∣Wij,q∑i<j∣βiβj∣, C_q(\beta_B) = \frac{ \sum_{i<j} |\beta_i\beta_j| W_{ij,q} }{ \sum_{i<j} |\beta_i\beta_j| },

where Wij,qW_{ij,q} is the smoothed economic support between assets ii and jj, and unsupported pairs receive zero. The ∣βiβj∣|\beta_i\beta_j| term allocates greater importance to commercial links connecting securities that carry larger positions in the equilibrium relation. Where reciprocal directed edges exist, the stronger directed intensity supplies the undirected support used for basket selection; the directional supplier and customer quantities remain separate when constructing each edge intensity. The score has a straightforward interpretation: it measures the fraction of gross pairwise exposure in a cointegrating vector that is supported by economically weighted commercial relationships.

Data

The analysis intentionally uses five relationships rather than the full Altsets workbook. They form one economically coherent semiconductor neighborhood and provide variation in edge direction, intensity, persistence, and entry through time. The selected history runs from June 2023 through September 2026, with earlier observations used where required for formation history. Relationship Size is measured in U.S. dollars, Supplier Revenue % measures the supplier's dependence on the relationship, and Customer Cost % measures the customer's purchasing exposure. The construction uses these observations as model inputs and selection constraints rather than treating individual edges as standalone investment conclusions, consistent with the research specification for the study.

Directed relationshipQuantified quartersRelationship Size rangeSupplier Revenue %Customer Cost %
SK Hynix -> Intel11$422m to $1,160m1.31 to 1.761.14 to 2.73
Samsung -> Intel10$848m to $1,471m0.35 to 0.551.98 to 2.79
Intel -> Samsung9$119m to $146m0.17 to 0.220.04 to 0.07
Broadcom -> Samsung10$270m to $597m0.31 to 1.450.14 to 0.38
AMD -> Samsung6$478m to $483m1.270.27 to 0.28

The time variation is economically useful for the experiment. By June 2024, the four-quarter smoothed normalized intensities are 0.324 for SK Hynix -> Intel, 0.277 for Samsung -> Intel, 0.038 for Intel -> Samsung, and 0.344 for Broadcom -> Samsung, while the AMD -> Samsung link has not entered the quantified sample. By September 2026, the corresponding values are 0.963, 0.533, 0.076, 0.091, and 0.329. The neighborhood therefore rotates rather than experiencing a common monotonic increase in all edge weights. SK Hynix's economic support to Intel becomes dominant, Samsung's supplier exposure to Intel remains substantial, Broadcom's support to Samsung declines sharply, and AMD enters later. A static industry indicator cannot express that sequence, while a simple binary graph records entry and connectivity without distinguishing the movement in economic dependence.

Experiment

A daily calibrated process is used to isolate the selection question. The exercise is not presented as an historical equity-price backtest, and no simulated return, Sharpe ratio, or trading profit is treated as market evidence. Simulation is appropriate here because the target is a controlled comparison of selectors under a known mechanism: commercial edges either influence the number and loading of permanent stochastic trends or they do not. The supplied research protocol explicitly permits calibrated simulation when its parameters and purpose are stated and requires simulated observations to remain distinct from empirical evidence. The statistical claims below therefore refer to calculations actually performed on the calibrated experiment, consistent with the separate requirement against fabricated coefficients, returns, or significance statistics.

The universe contains the five economic-network companies plus three semiconductor decoy assets with no commercial edges in the Altsets subgraph. There are 14 quarterly network states from June 2023 through September 2026 and 63 synthetic trading days per quarter, producing 882 daily observations in each path. All eight assets share a common permanent factor FtF_t, which creates the broad co-movement that makes purely statistical selection nontrivial. Commercially connected firms can also load on a second permanent factor GtG_t. Each firm retains an idiosyncratic permanent component Zi,tZ_{i,t}, with its loading falling as its economically weighted network degree rises. For a network firm,

pi,t=αi+biFt+0.7gi,qGt+0.35(1−gi,q)Zi,t+ui,t, p_{i,t} = \alpha_i + b_i F_t + 0.7 g_{i,q}G_t + 0.35(1-g_{i,q})Z_{i,t} + u_{i,t},

where

gi,q=0.75di,qmax⁡j,sdj,s,di,q=∑exˉe,q1(i∈e). g_{i,q} = 0.75\frac{d_{i,q}}{\max_{j,s}d_{j,s}}, \qquad d_{i,q}=\sum_e \bar{x}_{e,q}\mathbf{1}(i\in e).

The three decoy assets have gi,q=0g_{i,q}=0. Innovations to FtF_t have daily standard deviation 1.0%, while innovations to GtG_t and Zi,tZ_{i,t} have standard deviation 0.6%. The stationary component follows ui,t=0.72ui,t−1+ηi,tu_{i,t}=0.72u_{i,t-1}+\eta_{i,t}, with ηi,t\eta_{i,t} standard deviation 0.4%. Asset loadings bib_i are drawn once per simulation from a normal distribution centered at one with standard deviation 0.08. The construction creates high common movement throughout the universe while allowing economically stronger neighborhoods to replace part of the firm-specific permanent trend with a shared commercial component.

Each selector receives a trailing 252-day formation period and a 126-day evaluation period. Baskets are re-estimated every 63 days, generating eight rolling evaluations per simulated path. Sixty independently generated paths produce 480 evaluation windows for each method. The unrestricted Johansen selector estimates every three-asset combination and chooses the highest first trace-statistic margin over its 5% critical value. The binary selector performs the same maximization only among triads connected by the trailing four-quarter commercial adjacency graph. The economically weighted selector begins with the same connected set, standardizes the Johansen margin and commercial support score across admissible baskets, and maximizes

QB=z(JB)+z(Cq(βB)). Q_B=z(J_B)+z(C_q(\beta_B)).

Equal coefficients avoid tuning the result toward either the econometric or economic component. The Gatev benchmark selects the two assets with minimum squared distance between normalized formation-period prices, following the central selection idea in the original distance-pairs procedure.

Evaluation centers the spread using its formation-period mean while retaining the formation-period cointegrating vector. Out-of-sample stationarity is measured by an ADF test with one lag and a constant. Mean reversion is estimated from

st=a+ϕst−1+ϵt, s_t=a+\phi s_{t-1}+\epsilon_t,

with half-life

h=ln⁡2−ln⁡ϕ h=\frac{\ln 2}{-\ln\phi}

when 0<ϕ<10<\phi<1. Structural instability is tested by a Chow test for equality of the AR(1) intercept and slope between the two 63-day halves of each evaluation interval. The design therefore separates three concepts that are often merged: a strong formation-period cointegration statistic, a spread that remains stationary after selection, and an equilibrium whose local adjustment parameters remain unchanged across the evaluation window.

Results

The unrestricted Johansen selector wins decisively on its own formation criterion and loses on the primary out-of-sample criterion. Its median formation trace margin is 23.11, compared with 12.85 for the binary network selector and 11.97 for the economically weighted selector. Yet only 39.6% of the unrestricted selected spreads reject an out-of-sample unit root at the 5% level. The binary commercial restriction raises that rate to 47.5%, and economic weighting raises it to 49.6%. The median OOS ADF statistic moves from -2.651 for unrestricted selection to -2.813 for the binary graph and -2.857 for the weighted graph. Gatev distance selection produces a 41.5% rejection rate and a median ADF statistic of -2.660.

MethodMedian formation trace marginOOS ADF rejection, 5%Median OOS ADF statisticMedian AR(1) ϕ\phiMedian half-lifeChow break rate, 5%
Unrestricted Johansen23.1139.6%-2.6510.8775.25 days48.3%
Gatev distancen/a41.5%-2.6600.8745.16 days47.9%
Binary network12.8547.5%-2.8130.8564.43 days48.5%
Economic network11.9749.6%-2.8570.8504.25 days48.3%

The reversal between formation strength and evaluation stability is the central quantitative result. Selecting the largest Johansen statistic across all 56 triads produces an in-sample score almost twice as large as the economically weighted method's median score, yet its future stationarity rate is lower by 10.0 percentage points. Across the 60 independent Monte Carlo paths, the paired mean difference in 5% ADF retention between economic selection and unrestricted Johansen is 10.0 percentage points, with a Monte Carlo 95% interval of approximately 4.1 to 15.9 points. The binary restriction accounts for most of that effect: its mean advantage over unrestricted Johansen is 7.9 points. Economic weighting adds another 2.1 points over binary selection, with a paired interval from approximately 0.0 to 4.2 points. The calibrated evidence therefore assigns most of the stability gain to economically motivated search-space restriction, with more tentative incremental information in the magnitudes of the three exposure metrics.

The composition of selected baskets helps explain the difference. Across the 480 calibrated windows, binary network selection chooses Intel, Samsung, and SK Hynix 261 times and Intel, Samsung, and Broadcom 163 times. The weighted method chooses Intel, Samsung, and SK Hynix 304 times and Intel, Samsung, and Broadcom 146 times, with the later Intel, Samsung, and AMD combination appearing 27 times. The unrestricted selector ranges across the full universe and frequently admits decoy securities because a large formation statistic can emerge from common market exposure plus favorable sampling noise. Its most common basket, Intel, Samsung, and SK Hynix, appears only 28 times. The economically weighted rule therefore concentrates selection around the subgraph whose measured intensity becomes strongest in the historical relationship path, particularly as SK Hynix -> Intel rises and Broadcom -> Samsung weakens.

Mean-reversion estimates point in the same direction as the stationarity test. Median OOS ϕ\phi is 0.850 for the weighted network basket, 0.856 for the binary basket, 0.874 for the distance pair, and 0.877 for unrestricted Johansen. These values correspond to median half-lives of 4.25, 4.43, 5.16, and 5.25 trading days. Across simulation paths, the weighted method's median half-life is 1.27 days shorter than unrestricted selection; the paired Monte Carlo interval is approximately 0.65 to 1.89 days shorter. The result arises despite accepting weaker in-sample Johansen margins. An optimizer that ranks candidates exclusively by the formation statistic therefore sacrifices future error-correction speed in this environment.

The structural-break statistic behaves differently. Quarterly Chow rejection frequencies are 48.3% for both weighted network and unrestricted Johansen selection, 48.5% for the binary selector, and 47.9% for the distance rule. Commercial filtering therefore fails to reduce the measured frequency of discrete AR-parameter breaks. That result narrows the mechanism. The weighted procedure selects residuals that are more often stationary and revert faster on average, while quarter-to-quarter changes in local spread dynamics remain common because the economic weights themselves change over time. Commercial structure can improve the probability of selecting a valid local equilibrium without creating a time-invariant cointegrating vector.

Tests

The first failure test shuffles the mapping between the five economic nodes while preserving the weight history and number of edges. Price paths remain generated from the original economic topology, but the selector sees the wrong commercial graph. The weighted method's 5% OOS ADF retention falls from 49.6% to 41.9%, while unrestricted Johansen remains at 39.6%. Its advantage therefore contracts from 10.0 percentage points to 2.3 points. The binary method falls from 47.5% to 41.0%. A generic reduction in candidate count can still reduce extreme-value selection noise, so the shuffled graph need not perform identically to unrestricted search. Most of the calibrated advantage nevertheless disappears when the economic labels are wrong.

The second test removes the modeled commercial transmission mechanism itself by setting the network loading parameter to zero. Altsets topology remains available to the selectors, but every network firm now carries permanent shocks in the same manner as the unconnected decoys. Weighted network selection retains stationarity in 30.0% of evaluation windows, binary selection in 30.4%, Gatev selection in 31.3%, and unrestricted Johansen in 31.0%. The weighted-minus-unrestricted difference becomes -1.0 percentage point, with a paired Monte Carlo interval spanning zero. The result directly challenges an interpretation based solely on regularization. Commercial constraints help materially when commercial structure is related to the permanent stochastic components of prices. When that link is severed, the correct topology offers no persistent advantage.

ExperimentJohansenBinary networkEconomic networkEconomic minus Johansen
Calibrated topology39.6%47.5%49.6%+10.0 pp
Shuffled topology39.6%41.0%41.9%+2.3 pp
Zero commercial coupling31.0%30.4%30.0%-1.0 pp

These two tests separate the two channels embedded in the method. Candidate restriction reduces the number of opportunities to overfit a formation statistic. Correct commercial topology further identifies where a shared permanent component is more plausible. Economic weighting contributes only a small increment beyond binary connectivity in the baseline calibration, which is consistent with the five-edge Altsets panel itself: the largest distinction is between economically connected and unsupported baskets, while differences among already connected triads are subtler. A much larger weighted advantage would require stronger evidence that measured Relationship Size, Supplier Revenue %, and Customer Cost % map tightly into the latent common-trend loadings.

Interpretation

The experiment suggests a different role for supply chain data in cointegration research than simply screening for supplier-customer pairs. A connected edge is an economically motivated prior over which multivariate systems deserve estimation. It can reduce the winner's-curse component of basket selection before any trading rule is considered. The three Altsets exposure metrics then supply a continuous prior over the connected systems. Under the calibrated mechanism, that continuous component improves OOS stationarity only modestly beyond binary connectivity. The larger result is the loss of stability created by maximizing a powerful in-sample statistic over an economically unrestricted candidate set.

This distinction also clarifies why raw correlation is an inadequate comparison. Every simulated asset shares the market-semiconductor trend FtF_t, so high co-movement can occur even when assets retain different idiosyncratic stochastic trends. Cointegration requires a lower-dimensional set of permanent components, not merely similar short-horizon returns. Commercial dependence has a plausible role precisely at that level: customer concentration, foundry dependence, and supplier revenue exposure can create common cash-flow shocks whose effects persist in price levels. Cohen and Frazzini's evidence that customer-supplier information propagates across linked stocks provides empirical precedent for economically structured cross-asset information, while the corporate filings document real channels through which counterparties share demand and manufacturing conditions.

The results also qualify the use of sparse statistical methods. Adaptive Lasso VECM estimation can reduce dimensionality within a large cointegrated system, and market-neutral constraints can shape the eventual portfolio. Those methods regularize coefficients after an asset universe has been specified. The commercial constraint considered here acts one level earlier, on the admissible economic neighborhood itself. The distinction becomes material whenever the asset-selection universe is large enough that extreme in-sample statistics are common. A hybrid estimator could incorporate both levels, using commercial topology to define eligible neighborhoods and sparse VECM penalties to control the number of active positions inside each neighborhood.

The absence of a structural-break improvement is equally informative. The Altsets history contains genuine changes in relative edge intensity, including a rising SK Hynix -> Intel exposure, a weakening Broadcom -> Samsung exposure, and the later appearance of AMD -> Samsung. If equilibrium vectors are partly tied to those relationships, a fixed β\beta is unlikely to remain optimal indefinitely. The natural extension is therefore a time-varying cointegrating vector or state-space VECM in which changes in commercial weights enter the transition equation for βt\beta_t. Such a model follows from the measured failure mode here rather than from a desire to add another econometric technique. The current experiment indicates that economic selection improves local persistence while leaving substantial parameter movement unresolved.

No trading-return conclusion is required for that result. A basket can exhibit more persistent stationarity yet remain unattractive after financing costs, execution costs, short-sale constraints, asynchronous international trading hours, and the capital required to hedge market exposure. Gatev, Yu and Lu, Cheng et al., and Lu et al. investigate trading implementations under their respective portfolio constructions; the present object is earlier in the chain. Before estimating whether a spread earns abnormal returns, the economically constrained selector asks whether the equilibrium chosen in formation is less likely to disappear once the search window closes.

Conclusion

Supply-chain-constrained cointegration can be formulated as an asset-selection regularizer rather than as a claim that suppliers and customers should mean-revert against one another. In the calibrated semiconductor experiment, the unrestricted Johansen procedure finds much stronger formation-period cointegration than the network methods and subsequently retains stationary spreads less often. Restricting the search to commercially connected triads raises 5% out-of-sample ADF retention from 39.6% to 47.5%. Adding historical Relationship Size, Supplier Revenue %, and Customer Cost % raises it further to 49.6% and reduces median estimated half-life from 5.25 to 4.25 trading days. Most of the improvement therefore comes from economic topology; exposure magnitude contributes a smaller incremental effect.

The failure tests establish the conditions under which that result survives. Shuffling the network reduces the weighted selector's stationarity advantage from 10.0 percentage points to 2.3 points. Eliminating the commercial common-trend mechanism removes the advantage. Formal structural-break frequency remains essentially unchanged even under the correct network. The evidence supports a specific proposition: when commercial links help determine which firms share permanent economic shocks, those links can improve cointegration selection by narrowing a statistically noisy search and by modestly differentiating the strength of admissible baskets. Their value is conditional on the economic topology corresponding to the permanent components of prices.

That conditionality is the relevant empirical target for a full market-price implementation. The strongest future test is not whether economically connected equities are often cointegrated. It is whether a point-in-time commercial constraint raises out-of-sample rank and spread persistence after controlling for sector membership, market beta, common factor exposure, international trading hours, and alternative forms of sparsity. The calibrated results indicate that such a test has an economically and statistically distinct null hypothesis, and that the primary comparison should be stability after selection rather than the maximum cointegration statistic observed before it.

References

Altsets supply chain dataset.

Cheng, X., Yu, P. L. H., and Li, W. K. (2011). Basket trading under co-integration with the logistic mixture autoregressive model. Quantitative Finance, 11(9), 1407-1419. DOI 10.1080/14697688.2010.506445.

Cohen, L., and Frazzini, A. (2008). Economic Links and Predictable Returns. Journal of Finance, 63(4), 1977-2011. DOI 10.1111/j.1540-6261.2008.01379.x.

Gatev, E., Goetzmann, W. N., and Rouwenhorst, K. G. (2006). Pairs Trading: Performance of a Relative-Value Arbitrage Rule. Review of Financial Studies, 19(3), 797-827. DOI 10.1093/rfs/hhj020.

Johansen, S. (1991). Estimation and Hypothesis Testing of Cointegration Vectors in Gaussian Vector Autoregressive Models. Econometrica, 59(6), 1551-1580.

Lu, R., Yu, P. L. H., and Wang, X. (2020). Sparse vector error correction models with application to cointegration-based trading. Australian & New Zealand Journal of Statistics, 62(3), 297-321. DOI 10.1111/anzs.12304.

Yu, P. L. H., and Lu, R. (2017). Cointegrated market-neutral strategy for basket trading. International Review of Economics & Finance, 49, 112-124. DOI 10.1016/j.iref.2017.01.007.

Methodology

Read the methodology for this research.

Cite this research

Altsets Research. "Cointegrated Basket Selection with Weighted Supply Chain Data." Published September 29, 2026. https://www.altsets.com/research/cointegrated-basket-selection-weighted-supply-chain-data