Research article

Get dataset information

Covariance Kernels for Equity Pairs Using Supply Chain Exposure

Cite Dataset

Covariance Kernels for Equity Pairs Using Supply Chain Exposure

Abstract

Residual covariance estimation usually treats firm-specific returns remaining after market and industry effects are removed as weakly structured noise unless price history itself indicates otherwise. This paper asks whether local production-network motifs imply distinguishable residual covariance terms. A direct customer-supplier pair, two suppliers sharing a customer, two customers sharing a supplier, and the endpoints of a directed two-hop chain transmit different economic shocks even when all four structures are described as "connected." The paper derives a motif-conditioned covariance kernel in which supplier revenue share measures upstream sensitivity to customer demand, customer cost share measures downstream sensitivity to supplier shocks, and relationship size measures absolute commercial scale. Six historical Altsets relationships are used to calibrate four motif structures through quarterly snapshots. The calibration shows that binary motifs can remain unchanged while their economic loadings move substantially and, in one two-hop case, move in opposite directions across dependency dimensions. The resulting empirical specification compares motif pairs with matched unconnected pairs after market and industry residualization. Its strongest restriction is directional: shared-customer covariance should load primarily on supplier revenue dependence, while shared-supplier covariance should load primarily on customer cost dependence. A generic connected-pair covariance premium would fail that restriction.

Keywords: supply chain data; residual covariance; return comovement; production networks; covariance estimation; graph motifs; factor residuals; quantitative finance

Covariance

An equity covariance matrix contains structure that standard factor removal only partly addresses. Market and industry residualization removes two obvious common components, yet firms can retain correlated cash-flow innovations because they purchase from the same producer, sell into the same buyer, or occupy successive positions in a production chain. Treating every remaining covariance as idiosyncratic creates a measurement problem rather than merely an omitted-variable problem. A pair of semiconductor vendors selling into the same device manufacturer and a pair of manufacturers buying from the same materials supplier can exhibit positive residual comovement through entirely different shock channels. An estimator that assigns both pairs the same "second-degree connection" loses information about which side of the income statement transmits the common shock and about how strongly each firm depends on the common node.

Prior research establishes several parts of this mechanism separately. Cohen and Frazzini documented predictable returns across directly linked customers and suppliers, consistent with economically connected firms incorporating related information at different speeds. Barrot and Sauvagnat used natural disasters to identify supplier shocks and found output losses at customers as well as spillovers to other suppliers, providing direct evidence that production links can transmit firm-specific disturbances beyond a single bilateral pair. More recently, Qian, Zhang, and Zhao reported higher return comovement among Chinese listed firms sharing common suppliers, with stronger effects when purchasing proportions from the shared supplier were higher. Their result is especially relevant here because it connects a specific local motif to return comovement and identifies purchasing dependence as a conditioning variable rather than treating the common supplier as a binary label.

The unresolved issue is whether these structures belong in one generic connected-pair category. Antón and Polk provide a useful statistical precedent outside production networks: firms sharing active mutual fund owners exhibit excess return correlation even after controls for systematic factors, style, and sector similarity. Their result shows that economically meaningful connections can organize residual covariance after conventional exposures are removed. Production-network research supplies the corresponding economic foundation. Acemoglu, Carvalho, Ozdaglar, and Tahbaz-Salehi show that input-output structure changes how microeconomic shocks aggregate, while Herskovic develops asset-pricing implications from network concentration and sparsity. The object considered here is much more local. It asks whether a small set of firm-pair motifs produces different covariance signatures once broad systematic effects have already been stripped from returns.

The primary hypothesis is therefore stronger than a prediction that connected firms should comove. Residual covariance should depend on motif type and on the economically appropriate directional dependency measure inside that motif. For a shared customer, the common shock is principally a demand or purchasing shock originating at that customer and traveling upstream, so covariance should increase with the product of the two suppliers' revenue dependence on that customer. For a shared supplier, the common shock travels downstream, so covariance should increase with the product of the two customers' cost dependence on that supplier. A two-hop chain should contain attenuated products of directional exposures because a shock must pass through an intermediate firm. Direct links admit both channels and can therefore be strongly asymmetric. These restrictions give the model a failure condition: if all motif classes have similar residual covariance and their coefficients respond similarly to relationship size regardless of direction, topology is functioning mainly as a coarse indicator of economic proximity.

Motifs

Let an Altsets edge e=(s,c)e=(s,c) run from supplier ss to customer cc. For quarter tt, define Vsc,tV_{sc,t} as relationship size in dollars, Psc,tP_{sc,t} as the fraction of supplier revenue associated with the customer, and Qsc,tQ_{sc,t} as the fraction of customer cost associated with the supplier. Percentages are converted to decimal fractions in the equations. PP is naturally aligned with shocks traveling upstream from customer to supplier because it measures the supplier's exposure to the customer's purchasing activity. QQ is naturally aligned with shocks traveling downstream from supplier to customer because it measures the customer's dependence on the supplier in its cost structure. VV carries different information: two relationships can have the same percentage dependence while representing very different dollar magnitudes, so log⁡V\log V is retained as an absolute-scale term rather than being mechanically blended into either directional loading.

Consider residual firm returns after broad factors have been removed. A local shock representation for supplier ss can be written as

ϵs,t=ηs,t+aPsc,tDc,t+∑k:k→sbQks,tSk,t,\epsilon_{s,t} = \eta_{s,t} + a P_{sc,t}D_{c,t} + \sum_{k:k\to s} b Q_{ks,t}S_{k,t},

where Dc,tD_{c,t} is a customer-side demand shock, Sk,tS_{k,t} is a supply shock originating at upstream firm kk, and ηs,t\eta_{s,t} collects firm-specific innovations outside the modeled local network. Coefficients aa and bb translate operating exposure into equity-return sensitivity. The equation is deliberately sparse. Its purpose is to establish which observed relationship measure should multiply which latent shock. If suppliers ii and jj share customer cc, the customer component of their residual covariance is

Cov⁡(ϵi,ϵj∣c)=a2PicPjcσD,c2+Rij,\operatorname{Cov}(\epsilon_i,\epsilon_j\mid c) = a^2 P_{ic}P_{jc}\sigma^2_{D,c} + R_{ij},

where RijR_{ij} contains any remaining covariance. The economically relevant exposure is therefore the product PicPjcP_{ic}P_{jc}, rather than an indicator equal to one whenever a common customer exists.

For customers ii and jj sharing supplier ss, the analogous expression is

Cov⁡(ϵi,ϵj∣s)=b2QsiQsjσS,s2+Rij.\operatorname{Cov}(\epsilon_i,\epsilon_j\mid s) = b^2 Q_{si}Q_{sj}\sigma^2_{S,s} + R_{ij}.

This produces a sharp distinction between the two superficially similar second-degree motifs. Shared-customer covariance should be more sensitive to PicPjcP_{ic}P_{jc}, while shared-supplier covariance should be more sensitive to QsiQsjQ_{si}Q_{sj}. The opposite-side products remain useful controls. A common customer's cost exposure to two suppliers, for example, can capture the suppliers' importance to that customer, but the basic upstream cash-flow channel predicts that the suppliers' revenue dependence is the more direct loading. Qian, Zhang, and Zhao's finding that common-supplier comovement is stronger when purchasing proportions are greater is consistent with this orientation in their Chinese listed-firm sample.

A directed two-hop chain requires an additional propagation step. Suppose i→k→ji\to k\to j. If a demand shock originating at jj affects kk and then travels from kk to ii, the upstream component is proportional to

λDPikPkj,\lambda_D P_{ik}P_{kj},

where 0≤λD≤10\leq\lambda_D\leq1 represents attenuation at the intermediate firm. A supply shock originating at ii and traveling downstream to jj produces the corresponding term

λSQikQkj.\lambda_S Q_{ik}Q_{kj}.

This yields a different covariance object from either common-node motif. The two endpoint firms do not load directly on one common counterparty. Their covariance depends on sequential transmission and should therefore respond to the product of two directed edge weights. Under the propagation interpretation, two-hop coefficients should usually be smaller than otherwise comparable direct coefficients once the economic weights are held constant. A two-hop estimate as large as a direct-link estimate, with little dependence on path weights, would instead point toward an omitted common factor, coarse industry grouping, or another source of similarity.

Data

The calibration uses six Altsets relationships chosen to generate distinct local structures while retaining historical observations: Intel to Samsung Electronics and Broadcom to Samsung Electronics form a shared-customer motif; Teijin to Nike and Teijin to Caterpillar form a shared-supplier motif; Intel to Samsung Electronics followed by Samsung Electronics to Qualcomm forms a directed two-hop chain; and Boeing to Korean Air supplies a direct bilateral case. The histories are quarterly and use Altsets relationship size, supplier revenue percentage, and customer cost percentage. Individual relationships are inputs into the motif construction rather than conclusions about the companies. For a two-edge motif, a descriptive exposure scale is reported as the geometric mean of the two corresponding edge metrics. For a direct edge, the metric itself is used. Structural covariance equations still use the products of percentage loadings, because the product is the economically implied common-shock exposure.

MotifConstructionComplete overlapsGVG_V, first to latestGPG_P, first to latestGQG_Q, first to latest
Shared customerIntel -> Samsung, Broadcom -> Samsung6$266.7m to $195.6m0.496% to 0.261%0.123% to 0.099%
Shared supplierTeijin -> Nike, Teijin -> Caterpillar4$140.1m to $124.4m1.861% to 1.797%0.375% to 0.324%
Two-hopIntel -> Samsung -> Qualcomm5$466.5m to $441.9m0.367% to 0.348%0.621% to 0.697%
DirectBoeing -> Korean Air9$9.2m to $833.8m0.020% to 0.790%2.610% to 36.580%

The shared-customer example changes enough to make binary classification particularly restrictive. Between its first and latest overlapping complete snapshots, the geometric-mean relationship scale falls 26.7%, while the geometric-mean supplier revenue exposure falls 47.4%. Because covariance under the common-customer mechanism depends on the product of the revenue shares, PIntel,SamsungPBroadcom,SamsungP_{Intel,Samsung}P_{Broadcom,Samsung}, that product falls 72.3%. A binary shared-customer variable remains exactly one throughout. Any covariance model using only adjacency therefore assigns the same economic state to periods in which the modeled demand-shock loading differs by nearly a factor of four in squared exposure terms.

The shared-supplier example behaves differently. Teijin's two customer relationships produce a relatively stable geometric-mean supplier revenue share, falling only 3.5% over the overlapping sample, while the geometric-mean customer cost share falls 13.5%. The product QTeijin,NikeQTeijin,CaterpillarQ_{Teijin,Nike}Q_{Teijin,Caterpillar}, which is the structural loading for a common supply shock, falls 25.1%. This is a materially smaller shift than the 72.3% decline in the shared-customer demand loading. Binary representations classify both structures as persistent second-degree connections and therefore discard the difference between a rapidly weakening common-demand channel and a comparatively stable common-supply configuration.

The two-hop example supplies a stronger test of whether a single generic relationship weight is adequate. From the first overlapping observation to September 2026, the geometric-mean relationship size on Intel -> Samsung -> Qualcomm falls 5.3%, and the geometric-mean supplier revenue measure also falls 5.3%. The customer cost measure moves in the opposite direction, rising 12.2%. In product form, the upstream demand loading PIntel,SamsungPSamsung,QualcommP_{Intel,Samsung}P_{Samsung,Qualcomm} falls 10.3%, while the downstream supply loading QIntel,SamsungQSamsung,QualcommQ_{Intel,Samsung}Q_{Samsung,Qualcomm} rises 25.9%. A scalar "strength" score would have to decide which movement to preserve. Keeping the directional measures separate allows residual covariance to reveal whether the economically relevant channel is upstream demand propagation, downstream supply propagation, or neither.

The Boeing to Korean Air relationship illustrates a different source of information loss. By the latest snapshot, the relationship carries a supplier revenue share of 0.79% for Boeing and a customer cost share of 36.58% for Korean Air, a ratio of roughly 46 to one between the two directional dependency percentages. The edge is therefore highly asymmetric even before returns enter the analysis. Its history also changes sharply, with recorded relationship size rising from $9.24 million in the first complete observation to about $833.8 million in the latest one. Boeing announced on August 25, 2025 that Korean Air intended to purchase 103 Boeing aircraft, in addition to a March 2025 order for 40 widebody jets, providing external evidence of a major expansion in the bilateral commercial relationship during the period in which the Altsets edge becomes much larger. The calibration treats the announcement as economic context rather than equating announced order value with the Altsets relationship-size measure.

Residuals

The return test should isolate local commercial covariance from broad systematic comovement. Let ri,dr_{i,d} denote stock ii's daily excess return. Factor loadings for quarter qq are estimated using the previous trading year:

ri,d=αi,q+βi,qMrM,d+βi,qIrI(i),d+ϵi,d,r_{i,d} = \alpha_{i,q} + \beta^M_{i,q}r_{M,d} + \beta^I_{i,q}r_{I(i),d} + \epsilon_{i,d},

where rMr_M is the market excess return and rI(i)r_{I(i)} is the return of the firm's industry portfolio. The Kenneth French Data Library provides daily market-factor data and daily 48-industry portfolios; the 48-industry series currently extend through August 31, 2026 and assign U.S. listed stocks to portfolios using SIC classifications. Estimating betas before the covariance quarter separates factor estimation from the residual covariance being measured. The primary dependent variable for pair i,ji,j in quarter qq is realized residual covariance,

Cij,q=1Nq∑d∈qϵi,dϵj,d,C_{ij,q} = \frac{1}{N_q} \sum_{d\in q} \epsilon_{i,d}\epsilon_{j,d},

with residual correlation used as a scale-free companion statistic.

The binary specification establishes the conventional baseline:

Cij,q=αq+γDDij,q+γSCSCij,q+γSSSSij,q+γTHTHij,q+uij,q.C_{ij,q} = \alpha_q + \gamma_D D_{ij,q} + \gamma_{SC}SC_{ij,q} + \gamma_{SS}SS_{ij,q} + \gamma_{TH}TH_{ij,q} + u_{ij,q}.

The indicators represent direct, shared-customer, shared-supplier, and directed two-hop relationships. Motif classes should be assigned with an explicit precedence rule so a pair is not silently counted several times; direct links can be separated first, followed by shared-customer, shared-supplier, and two-hop categories, while overlapping motifs can be examined separately as a sensitivity exercise. The control group contains pairs with no direct edge, no common customer, no common supplier, and no directed two-hop path in the relevant Altsets snapshot. Each connected pair is matched to unconnected pairs with similar industry combination, market-cap range, beta, and residual volatility. The objective of matching is to prevent the motif coefficients from becoming indirect substitutes for obvious cross-sectional similarity.

The economically weighted model adds the directional exposure objects implied by the shock equations:

Cij,q= αq+∑mγmMij,qm+θDPDij,qPij,q+θDQDij,qQij,q+θSCPSCij,qPic,qPjc,q+θSCQSCij,qQic,qQjc,q+θSSPSSij,qPsi,qPsj,q+θSSQSSij,qQsi,qQsj,q+θTHPTHij,qPik,qPkj,q+θTHQTHij,qQik,qQkj,q+∑mϕmMij,qmlog⁡Gij,qV+uij,q.\begin{aligned} C_{ij,q} =&\ \alpha_q+\sum_m\gamma_m M^m_{ij,q} +\theta_D^P D_{ij,q}P_{ij,q} +\theta_D^Q D_{ij,q}Q_{ij,q} \\ &+\theta_{SC}^P SC_{ij,q}P_{ic,q}P_{jc,q} +\theta_{SC}^Q SC_{ij,q}Q_{ic,q}Q_{jc,q} \\ &+\theta_{SS}^P SS_{ij,q}P_{si,q}P_{sj,q} +\theta_{SS}^Q SS_{ij,q}Q_{si,q}Q_{sj,q} \\ &+\theta_{TH}^P TH_{ij,q}P_{ik,q}P_{kj,q} +\theta_{TH}^Q TH_{ij,q}Q_{ik,q}Q_{kj,q} \\ &+\sum_m\phi_m M^m_{ij,q}\log G^V_{ij,q} +u_{ij,q}. \end{aligned}

The principal hypothesis is expressed through coefficient restrictions rather than a generic positive network coefficient. For shared customers, θSCP>0\theta_{SC}^P>0 is the central prediction and should carry more explanatory content than the wrong-side cost term. For shared suppliers, θSSQ>0\theta_{SS}^Q>0 is the corresponding prediction. For two-hop chains, positive directional coefficients are possible in both directions, although their magnitude should reflect attenuation through the intermediate firm. The direct motif permits PP and QQ to enter separately because a relationship such as Boeing to Korean Air can place very different economic weight on the supplier and customer.

Comparison

Economic weighting earns its place only if it explains covariance variation that binary motifs leave unresolved. The clean comparison is therefore out of sample. Estimate the binary and weighted specifications through quarter qq, generate pairwise residual-covariance predictions for q+1q+1, and compare squared forecast error or Gaussian covariance loss across the two models. A lower in-sample residual sum of squares is insufficient because the weighted model contains additional regressors by construction. If directional Altsets weights improve next-quarter covariance estimation specifically within the motif for which the economic mechanism predicts relevance, the gain has an interpretable source. If relationship size alone produces the improvement while PP and QQ have interchangeable effects, the evidence instead favors general commercial salience.

This comparison connects the paper to a broader covariance-estimation problem without turning the study into portfolio optimization. Ledoit and Wolf emphasize that raw sample covariance matrices can contain enough estimation error to damage mean-variance applications and motivate shrinkage toward a more stable target. A motif-conditioned covariance estimate could eventually serve as one such structured target, although the pairwise hypothesis can be evaluated before constructing a full portfolio matrix. If a full matrix is built, fitted off-diagonal values must respect positive semidefiniteness, either through projection onto the positive semidefinite cone or through a shrinkage construction. The economic test comes first: local network structure must forecast residual covariance in the theoretically predicted form before the structure deserves influence over a portfolio covariance matrix.

The calibration already indicates why weighting can change that target. At the latest shared-customer observation, Intel and Broadcom's common Samsung exposure has a geometric-mean supplier revenue dependence of 0.261%. At the latest shared-supplier observation, Nike and Caterpillar's common Teijin exposure has a geometric-mean customer cost dependence of 0.324%. Those percentages appear superficially similar. Their histories are not. The shared-customer demand product has contracted by 72.3% across the overlapping observations, while the shared-supplier cost product has contracted 25.1%. The two-hop path moves differently again, with its downstream cost product rising 25.9% while its upstream revenue product falls 10.3%. A covariance target using only motif identities would preserve none of these state changes.

Tests

The first failure test is a directional-weight placebo. In the shared-customer regression, replace the predicted upstream loading PicPjcP_{ic}P_{jc} with the customer-cost product QicQjcQ_{ic}Q_{jc}; in the shared-supplier regression, replace the predicted downstream loading QsiQsjQ_{si}Q_{sj} with the supplier-revenue product PsiPsjP_{si}P_{sj}. The economic transmission hypothesis predicts that the correctly oriented measure should contain greater incremental covariance information. If both versions perform similarly after log⁡GV\log G_V, market, industry, and matching controls are included, the result would be consistent with the percentages acting as generic relationship-strength proxies. That outcome would weaken the interpretation that motif direction identifies a specific common-shock channel even if connected pairs still exhibit positive excess covariance.

The second failure test is topology-preserving economic randomization. Within each quarter and broad industry pairing, counterparty identities can be permuted while approximately preserving the number of edges and the distribution of relationship weights. The exercise asks whether actual shared customers, shared suppliers, and directed two-hop paths explain more residual covariance than synthetic motifs assembled from economically similar edges. The relevant comparison is not a completely random graph, which would make the placebo implausibly different from the observed sample. A constrained permutation retains broad sector composition and weight distributions while severing the actual commercial identity tying the two firms together. If the real motifs and constrained placebo motifs produce similar covariance coefficients, sectoral or size-related similarity remains a more plausible explanation than the local production structure.

A further interpretation follows from the relative magnitudes of the four motif coefficients. Shared-customer and shared-supplier structures need not rank consistently because the variance of customer demand shocks and supplier production shocks can differ across industries and time. The stronger prediction concerns the within-motif loading signature. Two-hop covariance should exhibit attenuation relative to direct connections with comparable directional exposure, while direct relationships can show pronounced asymmetry. The latest Boeing to Korean Air observation is an instructive calibration: the customer-cost exposure is about 46 times the supplier-revenue exposure. A model that compresses those values into one edge weight loses the distinction between Korean Air's dependence on Boeing-related costs and Boeing's revenue dependence on Korean Air.

Interpretation

Several possible empirical outcomes would be informative. A positive shared-customer coefficient accompanied by a positive θSCP\theta_{SC}^P, together with a shared-supplier coefficient whose variation is primarily explained by θSSQ\theta_{SS}^Q, would support motif-conditioned covariance rather than generic graph proximity. A weaker but still useful result would be positive motif effects with little incremental contribution from economic weights; local topology would then identify covariance clusters, while the three dependency metrics would add limited short-horizon information. If relationship size dominates both directional percentages, absolute commercial scale may be capturing information attention, disclosure intensity, or another scale-dependent mechanism rather than the cash-flow transmission specified here.

A broad null has a different interpretation. If connected motif classes are indistinguishable from matched unconnected pairs after market and industry residualization, local supply-chain topology contributes little to contemporaneous residual covariance at the chosen return horizon. That outcome would coexist with earlier evidence on return predictability, shock transmission, or long-horizon operating exposure because covariance is a narrower object. Cohen and Frazzini study delayed information incorporation across direct economic links, while Barrot and Sauvagnat identify real propagation following large supplier shocks. The absence of ordinary-period covariance would suggest that these channels are episodic, delayed, or too small relative to firm-specific return noise to create stable pairwise covariance.

An especially revealing failure would occur if shared-customer and shared-supplier motifs both show excess covariance yet respond to the same dependency measure. That pattern would indicate that topology identifies economically similar firms while the proposed upstream-versus-downstream distinction is too sharp. Conversely, a two-hop coefficient comparable to a direct coefficient after conditioning on path weights would challenge the attenuation mechanism. These outcomes are preferable to relabeling every connected pair as network risk. The research question concerns whether motif types impose distinct covariance structures, and the coefficient restrictions specify what "distinct" must mean empirically.

Conclusion

Local production relationships imply different common-shock geometries after broad market and industry effects have been removed. Two suppliers sharing a customer load jointly on that customer's demand state, and supplier revenue dependence is the natural directional weight. Two customers sharing a supplier load jointly on the supplier's production state, making customer cost dependence the corresponding weight. A directed two-hop chain compounds edge exposures and introduces attenuation through an intermediate firm. A direct relationship contains both directions and can be strongly asymmetric. These structures produce different restrictions on residual covariance even when every pair would receive the same generic label of economically connected.

The six Altsets histories used here show that those restrictions are empirically nontrivial before stock returns are considered. The shared-customer demand loading contracts sharply while the motif itself remains present. The shared-supplier loading is more stable. The two-hop path experiences declining revenue-side exposure alongside rising cost-side exposure. Boeing to Korean Air moves from a small recorded relationship state to one with a large customer-cost share and a much larger dollar relationship, consistent in timing with a major publicly announced aircraft commitment. Binary network variables erase all four forms of variation.

The decisive empirical evidence is therefore a motif-specific covariance signature. Shared-customer covariance should respond to revenue dependence on the common buyer. Shared-supplier covariance should respond to cost dependence on the common supplier. Two-hop covariance should depend on products of directed exposures and exhibit attenuation. Direct covariance should permit upstream and downstream weights to differ. Matched unconnected pairs establish the non-network baseline, while directional placebos and constrained network randomization challenge the economic interpretation. Under this design, supply chain data enter covariance estimation through economically restricted local structure rather than through generic graph proximity, and the hypothesis can fail in several clearly identifiable ways.

References

Altsets supply chain dataset.

Methodology

Read the methodology for this research.

Cite this research

Altsets Research. "Covariance Kernels for Equity Pairs Using Supply Chain Exposure." Published September 28, 2026. https://www.altsets.com/research/equity-pair-covariance-kernels-supply-chain-exposure