Effective Sample Size in Equity Tests with Supply Chain Dependence
Abstract
Cross-sectional finance often treats a sample of firms as containing approximately independent observations once market, style, and industry effects have been removed. That interpretation is fragile when residual firm outcomes remain correlated through commercial relationships. This paper derives a network-adjusted effective sample size that converts residual covariance among economically connected firms into a design effect, making the loss of independent statistical information directly comparable with conventional clustered inference. Under homoskedastic residuals, the adjustment is , which reduces to the familiar cluster design effect under equicorrelation and to for a sparse network with average correlated degree . Historical Altsets relationships are used to construct binary, relationship-size, supplier-revenue, and customer-cost representations of economic dependence for six directional links centered on Samsung Electronics and four semiconductor firms. The resulting network illustrates why industry clustering and binary supply-chain clustering need not approximate the same covariance structure. In the September 2026 snapshot, five of six focal-firm pairs share Samsung in the same economic direction, yet a directional percentage-weighted covariance kernel differs by more than three orders of magnitude between the strongest and weakest connected pairs. A calibration holding the strongest linked-pair residual correlation at 5 percent produces effective sample sizes of 3.48 under industry clustering, 3.56 under binary network clustering, 3.70 under relationship-size weighting, and 3.86 under directional dependency weighting for a nominal four-firm cross-section. The general result is that sample size is a property of the residual covariance graph as well as the count of securities.
Keywords: effective sample size; cross-sectional dependence; supply chain data; residual covariance; network econometrics; clustered inference; quantitative finance
Problem
Cross-sectional asset-pricing and corporate-finance tests commonly obtain precision from the number of firms in the sample. A regression with 500 firms may contain extensive controls for market beta, industry membership, size, value, momentum, profitability, investment, or volatility and still be interpreted as drawing information from roughly 500 separate economic units. The statistical premise behind that interpretation concerns the disturbances rather than the regressors. If residuals remain positively correlated across firms, additional securities contribute less independent information than their count suggests. Moulton's grouped-data result established the general econometric danger: disturbances shared within groups can cause conventional standard errors to be severely understated even when the number of micro observations is large. Conley's treatment of cross-sectional dependence subsequently generalized the problem by allowing covariance to decline with an economically meaningful distance rather than requiring observations to occupy discrete clusters.
A production network gives economic distance a concrete interpretation. Customer and supplier links can transmit demand, input, financing, and information shocks between firms that have already been stripped of broad market and industry components. Cohen and Frazzini document return predictability across economically linked customer-supplier firms, while Barrot and Sauvagnat show that idiosyncratic supplier shocks associated with natural disasters generate output and market-value effects at customer firms, particularly when inputs are difficult to substitute. Carvalho, Nirei, Saito, and Tahbaz-Salehi identify propagation both upstream and downstream after the Great East Japan Earthquake, including effects beyond immediately connected firms. These results provide an economic basis for residual dependence that follows commercial topology rather than conventional sector labels alone.
The statistical question is narrower than whether production links affect returns. Suppose the outcome of interest is a cross-sectional return, earnings surprise, forecast error, valuation residual, or estimated alpha, and conventional factors have already been removed. How much independent information remains if the remaining disturbances are correlated according to the supplier-customer graph? The relevant quantity is not the number of network edges or the average centrality of the firms. It is the contribution of network-conditioned covariance to the sampling variance of the statistic being estimated. Acemoglu, Carvalho, Ozdaglar, and Tahbaz-Salehi make a related aggregation point at the macroeconomic level: the rate at which idiosyncratic shocks diversify depends on network structure, and network sparsity by itself is insufficient to characterize aggregate volatility. The same logic can be applied to statistical information. Five hundred firms connected by weak, economically irrelevant links can still provide nearly 500 independent observations, while a considerably sparser graph carrying strong residual covariance can generate a much smaller effective sample.
Dependence
Let denote the excess return of firm on trading day or week . A practical first stage removes conventional systematic variation:
where contains market and standard style factors and is a leave-one-out return for firm 's industry. The leave-one-out construction prevents the firm's own return from mechanically entering its industry control. A Fama-French specification is a natural baseline because size, value, profitability, and investment factors are explicitly designed to absorb broad cross-sectional return patterns, while the industry term removes a separate source of within-sector covariance. The object retained for the network test is , so the hypothesis concerns dependence that survives these conventional controls.
For quarter , define the residual correlation of firms and over an evaluation window as
The primary hypothesis is that residual correlation remains positive for economically connected firms and declines as network distance increases, conditional on industry membership. The stronger version predicts heterogeneity within a fixed network distance: two firms connected through a common economically important customer or supplier should have greater residual covariance than two firms connected through a weak relationship. Evidence consistent with this distinction already exists outside the effective-sample-size setting. Federal Reserve Bank of Boston research on monetary-policy transmission reports large production-network effects on stock prices even with industry-demeaned returns and controls for common shocks, while Kim and Liu document increased customer-supplier stock-return comovement associated with firm-specific information and cash-flow news. The effective- problem asks what such residual dependence does to statistical precision.
The covariance model can be estimated in nested forms. An industry model assigns dependence according to . A binary network model replaces or augments that structure with direct and short-path indicators such as and , where is shortest directed or economically admissible network distance. An economic model uses continuous relationship weights. Comparing these specifications is essential because an industry cluster imposes equal dependence on every pair inside a category, while a binary network imposes equal dependence on every connected pair. Neither restriction follows automatically from shock propagation. Conley's concept of economic distance is especially relevant here: covariance can be structured by a measured distance without assuming that every observation inside an arbitrary group is equally related.
Data
The application uses six historical Altsets relationships centered on Samsung Electronics. The focal firms are Intel, AMD, Qualcomm, and Broadcom. Three relationships run from Samsung as supplier to a focal customer, and three run from a focal supplier to Samsung as customer. This produces two economically distinct sources of common exposure within the same small graph. A Samsung supply shock can jointly affect firms purchasing from Samsung, while a Samsung demand shock can jointly affect firms selling to Samsung. Intel and AMD appear on both sides, which permits the covariance structure to contain more than a single star-shaped channel.
| Directed relationship | First selected snapshot | September 2026 relationship size, USD m | Supplier revenue % | Customer cost % |
|---|---|---|---|---|
| Samsung to Intel | 2023-06 | 896.3 | 0.35 | 1.98 |
| Samsung to Qualcomm | 2025-06 | 1,378.0 | 0.55 | 6.94 |
| Samsung to AMD | 2026-06 | 2,400.0 | 0.96 | 11.82 |
| Intel to Samsung | 2023-06 | 141.7 | 0.22 | 0.07 |
| AMD to Samsung | 2025-06 | 482.6 | 1.27 | 0.28 |
| Broadcom to Samsung | 2023-06 | 269.9 | 0.31 | 0.14 |
The historical movement matters for covariance construction because a fixed adjacency matrix would treat every active link as economically unchanged. Samsung's relationship size with Intel was about $853.8 million in June 2023, reached about $1.47 billion in June 2025, and was about $896.3 million in the September 2026 workbook snapshot. Intel's customer-cost exposure associated with that supplier relationship moved from 2.15 percent to 2.79 percent and then to 1.98 percent over the same selected observations. The Samsung-to-Qualcomm relationship declined from approximately $1.83 billion and a 9.65 percent customer-cost share in June 2025 to about $1.38 billion and 6.94 percent by September 2026. On the other side of the graph, Broadcom's relationship with Samsung changed from approximately $597.5 million in June 2023 to $269.9 million in September 2026, while its supplier-revenue exposure declined from 1.45 percent to 0.31 percent. A binary link indicator is constant across movements of this kind once the edge is active.
The three Altsets metrics enter the covariance model according to their economic meaning. For a supplier serving two customers and , customer-cost percentages and measure the respective buyer-side dependency. Their common-supplier kernel is
For two suppliers and selling to common customer , supplier-revenue percentages and measure seller-side dependence on the shared customer:
Percentages are expressed as fractions in both equations. The directional exposure kernel is . This construction gives the product an economic interpretation: if a counterparty-specific shock creates return sensitivity proportional to the relevant exposure share, covariance induced by the common shock is proportional to the product of the two loadings.
Relationship size supplies a separate weighting scheme. For each common counterparty and direction, relationship dollars are normalized across the relevant edges,
with orientation changed appropriately for common suppliers. The size kernel sums across common counterparties. Normalizing locally prevents nominal firm scale from assigning near-unit covariance weight merely because a large company conducts large dollar transactions. The size kernel answers a different question from the percentage kernel. Dollar weighting identifies which pairs occupy large portions of the observed commercial flow around a shared node, while supplier-revenue and customer-cost shares identify how concentrated each participant is on the relationship. Keeping these objects separate allows the return data to determine which economic channel maps more closely into residual covariance.
Effective N
Consider the equal-weight cross-sectional mean of residual outcomes,
With firm-specific residual variances and pairwise residual correlations ,
Define as the number of independent observations with average residual variance that would generate the same variance for the sample mean. The resulting effective sample size is
Under homoskedasticity this simplifies to
The denominator is the network design effect,
This expression provides the link between conventional clustered inference and network dependence. If observations are divided into equal clusters of size , every within-cluster pair has correlation , and every cross-cluster pair has zero correlation, then
which is the familiar equicorrelated cluster design effect. The network version therefore generalizes clustered effective sample size rather than replacing it with an unrelated statistic. If each firm instead has an average of economically correlated neighbors and every relevant pair has residual correlation , the number of correlated pairs is , giving
That result produces an immediate scale calculation. For a nominal sample of 500 firms with ten residual-correlated network neighbors per firm, a 1 percent residual correlation corresponds to ; a 2 percent correlation gives approximately 417; and a 5 percent correlation gives approximately 333. With twenty correlated neighbors, the same three correlations imply effective samples of approximately 417, 357, and 250. A network can therefore halve the independent information in a 500-stock cross-section without requiring high pairwise correlation. Repeated small correlations accumulate because sampling variance depends on the sum of covariance terms.
The same principle extends to a regression coefficient. If is the cross-sectional design matrix and is the network-conditioned residual covariance matrix,
For coefficient , a parameter-specific effective sample size can be defined as
A single dataset can consequently have different effective sample sizes for different coefficients. A characteristic concentrated inside a commercially connected subgraph can suffer a larger precision loss than a characteristic spread across weakly connected firms, even though both regressions contain the same nominal . This distinction is particularly relevant for characteristic-sorted cross sections, event samples, and alternative-data signals whose exposures may themselves cluster in production networks.
Calibration
The four-firm Altsets subgraph provides a compact comparison among clustering assumptions. There are six possible focal-firm pairs. In the September 2026 snapshot, five pairs share Samsung in the same economic direction through at least one of the selected relationships; Qualcomm and Broadcom are the lone pair without such a shared directional channel. A broad semiconductor industry cluster would place all six pairs together. The binary network model therefore changes the covariance support only modestly relative to industry clustering. Economic weights change it substantially.
Using the directional percentage kernel, the strongest current pair is AMD-Qualcomm, with . Intel-AMD is next at 0.00236830, followed by Intel-Qualcomm at 0.00137412. AMD-Broadcom is 0.00003937 and Intel-Broadcom is 0.00000682. The strongest and weakest positively connected pairs differ by a factor of about 1,203 even though a binary two-step network indicator assigns both a value of one. The relationship-size kernel is less dispersed but still heterogeneous: the largest pair weight is approximately 3.85 times the smallest positive pair weight. This dispersion is generated directly by the observed relationship sizes and dependency percentages rather than by a statistical transformation of returns.
A useful structural calibration holds the maximum residual correlation among connected pairs fixed and changes only the assumed covariance geometry. Let denote the residual correlation assigned to the strongest connected pair. Industry clustering assigns to all six pairs. Binary network clustering assigns it to the five shared-direction pairs. The size and directional models assign . Under this normalization, the current relationship-size weights sum to 3.276 strongest-pair equivalents, while the directional percentage weights sum to 1.462. The resulting effective sample sizes are:
| Covariance structure | , | , | , |
|---|---|---|---|
| Naive independence | 4.000 | 4.000 | 4.000 |
| Industry cluster | 3.774 | 3.478 | 3.077 |
| Binary network | 3.810 | 3.556 | 3.200 |
| Relationship size | 3.873 | 3.697 | 3.437 |
| Directional percentages | 3.942 | 3.859 | 3.728 |
These figures are calibrated consequences of the observed graph and weights, rather than estimated residual correlations. Holding fixes the strongest pair across specifications and isolates the cost of treating weaker pairs as equally dependent. The industry model removes about 13 percent of the nominal information in the four-stock sample, while the directional exposure model removes about 3.5 percent under the same strongest-pair correlation. The gap arises because industry clustering spreads maximum covariance across six pairs, binary clustering across five, and the economic kernel concentrates most covariance on a small subset of economically stronger connections.
The historical network also changes the sufficient statistic entering the design effect. In June 2023, only one of the six focal pairs in this selected subgraph had a same-direction Samsung connection using the selected relationships. By June 2026, five pairs did. A binary approach therefore produces a stepwise increase in the network design effect. Economic weighting produces a different path. At a fixed 5 percent maximum-pair correlation, the September 2026 binary structure gives , the size structure gives 3.70, and the directional percentage structure gives 3.86. Earlier snapshots with a single active shared-direction pair produce approximately 3.90 under all three normalizations because there is no cross-pair heterogeneity to distinguish. The divergence between estimators appears precisely when network density expands and relationship intensities become unequal.
Tests
The empirical implementation should estimate rather than assume the mapping from network structure into residual covariance. For each quarter , factor and leave-one-out industry regressions generate firm residuals. Pairwise residual correlations for the subsequent return window can then be related to lagged network structure. A parsimonious specification is
The distance and weighted terms can also be estimated in separate nested specifications to reduce collinearity. The economically relevant predictions are when dependence decays with path length and when stronger directional dependence produces greater residual comovement. Theory supplies no requirement that the relationship-size coefficient dominate the percentage coefficient. Absolute transaction scale and proportional dependency describe different propagation mechanisms, so their relative explanatory content is an empirical question.
The first failure test is network randomization. Rewiring customer-supplier edges while preserving firm degree and, where feasible, industry composition creates placebo graphs with similar density but different economic counterparties. If the actual network and rewired graphs produce similar out-of-sample residual-covariance predictions, the supply-chain interpretation loses force. The second test exploits edge direction. For firms sharing a supplier, customer-cost percentage is the natural buyer-side exposure; for firms sharing a customer, supplier-revenue percentage is the natural seller-side exposure. A direction-swapped placebo uses supplier-revenue share for the common-supplier channel and customer-cost share for the common-customer channel. If the swapped kernel predicts covariance as well as the economically aligned kernel, the result is more consistent with generic firm similarity or scale than with directional shock transmission.
Industry clustering remains a necessary benchmark rather than a straw comparison. Semiconductor firms can share input costs, technology cycles, capital-spending conditions, and macro sensitivities without any relationship-specific transmission. A serious test therefore asks whether the network variables predict covariance among firms already within the same industry and whether connected firms across industries exhibit dependence that sector clustering misses. Barrot and Sauvagnat's stronger propagation through specific inputs and Carvalho and coauthors' evidence of multi-step upstream and downstream transmission provide reasons to expect economically meaningful heterogeneity among links. A null incremental network coefficient after industry residualization would instead imply that conventional sector controls absorb the relevant common variation for this purpose.
Interpretation
Effective sample size changes the interpretation of statistical precision without requiring a new return anomaly. Suppose a factor test reports a cross-sectional coefficient from 500 firms and conventional calculations treat those firms as independent after residualization. If fitted network covariance implies an average correlated degree of ten and an average residual correlation of 2 percent across those effective neighbors, the design effect is 1.20 and the sample contains the information equivalent of about 417 independent firms. If the network is denser or residual dependence is stronger, the adjustment can become much larger. The same calculation also permits the opposite outcome. Weak estimated network correlations or a rapidly decaying distance effect would leave close to the nominal count, providing evidence that standard cross-sectional intuition is adequate for the specification being studied.
The distinction between industry and network adjustment is therefore empirical rather than semantic. An industry cluster is a block covariance model. A production network is a sparse, overlapping, weighted covariance model in which a firm can participate in several economically distinct dependence channels at once. The selected Altsets semiconductor subgraph illustrates the difference sharply: all four focal firms are natural candidates for a broad common industry control, yet the observed directional exposure kernel assigns very different dependence intensity to their six pairings. Treating AMD-Qualcomm and Intel-Broadcom as exchangeable connected pairs discards information contained in the supplier-revenue and customer-cost shares. Treating Qualcomm-Broadcom as equivalent to the other five pairs because of industry membership imposes covariance where the selected Samsung-centered network supplies no same-direction connection.
The broader statistical implication follows from the design-effect equation rather than from any one industry. Nominal measures the number of securities. Effective measures the amount of independent residual information relevant to a specified estimator. When commercial links create residual covariance after factor and industry adjustment, these quantities separate. The difference can be represented with a single network statistic once covariance has been estimated: the weighted sum of pairwise residual correlations entering . Binary path length determines where dependence is allowed, while relationship size and directional dependency metrics determine how much each path should contribute. That construction converts production-network structure from an explanatory narrative into an explicit component of statistical inference.
Conclusion
Cross-sectional sample size should be interpreted through the covariance structure of the observations rather than through the security count alone. The network design effect
provides a direct translation from economically structured residual dependence into effective sample size. It nests conventional equicorrelated clustering, yields the sparse-network approximation , and extends naturally to coefficient-specific regression variance. Its empirical content comes from estimating residual covariance after conventional market and industry controls and then asking whether network distance and economic relationship weights explain what remains.
The six-relationship Altsets application demonstrates why the economic weights can alter that calculation even when the network topology is fixed. By September 2026, five of six focal semiconductor pairs share Samsung in a common directional relationship within the selected subgraph. Binary clustering treats those five connections equally. The directional dependency kernel, built from customer-cost exposure for common suppliers and supplier-revenue exposure for common customers, places the strongest and weakest positive pair weights more than three orders of magnitude apart. Relationship-size weighting produces a different intermediate covariance geometry. Under a common strongest-pair correlation calibration, those choices generate materially different effective sample sizes despite identical nominal .
The proposed empirical test is falsifiable. If factor- and industry-residualized covariance carries no incremental association with actual network distance or economic edge weights, the estimated network design effect converges toward one and nominal sample size remains an adequate description of information. If connected residuals remain correlated and the dependence is concentrated in economically strong or short network paths, conventional precision can be overstated even in large cross sections. The relevant adjustment is then neither a blanket penalty for using related firms nor a generic supply-chain correction. It is the amount of sampling information lost to the particular residual covariance structure generated by the observed commercial network.
References
- Acemoglu, D., Carvalho, V. M., Ozdaglar, A., and Tahbaz-Salehi, A. (2012). "The Network Origins of Aggregate Fluctuations." Econometrica 80, 1977-2016.
- Barrot, J. N., and Sauvagnat, J. (2016). "Input Specificity and the Propagation of Idiosyncratic Shocks in Production Networks." Quarterly Journal of Economics 131, 1543-1592.
- Carvalho, V. M., Nirei, M., Saito, Y. U., and Tahbaz-Salehi, A. (2021). "Supply Chain Disruptions: Evidence from the Great East Japan Earthquake." Quarterly Journal of Economics 136, 1255-1321.
- Cohen, L., and Frazzini, A. (2008). "Economic Links and Predictable Returns." Journal of Finance 63, 1977-2011.
- Conley, T. G. (1999). "GMM Estimation with Cross Sectional Dependence." Journal of Econometrics 92, 1-45.
- Fama, E. F., and French, K. R. (2015). "A Five-Factor Asset Pricing Model." Journal of Financial Economics 116, 1-22.
- Kim, D., and Liu, Y. (2020). "Aggregation of Idiosyncratic Shocks in the Customer-Supplier Network." Working paper.
- Moulton, B. R. (1990). "An Illustration of a Pitfall in Estimating the Effects of Aggregate Variables on Micro Units." Review of Economics and Statistics 72, 334-338.
Related research
Methodology
Cite this research
Altsets Research. "Effective Sample Size in Equity Tests with Supply Chain Dependence." Published September 27, 2026. https://www.altsets.com/research/effective-sample-size-equity-tests-supply-chain-dependence
