Do 1,000 Stocks Really Give a Supply-Chain Model 1,000 Independent Observations?

September 14, 2026

Altsets

Research by Altsets Research

Share

Connected companies can share customers, suppliers, industries, and shocks, so the row count in a cross-sectional network model can materially exceed the amount of independent statistical evidence.

Data used:Altsets Supply Chain Intelligence: 90k+ entities, 400k+ relationships, 20+ years of history.

Key findings

  • The supplied network contains repeated economic clusters such as Micron and SK Hynix around Nvidia and Shin-Etsu around several semiconductor manufacturers, so many stock rows can still reflect a smaller set of common shocks.
  • Network econometrics provides inference methods for observations connected through an observed graph, supporting network-aware or clustered treatment of standard errors rather than assuming every company observation is independent.

No. One thousand stocks can represent far fewer than one thousand independent observations when common customers, suppliers, industries, and bottlenecks make their returns and forecast errors move together. Effective sample size should reflect those clusters rather than raw row count.

The graph itself tells you why observations are dependent

The supplied Altsets network gives several concrete examples. Micron and SK Hynix both map to Nvidia as a customer, so the two suppliers share an obvious source of demand information. Shin-Etsu Chemical maps to Samsung Electronics, TSMC, and Intel, with displayed supplier revenue shares of 2.43%, 4.02%, and 1.79% respectively. Those are distinct customer relationships, but the customers also participate in the same semiconductor manufacturing cycle. A cross-sectional strategy that owns or studies all of these companies does not have a set of economically isolated bets.

This matters when a researcher reports significance from a factor regression or a long-short portfolio. If several stocks move because the same customer or industry shock reached them through the network, the sample contains clustered information. Adding more connected stocks can increase the row count without increasing the amount of independent evidence by the same amount. The issue is not unique to supply chains. The advantage of the graph is that it gives the researcher an observed structure for defining where some of that dependence may come from.

Network dependence can change the standard error even when the coefficient stays the same

Econometric research explicitly studies random variables whose dependence arises from an observed network. Kojevnikov, Marmer, and Song develop limit theory and network-HAC standard errors for settings where observations are connected and dependence decays through the network. The practical implication for a supply-chain quant is straightforward: ordinary heteroskedasticity-robust errors can be too optimistic if residuals remain correlated among economically connected firms.

A first diagnostic can be much simpler than a full network-HAC implementation. After fitting a cross-sectional model, calculate whether residuals remain more correlated among direct customer-supplier pairs, companies sharing a customer, or companies inside the same network cluster than among random matched pairs. If the residual dependence is strong, the researcher should be skeptical of significance estimates that assume independent cross-sectional errors. Clustered errors by industry may help, but they can still miss cross-industry relationships such as HPE and Nvidia both connecting to Microsoft.

Effective sample size should be thought of in shocks as well as stocks

Suppose a strategy tests supplier reaction to customer earnings across 200 supplier observations. If 70 of those observations come from suppliers tied to the same ten customers, the economically independent event count can be much closer to ten customer shocks than to seventy separate experiments. The exact effective sample size depends on the correlation structure and cannot be read directly from the graph, but the network can reveal when a headline observation count is misleading.

This becomes especially important in event studies. A single Nvidia earnings report can generate observations for Micron, SK Hynix, and other connected companies. Those supplier responses are valuable cross-sectional information about exposure differences, but they do not create several independent Nvidia earnings events. Statistical inference should recognize the event cluster even if the portfolio contains several securities. Counting stocks instead of shocks can make a small event sample appear much larger than it is.

Train-test splits can inherit the same dependence problem

Network dependence affects validation as well as standard errors. If companies connected to Nvidia appear in both training and test sets, the model can benefit from economic similarities shared across the split. That can be intentional when the deployment universe contains the same cluster, but it should not be confused with broad generalization. Likewise, a test period containing hundreds of stock-month observations can still be dominated by a small number of common macro or network shocks.

A stronger research report can therefore state several sample sizes: number of security observations, number of unique companies, number of unique relationship pairs, number of major customer or supplier events, and number of independent time periods or network clusters. None of those is a perfect effective-sample-size formula, but together they provide a much more honest picture than one giant row count. The researcher can then use clustered, block, or network-aware inference appropriate to the hypothesis.

The conclusion is that rows are not the same thing as independent evidence

Supply-chain data can expand a quant dataset dramatically because one event can connect to many firms. That is analytically useful, but it can also create false confidence if every connected row is treated as a new independent experiment. The graph should be used not only to build features, but also to identify where the observations themselves are statistically dependent. A model supported by many companies, many distinct network clusters, and many independent events is much stronger than one whose impressive sample size comes mostly from repeating the same underlying shock across connected firms.

The graph-aware validation guide explains how connected firms can leak economic similarity across train and test sets. The event-study guide explains why several supplier reactions to one customer event should still be recognized as one event cluster.

For relationship definitions and evidence limits, read the Altsets methodology.

Sources

Methodology

Read the methodology for this research.