How to Turn Supply-Chain Data Into Quantitative Stock Factors

July 18, 2026

Altsets

Research by Altsets Research

Share

Engineer point-in-time customer, supplier, overlap, and network features for systematic equity research while keeping relationship denominators separate and avoiding look-ahead bias.

Data used:Altsets Supply Chain Intelligence: 90k+ entities, 400k+ relationships, 20+ years of history.

Key findings

  • Apple at 8.75% of SK Hynix revenue and Tesla at 19.03% of LG Energy Solution revenue are comparable supplier-side observations, while the corresponding customer-cost percentages belong in a different feature family.
  • Supply-chain factors can use economic magnitude, network structure, and event linkage, but each feature still needs point-in-time construction, ordinary factor controls, and out-of-sample validation before it can be treated as a return signal.

Supply-chain data can become a quantitative stock feature, but the feature should be built from one economic question at a time. The useful raw material is not a generic "dependency score." It is a point-in-time set of directional customer and supplier relationships that can be transformed into reproducible cross-sectional variables.

The supplied Altsets data illustrates the difference. Apple is associated with 8.75% of SK Hynix revenue, while SK Hynix is associated with 1.76% of Apple's cost base. Those two numbers describe the same relationship from different sides. A quantitative model can use either one, but it should not average them merely because both are percentages.

The first factor can be customer dependence

A simple supplier-side feature is the largest customer revenue exposure for each stock. In the SK Hynix example, the Apple relationship contributes an 8.75% supplier-revenue observation. LG Energy Solution's mapped Tesla relationship contributes a 19.03% supplier-revenue observation for another company.

Those values are directly comparable because they use the same supplier-side denominator. A cross-sectional screen could rank companies by the largest disclosed or estimated customer share, top-three customer share, or another explicitly defined aggregation. The factor then has a clear economic interpretation: customer dependence.

A different feature can measure upstream dependence

Customer cost percentage belongs in another feature family. It can identify companies whose cost base appears concentrated in one supplier or group of suppliers. That may be useful for research on margin sensitivity, bottlenecks, pricing power, or event exposure.

The critical point is that upstream dependence should remain distinct from customer dependence. A model can include both features as separate columns. Combining them into one average removes the directional meaning that made the data useful.

Network structure creates features that financial statements cannot

The graph can also generate non-percentage features. Shared-customer counts, shared-supplier counts, number of economically material second-order paths, relationship concentration, and network centrality are all candidates for systematic research.

Feature familyExample constructionInvestment question
Customer dependenceLargest supplier-revenue percentageWhich stocks rely most on one buyer?
Upstream dependenceLargest customer-cost percentageWhich stocks rely most on one supplier?
Counterparty overlapNumber of shared customers or suppliersWhich stocks carry similar hidden economic exposure?
Network concentrationConcentration across mapped relationshipsIs the company dependent on a narrow network?
Event linkagePoint-in-time relationship joined to an eventWhich stocks are economically connected to the event company?

These are feature definitions, not claims that any one of them predicts returns.

There is real asset-pricing precedent for economically linked firms

Academic finance has studied return propagation through customer-supplier relationships for years. Cohen and Frazzini's 2008 Journal of Finance paper documented historical return predictability across economically linked firms and described a customer-momentum effect. More recent work has continued to study customer momentum and network spillovers, while some post-discovery evidence suggests the effect can weaken or depend on the sample and event type.

That history is useful for a quant researcher because it supplies hypotheses, not guaranteed alpha. A proprietary relationship dataset can improve coverage, direction, history, or economic weighting, but the resulting feature still has to survive a proper out-of-sample test.

Economic weighting can make a network feature more precise

A binary relationship variable treats a tiny customer and a dominant customer the same. Supplier revenue percentage can add economic magnitude to the customer side, while customer cost percentage can add magnitude to the supplier side.

The correct weighting depends on the hypothesis. A customer-momentum feature might weight customer returns by supplier revenue exposure. An upstream-shock feature might instead weight supplier events by customer cost exposure. The choice should be specified before the backtest rather than selected after seeing which version performs best.

Point-in-time construction is not optional

A quantitative feature is invalid if it uses relationships that were not known at the historical signal date. The same problem occurs if an analyst uses a current estimate to fill an earlier period automatically.

For every rebalance date, the model should use the relationship snapshot available at that date, resolve the relevant securities as they existed at the time, and calculate the feature only from information that the historical investor could have observed. This is where Altsets' monthly point-in-time snapshots become more important than a current relationship table.

Factor research needs ordinary quant controls too

A supply-chain feature can accidentally proxy for company size, industry, country, momentum, liquidity, or another well-known effect. A serious test should examine those correlations and determine whether the signal survives reasonable controls.

The purpose is not to strip away every economic connection. It is to understand what the factor actually measures. If a customer-dependence feature works only because the highest-exposure stocks are all small semiconductor companies, the investment interpretation is very different from a signal that persists across sectors and size buckets.

Do not confuse a useful feature with a tradeable factor

A feature can improve ranking, risk control, or explanatory power without producing standalone alpha. It can also work only in combination with valuation, earnings revisions, momentum, or another input. That still makes the data valuable.

The economic-basket guide focuses on grouping stocks around measurable exposure. This article addresses a different quantitative job: turning the relationship graph into columns that can enter a cross-sectional model.

The API is the natural interface once the hypothesis is fixed

A visual map and LLM are useful while the researcher is discovering possible features. Once the exact definition is fixed, a programmatic workflow is usually better because the same calculation has to be repeated across many securities and dates without changing the rules. The Altsets documentation covers the available data interfaces, and the point-in-time backtesting guide explains the historical discipline required before interpreting backtest results.

For relationship definitions and evidence limits, read the Altsets methodology. Browse Supply-Chain Data Use Cases for other systematic and event-driven investment research ideas.

Sources

Methodology

Read the methodology for this research.