How Should a Quant Model Missing Supply-Chain Data?

September 14, 2026

Altsets

Research by Altsets Research

Share

Missing relationships, structural-only edges, and missing economic metrics are different states, so forcing every blank to zero can create false exposures while complete-case filtering can bias the research universe.

Data used:Altsets Supply Chain Intelligence: 90k+ entities, 400k+ relationships, 20+ years of history.

Key findings

  • A structural relationship with no metric should remain distinguishable from no observed relationship, because encoding both states as zero turns unavailable magnitude into a false economic conclusion.
  • Missingness can correlate with size, geography, disclosure quality, and vendor coverage, so quant models should test whether coverage variables themselves are driving apparent predictability.

Missing values are unavoidable in supply-chain data, but treating every blank as zero creates a model that answers a different question from the one the researcher intended. A missing supplier revenue percentage does not mean the customer contributes zero revenue. A structural relationship without a relationship size does not mean the economic relationship has no value. For a quant, missingness is therefore part of the data-generating process and should be modeled deliberately rather than silently converted into an economic statement.

Separate missing relationship evidence from missing relationship metrics

There are at least two different forms of missingness in a supply-chain graph. In one case, the dataset knows that two companies are connected but lacks a reliable economic metric for the edge. In another case, the relationship itself may be absent because it was never disclosed, never observed, or did not exist. Those cases should not receive the same encoding because one says "known relationship, unknown magnitude" while the other says "no observed relationship in this dataset."

The supplied Tesla network makes the distinction concrete. LG Energy Solution has quantified directional metrics in the supplied view, while several other Tesla supplier relationships are structural without the same economic fields. Encoding every missing percentage on those structural edges as zero would make them look economically disproven rather than unquantified. A model can instead preserve an edge indicator, a metric-availability indicator, and the observed metric only where it exists, allowing the learning algorithm to distinguish topology from magnitude.

Missingness can be correlated with company size, geography, and disclosure behavior

A company is not randomly selected into a well-covered supply-chain dataset. Large public companies may disclose more information, certain jurisdictions may require different reporting, major customers may be named while smaller ones remain anonymous, and private companies can be much harder to map than listed firms. This means missingness can carry information about the observation process itself. If the researcher simply drops every company with incomplete network features, the surviving sample may become systematically larger, more transparent, or more developed-market oriented.

Asset-pricing research has shown that missing predictor data can materially affect cross-sectional estimation and that simplistic complete-case approaches can be inefficient or biased. Supply-chain data adds another layer because missingness can occur at the company, relationship, and metric levels simultaneously. The quant should inspect whether feature availability correlates with market capitalization, exchange, sector, country, age, liquidity, and future returns before deciding how to impute or filter. A model that appears predictive may otherwise be exploiting the coverage pattern rather than the economic relationship.

Imputation should respect what the variable means

Cross-sectional mean or median imputation can be reasonable for some standardized numerical predictors, especially when paired with an explicit missingness flag. It is much less defensible when the number represents a directional relationship percentage whose absence has a semantic meaning. Imputing the median supplier revenue share onto a structural edge can manufacture an exposure that was never observed. Likewise, filling every missing customer cost percentage with the industry average can make sparse companies look more completely measured than they are.

A safer design can keep structural and quantified feature families separate. For example, one model feature can count observed customer edges, another can summarize only the subset with quantified revenue shares, and a third can describe the fraction of observed edges that have metrics. The model is then allowed to learn from coverage without confusing "unknown" with "small." This also makes the feature easier to audit when a surprising prediction appears.

Historical coverage expansion can create an artificial signal

Relationship datasets often improve through time as more disclosures are processed, entities are resolved, and old records are backfilled. If later years contain denser graphs than earlier years, a raw network-degree feature can trend upward even when the true economic network did not change. A model can interpret that coverage growth as a temporal signal. The same problem can appear cross-sectionally if one group of companies receives deeper research coverage than another.

The solution is not necessarily to freeze the dataset at its weakest historical coverage. The researcher can normalize features by observable coverage, include coverage controls, restrict tests to stable subsets, or use only point-in-time information that would actually have been available. What matters is recognizing that graph density can change because the data collection process changed. A backtest should not reward the strategy for a vendor becoming better at finding relationships.

Missingness indicators should be tested as features, not assumed harmless

Sometimes the fact that a metric is disclosed can itself be informative. Companies may quantify material customers while leaving smaller relationships unnamed, or vendors may obtain better metrics for large economically important edges. A missingness flag can therefore acquire predictive power even when the underlying economic metric is unavailable. That is not automatically invalid, but the researcher needs to understand whether the model is learning economics, disclosure policy, coverage bias, or some mixture of all three.

One useful diagnostic is to train a simple model using only coverage and missingness variables. If those variables alone predict returns or portfolio membership surprisingly well, the full network model deserves closer inspection. The signal may still be legitimate, but it should not be described as customer or supplier economics until the researcher has separated the observation process from the economic process. This kind of baseline can prevent a sophisticated model from hiding a very simple data-quality artifact.

The conclusion is to preserve unknown as a state of its own

Quant pipelines prefer dense matrices, but supply-chain evidence does not naturally arrive that way. Forcing every blank into zero can turn uncertainty into false certainty, while dropping every incomplete observation can create a biased universe. A defensible model should distinguish no observed relationship, known structural relationship, known quantified relationship, and unavailable metric, then test whether coverage itself is influencing the result. That approach is less convenient than a fully filled matrix, but it keeps the meaning of the data intact.

The missing-metrics guide explains why an unquantified relationship should not be interpreted as zero. The trust-incomplete-data guide explains how evidence quality should constrain the strength of the investment claim.

For relationship definitions and evidence limits, read the Altsets methodology.

Sources

Methodology

Read the methodology for this research.