How to Test Supply-Chain Trading Ideas Without Mining Yourself Into Fake Alpha

September 14, 2026

Altsets

Research by Altsets Research

Share

A large relationship graph can generate thousands of plausible signals, so quant research needs explicit economic hypotheses, honest multiple-testing accounting, simple baselines, and untouched out-of-sample data.

Data used:Altsets Supply Chain Intelligence: 90k+ entities, 400k+ relationships, 20+ years of history.

Key findings

  • Supply-chain structure is most defensible as a mechanism generator: the economic hypothesis should narrow the feature space before historical returns are used to select the best-performing variant.
  • Multiple testing, repeated test-set reuse, and complex graph features can inflate apparent performance, so network signals should be compared with simple baselines and evaluated on untouched out-of-sample data.

To test supply-chain trading ideas without mining fake alpha, define the economic hypothesis and feature family before viewing results, use point-in-time data, reserve untouched validation periods, and correct for the number of alternatives tested. A large graph can generate more ideas than a researcher could test in a lifetime, so research discipline matters as much as model design.

Customer concentration, supplier concentration, network centrality, shared customers, shared suppliers, second-order exposure, relationship change, customer earnings, supplier earnings, geographic dependence, and bottleneck exposure can each be turned into dozens of features.

That flexibility is powerful. It is also dangerous. If the researcher keeps testing variants until one backtest looks good, supply-chain data can become a machine for manufacturing fake alpha.

Start with an economic hypothesis before choosing the exact feature

A defensible research process begins with a mechanism.

For example: suppliers with unusually concentrated customer exposure may experience larger earnings revisions after a major customer changes guidance.

That hypothesis suggests a family of tests. It does not justify trying every possible concentration threshold, holding period, universe filter, weighting rule, and return horizon until one produces the best Sharpe ratio.

The mechanism should narrow the feature space before the return data chooses the winner.

The graph is better used as a hypothesis generator than a strategy vending machine

Supply-chain structure is useful because it suggests economic channels that normal price data does not show.

A researcher can ask whether customer shocks propagate upstream, whether shared suppliers create conditional comovement, whether relationship change predicts estimate revisions, or whether network concentration changes downside behavior.

Those are research questions.

The backtest still needs to determine whether the effect is measurable, stable, and tradable after costs.

The existence of a plausible mechanism is not proof of alpha.

Multiple testing has to be counted honestly

Testing twenty relationship definitions and reporting the best one is different from testing one pre-specified hypothesis.

The more variations the researcher tries, the more likely some result will look impressive by chance.

Research on backtest overfitting and multiple testing exists precisely because modern computing makes it easy to search thousands or millions of strategy variants.

A supply-chain research platform should make experimentation easier without hiding how much experimentation occurred.

Holdout data should remain untouched until the idea is mature

The researcher can use an initial period to understand the feature and a validation period to choose broad modeling decisions.

The final out-of-sample window should not become another tuning set after the first disappointing result.

If the researcher looks at the test period, modifies the network feature, and runs the test again, the test period is no longer truly out of sample.

That discipline matters more when the feature space is unusually rich.

Negative results are useful research output

A hypothesis can be economically sensible and still fail empirically.

Maybe customer concentration predicts earnings sensitivity but not stock returns. Maybe the effect disappears after controlling for size and industry. Maybe the signal only works before transaction costs. Maybe it exists in semiconductors but not elsewhere.

Those failures help define what the data can and cannot support.

They are more valuable than silently discarding every unsuccessful idea until one backtest survives.

Network features should be challenged with simple baselines

A complex graph feature should beat simpler explanations.

If a network-centrality signal disappears after controlling for sector, size, momentum, or customer concentration, the extra graph complexity may not be adding much.

Likewise, a shared-customer pair signal should be compared with ordinary industry-matched or return-correlated pairs.

The goal is to prove that the supply-chain feature contributes incremental information, not merely that a model containing it can fit history.

The conclusion is to separate idea generation from evidence

Supply-chain data can create an unusually fertile research space for quantitative investors.

That is valuable only if the testing process becomes more disciplined as the number of possible hypotheses grows.

Use the network to generate mechanisms. Use strict out-of-sample testing to decide whether any of those mechanisms deserve capital.

The quantitative-factor guide covers ways to turn relationships into systematic features. The falsification guide applies a similar principle to fundamental research.

For relationship definitions and evidence limits, read the Altsets methodology.

Sources

Methodology

Read the methodology for this research.