If the Sources Are Public, What Makes Supply-Chain Data Proprietary?

September 14, 2026

Altsets

Research by Altsets Research

Share

The sources can be public while the dataset remains proprietary. The value comes from collecting fragmented evidence, resolving entities, normalizing directional relationships, preserving point-in-time history, and turning inconsistent disclosures into one comparable graph.

Data used:Altsets Supply Chain Intelligence: 90k+ entities, 400k+ relationships, 20+ years of history.

Key findings

  • Public-source access is not the same as having normalized customer-supplier data: raw evidence arrives across filings, PDFs, presentations, announcements, aliases, currencies, and inconsistent disclosure formats.
  • The supplied Shin-Etsu relationships illustrate the value of normalization because several customers can be compared inside one directional metric framework rather than researched document by document.

The sources can be public while the dataset is still proprietary. The difficult part is not discovering that filings, supplier lists, annual reports, and company announcements exist. The difficult part is turning fragmented disclosures into one normalized, point-in-time graph where the same company, relationship direction, economic metric, and security can be compared consistently across thousands of entities and many years.

Public documents do not arrive as one supply-chain database

The SEC provides more than twenty years of searchable EDGAR filings and APIs for submissions and XBRL financial-statement data. That is extremely useful raw material, but it does not provide a ready-made field saying that one company is a supplier to another, how economically important that relationship is on each side, whether the relationship was observable at a particular historical date, or how the companies should be resolved across aliases and securities.

Supplier information also appears outside one standardized filing system. Apple publishes a supplier list as a separate document. Shin-Etsu Chemical describes its semiconductor-material businesses in annual reporting. Other relationships can appear in press releases, presentations, contract descriptions, customer-concentration disclosures, or country-specific filings. The raw evidence is public in many cases, but the evidence format, naming, units, and disclosure style are not standardized.

Normalization is what makes two relationships comparable

The supplied Altsets data maps Shin-Etsu Chemical to TSMC, Samsung Electronics, and Intel with displayed supplier revenue shares of 4.02%, 2.43%, and 1.79% respectively. The investment value of that view is not merely knowing that Shin-Etsu sells into the semiconductor industry. The relationships have been placed into the same directional framework so the investor can compare how important those customers are to Shin-Etsu.

That comparison requires more than collecting documents. Company names need to resolve to stable entities, the supplier and customer direction has to remain correct, currencies and units have to be normalized where absolute values are compared, and missing economic metrics must remain missing instead of being silently interpreted as zero. A raw filing search can find evidence. A normalized relationship dataset makes the evidence comparable.

Historical reconstruction is harder than current-state research

Finding a relationship today is easier than determining what a researcher could have known in March 2014. Companies merge, delist, change tickers, rename subsidiaries, add new securities, and revise disclosures. A relationship can also be discovered later even though the underlying commercial tie existed earlier. A historical trading dataset has to decide whether it is representing economic truth or the information state that was actually observable at the time.

That point-in-time layer is a major part of the data product. Re-running a modern filing parser over twenty years of documents does not automatically recreate what the database looked like in each historical month. Later entity-resolution work and newly discovered evidence can leak backward unless the system preserves historical observation state.

Proprietary does not have to mean secret source documents

A dataset can be proprietary because of collection, normalization, reconciliation, entity resolution, historical versioning, and derived estimates even when many underlying source documents are publicly accessible. Market-data vendors have always created value from organizing information that investors could theoretically collect themselves. The relevant question is whether rebuilding the same clean, historical, cross-company structure would require meaningful time, engineering, and research effort.

This distinction is important for investors evaluating alternative data. A source being public does not mean every market participant has the same usable dataset. The informational edge can come from having the relationship graph already normalized, searchable, historically versioned, and connected to tradable securities while another researcher is still assembling documents company by company.

The conclusion is that the moat is the structured history, not the existence of filings

The sources can be public while the dataset remains proprietary. The value comes from converting inconsistent public evidence into a normalized relationship graph with stable entities, directional metrics, point-in-time history, security mapping, and explicit missingness. Anyone can open a filing. Recreating the same comparable network across thousands of companies and historical periods is a different task.

The filings-versus-data guide explains what structured relationship data adds after the investor has access to the original documents. The point-in-time feature-store guide explains why historical lineage and observation timing matter once the data is used quantitatively.

For relationship definitions and evidence limits, read the Altsets methodology.

Sources

Methodology

Read the methodology for this research.