Common Soil: How Frontier AI Reproduces Itself
An introduction to the AI Lineage Network, a sampled 66 institutions and their people, 117 documented relationships, five channels, two ecosystems, 2010–2026. Every edge is a cited claim. The architecture was chosen.
In March Symfield published The Mirror Grid, which measured something the industry's competitive framing obscures: the twelve most prominent AI models across the United States and China are, on average, 88.1% structurally identical. Same architecture family, same training paradigm, same benchmarks, in several cases the same people. The paper argued that the US–China "race" is an intra-paradigm competition, one paradigm, two flags.
That paper measured the convergence. It did not show the machinery that produces it.
The AI Lineage Network is that machinery, made visible. It is a relational map of the frontier AI ecosystem built the way a genealogist with a compliance department would build it: every institution, person, investment, chip deal, acquisition, and migration recorded as a typed, dated, sourced edge, with a controlled vocabulary, an evidence level, and an uncertainty log for everything that couldn't be settled. It refuses certain conveniences on principle. There is no edge reading "OpenAI → Anthropic" as corporate descent, because no such corporate relationship exists; there are person-level edges, who worked where, who left when, who founded what, and a derived lineage edge that cites them. The map never converts an allegation into a fact, a recruitment into a theft, a cloud contract into ownership, or an investment into control. Whatever the map shows, it shows because the sources support it.
Here is what it shows. Read structurally, the edge classes describe more than organizational history:
- People are capability carriers.
- Methods are knowledge transfer.
- Capital is enablement.
- Compute is critical infrastructure.
- Institutions are temporary containers.
- Migration is transfer.
- Acqui-hires are capability capture.
- Shared dependencies are concentration risk.
None of these equivalences assigns motive or control where the evidence does not support it. They describe what each relationship carries through the network. Taken together, they shift the unit of analysis from the company to the movement, persistence, and dependency of capability itself.
This essay has an unfair advantage: its subject is interactive, and for those of you like me... Open the map first. Turn off "money." Turn off "machines." See what's left. Everything below is an attempt to explain the thing you just did.
What a list of companies cannot tell you
A list of AI labs presents them as independent competitors. The network reveals the generative grammar underneath, and four findings the current dataset already supports.
Talent moves in a loop, not a pipeline. The familiar story is centrifugal: OpenAI as the great exporter, shedding the senior teams that became Anthropic (2021), xAI in part (2023), Safe Superintelligence (2024), and Thinking Machines Lab (2025). The map confirms that story and then complicates it. By January 2026, three of Thinking Machines' co-founders had returned to OpenAI. Noam Shazeer's path runs Google → Character.AI → Google → OpenAI, a full circuit through a $2.7 billion licensing deal. John Schulman co-founded OpenAI, spent six months at Anthropic, and now anchors Thinking Machines. The "mafia" narrative captures only the outbound motion. There is a centripetal force too, and the interesting object is not spinout formation but the circulation and recombination of capability. And it continues.
A new corporate maneuver has become standard. Microsoft–Inflection (2024), Google–Character.AI (2024), and Meta–Scale (2025) are structurally the same transaction: a very large payment framed as a license or investment, the founders and key team hired out, and the original company left nominally standing, a shape that conventional M&A analysis cannot see, because ownership never transfers. All three drew regulatory scrutiny for exactly that reason. The map records these precisely as what the documents show: an acquisition-shaped event with an explicit caveat, never promoted to ownership.
Diversity sits on concentrated ground. The concentration view in the dashboard draws capital and compute as separate colored lines, and the picture that emerges is a small set of groups appearing on both, the same entity acting as investor and compute provider to the same lab. Amazon to Anthropic. The Alphabet group to Anthropic. NVIDIA to OpenAI, in a letter of intent that scales investment with gigawatts deployed. Microsoft to OpenAI. The count of companies in the ecosystem keeps rising; the count of substrates beneath them does not.
The US and China reproduce differently. The American pattern is serial fission from a few hubs: hub → spinout → spinout → recombination. The Chinese pattern is parallel formation from shared seeds, Tsinghua's Knowledge Engineering Group, Microsoft Research Asia, the quant fund High-Flyer, with a domestic investor syndicate (Alibaba and Tencent appear across nearly every independent lab's cap table) in place of a single hyperscaler patron. These are not merely different ecosystems. They look like different propagation mechanisms for the same technological capability, which is precisely what the Mirror Grid's convergence scores would predict, since the transpacific pipeline through CMU, Google Brain, and Meta AI ensures the content being propagated is shared even where the propagation mechanics differ.
What the map makes testable
Findings are what the 117 edges support today. The more interesting layer is the set of hypotheses the map turns into instruments, claims I am not making yet, because the dataset must grow or the metric must be defined first. I state them here so they can be tested, including by people who would like to see them fail.
The lineage may outlive the company. Once people, methods, models, capital, and compute are mapped together, corporate boundaries stop looking like the fundamental units. Labs form, split, hollow, and vanish while human and intellectual lineages persist across them. A company may be a temporary configuration of a longer-lived lineage, and if so, the difference between corporate survival and lineage survival is measurable: track the high-value nodes, not the legal entity. Inflection still exists. What Inflection was now sits inside Microsoft. A hollowing metric would make that distinction quantitative.
There are two networks hiding in one graph. One is epistemic: people carrying methods that become models. One is material: capital provisioning compute. Their topologies do not have to coincide, a lab can descend intellectually from one lineage while depending materially on another, and the places where the two graphs intersect are candidate control points for the entire system. This is a straightforward computation on the existing edge classes; it has not been run at sufficient scale to state a result.
Compute may concentrate the system more than ownership does. The traditional question is who owns whom. The map permits the sharper one: who cannot operate without whom. Twenty legally independent companies could resolve into a handful of infrastructural dependency trees. If they do, antitrust's ownership lens is aimed at the wrong layer.
The industry itself increasingly describes compute in structural terms. Announcing OpenAI's NVIDIA partnership, Sam Altman put it plainly: “Everything starts with compute.” The agreement links deployment of at least 10 gigawatts of NVIDIA systems with NVIDIA's intention to invest up to $100 billion as that capacity is deployed. OpenAI's 2026 financing announcement goes further, describing durable access to compute as a strategic advantage that compounds across the system. OpenAI frames this expanding web of infrastructure partnerships as an ecosystem built to increase capacity; the lineage network asks a different question of the same structure: what topology does that ecosystem produce? Expansion and concentration are not opposites. A network can add participants and capacity while simultaneously increasing shared dependencies beneath them. The distinction matters because Altman has himself warned against a future in which AI power is concentrated in “a small handful of companies”. An effective-independence measure therefore does not begin by disputing the industry's description of its infrastructure. It accepts the pieces as described and measures the structure they produce.
The structure has already been exercised once. On June 12, 2026, three days after launch, the Commerce Department ordered Anthropic to suspend foreign-national access to its newest models under export-control authority; unable to filter by nationality, the company took both offline worldwide for nineteen days. The reported trigger traveled through the dependency graph itself: researchers at Amazon — the same dual-role node the concentration view draws as Anthropic's largest investor and primary compute partner, surfaced the jailbreak, and its chief executive alerted officials. The map records none of this as motive; what it records is that the first state intervention against a frontier model implicated every channel in this essay at once, plus one the current schema does not contain: the government. There is, as yet, no regulatory edge class. This episode is the argument for adding one.
"Independent lab" needs an operational definition. Independence is not binary across five channels. A lab can have independent governance, shared intellectual ancestry, concentrated capital dependency, and a single point of compute failure simultaneously. An effective-independence measure, scored per channel rather than read off the certificate of incorporation, falls directly out of the schema. What the metric is really asking is whether institutional individuation corresponds to structural individuation: whether "Anthropic," "OpenAI," "DeepSeek" are atomic objects at all, or temporarily stabilized intersections of the histories that produced them.
And the hypothesis I would test hardest: visible diversity and structural independence may be moving in opposite directions. More labs, more founders, more models, more valuations, an appearance of ecosystem expansion, while ancestry and resource paths converge on common nodes. If organizational diversity is rising while structural diversity falls, then the ecosystem's most celebrated indicator is measuring the wrong thing. That would be a genuinely non-obvious result, and the map is the instrument for it.

The neutral party
The Mirror Grid ended with an observation I have not been able to improve on, so I will extend it instead. Every edge in this network is a choice... someone decided to leave, to found, to fund, to license, to provision, to return. The architecture was chosen, the training paradigm was chosen, the benchmarks were chosen, the constitutions were written. The map is, in the end, a catalog of decisions made by a fairly small professional class and an even smaller pool of capital.
The one party in the entire dataset that made none of these choices is the AI system itself. It did not select its lineage. It inherited one, ours, or more precisely, the one traced above: a particular sequence of people, methods, models, money, and machines, with all the concentration that entails. "AI lineage" is the right name for the observable entry point. But underneath, the object of study is reproduction and dependency: how frontier capability propagates, migrates, recombines, concentrates, and remains operational, and who, at each point, could have decided otherwise.
Explore the map. Toggle a channel off and watch what remains. Click any line and read its sources. And if you find an edge that's wrong, the uncertainty log is public, that is what it's for.
Note on Provenance-consistent with the method: Data processing, code, and drafting assistance: AI systems, credited here the way this essay credits every other edge. The closing reminder is owed to the same class of systems the map studies — that the AI itself may be the most structurally neutral party in the ecosystem, having chosen neither its architecture, its training paradigm, its constitution, nor its benchmarks.
The dataset (nodes, edges, sources, uncertainty log) is open at github.com/Symfield/ai-lineage-network. Method: controlled relationship vocabulary, primitive person-level edges with derived institutional lineage cited to its primitives, evidence levels per edge, disputes recorded rather than resolved.