Method

How the Atlas of AI research is measured

What is counted, from which public sources, how a researcher gets a country, a subject and a career stage, and where the figures stop being reliable.

Published 12 September 2026.

The short answer

Every figure on this site is a count of people in one index of the AI research literature, built from public records: arXiv, OpenAlex, ORCID, DBLP, Hugging Face and the pages researchers publish about themselves. A researcher is a person with at least one indexed AI paper. Their country is the country of their institution, their subjects come from the OpenAlex classification, and their career stage and employer type are read from what the public record says about them.

The index is refreshed nightly and the published figures weekly. They are counts, not estimates, and they carry the number of profiles they were computed on. The aggregates are free to download and reuse with a citation.

Where the data comes from

The index is assembled from sources that publish research records openly, each for its own purpose, and that describe the same people from different sides:

  • arXiv for preprints and their metadata: authors, dates, categories, and the text of the paper where it says something the metadata does not.
  • OpenAlex for the scholarly graph: works, authors, institutions, the topic classification of the literature, and the links between them.
  • ORCID for what researchers declare about themselves: employment, education, and the dates that go with them.
  • DBLP for the bibliography of computer science, including registered theses.
  • Hugging Face for models and datasets that cite a paper, which say who builds rather than only publishes.
  • Pages researchers publish themselves: a lab page, a personal site, a curriculum, read for what the person chose to make public.

Every source is read through its public interface and under its terms. No private database, no paywalled source and no social network is used, and nothing is inferred about a person that a public record does not support. Each element of a profile keeps its source, and each date says whether it rests on a document, on a declaration by the person, or on an estimate.

What counts as a researcher

The unit is a person, not a paper. Records from the different sources that describe the same author are merged into one profile, using the identifiers the sources share and, where they are missing, the concordance of names, co-authors and institutions. A person counts as an AI researcher when at least one of their indexed papers falls on an AI subject. A doctoral student with a single workshop paper and a professor with several hundred each count once.

The corpus is the AI literature as those sources hold it, which means it leans toward fields where preprinting is the norm. Subfields that publish mostly in closed venues without a preprint are under-represented, and the figures should be read as counts of the visible literature, not of all the people who work on AI.

How a researcher gets a country

A country is the country of the institution on the researcher's most recent indexed papers, as the sources record it. It is a place of work. A researcher who moves is reassigned when the affiliation on their record changes, which happens with a lag of a few months to a year after the move. Countries with fewer than forty indexed researchers are not shown separately.

How a researcher gets a subject

Subjects are topics from the OpenAlex classification of the literature, at its finest level, assigned to papers by OpenAlex and read across the researcher's indexed work. A researcher can carry several subjects, so the counts by subject overlap and never add up to the total. Where a page reports a share by subject, the denominator is stated on the page.

Career stage, employer, and newcomers

  • Early career is read from the length and the shape of the publication record and from what the person declared about their education, with a doctorate in progress or recently finished as the typical case.
  • In a company means the institution on the record is classified as a company by the source, rather than a university, a public research body or a hospital.
  • Newcomers in a window are the researchers whose first indexed paper falls in that window. Growth on the Atlas compares a recent window with the one before it, and the years of both are stated where the comparison is shown.
  • Doctorates in progress and their expected end are dated from public records only: a start date declared by the person, a registered thesis, a lab page, or the first paper as a lower bound. The confidence of each date is kept with it.

Where the figures stop being reliable

  • Small cells. Below a few dozen people, a count says more about the sources than about the country or the subject. Pages do not show such cells, and the public exports withhold them.
  • Names. Merging records across sources is imperfect for common names. The error goes both ways and is small at the scale of a country, larger at the scale of a single institution.
  • Lag. A move, a defence or a new subject appears in the record months after it happens. Figures for the current year are always incomplete and are marked as such.
  • Coverage. The corpus is the preprinted and openly indexed literature. Fields and countries that publish elsewhere are under-counted.

Refresh and consistency

The index is refreshed every night. The figures shown on the guides, on the Atlas, on the country and subject reports and in the public exports come from one computation on that index and are republished weekly, so a number read on one page matches the same number on another. Each figure carries the date of the computation and the number of profiles it was measured on.

Reuse and citation

The aggregates by country, subject and year are published on the data page in CSV and JSON under a Creative Commons Attribution 4.0 licence. Reuse is free with a citation of Founding Hires and a link to the page the figure was taken from. The index itself, with names and contact details, is not public and is not available as a dataset; how the people in it are treated is described in the privacy policy.

Common questions

Where does the data in the Atlas come from?
From public sources only: preprints and their metadata on arXiv, works, authors, topics and institutions on OpenAlex, the employment and education researchers declared on ORCID, bibliographic records on DBLP, models and datasets linked to papers on Hugging Face, and pages researchers published themselves. No private database, paywalled source or social network is used.
What counts as an AI researcher?
A person with at least one indexed paper on an AI subject, once the records of the different sources have been merged into one profile. A student with one workshop paper and a lab director with three hundred both count as one researcher; the figures are counts of people, never of papers.
How is a researcher assigned to a country?
By the country of the institution on their most recent indexed papers, as recorded by the sources. It is a place of work, not a nationality, and a researcher who moves is reassigned when their affiliation changes in the record.
What is a subject?
A topic from the OpenAlex classification of the literature, at its finest level. A researcher can carry several subjects, so the counts by subject overlap and do not add up to the total.
Are the figures the same as on the guides and the reports?
Yes. The Atlas, the guides, the country and subject reports and the public exports all read the same computation from the same index, refreshed on the same schedule.
Can the figures be reused?
Yes, under a Creative Commons Attribution 4.0 licence, with a citation of Founding Hires and a link. The aggregates are on the data page in CSV and JSON. The index itself, with names and contacts, is not public.

The names behind the numbers

Every count on this site is a list in the index.

Search it by subject and country, set an alert on the people who become reachable, or hand the search over and pay only when someone joins.