Blog / Common Crawl's own researchers put its persistent core near 40 percent

Common Crawl's own researchers put its persistent core near 40 percent

A paper from three Common Crawl Foundation authors, posted to arXiv on 15 July 2026, fits 51 monthly crawls from 2020 to 2025 and reports a persistent core fraction of about 0.4 at domain granularity, with the rest of the domain population churning through a shell.

In short

  • A paper titled Measuring What the Crawler Sees, by Michael Paris, Hande Celikkanat and Luca Foppiano of the Common Crawl Foundation, was submitted to arXiv as 2607.13636 on 15 July 2026 and runs to 16 pages with 4 figures.
  • Fitting two independent projections of the same sampling process to a 51 crawl window of Common Crawl from 2020 to 2025, at domain granularity, the paper reports both fits landing on a persistent core fraction of about 0.4, written in the paper as an approximation rather than to further digits.
  • The paper reports a per round coverage fraction of about 0.78 for Common Crawl at domain granularity, against about 0.51 for the German Academic Web at URL granularity, and states that coverage is the parameter an operator tunes while survival is supplied by the web.
  • The same procedure on the German Academic Web, a closed Heritrix crawl of 2013 to 2019 seeded from roughly 150 German academic homepages, converges on a triple of about 0.06, 0.62 and 0.51 for core fraction, survival and coverage.
  • The paper publishes no numeric survival rate for Common Crawl at domain granularity and no numeric shell parameters, and it states that a residual disagreement on shell coverage means the shell is not homogeneous, so nothing in it identifies whether any individual domain sits in the core or the shell.

Three researchers at the Common Crawl Foundation have published a measurement of their own crawler, and the finding worth carrying out of it is a mass split. Treating a longitudinal crawl as repeated sampling from an urn, and fitting two independent views of that process to 51 monthly Common Crawl archives from 2020 to 2025, they report that the domain population does not behave as one uniform pool. About four tenths of it sits in a persistent core that the crawl keeps resolving. The rest churns through a shell, appearing and disappearing at a rate the two views of the data still disagree about.

Lantad measured none of this. Everything below is a reading of a paper on arXiv and of two pages of Common Crawl's own site, opened and checked on 1 August 2026, with the figures taken from the paper's own results and discussion sections rather than from anyone's summary of them. The part this site can speak to comes at the end, and it is narrower than the paper: none of these parameters tell you where your domain sits, and the one thing you do control is what your server hands an AI crawler on the round it does arrive.

Pairwise containment

  • Fraction of one crawl's URLs seen again later
  • Uses one crawl pair at a time
  • Described as recurrence weighted
  • The paper says it reads the persistent core
  • Pulls survival high and coverage moderate

Discovery curve

  • Distinct items seen across a window of crawls
  • Uses the full sequence of crawls
  • Described as fresh discovery weighted
  • The paper says it reads the ephemeral shell
  • Pulls both parameters lower
The two projections the paper fits, as described in its own method and results sections. A restatement of published research, not a measurement of any site.

What the paper measured, and who wrote it

The document is Measuring What the Crawler Sees on arXiv, filed as 2607.13636 under Physics and Society with cross lists to Digital Libraries and Information Retrieval, submitted on 15 July 2026, and carrying the comment 16 pages, 4 figures, web metrics. Its three authors, Michael Paris, Hande Celikkanat and Luca Foppiano, give their affiliation as the Common Crawl Foundation and their contact addresses on the commoncrawl.org domain. That provenance is most of why the paper is worth a post. This is not an outside audit of a crawler, and it is not a vendor blog claiming a result. It is the operator of the largest open web corpus publishing a model of what its own sampling process does, in a venue where the derivation can be checked.

The method starts from a deliberately plain picture. A crawl is treated as repeated draws from an urn holding some number of unique items. Each round draws a fraction of the urn, called coverage, and retains a fraction of it, called survival, replacing what it does not retain with items never seen before. Under that model every item in a round meets one of three fates: it is churned out, it survives and is sampled into the crawl, or it survives and is missed, persisting unseen into the next round. Those three probabilities sum to one, and the third is the one that gives the system its memory.

From there the paper does something that makes the result falsifiable rather than merely fitted. Two different observables are functions of the same two parameters. Pairwise containment asks what fraction of one crawl's items reappear in a crawl some number of rounds later. The discovery curve asks how many distinct items accumulate over a sliding window of consecutive crawls. If the population really were uniform, fitting each observable independently would recover the same pair of numbers. The paper's phrasing for what happens when they do not is the sentence to keep: any disagreement is itself a measurement.

The Common Crawl inputs are public. The paper states that Common Crawl publishes a per crawl URL count and a family of rolling window unique URL counts for windows of 1, 2, 3, 4, 6, 9 and 12 crawls, and it pivots those across every valid combination of starting crawl and window length to trace the discovery curve. It reduces both projections to the same 2020 to 2025 monthly window and aggregates to the domain population, which it argues is the granularity at which the operator's host budgets and centrality thresholds actually act.

ItemWhat the paper statesWhat it does not state
Identifier and datearXiv 2607.13636, submitted 15 Jul 2026, 16 pages, 4 figuresAny journal acceptance or peer review outcome
AuthorsMichael Paris, Hande Celikkanat, Luca Foppiano, Common Crawl FoundationAny statement that this is official operator policy
Common Crawl window51 crawls, 2020 to 2025, aggregated to domain granularityResults at URL granularity for Common Crawl
Second archiveGerman Academic Web, URL granularity, 2013 to 2019, biannualA third archive, or any commercial search index
Core fraction fittedAbout 0.4 at domain granularity, from both projectionsA confidence interval, or digits beyond the approximation
Shell parametersReported as tightening, with a residual gap on shell coverageNumeric values for shell survival or shell coverage
Facts stated in arXiv 2607.13636, read on 1 August 2026. Every row is a reading of the paper, not an independent measurement.

What a persistent core of about 0.4 actually says

Run the two fits under the assumption that the domain population is uniform and they do not agree. The paper describes containment as pulling survival high and coverage moderate, because it is dominated by the recurring mass, while the discovery curve pulls both lower, because it is dominated by churn driven fresh discoveries. Neither single pair of numbers explains both views at once, and the size of that gap is what the paper calls the homogeneous fit failure on this window.

The minimal repair is to stop treating the population as one pool. The paper splits it into a persistent core, modelled as fully sampled and not churning, holding a fraction it calls kappa, and a shell with its own survival and coverage. Refitting both projections under that two component model, the paper reports that both land on a common core fraction of about 0.4 at domain granularity, and it presents that agreement as the point: a mass split that the single population reading could not see at all.

Two cautions belong immediately next to that number, and the paper supplies both. First, it is written as an approximation. The value appears as about 0.4 in the figure caption and again in the discussion, with no interval and no further digits, so quoting it as 40 percent with a decimal place would be adding precision the source does not carry. Second, the repair is not complete. A residual disagreement between the two projections survives on shell coverage, which the paper reads as a signal that the shell itself is not uniform, with containment anchoring on the inner shell and the discovery curve dragging toward the outer one. The scalar core fraction is offered as an entry point to a rank resolved treatment that the authors explicitly leave to follow up work.

The core is also not the same thing as the crawl. The paper's own framing of survival is worth reproducing because it is the sort of caveat that gets stripped in retelling: the urn has no rebirth mechanism, so a domain absent for several crawls and reappearing later folds in either as a long run of survived but missed events or as a fresh item that happens to share an identifier, and the two are observationally indistinguishable. The recovered survival rate is therefore described in the paper as not a strict web survival rate but the survival rate the urn sees. That is a measurement of the crawl, not of the web, and the paper says so.

  • Persistent core, both projections agree about 0.4 Modelled as fully sampled and not churning. Written in the paper as an approximation, with no interval given.
  • Shell, the remainder about 0.6 Carries its own survival and coverage. The two projections still disagree on shell coverage, which the paper reads as internal structure.
The domain population split reported in arXiv 2607.13636 for Common Crawl, 2020 to 2025, 51 crawls. The core figure is the paper's approximation; the shell figure is its complement and is not separately fitted in the paper.

Coverage is the operator's lever, survival is not

The discussion section makes a distinction that matters more for a site owner than the core fraction does. Of the parameters recovered, the paper reads coverage as the per round fraction the operator tunes, through host budgets, centrality rank thresholds and revisit cadence, and survival as the rate the operator does not pick, because the web supplies it through the persistence of the eligible set. The core fraction is described as a joint readout of both: which items the policy keeps eligible, and how stable that eligible set turns out to be.

The empirical support offered for that split is a comparison across two crawlers with opposing designs. Common Crawl is described as Nutch based and recursive, with each monthly fetch list derived from the previous crawl database and ranked by harmonic centrality over a webgraph the paper puts at 200 billion edges. The German Academic Web is a closed Heritrix crawl, breadth first, seeded from roughly 150 German academic homepages, running biannually rather than monthly over 2013 to 2019 at URL rather than domain granularity. The paper reports coverage of about 0.78 for Common Crawl at domain granularity against about 0.51 for the German Academic Web at URL granularity, and reads that difference as tracking the operating point rather than the population. It is careful to label the split a conjecture rather than a theorem, noting that pinning it down would require holding one side fixed while varying the other.

The second archive is doing real work here beyond a headline figure. Its role is to test whether the structural prediction, that two projections converge on one triple, is a property of the framework or an artefact of Common Crawl's pipeline. On that archive the paper reports a fitted triple of about 0.06, 0.62 and 0.51 for core fraction, survival and coverage, numerically unlike the Common Crawl values, with the structural signature holding anyway. A core fraction of 0.06 against 0.4 is a large difference, and the paper attributes it to the archives sampling different populations at different cadences rather than to one crawler being better.

One more detail from the figure captions deserves carrying, because it constrains how any of this converts into calendar time. Common Crawl's cadence is not evenly monthly across the window. The paper notes it ranges from biweekly to occasional gaps of one to two months in 2021 and 2022, so a twelve crawl window can span anywhere from 0.9 to 2.1 years. Per round parameters are per crawl, not per month, and anyone translating them into how long a domain stays visible needs that conversion stated rather than assumed.

PropertyCommon CrawlGerman Academic Web
Crawler and strategyNutch based, recursive, ranked by harmonic centralityHeritrix based, breadth first, fixed seed hopper
Seed or fetch listDerived from the previous crawl databaseRoughly 150 German academic homepages
Window analysed2020 to 2025, 51 crawls2013 to 2019, biannual
Granularity fittedDomainURL
Coverage per roundAbout 0.78About 0.51
Core fractionAbout 0.4About 0.06
Survival per roundNot published numericallyAbout 0.62
The two archives compared in arXiv 2607.13636, with the fitted values the paper reports in its discussion section. Figures are the paper's, at the granularity each was fitted.

None of this tells you whether your domain is in the core

This is the part most likely to be lost if the finding travels, so it is worth stating before anything else is built on it. The core fraction is a population parameter recovered from aggregate published statistics: per crawl counts and rolling window unique counts. It is fitted without any per domain data at all. Nothing in the paper identifies which domains occupy the core, nothing ranks them, and nothing offers a test a site owner could run to find out. The authors are explicit that promoting the scalar to a per domain rigidity profile, along the operator's centrality rank axis, is follow up work that has not been done yet.

That limitation is not a footnote to be waved at. It is the difference between a fact about a corpus and a fact about a website, and this blog has been strict about that distinction before, most recently in noting how a survey of 45 GEO studies found the crawling stage is the one few studies observe. A number describing how a sampling process behaves in aggregate cannot be converted into advice for one site without an extra measurement nobody has published. Anyone who tells you your domain is in Common Crawl's persistent core is asserting something this paper does not contain.

Common Crawl's own documentation says the same thing in plainer language, and from a different direction. Its frequently asked questions page states that the dataset is a sample of the web, and that it does not generally archive any entire website but a randomly selected subset of it. Read that alongside a core fraction near 0.4 and the two are consistent: partial coverage of the population each round, with some of the population far more reliably resolved than the rest. Neither statement promises any particular site anything.

There is a familiar shape here that shows up whenever a platform level statistic meets a single site. The same problem appears when counting how many prompts an AI visibility measurement needs before its numbers mean anything, which an earlier post worked through at length, and it is why this site withholds a grade rather than converting a partial observation into a confident score. Aggregate parameters are real and useful. They are not per site verdicts, and dressing them up as one is the failure mode.

  • The domain population is not uniform Supported Two independent projections disagree under a single population model, and the paper presents that disagreement as the measurement.
  • A core fraction near 0.4 at domain granularity Supported, approximate Both fits land on it. Written as an approximation with no interval, so it should be quoted that way.
  • The shell is internally structured Indicated, not resolved A residual disagreement on shell coverage survives the two component fit. The paper leaves the rank resolved treatment to follow up work.
  • Your domain is in the core Not supported No per domain data enters the fit. The paper identifies no members of either component and offers no test for one site.
What arXiv 2607.13636 supports and what it does not, based on its stated method of fitting aggregate published crawl statistics.

Why crawl coverage reaches a model at all

The paper's own motivation for caring about any of this is stated in its related work section, and it is the reason a scanner's blog is covering a physics preprint. Common Crawl is described as the foundational corpus for most open and commercial pretraining datasets, with C4, The Pile, RefinedWeb, RedPajama, Dolma, FineWeb, DCLM and Nemotron-CC named as direct derivatives, and the web portions of several well known model families sitting on that stack. Each derivative competes on filtering: language identification, deduplication, quality classification, perplexity scoring.

Then comes the sentence that does the work. Filter quality, the paper argues, is bounded by fetch coverage, because if a URL is never fetched, no downstream filter can recover it. That is an ordering claim rather than a magnitude claim, and it is the same ordering this site keeps arriving at from the outside: what happens at the crawl layer constrains everything after it, and no amount of care further down the pipeline reaches back to fix a page that was never read. It is also the ordering that generative engine optimization advice tends to skip, spending its attention on the ranking or citation stage while assuming the fetch succeeded.

Two limits keep this from becoming a bigger claim than it is. Presence in a pretraining corpus is not the same as being cited in a live answer, and this blog has published the measurement showing those are different questions: counting citations and measuring how much of a cited page reaches the answer text ran in opposite directions across platforms. Nothing in the crawl coverage literature says a page in the core gets quoted more. It says a page never fetched cannot be in the corpus at all, which is a floor, not a lever.

The second limit is what the archive records when it does arrive. Common Crawl's overview page describes the corpus as containing petabytes of data, regularly collected since 2008, while the paper counts over 90 archives since 2013 at roughly 2 to 3 billion URLs each. What those archives hold of your site is whatever your server returned, because Common Crawl's FAQ states that JavaScript is not executed and cookies are not used, which is the subject an earlier post here covered against the July archive. A client rendered page contributes its loading shell, on whichever rounds it is sampled.

The ordering the paper argues for, from fetch coverage to downstream corpora, as stated in its related work section. A restatement of the paper's argument, not a measured pipeline.

What a site owner can check, given none of this is about one site

Take the paper at its word and most of what it measures is outside any site owner's reach. Coverage is the operator's knob. Survival is what the web supplies. The core fraction is a joint readout of both, fitted across a population, with no per domain resolution yet. There is no configuration change that moves a domain into the core, and anyone selling one is ahead of the published evidence.

What is left is the round itself. Whenever a crawler does arrive, what it takes away is fixed entirely by what the server returns, and that part is auditable from outside in a few minutes. Start with whether the served HTML carries your words rather than a shell, which is what prose parity names and what fetching a page the way a crawler does will answer. The crawler view tool does that fetch, and the stack guides cover the usual causes, including the Next.js case where content arrives only after hydration.

Then check that your rules permit the fetch you think they permit. A robots.txt file is evaluated per user agent, so an answer for one token says nothing about another, which is why the robots.txt tester evaluates a path against each published crawler separately. Two failure modes are worth naming because they are silent. A file that returns the wrong status flips its meaning entirely, since a missing file allows everything and a broken one obliges a crawler to assume complete disallow, which an earlier post traced through the specification. And the file is only one of two layers, because an edge rule can refuse a request the file allows, as documented here before.

Last, be precise about which crawler you are reasoning about. CCBot is one token among the fifteen in this scanner's registry, most vendors publish only one token each, and Common Crawl's archive is a training substrate rather than a live answer engine, so a rule aimed at it is not a rule aimed at a chat product. What a scan can tell you is bounded in the same way this paper is bounded, and the methodology page says where those bounds sit. Lantad identifies itself when it fetches, and its own crawler page documents that, because a tool that argues about crawler behaviour should be legible as one.

  • Prose present in the served HTML Fully in your control. Common Crawl's FAQ states its crawler executes no JavaScript, so a client rendered page contributes what the server returned.
  • robots.txt permits the fetch for the token you mean Fully in your control, and evaluated separately per user agent rather than once for all crawlers.
  • robots.txt returns a status that means what you intend Fully in your control. A missing file and a failing file carry opposite meanings under the specification.
  • Edge or firewall rules not refusing the request Fully in your control, and independent of the file. The two layers can disagree without any warning.
  • Coverage per round The operator's lever, described in the paper as tuned through host budgets, centrality thresholds and revisit cadence.
  • Survival of the eligible set Not chosen by anyone. The paper describes it as supplied by the web, and as the rate the urn sees rather than a strict web survival rate.
  • Whether your domain sits in the persistent core Not determinable from this paper or from any external scan. No per domain data enters the fit and no membership test is published.
What is inside a site owner's control and what belongs to the crawl operator, following the parameter split stated in arXiv 2607.13636.

Related

Common questions

What does a persistent core fraction of 0.4 mean for Common Crawl?

It is a fitted model parameter, not a count of websites. In arXiv 2607.13636, submitted 15 July 2026, the domain population of 51 Common Crawl archives from 2020 to 2025 is modelled as a persistent core that is fully sampled and does not churn, plus a shell with its own survival and coverage. Both of the paper's independent projections land on a core fraction of about 0.4. The paper writes it as an approximation, gives no interval, and identifies no individual domains as belonging to either component.

Can I find out whether my site is in Common Crawl's persistent core?

Not from this paper. The fit uses only aggregate published statistics, namely per crawl URL counts and rolling window unique URL counts, so no per domain information enters it. The authors describe a rank resolved version, which would attach the split to the operator's centrality rank axis, as follow up work that has not been published. Any tool claiming to report your core membership today is going beyond the available evidence.

Does being in Common Crawl mean an AI model will cite my site?

No, and the two questions are separate. Common Crawl is a pretraining substrate, and the paper's own argument is a floor rather than a lever: if a URL is never fetched, no downstream filter can recover it. Live citation in a chat answer is decided by different systems fetching under different crawler tokens, and published measurement has found citation counts and how much of a cited page reaches the answer text moving in opposite directions across platforms.

What can a site owner actually control here?

What the server returns on the round a crawler does arrive. That means the prose being present in the served HTML rather than assembled by JavaScript, robots.txt permitting the fetch for the specific token you mean, robots.txt returning a status whose meaning matches your intent, and no edge rule refusing a request the file allows. Coverage per round and survival of the eligible set are the operator's and the web's respectively, and the paper is explicit that a site owner picks neither.

See what AI can read on your site

Run a free scan and get a graded report of exactly what AI crawlers can and cannot read, with ranked fixes.