A client forwarded me a LinkedIn post last month with the subject line “are we tracking this?” The post was about harmonic centrality, framed as a metric quietly deciding which brands show up in AI answers. He runs a mid-size SaaS company, has a marketing team of three, and was ready to brief them on a new ranking factor by Monday.
I had to tell him the truth: harmonic centrality is not new. It is not even new to SEO. What is new is that Common Crawl, the archive a large share of AI models train on, uses it to decide how often your site gets crawled. That is a real and specific reason to care. It is just not the reason most of the posts he was reading made it sound like.

Ai Search
Key Takeaways
- Harmonic centrality measures how close a page or domain sits to the rest of the web’s link graph, counted in link hops rather than link authority.
- Network scientists Massimo Marchiori and Vito Latora introduced the metric in 2000, and SEO writers were already debating its relevance to Google back in 2019.
- Common Crawl now uses harmonic centrality to decide which domains get crawled most often, and Common Crawl supplies a large share of the training data behind today’s large language models.
- A site’s harmonic centrality score correlates strongly with PageRank at the very top of the web, but the two diverge for everyone outside the top 100 domains, which is where almost every Wild Creek client actually sits.
- The fix is not a new tactic. It is the same structural work that has always mattered: earning links from well-connected sites and building an internal link structure that puts your important pages within a few clicks of everything else.
Here is why that gap between “new metric” and “new stakes” actually matters for your traffic and your AI visibility.
What Is Harmonic Centrality?
Harmonic centrality is a graph theory measure of how easily one node in a network can be reached from every other node, calculated by summing the inverse of the shortest-path distance between that node and each of the others. Applied to the web, a domain’s harmonic centrality score describes how many link hops separate it from the rest of the internet. A page three hops from most of the web scores higher than a page nine hops away, regardless of how prestigious the sites doing the linking are.
Harmonic centrality is a network science metric that scores how structurally close a node is to every other node in a graph, based on shortest-path link distance rather than link authority. Applied to the web, it measures how discoverable a domain is through link traversal, independent of who is linking to it or how much authority those links carry.
That last distinction is the one worth sitting with. Harmonic centrality does not ask who vouches for you. It asks how many clicks it takes to get to you from anywhere else on the web. Those are different questions, and conflating them is where most of the recent coverage goes wrong.
Where Did Harmonic Centrality Actually Come From?
Harmonic centrality was introduced by Massimo Marchiori and Vito Latora in their 2000 paper “Harmony in the Small World”, published in Physica A and later posted to arXiv, as a fix for a real problem in closeness centrality: on a network with disconnected clusters, closeness centrality breaks down, because averaging distances to unreachable nodes produces nonsensical results. Harmonic centrality sidesteps that by summing the inverse distances instead of inverting the average, so an unreachable node simply contributes zero rather than corrupting the whole score.
The measure got its mathematical rigor over a decade later, when Paolo Boldi and Sebastiano Vigna of the University of Milan formalized it alongside PageRank and other centrality measures in their 2014 paper “Axioms for Centrality”, published in Internet Mathematics. That is not a new paper by AI-hype standards. It predates most of the SEO industry’s current interest in AI search by ten years.
And SEO writers had already found it before that hype cycle started. In January 2019, Aysun Akarsu wrote a piece for Search Engine Journal asking whether harmonic centrality could be “the new PageRank,” walking through the same math, the same comparison, and largely the same conclusion this article reaches. That was seven years before Common Crawl’s own team gave the metric fresh relevance in a blog post in January 2026.
“The math didn’t change. The audience did. Harmonic centrality has been sitting in network science papers and one 2019 SEO article for two decades. What changed is that Common Crawl started publishing it, and Common Crawl feeds the models everyone suddenly cares about impressing.”
None of this makes the metric irrelevant. It makes it exactly what the “old wine, new bottle” framing suggests: a genuinely useful, decades-old tool that has resurfaced because the thing consuming its output changed, not because the tool itself did.
How Is Harmonic Centrality Different From PageRank?
PageRank is the algorithm most people already have some intuition for. It estimates a page’s importance by modeling how link authority flows through the web’s link graph, weighting each incoming link by how much authority the linking page itself has. Harmonic centrality ignores authority entirely and asks a narrower structural question: how many hops does it take to reach this page from the rest of the web.
| Question | PageRank | Harmonic Centrality |
|---|---|---|
| What does it measure | Authority flow through weighted links | Structural closeness, measured in link hops |
| Computation | Iterative, requires the graph to converge | Non-iterative shortest-path calculation |
| Handles disconnected sections of the web | Can distort scores in sparse or isolated clusters | Unreachable nodes simply contribute zero |
| Resistance to manipulation | Vulnerable to link farms and reciprocal schemes | Harder to game because hop count matters more than link volume |
The table explains what each metric measures. What matters more is how closely the two actually agree once you leave the very top of the web.
In an 87-million-domain Common Crawl graph analysis published by Search Engine Journal’s Aysun Akarsu in January 2019, PageRank and harmonic centrality rankings correlate at 0.948 among the top 100 domains, but that correlation falls to 0.317 once the comparison widens to the top 100,000 domains. A separate run on the smaller Stanford Web Graph dataset in the same analysis showed the same pattern: strong agreement near the top, near-zero agreement further down.
The two metrics agree closely at the very top of the web, the Wikipedias and Googles of the world, but they tell increasingly different stories the further down the rankings you go. If your business is not competing with Wikipedia, and it is not, that divergence is the part that applies to you.
Why Is Everyone Talking About Harmonic Centrality Again?
Everyone is talking about harmonic centrality again because Common Crawl, the open web archive that supplies a large share of AI training data, uses the metric to decide which sites it crawls most often, and crawl frequency shapes how much of your content ends up in the datasets large language models are trained on.
Common Crawl uses harmonic centrality to determine crawl priority, and “sites with higher HC scores are crawled more frequently,” according to Common Crawl’s own blog, published in January 2026. More frequent crawling means more appearances in the monthly web archive snapshots that AI models draw from.
A February 2024 Mozilla Foundation report that analyzed 47 large language models released between 2019 and October 2023 found that “at least 64% of these models (30) used at least one filtered version of Common Crawl for their pre-training,” and that “Common Crawl made up more than 80% of the tokens in OpenAI’s GPT-3.”
That is the actual chain worth understanding, and it runs through several links, not one: a higher harmonic centrality score gets you crawled more often by Common Crawl, more crawling means more of your content lands in the archives that feed AI training pipelines, and more representation in that training data is one input, among many, into whether a model has anything to say about your brand at all. It is a plausible mechanism. It is not proof that improving one metric alone changes whether AI systems cite your brand, and nobody running a credible study has claimed otherwise.
Two caveats worth stating plainly. First, Common Crawl is not the only corpus behind these models. Most labs blend it with licensed data, curated academic sources, and their own proprietary crawls, so a low harmonic centrality score does not automatically mean a model has never encountered your brand, and a high score does not guarantee it has. Second, crawl frequency is necessary but not sufficient. Being crawled often gets your content into the archive. Whether a model actually retains, weights, and surfaces that content in an answer depends on filtering, deduplication, and training decisions Common Crawl has no say in.
Does Harmonic Centrality Actually Affect Your AI Search Visibility?
Harmonic centrality is one structural input into whether AI systems have training data about your brand, but it is not a direct AI-visibility ranking factor you can chase in isolation, because the domains with the highest harmonic centrality scores, such as Wikipedia, Reddit, and YouTube, earned that position through decades of link accumulation that has nothing to do with a deliberate optimization campaign.
I want to be direct about where this sits in your priorities. If you are a founder with a three-person marketing team, or a CMO defending a budget to a board that just read the same LinkedIn post my client did, harmonic centrality is not a lever you pull this quarter and see AI citations jump next quarter. It is a slow, structural signal that responds to the same work good SEO has always required: more people linking to you from sites that are themselves well connected, and an internal structure that does not bury your best pages nine clicks deep.
“A metric with a name you’ve never heard is not the same thing as a strategy you’ve never tried. Most of what harmonic centrality asks for is the same homework I’ve been giving clients since before anyone said the word ‘GEO.'”
Here is roughly where that structural signal sits relative to how mature a site’s AI visibility work actually is:
| Stage | What it looks like | Harmonic centrality implication |
|---|---|---|
| Early-stage AI Inclusion | Site is indexed but rarely crawled by Common Crawl, thin link profile, isolated internal linking | Low HC score, infrequent crawl passes, minimal presence in training snapshots |
| Mid-stage AI Discoverability | Growing backlink profile from relevant sites, cleaner internal link structure, some AI citations appearing | Rising HC score, more frequent crawl inclusion, occasional model recognition |
| Mature AI Visibility | Consistent inbound links from well-connected sites, flat internal architecture, regular AI Overview and assistant citations | Stable, comparatively high HC score, reliable crawl frequency, established presence across training snapshots |
That framework is Wild Creek’s own proprietary way of thinking about AI-era visibility maturity, not a neutral industry standard, and other agencies will draw the stages differently. It is where harmonic centrality fits in our thinking: as one signal inside a larger structural picture, not a standalone dashboard number to obsess over. You can read more about how we apply it to GEO work for clients.
One edge case the maturity table glosses over: a site can be genuinely well connected within a tight niche community, such as a regional trade association or a small professional network, and still show a low harmonic centrality score, because that measure runs against the whole web graph rather than against your specific competitive set. A well-linked site in a small niche is not the same thing as a well-linked site in Common Crawl’s sampling, and the two can diverge.
What Should You Actually Do About Harmonic Centrality?
You cannot see your own harmonic centrality score on demand the way you can check a PageRank estimate or a domain authority number, since Common Crawl does not publish a self-serve lookup tool for it, so the practical response is to work on the two things that actually move it: earning links from sites that are themselves well connected, and shortening the path between your homepage and your most important pages.
A few mistakes I see agencies and in-house teams make once they hear about a metric like this one:
They chase the number instead of the mechanism. There is no dashboard where you type in a domain and get a harmonic centrality score back from Common Crawl, so anyone selling you a “harmonic centrality report” is estimating it from a proxy, usually an open dataset like the one referenced in Common Crawl’s own post. Treat those estimates as directional, not gospel.
They ignore internal linking because it feels less important than backlinks. Internal link structure directly shortens or lengthens the hop count between your pages and the rest of your site’s connections to the wider web. A flat structure where every important page is two or three clicks from the homepage does more for this metric than most people assume.
They treat this as a reason to panic rather than a reason to keep doing the work. This is the “old wine, new bottle” point in practice: the fix for a low harmonic centrality score is the same fix that has applied to SEO for twenty years, earn real links and build a sane site structure. Nothing about AI search changes that underlying mechanic, even though it changes who is now paying attention to it.
This week, run a short structural audit rather than waiting on a harmonic centrality tool to tell you what to do:
- Do your highest-priority pages sit within two or three clicks of your homepage, or are you burying them under category pages nobody navigates through?
- Of your last ten backlinks, how many came from sites that are themselves reasonably well linked, rather than from directories nobody else links to?
- Does your internal linking route visitors and crawlers toward your best content, or toward whatever page happened to get published most recently?
- If a new visitor landed on your least-connected page, how many clicks would it take them to reach your most important one?
None of those four questions require a harmonic centrality score to answer. That is the point.
We built our Human Algorithm approach around exactly this tension: the data tells you a metric matters, and judgment tells you how much of your limited time it actually deserves this quarter. For most of our clients, that answer is “keep building links and fixing internal structure the way you should be anyway,” not “build a new campaign around a 25-year-old graph theory paper.”
Frequently Asked Questions
Is harmonic centrality a Google ranking factor?
Harmonic centrality is not a confirmed Google ranking factor. It is a metric Common Crawl uses to decide crawl priority for its own open web archive, and while Google has its own internal signals for crawl and index decisions, no public source ties harmonic centrality directly to Google’s search ranking algorithm.
How is harmonic centrality different from domain authority?
Harmonic centrality measures structural link distance in hops, while domain authority is a third-party score, developed by Moz, that estimates overall ranking strength from a mix of link quantity and quality signals. The two can move in similar directions since both respond to a stronger link profile, but they are calculated differently and neither one substitutes for the other.
Can I check my website’s harmonic centrality score?
There is no official self-serve tool from Common Crawl to check a live harmonic centrality score for a specific domain. Some third-party tools estimate it from Common Crawl’s published web graph datasets, which is useful directionally but should be treated as an estimate rather than an authoritative figure.
Does improving harmonic centrality guarantee better AI search visibility?
No single metric, including harmonic centrality, guarantees better AI search visibility on its own. It is one structural input into how often a site gets crawled and represented in AI training data, alongside factors such as content quality, entity clarity, and citation-worthy structure, and treating it as the whole strategy would overstate what any one signal can do.
Why is harmonic centrality getting attention now instead of years ago?
Harmonic centrality is getting attention now because Common Crawl, which supplies a large share of the data used to train large language models, published details in January 2026 about using the metric to prioritize crawling, connecting a two-decade-old network science measure to today’s AI search conversation for the first time at scale.
Sources & Further Reading
- Marchiori, M. and Latora, V., “Harmony in the Small World,” Physica A 285 (2000), pp. 539-546. arxiv.org/abs/cond-mat/0008357. Original paper introducing the metric.
- Boldi, P. and Vigna, S., “Axioms for Centrality,” Internet Mathematics 10 (2014), pp. 222-262. arxiv.org/abs/1308.2140. Formalized harmonic centrality’s mathematical properties alongside PageRank.
- Akarsu, A., “Can Harmonic Centrality Be the New PageRank?” Search Engine Journal, January 2019. searchenginejournal.com/harmonic-centrality-pagerank/283985. Source of the 0.948 (top 100) and 0.317 (top 100,000) correlation figures, computed on an 87-million-domain Common Crawl graph.
- Common Crawl, “How SEOs Are Using Common Crawl’s Web Graph Data for AI Ranking Signals,” January 2026. commoncrawl.org/blog/how-seos-are-using-common-crawls-web-graph-data-for-ai-ranking-signals. Source of the crawl-priority mechanism described above.
- Mozilla Foundation, “Training Data for the Price of a Sandwich: Common Crawl’s Impact on Generative AI,” February 2024. mozillafoundation.org/en/research/library/generative-ai-training-data. Source of the 64 percent and 80 percent figures on LLM training data.



