link-infrastructure
PBN Footprint Scanner: A Technical Investigation Playbook
Short answer
A PBN footprint scanner does not look at one domain at a time. That is the first operating principle, before any tool or checklist enters the picture: evidence that a cluster of linking domains is a private blog network lives in the relationship between domains, not in any single site's own metrics. A thin, oddly-registered site is just a low-quality website until you can show it shares a hosting block, an analytics ID, a registrar, or a set of money pages with a dozen other "independent" sites that all happen to link to the same client. This playbook walks through that investigation in order: pull the full list of domains linking to a target, then run hosting, tracking-code, registration, content, and outbound-link checks against that list as a single dataset, so each step either strengthens or narrows the case before you draw a conclusion. It is a hands-on, procedural companion to the broader framework laid out in Backlink Intelligence Explained — this post is the literal, step-by-step version of the "evidence over a single score" idea that pillar describes in the abstract.
What a PBN footprint scanner is actually checking for
A PBN footprint scanner is checking for a network, not a single bad site. Ahrefs' own glossary defines the target plainly: a private blog network (PBN) is "a network of websites created solely to link out to another website and improve its organic search visibility." The operative word is network — a PBN's entire value to whoever built it depends on the sites looking independent while actually being commonly controlled and commonly purposed. That's also the exact behavior Google's own spam policies define as link spam: "the practice of creating links to or from a site primarily for the purpose of manipulating search rankings."
The tell, as a result, is structural rather than visual. A single PBN site, viewed alone, might have a clean design, a plausible author bio, and content that reads fine at a skim. What it can't fake in isolation is its relationship to the rest of its own network — and that relationship only becomes visible once you've pulled every domain linking to the target and are comparing them against each other, not reading them one at a time.
That's also why a PBN footprint scanner is a process more than a piece of software. Automated tools speed up several of the steps below — bulk WHOIS lookups, bulk IP resolution, source-code scraping for tracking IDs — but the underlying logic is a manual investigative discipline first: build the full list, then look for repetition across it. Backlink Intelligence Explained covers why evidence should outrank any single score when judging backlink quality generally; this playbook applies that principle narrowly to one specific, well-documented pattern of link manipulation.
Step 1: Pull the full list of linking domains and treat it as one dataset
Before checking anything else, export every domain currently linking to the target — not a sample, not the ten sites you happen to recognize. The mechanics differ by platform: Ahrefs and Moz both expose a referring-domains report with export built in, DataForSEO's Backlinks API returns the same data programmatically, and vendor-vetting platforms like bklink pull the same list before a listing goes up for evaluation. The specific tool matters less than the discipline: none of the checks in steps 2 through 6 mean anything applied to a single domain. A shared IP address is unremarkable for one site and a strong signal once it repeats across fifteen "unrelated" ones. A shared analytics ID is meaningless in isolation and diagnostic once it shows up twice. Every subsequent step in this playbook depends on having the complete set in front of you first, which is exactly why this is a scanner run across a set, not a lookup run against one page.
Ahrefs' own guidance for spotting suspect referring domains reflects the same set-based logic: check a target's referring-domains report and sort by organic traffic in ascending order, because genuine PBN feeder sites consistently cluster at the bottom of that sort — sites that "generally not look like the kind of site you'd ever visit or read as a user." That sort only works as a filter because you're looking at the whole list at once, ranked against itself, not vetting domains individually as they happen to come up in an outreach inbox or a disavow request.
Step 2: Check for shared IP addresses and hosting infrastructure
With the full domain list in hand, resolve every domain to its hosting IP and look for repetition — either exact IP matches or domains clustered in the same narrow subnet, which is how shared-hosting providers typically allocate addresses to customers on the same plan. Search Engine Land's guide to spotting PBNs puts the mechanics plainly: "having the same IP address doesn't always indicate a PBN. But it's pretty easy to spot a bunch of domains using the same IP address." That first sentence matters as much as the second — shared hosting is also just cheap hosting, and plenty of unrelated small sites end up on the same budget server or the same CDN-fronted range by coincidence, not conspiracy. What separates a coincidence from a footprint is what else clusters alongside the IP overlap: if the same group of domains sharing an IP also links overwhelmingly to the same money site and nothing else of substance, the hosting overlap stops looking incidental.
This is a genuinely old technique. Google's public actions against link networks have specifically targeted operations where, as Search Engine Land's earlier reporting on PBN penalties describes the underlying tell, "the key to identifying a PBN is the cross-site 'footprint' where much of the technical data on the sites are the same" — starting with whether the sites in question are "all on the same IP." That same reporting notes Google's Penguin algorithm "now runs in real time as part of the core ranking algorithm" and can devalue rankings once this kind of scheme is detected, which is precisely why serious PBN operators today spread their sites across different registrars and hosting accounts specifically to avoid this check. A clean result at this step narrows your suspect list; it does not clear it, since a sophisticated network is built to pass this exact test.
Step 3: Check for reused analytics and ad-network IDs in the page source
Pull the raw HTML of every domain in your set and extract any Google Analytics property ID, Google Tag Manager container ID, AdSense publisher ID, or other ad-network identifier present in the source. These identifiers exist to tell an ad or analytics platform which account owns a given pageview, which means they double, unintentionally, as an ownership fingerprint anyone can read directly out of the page source. Search Engine Land's PBN guide describes exactly this mechanic: "when multiple sites use the same tracking codes in their links and HTML source code, it becomes fairly easy to identify those that belong to the same PBN."
This check has been formalized enough that dedicated tools exist for it. HackerTarget's Reverse Analytics Search tool identifies common ownership across sites by searching a crawled dataset for "a common Google Analytics or Adsense ID," on the logic that "linking of Analytics or Adsense identifiers can reveal associated web sites, virtual hosts, and IP address ranges" that would otherwise look unconnected. The tool's own results are pulled from public crawl data — the Alexa Top 1 Million and Common Crawl datasets — which is a reasonable proxy for the same lookup you're running by hand against your own referring-domain list; the goal in both cases is to find the same identifier repeating where it has no legitimate reason to repeat.
Treat this, like hosting, as an asymmetric signal. A shared ID across nominally independent domains is close to conclusive — there's no innocent explanation for two sites that would never otherwise be mentioned in the same sentence sharing an AdSense publisher ID. But the absence of shared IDs proves nothing, since avoiding this exact footprint — a separate analytics account per site, or no tracking at all — is one of the first things a competent PBN operator learns to do.
Step 4: Check WHOIS and registration data for clustering
Pull WHOIS or RDAP records for every domain in the set and compare four fields across the list: registrar, registration (creation) date, nameservers, and registrant organization where it's visible. None of these is meaningful alone — plenty of legitimate, unrelated site owners happen to use the same large registrar. What's meaningful is clustering: a disproportionate share of your candidate list registered through the same registrar within the same narrow window. Domain-investigation guidance on this exact pivoting technique — developed for tracing infrastructure in security investigations, but built on the same reverse-WHOIS logic that applies to any domain cluster — treats tight registration timing as a standard indicator: "domain creation dates within 7 to 30 days of a known incident are used as standard indicators across threat intelligence workflows," and separately notes that "a domain created one to three days before the campaign launched confirms purpose-built infrastructure." The same source flags nameservers as an even more durable pivot than registration timing, since "threat actors change nameservers less frequently than registrant details because DNS changes require propagation time and carry operational risk" — the same operational inertia applies to anyone running a PBN, who has no ongoing reason to touch nameservers once a network is stood up.
Privacy-shielded ownership is a modifier, not a verdict
A meaningful share of all domains, PBN or not, use a WHOIS privacy or proxy service. ICANN's own background material on WHOIS explains that these services exist precisely so a registrant's real contact details aren't published for anyone running a free WHOIS lookup to see, substituting the provider's own information instead — because public disclosure of registrant data "may also expose you to potential spam, phishing, and identity theft." That's a legitimate, common consumer choice, not inherently a red flag on a single domain. It becomes relevant to a PBN investigation only in combination with the rest of the cluster: privacy-shielded ownership repeating across every domain in a candidate set, layered on top of a shared registrar and a tight registration-date window, is a very different fact pattern from one privacy-shielded domain sitting among otherwise-public registrations. Search Engine Land's older reporting on PBN detection treats hidden WHOIS data the same way — as one input among several, not disqualifying on its own, noting plainly that "having hidden WHOIS data is a red flag" alongside the hosting and hidden-ownership checks discussed above.
Step 5: Check content patterns across the cluster
Read the content on every domain in the set with one question in mind: does this page exist to serve a reader, or to hold a link? Look for thin pages with little unique information, content duplicated word-for-word (or lightly reworded) across multiple domains in the cluster, spun text that reads grammatically but says nothing specific, or AI-generated pages with no evidence of editorial review layered on top. Google's own spam policies give this exact pattern a name and a definition — scaled content abuse, described as when "many pages are generated for the primary purpose of manipulating search rankings and not helping users," with "using generative AI tools or other similar tools to generate many pages without adding value for users" listed as a direct example. That policy language is useful here for a specific reason: it describes exactly the evidentiary bar a PBN page needs to clear to count as a real, independent publication rather than a page that exists to hold a link, which is the question this whole step is trying to answer.
Layer an audience-signal check alongside the content read itself: real publications accumulate messy, organic signs of a real audience over time — comments, even sparse ones, social shares, a visible history of posts that don't all point the same direction, the occasional off-topic piece a genuine blogger writes because something caught their interest. A PBN page usually has none of this, because no one was ever meant to read it beyond a crawler. Template and design repetition compounds the same evidence: Search Engine Land's PBN guide notes that networks "often use the same or slightly modified templates and design elements," which becomes a fast visual tell once you're comparing ten or twenty sites side by side in a spreadsheet rather than judging each one's design in isolation.
Step 6: Check outbound link concentration from the suspected network
For every domain in your set, pull its own outbound link profile — not who links to it, but who it links to. The question that matters is concentration: do these "different" sites, ostensibly covering different topics with different authors, funnel their substantive outbound links to the same small handful of money pages, with nothing else of real weight going anywhere else on the web? A genuine independent blog cites a wide, idiosyncratic range of sources over its lifetime, reflecting whatever its author actually read and cared about. A PBN page exists to point at a client, so its outbound profile tends to be almost entirely money-site links plus a thin layer of filler links added to make the page look less obviously purpose-built.
Anchor text is the fastest surface check here. Ahrefs' own glossary flags "unnaturally placed links with exact-match anchor texts" as one of the clearest tells that a link exists to manipulate rankings rather than to help a reader navigate somewhere useful. Google treats this specific pattern — a site's own outbound links, not just its inbound ones — as independently evaluable: Search Console's manual actions documentation describes a dedicated "unnatural links from your site" action, separate from the inbound version, triggered when Google has "detected a pattern of unnatural artificial, deceptive, or manipulative outbound links" originating from a site. That a site's own outbound behavior can draw a manual action in its own right is a useful frame for this step: you're not just evaluating whether these domains deserve to pass authority to your target, you're checking whether their whole outbound pattern is the kind Google's own policies already treat as a violation on its own terms. At scale, this is the same thing search engines do algorithmically — Search Engine Land's guide describes Google modeling the link graph itself to identify closely affiliated sites, which is the automated version of the same cross-domain outbound comparison this step runs by hand.
Building the evidentiary case: why no single step is proof
Every check in this playbook produces a partial signal, and every partial signal has an innocent explanation in isolation. Shared hosting happens by coincidence on cheap infrastructure. Registrar overlap happens because a handful of registrars dominate the market. A privacy-shielded WHOIS record is a normal consumer choice. Thin content is also just bad writing. None of that changes once you're looking at a full set of domains instead of one — but the pattern of overlaps across the set is what stops being explainable by coincidence. Search Engine Land's framing of this is the right one to carry through all six steps: "spotting a PBN isn't about just one red flag; it's about connecting the dots." A cluster that shares an IP block, a registrar, a registration week, an analytics ID, thin near-duplicate content, and a closed outbound loop to the same three money pages isn't six weak signals — it's one strong one, made of six independent confirmations that would each need a separate, unrelated explanation to be innocent.
That's also the reason Ahrefs' own advice on PBNs is as blunt as it is, aimed as much at anyone evaluating whether to buy a link from a suspect network as at anyone building one: "do not use them at all. Doing so is extremely risky and, arguably, unethical" — because once a network gets deindexed, every link inside it stops passing anything at once, and the buyer who never built the network still absorbs that loss. How much of this evidence is actually enough to act on — where the line sits between "worth a second look" and "confirmed," and how to avoid false-positiving a small, legitimately independent publisher that happens to share a hosting provider with a bad neighbor — is a judgment call this playbook deliberately doesn't resolve on its own. That question is the subject of PBN Spam Detection: Footprints, False Positives, and Evidence, which picks up the evidence-standards side of this work exactly where this technical sequence leaves off.
The PBN footprint scanner checklist
Run the six steps in this order and log what each one actually establishes — not just a yes or no, but whether the result is strong enough to stand alone or only meaningful in combination with the rest of the set.
| Step | What you're checking | How to check it | Signal strength alone |
|---|---|---|---|
| 1. Pull the full domain list | Every domain currently linking to the target, treated as one set | Referring-domains export from a backlink data tool or API | Not evidence on its own — the prerequisite for every later step |
| 2. Shared hosting and IP | IP address or subnet overlap across the domain set | Bulk reverse-IP and hosting lookups | Weak alone; cheap shared hosting causes false positives |
| 3. Reused analytics or ad IDs | Repeated Google Analytics, Tag Manager, or AdSense identifiers in page source | View-source check or automated crawl across the set | Strong when found; proves nothing when absent |
| 4. WHOIS and registration clustering | Shared registrar, clustered registration dates, shared nameservers | Bulk WHOIS or RDAP lookups compared across the set | Moderate; privacy-shielded WHOIS alone is not suspicious |
| 5. Content patterns | Thin, duplicated, spun, or AI-generated content with no real audience signals | Manual read plus a duplicate-content check across the set | Moderate; must be read alongside the rest |
| 6. Outbound link concentration | Whether the "different" sites all link to the same small set of money pages | Outbound-link export per domain, compared across the set | Strong when combined with steps 2 through 5 |
Treat the table itself as a record of what was checked and what each check did or didn't establish — not a scorecard where any single yes/no should ever, on its own, decide the outcome.
Where investigations go wrong
- Stopping at the first positive signal. A shared IP address feels conclusive the first time you find it. It isn't, on its own — see step 2.
- Only checking the site you were pitched. PBN evidence is cross-domain by definition; checking one domain in isolation, however thoroughly, will never surface it.
- Reading a clean result on steps 2 through 4 as clearing a network. A sophisticated operator's entire design goal is to pass exactly those checks. Content and outbound-link patterns (steps 5 and 6) are harder to fake at scale precisely because they cost real time and money to fix, and deserve equal weight even when the technical footprint looks clean.
- Treating WHOIS privacy or shared hosting alone as disqualifying. Both have entirely ordinary explanations outside a PBN context, and treating either as a standalone verdict produces false positives against small, legitimate publishers.
Related reading
- Backlink Intelligence Explained — the broader framework this playbook applies: evaluating links as evidence, not as a single score.
- PBN Spam Detection: Footprints, False Positives, and Evidence — the companion piece on how much of this evidence is enough to act on, and how to avoid false positives.
Key takeaways
- PBN evidence is cross-domain: pull the full list of domains linking to a target before checking anything else, then evaluate it as one set.
- A shared IP address or hosting block is a real but weak signal alone - cheap shared hosting produces the same overlap among entirely unrelated sites.
- Reused Google Analytics, Tag Manager, or AdSense IDs across nominally independent sites are close to conclusive when found, but their absence proves nothing.
- WHOIS clustering (same registrar, same registration window, same nameservers) matters as a pattern across many domains, not as a flag on any single one - WHOIS privacy alone is not suspicious.
- Thin, duplicated, spun, or AI-generated content with no real audience signals is what Google's own spam policies define as scaled content abuse.
- The strongest tell is outbound link concentration: whether the different sites in a cluster all funnel links to the same small set of money pages.
- No single check is proof - the case is built by connecting multiple independent signals across the same domain set, not by any one red flag.
Frequently asked questions
What is a PBN footprint scanner?
It is the process of pulling every domain linking to a target and checking that set for cross-domain patterns - shared hosting, reused tracking IDs, WHOIS clustering, content patterns, and outbound link concentration - that indicate common ownership rather than independent endorsement. It is a procedure more than a single tool, since the evidence only appears once domains are compared against each other.
Is a shared IP address alone proof that sites belong to a PBN?
No. Shared or budget hosting regularly puts unrelated sites on the same IP address or subnet, so IP overlap on its own is a weak signal. It becomes meaningful only when it lines up with other patterns in the same domain set, such as those sites also linking overwhelmingly to the same money page.
Does WHOIS privacy automatically mean a domain is hiding something?
No. WHOIS privacy and proxy services are a normal, common consumer choice used by many legitimate site owners specifically to keep personal contact details out of a publicly queryable database. It only becomes relevant to a PBN investigation when privacy-shielded ownership clusters across many domains that also share a registrar and a tight registration window.
Can a well-built PBN avoid every check in this playbook?
It can avoid the technical footprints in steps 2 through 4 - spreading sites across different hosts, registrars, and analytics accounts is straightforward. Content and outbound-link patterns (steps 5 and 6) are harder to fully disguise at scale, since they require ongoing, genuine editorial effort and a naturally diverse set of outbound citations, which is exactly what a link-only site does not have.
What is the fastest single check to run first?
Pulling the complete list of linking domains, since every other step depends on comparing values across that set rather than checking domains one at a time. After that, reused analytics or ad-network IDs tend to be the quickest technical check to run and the hardest to explain away when found.
How many domains need to share a pattern before it is worth investigating further?
There is no fixed threshold - a handful of domains sharing multiple independent signals, such as hosting, registration timing, and outbound concentration together, is more telling than dozens sharing only one weak signal like a common registrar. Running all six steps is about seeing how many independent confirmations stack up, not hitting a specific count.
Where can I read more about how much of this evidence is enough to act on?
That evidence-standard question, including how to avoid false-positiving a small, legitimate publisher, is covered in the companion piece 'PBN Spam Detection: Footprints, False Positives, and Evidence,' which focuses on the judgment calls behind the technical checks in this playbook.
Sources
- 1. Ahrefs SEO Glossary - Private Blog Network (PBN)
- 2. Google Search Central - Spam Policies for Google Web Search
- 3. Search Engine Land - What Are PBNs? Risks, Rewards and SEO Implications Explained
- 4. Search Engine Land - Private Blog Networks: A Great Way to Get Your Site Penalized
- 5. HackerTarget - Reverse Analytics Search
- 6. WhoisFreaks - Mastering WHOIS OSINT for Effective Domain and IP Investigations
- 7. ICANN At-Large - Background: WHOIS
- 8. Google Search Console Help - Manual Actions Report
Part of series
Backlink Intelligence & Infrastructure
Related articles
Backlink Intelligence Software: Category Explained
Backlink intelligence software gets used as a marketing label more often than a defined category, but three capabilities actually set it apart from Ahrefs, Semrush, or Moz: infrastructure and footprint detection built into the product, workflows that verify a link is live right now rather than just once-indexed, and evidence-first reporting that shows the signal behind a score. This piece compares the category directly against mainstream backlink checkers and covers who genuinely needs dedicated intelligence tooling versus who is already well served by a standard audit tool.
Backlink Footprint Audit: Finding Shared Infrastructure Across Sites
Shared hosting alone rarely proves a backlink network is fake, since plenty of legitimate small sites share the same budget host. This guide walks through the four infrastructure signals worth checking across a set of linking domains, and why real evidence only appears once several signals stack on the same cluster.
IP and ASN Resolution for Backlink Investigations
Two backlink domains can use completely different IP addresses and still sit inside the same hosting company's network once those IPs are resolved back to an ASN. This piece covers what an ASN actually is, how IP-to-ASN resolution works in practice, and why a shared ASN means very different things depending on whether the host behind it serves millions of unrelated sites or a few thousand.
