Method · Two passes, 35+ signals, three scores

Read every website. Extract the same signals. Keep the evidence.

Every candidate website is read against your written thesis — first in a fast triage pass, then in a deep extraction pass. Every claim in the deliverable traces back to a sentence a company published about itself.

2-passscreening architecture
15signals per company
3weighted composite scores
100%verbatim, sourced evidence
120M+ classified domains
24.7M business & finance sites
700+ industry categories
300+ organisations run on our data
Evidence for every call
The starting line

Screening begins with the entire category, not a curated list

We operate several AI platforms and maintain large-scale specialized datasets. More than 300 enterprise organisations run on our data — among them one of Europe's largest telecom operators, adtech and cybersecurity platforms, and a leading airline metasearch.

That classification of 120M+ active domains is where every engagement starts. Your thesis maps to categories; we pull every domain in scope. Nobody pre-filtered the list, so nothing was silently missed before you arrived.

Full-universe screening, defined

Starting from every classified domain in a category — not from a database of companies somebody previously found, tagged, and decided was worth keeping.

One real category, end to end

Domains in the Mechanical & Industrial Engineering category, worldwide367,766
Live, operating companies after triage (estimate, full category)~254,000
Eligible independents across ten subverticals (estimate, full category)~50,000

Estimates extrapolate the measured triage rates to the full category. Every drop between rows is logged with a reason — nothing exits silently.

The pipeline

Two passes, five stages, no black box

A fast pass over everything, a deep pass over survivors. That split is what makes reading tens of thousands of websites economically sane — without ever sampling or skipping.

1

Universe pull

Your thesis maps to categories in the 120M+ domain classification. Every domain in scope is pulled — typically tens of thousands.

2

Triage pass

An LLM reads every homepage: operating company or directory? Independent or branch? Which subvertical, really? Dead domains fall away, with reasons.

3

Deep extraction

Survivors get a full crawl — about, team, history, services, certifications, careers — and structured extraction of all 35+ signals (20 baseline + the vertical-specific set).

4

Scoring

Three weighted composite scores rank the universe against your thesis. Group-owned companies zero out of outreach, whatever their fit.

5

Analyst QC

Evidence is re-verified against site text, exclusions are audited, and the package ships: ranked CSV, report, exclusion log.

Pass one — triage reads

Is this an operating company at all — or a directory, a marketplace, a parked domain, a blog?
Is it independent, or visibly a branch, subsidiary, or group member?
Which subvertical does the work actually belong to, keywords aside?
Which country does the page itself say the company operates from?

Pass two — extraction reads

The full site: about pages, team pages, history, services, case studies, certifications, careers.
All 35+ signals, each recorded with a verbatim snippet and its source URL.
Subvertical-specific signals defined for your engagement — accreditations, stamps, authorizations.
A "not visible" record wherever the site is silent — never a guess.
Why not one deep pass over everything? Because most of a raw category is noise — directories, distributors' microsites, dead pages. Triage spends seconds on those so extraction can spend real attention on companies that might actually matter to you.
One sample run, in numbers

The scale the method is built to survive

The Mechanical & Industrial Engineering category, the same run the sample report is drawn from — full-category figures, with downstream numbers extrapolated from measured triage rates.

0
Domains in the category, worldwide
~0
Live operating companies (estimate, full category)
~0
Eligible independents, ten subverticals (estimate, full category)
0
Signals evidence-backed, every company
Ranked CSV, all signals Sample report Verbatim evidence log Documented exclusion log Custom ICP re-runs
The framework

The 15-signal framework

Every signal is extracted only from what a company publishes about itself. Where the site is silent, we record "not visible" — never an estimate, never an inference from a name or a photo.

01

Founder-led / family-led association

Reads: explicit founder or family language — "second generation", "founder-led", "family-owned" — captured as the exact phrase on the page.

Why it matters: roughly half of confirmed industrial fits carry this evidence, and outreach to an identifiable owner is a different conversation entirely.

02

Visible leadership bench depth

Reads: how many principals are actually named on the site, and whether the visible bench extends beyond one or two people.

Why it matters: a shallow published bench shapes transition context; a deep one signals professionalized management worth meeting.

03

Operating history & continued independence

Reads: stated founding year, years in operation, and independence language, taken from the company's own telling of its history.

Why it matters: long-established independents are the population most theses actually target — and the hardest to find in profile databases.

04

Strategic fit to thesis

Reads: subsector, services, and customer types, matched line by line against the buyer's written thesis.

Why it matters: this is the principal signal — it drives the Mandate Fit score that carries 70% of the ranking weight.

05

Geographic & branch footprint

Reads: headquarters, branch locations, and the stated service area — with owned locations distinguished from partner mentions.

Why it matters: footprint decides whether a company is a platform, an add-on, or out of territory before anyone books a call.

06

Service-led vs product-led model

Reads: whether revenue visibly comes from field service, manufacturing, distribution, or a hybrid of them.

Why it matters: two companies with identical keywords can have opposite business models — this signal is how a thesis avoids buying the wrong one.

07

Recurring-offering indicators

Reads: service contracts, maintenance agreements, scheduled programs, consumables — the published language of repeat revenue.

Why it matters: recurring language on the site is the closest website-visible proxy for revenue quality a screen can honestly offer.

08

Vertical specialization & end-market exposure

Reads: customer industries evidenced by case studies and named end markets — not guessed from homepage keywords.

Why it matters: documented aerospace, food, or utility exposure is what separates a niche specialist from a generalist with a good copywriter.

09

Acquisition-program / roll-up readiness

Reads: signs the company already belongs to a group or runs its own acquisition program — press pages, "part of the X family" lines, investor tabs.

Why it matters: this is usually a disqualifier. About one in ten keyword-perfect candidates fails here and becomes a documented exclusion.

10

Management professionalization

Reads: visible non-founder functions — finance, operations, HR, marketing roles named on the site.

Why it matters: named second-layer management tells you how much of the company walks out the door with the founder.

11

Hiring posture & functional investment

Reads: open roles and their types, from careers pages and posted listings on the site itself.

Why it matters: what a company hires for is a candid window into growth, capacity constraints, and operational maturity.

12

Website / news activity trajectory

Reads: the most recent dated content anywhere on the site — news posts, project updates, certifications with issue years.

Why it matters: visibly active and visibly dormant companies deserve different outreach priority, and the difference is checkable.

13

Partner & channel ecosystem position

Reads: named OEM partnerships, distributorships, buying groups, and association memberships published on the site.

Why it matters: channel position often is the moat in industrial services — and it transfers, or doesn't, in an acquisition.

14

Compliance & regulated-market readiness

Reads: explicit certifications only — ISO, AS9100, ITAR, ISO/IEC 17025, ASME — captured as the exact claim text.

Why it matters: certifications gate entire end markets. A claim recorded verbatim can be checked; a checkbox in a database cannot.

15

Digital-commercial maturity

Reads: online quoting, customer portals, e-commerce, pricing transparency — how the company sells through its own site.

Why it matters: digital maturity is a cheap, honest indicator of how investable the operation is beyond its machines and vans.

Beyond the fifteen

Subvertical-specific signals, added per engagement

The fifteen common signals are the floor, not the ceiling. Each engagement adds signals that only make sense in your niche — the accreditations, stamps, and authorizations that decide who is real in that trade.

Automation & control panel builders

Industrial Automation & Controls Integrators
UL 508A panel shop listing Named PLC platform partnerships SCADA / HMI integration scope In-house panel fabrication Robotics integration capability In-house controls engineering team Field commissioning & startup services Ongoing service & support contracts Hazardous-area (Class I Div 2) experience End-market focus (water, F&B, pharma)

Calibration & metrology labs

Calibration, Metrology & Industrial Testing
ISO/IEC 17025 accreditation Published accreditation scope Accreditation body named Stated turnaround commitments On-site vs in-lab service mix Number of accredited disciplines Regulated-industry capability (medical, aero) Asset-management software offering Multi-lab / multi-state footprint Recurring calibration-contract language

Boiler, burner & steam services

Boiler & Steam System Services
ASME code stamps (S, R, U) R-stamp repair authorization Combustion tuning programs 24/7 emergency coverage Code-welding capability Service-contract base indicators Rental boiler fleet Water-treatment program tie-ins Multi-state service territory Named boiler OEM partnerships

Coatings & surface finishing

Industrial Coatings & Surface Finishing
NADCAP accreditation AS9100 / aerospace approvals OEM process specifications In-house testing capability Process breadth (anodize, plate, powder) ITAR registration ISO 9001 certification Prototype-to-production range Quoted turnaround commitments Defense / aerospace / medical mix

Equipment dealers & material handling

Material Handling & Lifting Equipment Services
OEM dealer authorizations Factory-trained technicians Rental fleet on site Planned maintenance programs Multi-branch geographic footprint In-house service department Emergency / 24-7 service posture Parts inventory / distribution arm Full-maintenance lease language Industry association memberships

Each added signal follows the same discipline as the core fifteen: exact claim text, source URL, and a "not visible" record when the site doesn't say. The signal list is agreed with you before extraction starts — it is your thesis, operationalized.

The ranking

Three scores, one deliberately lopsided weighting

Fifteen signals collapse into three composite scores. The weighting is intentionally unbalanced: fit does the work, suitability gates the list, and context is capped so it can never flatter a weak match.

70

Mandate Fit — 70%

How closely the company matches the written thesis: subsector, service model, recurring offering, footprint, compliance posture. The score that does the heavy lifting, and the one your ICP re-runs re-weight.

Principal score
20

Outreach Suitability — 20%

Is there an identifiable decision-maker and an independent owner to talk to? Founder or family association, visible principals, and no existing group ownership feed this score.

Gating score
10

Transition Context — 10%

Long operating history, limited visible bench, self-described generational ownership — website-visible context only, and deliberately capped at a tenth of the total.

Capped score

The zero-out rule: a company already owned by a group or consolidator scores zero on outreach, whatever its fit. Roughly one in ten otherwise-perfect candidates fails exactly here — and each becomes a documented exclusion, not a silent omission.

Evidence discipline

Every claim traces to a sentence a company wrote

Deliverables don't say "founder-led: yes". They quote the line, name the page, and link the URL — so your associate can check our work in one click. Below, real entries from the published sample, anonymized for this page.

Target M-01Precision machining
"The company is a fourth-generation, family-owned precision CNC machining company."Industries page · signal: founder / family association
"Certified to AS9100D and registered under ITAR for defense work."Quality page · signals: compliance & certifications, end-market exposure
"Serving aerospace, medical device and industrial customers since 1968."About page · signals: operating history, strategic fit
"Our in-house inspection department operates three coordinate measuring machines."Capabilities page · signal: subvertical-specific capability
16/16 evidence snippets verified against site text
Target C-01Calibration & testing
"The company has been offering 17025-accredited calibration services you can trust since 1982."Homepage · signals: compliance + operating history
"Technicians perform calibrations on-site at your facility or in our climate-controlled lab."Services page · signal: on-site vs in-lab service mix
"Annual calibration agreements keep your instruments on schedule automatically."Services page · signal: recurring-offering indicators
14/14 evidence snippets verified against site text
Target A-01Automation integration
"UL 508A — Industrial Control Panels Certification for USA."Certifications page · signal: subvertical-specific compliance
"Our engineers program all major PLC platforms and build every panel in our own shop."Capabilities page · signals: in-house fabrication, platform partnerships
"A second-generation firm led by the founder\'s sons, with the same project managers for over a decade."About page · signals: founder / family association, bench depth
15/15 evidence snippets verified against site text
Excluded M-01Documented exclusion
"The company is a wholly owned subsidiary of a global industrial group."About page · rule: group ownership zeroes outreach
Excluded with reason — logged, not deleted
The "not visible" rule

If a signal is not on the website, the field says "not visible". It is never estimated, never borrowed from a third-party profile, and never inferred from a name, a photo, or a hunch.

Quality control

What happens between extraction and your inbox

LLM extraction at scale is powerful and fallible in equal measure. The QC pass exists because we assume errors, then go looking for them.

1

Snippet re-verification

Every quoted snippet is matched programmatically against the crawled page text. The verification count ships in the report — "16/16 verified", or honestly, "11/13". A snippet that can't be found is removed, and the signal reverts to "not visible".

2

Analyst review of top ranks and all exclusions

A human analyst reads the top of the ranked list and the complete exclusion log. Misclassified subverticals, mistaken independence calls, and over-generous fits get corrected before anything ships.

3

Zero-out audit

Every company flagged as group-owned is spot-checked against its own site, because a false ownership call costs you a real candidate. The one-in-ten exclusion rate is measured, not assumed.

4

Package assembly

Ranked CSV, sample report — 8 top fits, 5 keyword-missed fits, 5 documented exclusions, 2 insufficient-evidence cases — plus the full evidence and exclusion logs, delivered CRM-ready.

Honest edges

What this method cannot do

A method you can trust has edges you can see. These are ours, stated plainly rather than discovered mid-engagement.

It only sees what companies publish

No financials, no revenue estimates, no valuation guesses. A website shows what an owner chose to say — the method refuses to pretend otherwise, which is why the standards page exists.

It never detects intent

We do not claim to know whether any owner would entertain a conversation. No website signal supports that claim, and vendors who sell it are selling a guess.

Thin sites limit extraction

A two-page website yields two pages of evidence. Those companies are flagged "insufficient evidence" — two per sample report — rather than scored on imagination.

Websites drift

Sites change after we read them. A screen is a snapshot; the monitoring tier exists precisely because universes go stale. Re-screens catch what single passes cannot.

Related reading: the claims we refuse to make and the verticals we decline to serve are published in full on our standards page. The limits are the product.
Method FAQ

Questions deal teams ask about the pipeline

The sample report answers most of these with worked examples — these are the short versions.

How does a written thesis become a screen?
You send the thesis as you'd write it for your IC — subsectors, service models, geographies, must-haves, dealbreakers. We translate it into category scope for the universe pull, prompts for both passes, and the weighting inside Mandate Fit. You review that translation before extraction starts, so the screen tests your thesis rather than our paraphrase of it.
What happens when a signal isn't visible on a site?
The field records "not visible" and the composite scores treat absence as absence — not as a soft negative, and never as an invitation to guess. This costs us tidy-looking completeness and buys you a dataset where every populated field is checkable. In practice, the pattern of what a company chooses not to publish is often informative in itself.
How do you keep an LLM from inventing evidence?
By never trusting it. Every snippet the model attributes to a page is re-matched against the crawled text of that page, and the match rate is printed in the deliverable — 16/16, 14/14, or honestly lower. Snippets that fail matching are dropped and their signals revert to "not visible". Hallucination isn't argued away; it's filtered out mechanically.
Why two passes instead of one deep read of everything?
Cost and attention. A raw category is mostly noise — directories, marketplaces, parked domains, branch pages. Deep-reading all 367,766 domains in the sample category would spend most of the budget on sites that fail in the first sentence. Triage spends seconds each on those; extraction then gives real attention to the minority that survives — an estimated ~254,000 live operating companies across the full category.
Can the signal set be customized for our niche?
Yes — that's standard, not an upgrade. The fifteen common signals stay fixed so results are comparable across runs, and each engagement adds subvertical-specific signals: UL 508A for panel builders, ISO/IEC 17025 scope for calibration labs, ASME stamps for boiler services, NADCAP for finishing, OEM authorizations for dealers. The added signals follow the same evidence rules.
How current is the underlying domain data?
The 120M+ domain classification is updated continuously as part of the data business that 300+ organisations already rely on — it is not rebuilt per engagement. Site content is crawled fresh at extraction time, so signals reflect the website as it stood during your run, and the monitoring tier re-screens quarterly with monthly deltas.

See the method applied to a real category

One email. We send the sample report the same day — every company scored and ranked with full signal transcripts, including the ones we refused to include and why.

Request a Free Pilot