Cookieless Audiences
Home Database API Docs Pricing Live Demo Taxonomy
Use Cases
Media Planning by Persona Inventory Curation & Deal Packaging Seller-Defined Audiences CDP & Analytics Enrichment ABM Account Profiling
Industries
SSPs DSPs Publishers Agencies Curation Platforms
Company
Contact Login
Try Live Demo
Industry · Consultancies

Audience Evidence That Survives the Data Room

Strategy and diligence engagements run on evidence that has to hold up — to a partner review, to a client's board, to the other side's advisors. Cookieless Audience gives consulting teams a deterministic audience profile for any of 102 million domains: coded demographics, 285 sub-interests, 283 purchase-intent segments, B2B firmographics and 1,667 personas, every value drawn from fixed v1.0 vocabularies aligned with IAB Audience Taxonomy 1.1. The same query re-run by anyone, on the same file, returns the same figure — which is what “auditable” actually means. And it is available by instant download, on engagement timelines, not procurement timelines.

102MProfiled domains
100%Deterministic classification
v1.0Fixed, versioned vocabularies
0PII in the pipeline
The problem

Diligence evidence is usually someone else's homework

When an engagement touches a media property, a content business or anything whose value rests on “who the audience is”, the evidence base gets thin fast. Management presentations assert an audience (“affluent, professional, decision-makers”) that was itself assembled for persuasion. Expert calls give directional color that cannot be quantified or reproduced. Traffic estimates say how many but never who. And survey work — the rigorous option — takes longer than the exclusivity window of most deals and the sprint cadence of most strategy studies.

The structural issue is reproducibility. A diligence finding is only as strong as the answer to “how do you know, and could we check?” Probabilistic audience estimates from tracking-based vendors change between pulls, depend on opaque models, and degrade precisely where measurement is weakest — the roughly 40%+ of traffic on Safari, Firefox and iOS where third-party cookies are already blocked (they remain supported on Chrome). An analysis a client cannot re-run is an opinion with a chart.

Content-derived, deterministic classification changes the character of the evidence. Each domain's profile is computed from what it publishes, every attribute takes a value from a fixed vocabulary — documented in full on the taxonomy page — and carries a banded confidence level (low / medium / high). The workpaper can state the exact rule applied; the reviewer can apply it again and reconcile to the digit.

Engagement types

Where the data earns its place in the workplan

Any workstream that needs a defensible answer to “who does this property, partner or market actually reach?”

Commercial due diligence

Independently verify a target's audience claims — publisher, content network, affiliate portfolio, community — against coded profiles rather than the CIM's own narrative.

Market entry & growth strategy

Map a category's media landscape — who serves which audience segment, where the white space is — from a census-style corpus instead of a hand-built site list.

Marketing & media effectiveness

Audit where a client's spend and partnerships actually land: profile the domains on the media plan and score them against the client's own target-audience definition.

Post-deal value creation

Re-run the diligence queries each quarter on refreshed data to track whether the audience thesis is playing out — same vocabulary, same rules, comparable numbers.

Workflow

A diligence sprint on coded audience data

Built to fit inside a two-to-six-week workstream. No integration project, no panel fieldwork — a file, a join, and a rules pass.

STEP 1

Code the hypothesis

Translate the claim under test into vocabulary terms. “Affluent professional audience” becomes income_level: high|affluent, life_stage: established_professional, b2b_seniority: director+ — testable, not rhetorical.

STEP 2

Acquire the data

Instant-download the top-1M file ($1,990) or a vertical slice ($190–$490) the day the engagement starts. Spot-check specific properties and sections with the real-time API where page-level granularity matters.

STEP 3

Test against evidence

Join the target's domains (or the market's) to the file and apply the coded rule at a stated confidence threshold. Report the strict cut (high confidence only) and the broad cut (medium+) side by side.

STEP 4

File the workpaper

Deliverable cites vocabulary version, selection rule and threshold. Anyone with the same file reproduces every exhibit — the finding survives partner review, client challenge and the other side's advisors.

Worked example

Testing an audience claim in a CDD workstream

Illustrative scenario: a client is acquiring a personal-finance content group. The information memorandum claims “an affluent, investment-active professional audience.” Left, the claim coded as a testable rule; right, the flagship property's profile from the data.

The claim, coded

source: information memorandum, p.14

Rule under test

High / affluent income Established professionals Investment intent present

As vocabulary codes

income_level: high | affluent life_stage: established_professional PI.finance_insurance.stocks_and_investments threshold: medium+ vocab: v1.0

Test

Applied to all 14 domains in the target portfolio; exhibit reports the share of properties satisfying the full rule at the stated threshold, plus each property individually.

Flagship property, from the data

example-wealth-daily.com

Demographics

45–54 high Upper-middle income high Established professional med Postgraduate med

Interests & intent

Personal Investing high Stocks and Investments high Retirement Planning med

Finding for the exhibit

Partially substantiated. Investment intent and professional profile confirm at high/medium confidence — but income resolves to upper_middle, not high | affluent as claimed. Across the portfolio, 9 of 14 domains pass the full rule. Material for pricing the advertising-revenue thesis; the sell-side's “affluent” language overstates the modal audience.

Note what the exhibit does not depend on: no panel extrapolation, no vendor black box, no expert's impression. The client's own analysts can re-run the rule after close — and each quarterly refresh turns the same query into a value-creation tracking metric.

Comparison

Evidence sources in a diligence sprint

Each source answers a different question. The coded corpus is the only one that is simultaneously fast, quantitative and reproducible.

SourceSpeedReproducible?Best used for
Management materials / data roomImmediateAsserted, not testableThe claims to be tested, not the test
Expert interviewsDaysDirectional, unquantifiedContext, hypotheses, sanity checks
Commissioned surveysWeeks–monthsPartially (with full methodology)Demand-side questions when the timeline allows
Traffic / clickstream estimatesImmediateModel-dependent, varies between pullsScale ranking; says how many, not who
Coded domain-level audience dataInstant downloadDeterministic, versioned, confidence-bandedTesting audience claims; mapping markets; quarterly tracking

For firm-wide use — multiple engagement teams querying the corpus, larger extracts from 5M up to the full 102M domains, or feeds into internal benchmarking tools — custom licensing is available (from $15,000/year); contact us for a quote. Single engagements are usually covered by self-serve tiers.

FAQ

Frequently asked questions

Can we get the data inside a deal timeline?

Yes — the top-1M file ($1,990) and the top-100k file ($490) are instant card checkout with immediate download, and vertical or country slices ($190–$490) ship the same way. A team can be joining target domains against coded profiles on day one of the workstream. Only large custom extracts (5M+ domains) go through a quoted process. Individual properties can be checked immediately, for free, in the live demo.

Will the analysis stand up to scrutiny from the other side's advisors?

That is the design goal. Classification is deterministic — the same domain and data version always yield the same profile — and every attribute value comes from a fixed, versioned vocabulary (v1.0, aligned with IAB Audience Taxonomy 1.1) with a banded confidence level. Your workpaper states the selection rule, threshold and vocabulary version; any party with the same file reproduces the exhibit exactly. Disagreement then has to be about the rule, which is a substantive argument, not a data-quality one.

Do we need engineering support to use it?

No. The database is a flat, documented file keyed on domain — a join and a filter in Excel, SQL, R or Python, well within a normal analyst's toolkit. The coded fields (age_bracket, income_level, INT.*, PI.*, personas, firmographics) are enumerated on the taxonomy page. The real-time API (plans from $99/month for 10,000 credits) is there when a workstream needs page-level profiles of specific URLs.

How does this handle the cookieless share of the web?

Profiles are derived from each domain's published content, not from tracking users, so they are identical for Chrome, Safari, Firefox and iOS audiences. Tracking-based estimates under-observe the roughly 40%+ of traffic where third-party cookies are already blocked (Safari, Firefox, iOS — Chrome still supports them); a content-derived profile has no such blind spot, and no PII is processed at any point — which shortens the client's own privacy and procurement review.

Related pages

Adjacent workflows and industries

The same corpus powers the neighboring diligence and research workflows.

Bring reproducible audience evidence to the next engagement

Check any target domain in the live demo now — then license the file that covers the workstream and hand the client an analysis they can re-run themselves.

Stay in the loop

You are on the list!

We will send you updates that matter — no spam.