Cookieless Audiences
Home Database API Docs Pricing Live Demo Taxonomy
Use Cases
Media Planning by Persona Inventory Curation & Deal Packaging Seller-Defined Audiences CDP & Analytics Enrichment ABM Account Profiling
Industries
SSPs DSPs Publishers Agencies Curation Platforms
Company
Contact Login
Try Live Demo
Industry · Market research firms

Market Maps and Category Sizing from Domain-Level Audience Distributions

Market research firms use Cookieless Audience as a census-style input: a pre-computed audience profile for each of 102 million domains, coded in fixed, versioned vocabularies aligned with IAB Audience Taxonomy 1.1. Instead of extrapolating a media landscape from a panel that only resolves the largest sites, you aggregate coded demographics, 285 sub-interests, 283 purchase-intent segments and 1,667 deterministic personas across every domain in a category — and every number in the deliverable traces back to a versioned code a client can audit.

102MDomains in the corpus
v1.0Versioned vocabularies
29 / 285Interest groups / sub-interests
1,667Deterministic personas
The problem

Panels stop where most of the web begins

The standard inputs for a media landscape study — consumer panels, syndicated audience measurement, survey recall — share one structural limit: they only resolve the head of the web. A panel of tens of thousands of people produces stable audience estimates for a few thousand large sites; below that, the sample thins out and the long tail of a category becomes statistically invisible. Yet in most categories the long tail is precisely where the interesting structure lives: the specialist publishers, the niche communities, the emerging players a market map is supposed to surface.

Tracking-based alternatives have their own problem: they depend on third-party cookies and device identifiers, and roughly 40%+ of web traffic — Safari, Firefox and iOS — is already cookieless (Chrome still supports third-party cookies, but privacy regulation keeps adding pressure). A measurement approach anchored to tracking under-observes an increasingly large slice of exactly the audiences a study is meant to describe.

Content-derived audience classification takes the opposite route. Every domain is profiled from what it publishes, using the same fixed vocabularies — 8 age brackets, a 5-point gender skew, 6 income bands, 7 education levels, 14 life stages, household composition, employment, home ownership, urbanicity, plus coded interests (INT.*), purchase intent (PI.*) and B2B firmographics. Coverage is uniform across the head and the tail, no panelists are involved, and no PII enters the pipeline at any point. For a research firm that means one thing above all: the denominator of a category study can finally be the whole category. The full field list is documented on the audience segmentation taxonomy page.

Deliverables

What research teams build with it

The dataset behaves like a structured census of web properties — so it slots into the study types you already sell.

Market maps

Enumerate every domain serving a category, then cluster by persona, interest mix and demographic profile to show who competes for which audience — head to tail, not just the top 50 sites a panel can see.

Category sizing

Count and segment the supply side of a content category: how many properties address a given intent segment, at what confidence, with what demographic tilt — a defensible structural size for the media landscape.

Media landscape studies

Describe how a market's attention is distributed: which audience segments are over-served or under-served by existing publishers, and where white space exists for a client's product or content play.

Tracker studies

Quarterly refreshes on stable v1.0 vocabularies mean the same query re-run next quarter measures the same thing — category drift becomes observable instead of being an artifact of methodology churn.

Workflow

From coded corpus to citable finding

A landscape study on this data is a sequence of reproducible queries, not a sequence of judgment calls. Four steps, whether the universe is a vertical slice or the full corpus.

STEP 1

Define the universe

Select domains by coded attribute — e.g. every domain carrying INT.home_garden.home_improvement or PI.home_garden_services.home_improvement_and_repair — optionally filtered by country slice. The selection rule is itself part of the methodology section.

STEP 2

Aggregate distributions

Cross-tabulate the universe on any field: age_bracket, income_level, life_stage, urbanicity, persona. Confidence bands (low / medium / high) let you report a strict cut and a broad cut side by side.

STEP 3

Map the structure

Cluster domains by attribute similarity to reveal sub-markets, position named players on the map, and quantify white space — audience cells with demand signals but few dedicated properties.

STEP 4

Publish with citations

Every figure cites its vocabulary code and version. A client — or a peer reviewer — can re-run the query on the same file and get the same number. Deep-dive individual properties with the real-time API where a page-level view is needed.

Worked example

Mini landscape: the home-improvement content category

Illustrative scenario: a firm is mapping the home-improvement media landscape for a tools manufacturer. Left, the aggregate distribution across the selected universe; right, one domain from the map, as it appears in the data.

Category aggregate

universe: domains with INT.home_garden.home_improvement (medium+ confidence)

Age distribution of category audience (share of domains, primary bracket)

35–44
34%
45–54
28%
25–34
18%
55–64
13%
Other
7%

Modal attributes across the universe

Owner-occupied Suburban Upper-middle income Family life stages

Codes cited in the methodology

home_ownership: owner urbanicity: suburban income_level: upper_middle life_stage: family_young_children vocab: v1.0

One property on the map

example-renovation-magazine.com

Interests

Home Improvement high Remodeling & Construction high Interior Decorating med

Purchase intent

Home Improvement and Repair high Carpeting and Flooring Services low

Demographics

35–44 high Owner high Suburban med

Reading for the report

This property sits in the category's densest cell — suburban owner-families, 35–44. The map's white space is elsewhere: the aggregate shows an 18% share of domains skewing 25–34, mostly renters (home_ownership: renter), yet few properties serve that cell with dedicated rental-friendly improvement content — a findable, codeable gap the client can act on.

The point of the exercise is not any single profile — it is that both panels of this example come from the same file, same vocabulary, same version. The aggregate and the property-level view reconcile by construction, which is what makes the finding citable rather than anecdotal.

Comparison

Research inputs compared

Domain-level audience data does not replace panels or surveys — it covers the questions they structurally cannot answer, and it is the only one of the four that scales to the full web.

InputStrengthStructural limitRole in a landscape study
Consumer panelsPeople-level behavior on large sitesSample thins below the head; long tail invisibleReach validation for the largest mapped properties
SurveysAttitudes, awareness, stated preferenceRecall-based; cannot enumerate a media landscapeDemand-side context layered onto the map
Clickstream / traffic estimatesVolume ranking of sitesSays how many, not who; tracking-dependentWeighting the map by scale
Coded domain-level audience dataUniform audience attributes across 102M domains, versioned and reproducibleDescribes properties and their audiences, not individualsThe census frame: universe definition, segmentation, white-space analysis

Licensing follows study scope: the top-100k file ($490) covers head-of-market studies, the top-1M file ($1,990, instant checkout) covers most national landscapes, vertical and country slices run $190–$490, and multi-country or full-corpus programs — 5M up to all 102M domains — are quoted individually via contact. Quarterly refreshes keep trackers on a consistent baseline.

FAQ

Frequently asked questions

How is this different from the panel data we already license?

It answers a different question with different machinery. Panels observe a sample of people and estimate audiences for the sites large enough to register in that sample. This dataset profiles the properties themselves — every one of 102 million domains gets the same coded attribute set derived from its content — so coverage is uniform from the head to the long tail. Most research teams use both: the coded corpus as the census frame that defines and segments the universe, panel data to validate reach on the largest properties within it.

Can we cite this data in client deliverables and published studies?

Yes, and it is built for exactly that. Every attribute value comes from a fixed, versioned vocabulary (v1.0, aligned with IAB Audience Taxonomy 1.1), every attribute carries a confidence band, and classification is deterministic — the same input yields the same output. A methodology section can state the universe-selection rule, the vocabulary version and the confidence threshold, and anyone with the same file can reproduce every figure. The full vocabulary is public on the taxonomy page.

What license scope does a typical study need?

A single-category study is usually covered by a vertical slice ($190–$490) or the top-1M file ($1,990, instant download), which resolves most nationally relevant properties. Multi-market programs, syndicated products or trackers that need the deeper tail license larger extracts — 5M domains up to the full 102M corpus — under custom terms, quoted individually. Quarterly refresh options ($190–$590 depending on tier) keep longitudinal studies on a stable baseline.

Is there any PII or panel-derived personal data involved?

No. Profiles are derived from the published content of each domain, not from tracking individuals — no cookies, no device IDs, no panelists, no PII anywhere in the pipeline. That also means coverage does not degrade in cookieless environments (Safari, Firefox, iOS — roughly 40%+ of traffic), which increasingly distort tracking-based inputs. For research firms this simplifies both procurement review and the privacy statement in the final report.

Related pages

Adjacent workflows and industries

The same coded corpus supports the neighboring study types — and the industries that commission them.

Put a census frame under your next study

Open the demo and read the coded profile of any domain in your category — then license the slice that covers your universe and run the whole map in one pass.

Stay in the loop

You are on the list!

We will send you updates that matter — no spam.