Cookieless Audiences
Home Database API Docs Pricing Live Demo Taxonomy
Use Cases
Media Planning by Persona Inventory Curation & Deal Packaging Seller-Defined Audiences CDP & Analytics Enrichment ABM Account Profiling
Industries
SSPs DSPs Publishers Agencies Curation Platforms
Company
Contact Login
Try Live Demo
Industry · Curation platforms

The Data Substrate for Curated Marketplaces

A curation business lives or dies on one question: for any Deal ID you sell, can you say — credibly, at scale, across every SSP you touch — who the audience behind those domains is? Cookieless Audience answers it as infrastructure: pre-computed audience segmentation for 102 million domains, every attribute drawn from fixed vocabularies aligned with IAB Audience Taxonomy 1.1, plus a real-time API for page-level segmentation of individual URLs. It is the substrate under a curated marketplace — the layer that turns supply footprints into audience-defined deals, on cookied and cookieless traffic alike.

102Mdomains, head of market to long tail
285 / 283interest and purchase-intent segments
1,667deterministic personas for deal themes
Quarterlyrefresh cadence on refresh plans
Aligned with IAB Audience Taxonomy 1.1 Banded confidence per attribute No PII anywhere in the pipeline Fixed, versioned vocabularies (v1.0)
The curation problem

Deal IDs scale. Audience knowledge usually doesn’t.

Building one good curated deal is a research project: pick a theme, find the domains, justify the audience claim, keep the list current. Building a marketplace of hundreds of deals across tens of thousands of domains — and defending every audience claim to a buyer — is a data problem no ad-ops team can solve by hand.

The Cookieless Audience database makes the knowledge side of curation a query instead of a project. Every domain carries coded demographics (8 age brackets, 6 income bands, 14 life stages, household, employment, urbanicity), 29 interest groups with 285 sub-interests, 34 intent groups with 283 purchase-intent segments, B2B firmographics and one of 1,667 deterministic personas — each attribute with banded confidence. Deal construction becomes boolean logic over those fields, repeatable across every SSP integration you operate.

And because there is no PII and no identifier anywhere in the pipeline, the deals you build work identically on the 40%+ of traffic where Safari, Firefox and iOS block third-party cookies — cookies remain on Chrome, but your product no longer depends on them anywhere.

// A deal definition, as segment logic
// Theme: "Affluent frequent travelers"
{
  "deal_theme": "affluent_frequent_travelers",
  "rules": {
    "purchase_intent": [
      "PI.travel.hotels_and_resorts",
      "PI.travel.air_travel"
    ],
    "income_level": ["high", "affluent"],
    "confidence_min": "medium"
  }
}
// Run against the licensed file -> domain set.
// Attach to Deal IDs in each SSP seat you curate.
What the substrate powers

Three layers of a curation product, one dataset

Curation platforms use the data at catalog level, deal level and buyer level — the same coded attributes flow through all three.

Catalog-scale classification

Join the domain file to the total supply footprint reachable through your SSP and exchange integrations. Every reachable domain gets a consistent audience profile — including the long tail where no publisher declaration or first-party data will ever exist.

Deal construction at scale

Express each marketplace product as attribute logic — personas, INT.*, PI.*, demographics, firmographics — and regenerate domain sets programmatically. A hundred deals become a hundred queries, not a hundred research projects. Mechanics: inventory curation.

Buyer-facing evidence

Every deal ships with its rule set and confidence bands — a one-sheet a buyer can interrogate. Crosswalk attributes to IAB Audience Taxonomy 1.1 nodes and the same logic backs seller-defined audience declarations on the curated supply.

Workflow

From licensed file to live curated marketplace

The substrate is a file you own and query — every downstream artifact regenerates from it.

License the corpus slice

Top 100k or top 1M domains self-serve, vertical and country slices, or 5M up to the full 102M corpus for marketplace-scale operations — see pricing.

Join to reachable supply

Intersect with the domains your SSP and exchange integrations can actually transact, so every deal is buildable on day one.

Define deal logic

Write each product as boolean rules over coded attributes, gated on confidence bands. Version the rules — the vocabularies are fixed at v1.0, so logic stays stable.

Generate & sync deals

Materialize domain sets per deal and push them to Deal IDs across your seats. Regeneration is a re-run, not a rebuild.

Refine with the API

For large multi-topic publishers, use the per-URL API to curate at section level — the finance vertical of a news site can sit in a different deal than its sports vertical.

Substrate numbers

Sized for marketplace operations

102Mdomains in the full corpus — slices from 100k up
568interest + intent segments to build deal themes from
4segtax ID for Audience Taxonomy 1.1 declarations on curated supply
40%+of traffic is cookieless — your deals work there anyway
Worked example

A curation desk builds an “Affluent frequent travelers” deal

A curation platform wants a travel-audience product for luxury-hospitality and airline buyers, spanning every SSP it curates on. The deal is defined once, as attribute logic, and materialized everywhere.

Rule (coded values)Reads asWhy it’s in the logic
PI.travel.hotels_and_resorts or PI.travel.air_travelIn-market for hotels, resorts or flightsThe commercial core — intent is what the buyer pays a premium for
INT.travel.adventure_travel (optional boost)Adventure-travel interestSub-theme for an experiential-travel variant of the deal
income_level ∈ {high, affluent}High-income readershipMatches the luxury price point of the advertisers
audience_type = b2cConsumer-facing propertiesExcludes trade and B2B travel media from a consumer deal
confidence ≥ medium, high for the flagship tierEvidence gateTwo deal tiers: broad reach at medium, flagship at high confidence
Outcome: the query yields the qualifying domain set, split into a flagship tier (high confidence) and a reach tier (medium). The desk attaches each tier to Deal IDs in every SSP seat, publishes a one-sheet with the exact rules and confidence gates, and crosswalks the attributes to Audience Taxonomy 1.1 nodes so the curated supply can carry segtax: 4 declarations. Next quarter’s refresh re-scores the corpus; re-running the same logic updates every deal in the marketplace at once.
Comparison

Ways to source a curation substrate

 Build your own classifierPublisher-declared / first-party dataLicensed domain-level dataset (this one)
Time to first dealMonths of crawling, modeling, QAPer-publisher negotiationDays — the file arrives pre-computed
CoverageWhatever you crawl and maintainOnly cooperating publishers; long tail absent102M domains, uniform schema, long tail included
Taxonomy alignmentYours to design and defendVaries per publisherBuilt aligned with IAB Audience Taxonomy 1.1, fixed v1.0 vocabularies
FreshnessYour pipeline, your ops burdenPublisher-dependentQuarterly refresh; per-URL API between refreshes
Cost shapeEngineering headcount, ongoingRevenue shares, integrationsOne-time license + quarterly refresh; instant-buy tiers from the pricing page
Scope note: the substrate informs deal construction and audience evidence. It is not an impression-level pre-bid classifier — Deal ID targeting, packaging and any SDA declarations execute in your own and your SSP partners’ systems. Full field reference: audience segmentation taxonomy.
FAQ

Curation platforms — common questions

How much of the corpus does a curation platform actually need?

Most desks start with the top 1M domains — available as an instant card purchase with immediate download — because it covers the supply that transacts meaningfully in most markets. Platforms curating deep long-tail or international supply license 5M up to the full 102M corpus under a custom agreement. Vertical and country slices exist for specialist marketplaces. Details are on the pricing page.

Can deals be finer-grained than whole domains?

Yes. The domain file is the substrate for property-level deals; the real-time API applies the same v1.0 vocabularies to individual URLs, so you can curate sections of large publishers separately — putting a news site’s personal-finance vertical in an investing deal without dragging in its celebrity coverage. Domain-level and URL-level outputs share one schema, so mixed-granularity deals stay coherent.

How do buyers verify the audience claims behind a deal?

Every attribute is a coded value from a fixed, published vocabulary with a banded confidence score, so a deal one-sheet can state its exact inclusion logic — not a vague audience description. Buyers holding their own copy of the data can re-run the logic and check the domain list themselves. That inspectability is a sales asset: curated products survive procurement when the evidence is reproducible.

Does this classify impressions in the bidstream for our deals?

No. The dataset is planning and construction infrastructure: it determines which domains belong in a deal and documents why. Once a Deal ID exists, targeting and delivery run entirely in the SSPs and DSPs transacting it. We deliberately make no impression-level pre-bid claims — the product is the substrate, not the auction.

Related pages

Build your next hundred deals from one file

Test the attribute depth on any domain in the audience demo, then license the slice of the corpus your marketplace needs.

Open the audience demo See database pricing
Stay in the loop

You are on the list!

We will send you updates that matter — no spam.