In a media or content deal — a publisher roll-up, a newsletter group, an affiliate portfolio, a community business — most of the enterprise value is the audience, and most of what the deal team knows about that audience comes from the seller. Cookieless Audience gives investment teams an independent read: a deterministic, coded profile for any of 102 million domains — demographics, 285 sub-interests, 283 purchase-intent segments, B2B firmographics and 1,667 personas, all from fixed v1.0 vocabularies aligned with IAB Audience Taxonomy 1.1. Instant download, reproducible by anyone on the deal, refreshed quarterly for the hold period.
A content business is bought on a monetization thesis: this audience can carry higher CPMs, more commerce revenue, a subscription tier, a better sponsor roster. Every one of those levers depends on who the audience is — income, life stage, purchase intent, professional composition. Yet the evidence in the data room is the seller's own deck, a traffic chart that says how many but not who, and perhaps a survey the target commissioned about itself.
Independent checks are harder than they look. Panel-based measurement only resolves the largest properties — useless for the mid-tail sites that make up most sponsor-backed roll-ups. Tracking-based estimates are probabilistic, vary between pulls, and systematically under-observe the roughly 40%+ of traffic on Safari, Firefox and iOS where third-party cookies are already blocked (Chrome still supports them). Neither survives the question a smart IC asks: “if we run this again, do we get the same number?”
Content-derived classification does. Each domain is profiled from what it publishes, deterministically: same input, same output, every attribute a value from a fixed, versioned vocabulary with a banded confidence level — the full field list is public on the taxonomy page. The audience section of the IC memo becomes a set of queries anyone on the deal — or the lender's advisor — can re-run and reconcile.
The same coded corpus serves sourcing, diligence, valuation and the hold period — with numbers that stay comparable across all four.
Scan a thesis space — every domain serving a persona or intent segment — to build target lists beyond the obvious names, including the long tail no panel resolves.
Test the CIM's audience claims against coded profiles of the target's domain footprint — property by property, at a stated confidence threshold.
Benchmark the target's audience against comparable properties on identical attributes — does this audience justify the multiple relative to what similar audiences trade for?
Re-run the diligence queries on each quarterly refresh: audience drift, thesis tracking, and evidence for the exit story — on the same vocabulary you bought with.
Runs inside an exclusivity window: the files are instant download, and the analysis is a join plus a rules pass an associate can own.
List the target's domains — flagships, sub-brands, acquired sites — and a comp set of 5–15 similar properties. This is the universe every exhibit will reference.
Join footprint and comps to the database — age_bracket, income_level, INT.*, PI.*, personas, firmographics, each with a confidence band. Use the real-time API for section-level views of key properties.
Code the monetization thesis as attribute rules and score every property against them. Report strict (high-confidence) and broad (medium+) cuts so the IC sees the sensitivity.
Position the target against comps on the attributes the thesis depends on. Exhibits cite vocabulary version, rule and threshold — fully reproducible by any party to the deal.
Illustrative scenario: a PE fund is evaluating an outdoor-recreation content group. The thesis: the audience can support a step-up in gear-commerce and affiliate revenue. Left, the flagship's coded profile; right, the comp benchmark on the thesis-critical attributes.
PI.sporting_goods.outdoor_recreation_equipment PI.travel.camping income_level: upper_middle gender_skew: male_lean vocab: v1.0
| Property | Gear intent | Income band | Rule |
|---|---|---|---|
| Target flagship | High | Upper-middle | Pass |
| Target site 2 of 6 | High | Middle | Pass* |
| Comp A (trade sale 2024) | High | Upper-middle | Pass |
| Comp B (larger, ad-funded) | Medium | Middle | Fail |
| Comp C (general lifestyle) | Low | Middle | Fail |
5 of the target's 6 domains pass the thesis rule — the audience carries the gear-commerce intent the model assumes, and compares favorably with Comp A, the relevant transaction benchmark. The one caveat is coded, not anecdotal: income resolves to upper_middle, not premium — which caps the price-point assumption in the commerce plan, and is exactly the kind of adjustment diligence exists to find.
*Flagged in the exhibit: passes on intent, misses the income leg — reported, not hidden, because the strict and broad cuts are both stated. Post-close, the same rule becomes the monitoring metric: each quarterly refresh answers “is the audience thesis still true?” with the same query.
The usual sources are not wrong — they are partial. The coded corpus is the piece that makes the audience section of the memo testable.
| Deal question | Traditional answer | Weakness | With coded audience data |
|---|---|---|---|
| Who is the audience? | CIM narrative, target's own survey | Seller-produced; built to persuade | Independent coded profile per domain, confidence-banded |
| Is the audience claim true? | Expert calls, spot checks | Directional; not quantified or reproducible | Claim coded as a rule and tested across the footprint |
| How does it compare with comps? | Traffic rankings, multiples tables | Volume-only; audience quality absent | Target and comps scored on identical attributes |
| Is the thesis on track post-close? | Revenue actuals (lagging) | Explains outcomes only after they happen | Quarterly re-run of the diligence rule — a leading indicator |
Licensing matches deal cadence: single-deal work is covered by the top-1M file ($1,990, instant checkout), the top-100k file ($490) or a vertical slice ($190–$490). Funds running a screening program across a thesis space, or monitoring a portfolio, take larger extracts — 5M up to the full 102M corpus — under custom licensing (from $15,000/year); contact us for a quote.
Yes. The self-serve files are instant card checkout with immediate download — a deal team can have the target's footprint and a comp set profiled on the first day of confirmatory diligence. The analysis itself is a domain join plus attribute filters, sized for an associate with Excel or SQL rather than a data-engineering workstream. Individual domains can be checked immediately in the live demo.
Indirectly but concretely: the monetization levers in the model — premium CPMs, commerce conversion, subscription propensity, sponsorship rates — each presuppose audience attributes (income bands, purchase-intent segments, professional composition). Coding those presuppositions as rules and testing them across the footprint tells you which revenue lines rest on audience facts and which on audience hopes. Benchmarking against comps on the same attributes then disciplines the multiple. See audience-based domain valuation for the pattern in detail.
Classification is deterministic and the vocabularies are fixed and versioned (v1.0, aligned with IAB Audience Taxonomy 1.1), so any co-investor, lender or advisor with the same file reproduces every exhibit exactly — the discussion moves to the rule you chose, not the data you got. There is no panel sampling to challenge and no probabilistic model to dispute, and no PII is processed anywhere, which keeps the workstream clean for compliance review.
That is precisely where content-derived profiles hold their value. Roughly 40%+ of web traffic — Safari, Firefox, iOS — already blocks third-party cookies (Chrome still supports them), so tracking-based audience estimates under-observe those users and understate exactly the mobile, often affluent audiences many content deals are priced on. A profile computed from the domain's content applies to all visitors identically, whatever browser they arrive on.
The same corpus, applied to the neighboring questions a deal raises.
Profile the target's flagship in the live demo now — then license the file and run the footprint, the comps and the thesis rule before the IC meets.