Skip to main content
JoseePMM
Adobe Employee
Adobe Employee
September 9, 2026

Data Distiller - From Audiences to Insight: Identifying Your Most Valuable Customers 

  • September 9, 2026
  • 0 replies
  • 4 views

From Audiences to Insight: Identifying Your Most Valuable Customers 

If you're running Adobe Real-Time CDP, you've already done the hard part. Your data sources are connected. Identities are stitched. Profiles are unified, and your team can build an audience and activate it across channels in minutes.

So, here's the question worth sitting with: open your audience builder right now. Can you segment on how much each customer is actually worth to your business, ranked, current, and ready to activate today?

Most Real-Time CDP teams can't. And the reason isn't a gap in Real-Time CDP. It's that the attribute you'd need to segment on doesn't exist on the profile yet.

The limits you've probably already hit

Real-Time CDP gives you computed attributes, a genuinely useful way to aggregate event-level data up to the profile. If you need a count of purchases in the last 30 days, that's exactly the right tool.

But computed attributes work within the boundaries of the Profile Store. They operate on a pre-defined set of aggregation functions, over the data available in profile, bounded by your profile data time-to-live (TTL). That's fine for straightforward counts and sums.

It's not enough when the signal you want requires:

  • Historical depth beyond your profile TTL: multi-year purchase behavior, not just the last 90 days
  • Multi-dataset logic: joining transactions, loyalty records, and web behavior that don't all live in profile
  • Custom ranking: sorting your entire base relative to each other, which isn't an aggregation at all
  • Business logic your team defines: weighted scores where spend, frequency, and engagement each count differently

So what happens instead is familiar. An analyst exports data, builds the score in a spreadsheet or an external warehouse, then hands back a CSV. Someone ingests it. The campaign launches against a ranking that was stale before it shipped. Next quarter, the whole thing starts over because it was a favor rather than a process.

Meanwhile the raw ingredients have been sitting in your Experience Platform Data Lake the entire time.

Data Distiller: the data transformation engine behind your audiences

Data Distiller is Experience Platform's engine for scalable data transformation. For a Real-Time CDP team, the simplest way to think about it is this: it derives, from your existing data, the profile attributes your audience builder doesn't have yet.

Take the value-ranking example. Data Distiller runs your scoring logic in SQL against the full Data Lake, not just what's in profile, and sorts your entire customer base into ten ranked tiers, or deciles. Decile 1 is your top 10%; Decile 10 is the bottom. Every customer lands somewhere.

The inputs are yours to define: total spend, order frequency, recency, average order value, loyalty activity, engagement trajectory. Weighted however your business actually thinks about value.

Data Distiller writes the result back as a derived dataset and makes each customer's value rank available as a derived attribute on the profile.

The distinction that matters here: derived attributes aren't computed attributes. Computed attributes operate within profile constraints. Derived attributes can draw on the entire Data Lake, use custom transformation logic, and carry whatever look-back window you decide is right.

Then it's just an audience

This is the part that lands for Real-Time CDP users. Once the value rank is on the profile, it behaves like every other attribute you already segment on. No import step, no CSV, no new tool for marketers to learn. Your team opens the audience builder and selects it.

From there the tiering does real work across channels:

  • Deciles 1–2: early access, concierge treatment, genuine recognition. Retaining this group is worth more than acquiring several new customers.
  • Deciles 3–5: growth offers designed to move customers up a tier. This is where incremental spend does the most work.
  • Deciles 8–10: low-cost nurture, or suppression from expensive paid destinations entirely.

That last one is often the fastest ROI story. A real share of most activation budgets goes to customers who will never return the spend, and destination costs scale with the audiences you push. Knowing precisely who to exclude is worth as much as knowing who to prioritize.

And because the attribute lives on the profile, Journey Optimizer picks it up directly, so the same tiering drives personalized messaging in email, push, SMS, and in-app without anyone rebuilding the logic.

Beyond value ranking

Once Data Distiller is in place, the same pattern unlocks a set of things Real-Time CDP teams routinely ask for:

  • Advanced audience segmentation: author complex, multi-condition audiences in SQL directly on the Data Lake, register them in the Audience workspace, and activate them like any other audience
  • Profile enrichment and attribute engineering: compute and join new attributes such as CLTV, loyalty status, churn likelihood, and engagement scores
  • Attribute extraction for personalization: pull product, behavioral, or demographic attributes forward specifically for Journey Optimizer campaigns
  • Data quality and reduction: filter, deduplicate, and flatten datasets before they map to profile, so you're not carrying noise into the Profile Store
  • Reporting model extension: build custom dashboards and KPIs beyond the standard Real-Time CDP reports
  • Operational auditing: schedule snapshots of profile or audience changes for trend analysis and governance

Several of those have a direct efficiency angle worth flagging to whoever owns your contract: cleaning and minimizing datasets before they reach profile helps you manage what actually lands in the Profile Store, and activating refined audiences means pushing high-value segments rather than broad ones.

It stays accurate on its own

The difference between this and the spreadsheet version isn't the math. It's that this one keeps running.

Transformations can be scheduled on whatever cadence your business needs, and processed incrementally, so only new or changed data is recomputed rather than reprocessing full history every run. The ranking refreshes itself.

Because the logic lives in templated SQL, it's written once, versioned, and reused across teams. When the business decides loyalty engagement should outweigh raw spend, that's an edit to a query, not a new project. And query auditing tracks history and lineage, so "how was this customer scored" has an answer.

"We already have a data warehouse"

Many Real-Time CDP customers do, and it isn't going anywhere. Data Distiller doesn't replace it.

The difference is timing. When a new signal is needed in the final stretch before activation, exporting to an external warehouse, transforming, re-importing, and waiting for profiles to update introduces delay and friction at exactly the wrong moment. Worse, in many organizations, the team that owns the enterprise warehouse isn't the team that owns Experience Platform. Your data team ends up waiting in someone else's queue, behind someone else's priorities, for a transformation they could have owned themselves.

Data Distiller applies the transformation inside Experience Platform, in place, right before activation, within Experience Platform's governance framework and without unnecessary data movement. This closes the last-mile gap and eliminates the process friction that not having Data Distiller may introduce.

Is this you?

Data Distiller isn't required for every Real-Time CDP deployment. If computed attributes and out-of-the-box segmentation cover your use cases, they cover them.

It's worth a conversation if any of this sounds familiar:

  • Your audiences don't reflect true customer behavior, because the underlying attribute doesn't exist
  • Campaigns wait on an analyst to produce a score or a list
  • You're ingesting externally computed attributes back into Experience Platform on a recurring basis
  • The segmentation logic you actually want is too complex for rule-based audience building
  • You're carrying more into the Profile Store than you need because nothing filters it first
  • You find yourself having to export and re-import data to create new attributes regularly due to last-minute changes.

The bottom line

Real-Time CDP is exceptional at activating the profile you give it. Data Distiller ensures the profile contains exactly what you need.

If your team has ever built an audience and quietly known it was a rough approximation of what they actually meant, that gap is the one worth closing.

 

See the full implementation in the decile-based derived datasets use case on Experience League, or the customer lifetime value guide for tracking value signals over time. To scope what this looks like against your data, talk to your Adobe account team.

 
Partners? Access the newest resource to learn more - Data Distiller Turn Customer Data into Activation Ready Intelligence