Deterministic Versus Probabilistic Identity Resolution – A Brief Tour

HomeInsightsBlogs | Last Updated March 14, 2022 - by brett crawford under data science & analytics

Published onMarch 14, 2022

Once upon a time, getting to know your customers meant pleasantries exchanged across a counter at the time of sale. In today’s world, the number of potential customer touchpoints, both online and offline, has grown exponentially beyond these humble beginnings. These touchpoints provide behavioral signals that are critical to building a comprehensive picture of your customers and how they interact with your brand. In marketing-speak, this is often referred to as a ‘single’, ‘unified’, or ‘360 degree’ view of the customer.

The jargon however belies the practical difficulties of tying together all a customer’s touchpoints into a single record. This post seeks to answer some common questions around the subject of identity resolution.

What is identity resolution?

Identity resolution is the process of attributing all a customer’s behavior to a single, unified customer profile. Practically speaking, this involves collecting all your customer’s datapoints from applicable online and offline systems, and then merging those datapoints into a singular customer record.

Okay, but why do I care?

The ability to tie together all your customer interactions gives you a more holistic view of how your customers interact with your brand. This in turn allows you to craft more personalized and valuable experiences for your customers.

So, I’ll throw in a ‘join’ statement, what’s the big deal?

While a relatively simple concept, identity resolution is difficult to implement in practice because of the number of possible customer touchpoints, both online and offline. The impending demise of third-party cookies also adds an additional wrinkle.

How do I do it then?

There are two main branches of approaches used to tackle the identity resolution problem: deterministic and probabilistic.

What is deterministic identity resolution?

Deterministic identity resolution uses a set of rules to tie records back to a single customer. These rules are based off facts that you know to be true. Data points that are used in the rules are usually based off first-party personally identifiable information (PII) that is either authenticated by a user or is a unique system generated ID. Examples of this include phone numbers, email addresses, and cookie IDs. Essentially, this functions like a ‘join’ statement if the datapoints and relations between them are clean and well defined. The operative word being “if.”

What is probabilistic identity resolution?

Probabilistic identity resolution refers to the use of predicative algorithms to find the probability that any two records belong to the same individual. This is sometimes referred to as fuzzy matching and is used to match records that are probably related based on an accepted level of statistical confidence. Examples of datapoints used in probabilistic approaches are device operating systems, IP addresses, and time stamps.

But wait, isn’t this just fingerprinting?

Fingerprinting relies on using similar soft signals such as a browser’s version to try and predict a user’s web browser to be unique. This is essentially a question of semantics, but fingerprinting connotes passive probabilistic identification without a user’s knowledge, consent, or control.

Which one is better?

The choice between deterministic and probabilistic approaches is a tradeoff in precision versus reach. Because deterministic approaches make no assumptions or inferences to match records, these yield the highest degree of confidence in making a match. The caveat is that this approach breaks down when data is not clean or PII is not available, which unfortunately can rule out large amounts of data. Probabilistic approaches have the potential to find matches at a greater scale and in messier data than a deterministic approach but sacrifice a degree of accuracy to achieve that scale.

This question ultimately boils down to your use case and what you value more-accuracy or scale.

How can Softcrylic help?

Softcrylic has extensive expertise in stitching customer behaviors. For example, our Analytics Shift product allows you to join online customer data captured in Adobe Analytics with offline enterprise data.

Identity resolution is also a prerequisite for any audience modeling project. Our AI-based audience segmentation tool, AudienceMatch, combines all of your first party demographic and digital behavioral data to provide a holistic view of your customers, giving you the ability to actively target customers with messaging and strategy tailored to their unique characteristics.

Additionally, our Data Science and Analytics team can consult with you to build a custom identity resolution solution to specifically address the needs of your business.

If you have any questions regarding identity resolution for your business, reach out today!

For more information on Data Science & Analytics services fill the form below.

    Brett Crawford

    Brett is a consultant in our Data Science and Analytics practice. He enjoys applied mathematics and using data to solve real-world problems.

    Contact Us

    We're not around right now. But you can send us an email and we'll get back to you, asap.

    Not readable? Change text. captcha txt

    Start typing and press Enter to search