How I approached cross-company data strategy to support a major grocery integration
The situation
Combined customer data was the prerequisite for everything on the post-close roadmap. The prerequisite itself was ambiguous, contested, and politically ownerless.
- The business problem. Combined data was the unlock for cross-company loyalty, personalization, and alternate-profit revenue. The pressure to promise fast value was enormous, and so was the temptation to overclaim what we were ready for.
- The org problem. Customer data lived in a web of legacy systems with multiple sources of truth and fragmented ownership, across two companies that had never worked together and used the same words to mean different things.
- The customer problem. Any combination of data creates collision, because real people already exist in both systems. Get identity resolution wrong and you either fracture someone’s experience by failing to recognize them, or violate their trust by merging them with a stranger.
My leaders asked how we’d merge the data and how long it would take. I answered a different question: how do we make defensible identity decisions when the information is incomplete, the timeline keeps moving, and nobody owns the ambiguity?
What I built
The instinct in the room was to chase the promised value and let the messy details sort themselves out. I went after the mess first, because the mess was the actual product.
Exists in one system, never touched the other.
The mirror case.
One person, two real accounts, identifiers line up. The most upside.
Two different people colliding on a shared or system-generated identifier. Merging them is a trust violation.
One person in both systems, invisible to us because they signed up with different details. We fail to recognize someone we already know.
A merge is high stakes and hard to undo. Every path above requires confidence first, and the customer’s own input wherever we could get it.
How I operated
- I matched signal to confidence to action. Records would be matched on a tiered set of unique identifiers, corroborated by secondary confidence signals, with each confidence level tied to a specific resolution path. The trade-off I insisted on: no automatic merges on attribute matches. Merging is hard to reverse, so the system required confidence and, wherever possible, the customer’s own input.
- I turned “collision” into a taxonomy people could act on. Genuine matches, false matches, missed matches, and single-company customers each carry different customer impact and compliance risk. Naming them separately gave two companies a shared language for something they couldn’t previously frame consistently.
- I chose honesty over an impressive-looking plan. For every workstream I documented what was decided, what needed a business owner, and what depended on information we didn’t have. Governance, consent methodology, and volume forecasting went on the board as open decisions needing owners. I mapped how the foundation enabled downstream value while pushing back on pressure to promise experiences the data couldn’t support yet.
What happened
The transaction that prompted this work didn’t proceed, so the strategy was never applied in its original context. The value was real anyway.
- It gave a leaderless problem a shared language. Two organizations that couldn’t frame identity resolution consistently could now discuss it, disagree about it, and decide against a common taxonomy.
- It exposed a weak matching foundation. Profiling the existing customer base showed more than a third of records were missing the very fields the match logic depended on, and roughly half had no digital account at all. This was a data-quality problem as much as a merge problem, and both had to be solved at once. It’s exactly why I refused to accept automatic merging on attribute matches.
- It named what couldn’t be known. Collision volume couldn’t be forecast before day one with the data we had. Rather than model a number I couldn’t defend, I documented it as an open decision needing an owner. Leaders planned against a stated unknown instead of false precision.
- The thinking outlived the event. The identity-resolution and collision frameworks now inform how the enterprise approaches combined customer data inside its own systems. Same core problem, and the framework transferred cleanly.
- It set a pattern for working under scrutiny. Separate what you know from what you don’t, refuse to overclaim, and treat customer trust as a constraint you design around.
Who owned what
A small, senior tiger team across two companies. Legal and privacy owned the boundaries, including what could be discussed at all. Data and analytics owned system profiling and match feasibility. Loyalty owned the customer-facing implications. Leadership on both sides owned the go/no-go.
I owned identity resolution and collision strategy end to end: framing the problem, building the frameworks, and driving alignment across all of those groups. My influence was lateral and upward, creating shared language where none existed and converting ambiguity into decisions senior leaders could own.
What I’d do differently
- Push for forecasting sooner. Collision volume was a persistent unknown that made every downstream plan harder. I’d argue earlier for the analysis to size it, even directionally, so decisions weren’t made in the dark.
- Name the governance owner on day one. No clear governance and consent owner was the single biggest source of drift. I now treat “who owns this decision” as a prerequisite.