← All case studies Case study

How I approached cross-company data strategy to support a major grocery integration

Enterprise Data Strategy & Cross-Functional Product Leadership · Kroger

The situation

Combined customer data was the prerequisite for everything on the post-close roadmap. The prerequisite itself was ambiguous, contested, and politically ownerless.

  • The business problem. Combined data was the unlock for cross-company loyalty, personalization, and alternate-profit revenue. The pressure to promise fast value was enormous, and so was the temptation to overclaim what we were ready for.
  • The org problem. Customer data lived in a web of legacy systems with multiple sources of truth and fragmented ownership, across two companies that had never worked together and used the same words to mean different things.
  • The customer problem. Any combination of data creates collision, because real people already exist in both systems. Get identity resolution wrong and you either fracture someone’s experience by failing to recognize them, or violate their trust by merging them with a stranger.
The real question

My leaders asked how we’d merge the data and how long it would take. I answered a different question: how do we make defensible identity decisions when the information is incomplete, the timeline keeps moving, and nobody owns the ambiguity?

What I built

The instinct in the room was to chase the promised value and let the messy details sort themselves out. I went after the mess first, because the mess was the actual product.

Identity resolution
Five kinds of collision, and what each one costs
No overlap
Company A only

Exists in one system, never touched the other.

What we doNothing needed.
No overlap
Company B only

The mirror case.

What we doLoad it in, leave it alone.
Overlap we can see
Genuine match

One person, two real accounts, identifiers line up. The most upside.

What we doPrompt an active merge — or an active choice to stay separate.
Overlap that isn’t real
False match

Two different people colliding on a shared or system-generated identifier. Merging them is a trust violation.

What we doBuild confidence, then split. The compliance-critical case.
Overlap we can’t see
Missed match

One person in both systems, invisible to us because they signed up with different details. We fail to recognize someone we already know.

What we doSelf-solicited passive match, customer-initiated.
The line that held
Never merge on matching attributes alone

A merge is high stakes and hard to undo. Every path above requires confidence first, and the customer’s own input wherever we could get it.

Before this, “collision” was one word standing in for five problems with five different risk profiles. Splitting them apart gave two organizations something they could actually disagree about, and then decide. Volume stayed unknown — flagged as an open decision needing an owner, not dressed up as solved.

How I operated

  • I matched signal to confidence to action. Records would be matched on a tiered set of unique identifiers, corroborated by secondary confidence signals, with each confidence level tied to a specific resolution path. The trade-off I insisted on: no automatic merges on attribute matches. Merging is hard to reverse, so the system required confidence and, wherever possible, the customer’s own input.
  • I turned “collision” into a taxonomy people could act on. Genuine matches, false matches, missed matches, and single-company customers each carry different customer impact and compliance risk. Naming them separately gave two companies a shared language for something they couldn’t previously frame consistently.
  • I chose honesty over an impressive-looking plan. For every workstream I documented what was decided, what needed a business owner, and what depended on information we didn’t have. Governance, consent methodology, and volume forecasting went on the board as open decisions needing owners. I mapped how the foundation enabled downstream value while pushing back on pressure to promise experiences the data couldn’t support yet.

What happened

The transaction that prompted this work didn’t proceed, so the strategy was never applied in its original context. The value was real anyway.

  • It gave a leaderless problem a shared language. Two organizations that couldn’t frame identity resolution consistently could now discuss it, disagree about it, and decide against a common taxonomy.
  • It exposed a weak matching foundation. Profiling the existing customer base showed more than a third of records were missing the very fields the match logic depended on, and roughly half had no digital account at all. This was a data-quality problem as much as a merge problem, and both had to be solved at once. It’s exactly why I refused to accept automatic merging on attribute matches.
  • It named what couldn’t be known. Collision volume couldn’t be forecast before day one with the data we had. Rather than model a number I couldn’t defend, I documented it as an open decision needing an owner. Leaders planned against a stated unknown instead of false precision.
  • The thinking outlived the event. The identity-resolution and collision frameworks now inform how the enterprise approaches combined customer data inside its own systems. Same core problem, and the framework transferred cleanly.
  • It set a pattern for working under scrutiny. Separate what you know from what you don’t, refuse to overclaim, and treat customer trust as a constraint you design around.

Who owned what

A small, senior tiger team across two companies. Legal and privacy owned the boundaries, including what could be discussed at all. Data and analytics owned system profiling and match feasibility. Loyalty owned the customer-facing implications. Leadership on both sides owned the go/no-go.

I owned identity resolution and collision strategy end to end: framing the problem, building the frameworks, and driving alignment across all of those groups. My influence was lateral and upward, creating shared language where none existed and converting ambiguity into decisions senior leaders could own.

What I’d do differently

  • Push for forecasting sooner. Collision volume was a persistent unknown that made every downstream plan harder. I’d argue earlier for the analysis to size it, even directionally, so decisions weren’t made in the dark.
  • Name the governance owner on day one. No clear governance and consent owner was the single biggest source of drift. I now treat “who owns this decision” as a prerequisite.