AI Requires a Strong Data Foundation

ROI Solutions | AI Requires a Strong Data Foundation

Excitement around AI-powered analytics is everywhere.

AI copilots, analytics agents, natural language dashboards, and automated insights are transforming how organizations interact with data. The vision is compelling: instead of learning where data lives, understanding a reporting tool, or waiting for an analyst, someone can simply ask a question and receive an answer.

That’s the dream, right? 

Recently, Anthropic, the company behind Claude, put the dream to the test by asking its own LLMs business analysis questions about its raw business data.

The accuracy rate – only 21%!

So where did things go wrong? It wasn’t the LLM. It was the data.

The AI could absolutely do the analysis it was being asked to do. Anthropic concluded that “the complexity mainly lies in the ambiguity of the data.” How Anthropic enables self-service data analytics with Claude.

That observation should sound familiar to anyone who works with nonprofit data. It also connects directly to a conversation we’ve been having about shared language, connected data, and what organizations need to have in place before AI can deliver on its promise.

The good news is that by the end of the experiment, Anthropic raised its accuracy rate to 95%-99%! What made the difference? Let’s dig in a bit.

The Real Problem Isn’t Access to Data

“I just need to see all my data in one place!” This has been the mantra for years. Get my data in one place, and I can query it, build dashboards on it, and all my problems will be solved. Right?

Consider a seemingly simple question: How many active donors did we have last quarter?

Before we even think about the data sources involved in answering this question, we must first have a standard, shared definition in our organization of what an active donor means. 

Is it based on a single gift in a certain time frame? Should the time frame be absolute or rolling? What’s the minimum amount? Is it a single gift or a cumulative amount? Do we calculate it for an individual or a household? How do soft credits factor in? Are there certain types of gifts we shouldn’t include in our calculations? 

This isn’t a new problem, and we have been dealing with it as humans working with data since we started putting ones and zeros into machines.

Recently, my colleague Derek Drockelman wrote about this in Why Shared Language Is the Foundation for Audience-First Organizations. He argued that connecting technology isn’t enough if the people using it haven’t agreed on what their information means. Terms like engaged, active donor, major donor, and lapsed can sound universal while carrying very different definitions across departments.

AI doesn’t make this semantic ambiguity disappear. If anything, it makes defining them even more important. A traditional report built around an ambiguous definition might give someone a questionable number. An AI analytics environment can make that same questionable definition available to anyone who asks, returning an answer quickly and confidently, without anyone really digging in to verify. 

The easier we make it to ask questions of our data, the more important it becomes that we’ve done the work of determining what that data means.

More Data, More Ambiguity

These settled metric and semantic definitions are just the beginning. Let’s say we settled on the definition of an active donor as something like, “Any individual in a household who has donated at least one single $15 donation to our 501c(3) in the current or prior fiscal year.”

Now let’s say we have data from our CRM solution with in-house major gifts data, data from our direct mail program in another database, and, of course, the online donation forms on two different systems.

Put all that data, in its raw form (schemas, field names), from the source systems together, connect it to an LLM, and throw the above definition at it, and you can only imagine the results. 

How does it know who an individual is and what household they belong to? The same person can definitely be in multiple source systems. Where does it look to determine if the donation is a 501c(3) vs. 501 (c) (4) vs. PAC? What’s the definition of a fiscal year?

In Maybe We Should Stop Calling It a Donor Journey, Derek explored changes when nonprofits can see more of a constituent’s relationship with the organization. A donor record might tell us someone made a $50 gift. Connected data might tell us the same person had been consuming content, attending events, volunteering, advocating, or otherwise engaging with the organization for years before that gift arrived.

That’s enormously valuable context. But bringing more data together also means reconciling more definitions, structures, identities, and relationships.

Connecting those sources allows us to see more of the relationship. It doesn’t automatically mean we understand everything we’re seeing. AI inherits, and can even amplify, that challenge.

Data Foundations Control the Ambiguity

This is what I found particularly interesting about Anthropic’s experience. The answer wasn’t simply more data or better LLMs.

To go from 21% accuracy to 95-99%, Anthropic put a data foundations layer between the raw business data and the LLM. I like to think of the data foundation as a magic decoder ring that takes standard business definitions (“active donor”) and directs the model to the right bits of data and logic to produce the right answer. 

Otherwise, AI agents don’t even know where to look!

“By far the most common failure is that the agent can’t map a concept (“revenue for product X”) to the single correct table, column, and metric definition, usually because there are multiple plausible candidates with subtly different implementations. The fix is fewer, more heavily governed logical models… The goal is that when an agent searches for a concept, it finds a single governed answer”

That’s exactly what ROI Solutions is doing with the Common Data Model in Unite Analytics, our constituent data platform for nonprofit organizations.

Bringing disparate sources together is only part of the problem. Constituent, donation, campaign, interaction, engagement, and other data may originate in systems that represent the same concepts in very different ways. Unite’s Common Data Model creates a consistent analytical structure for that information, while identity resolution determines when records and activity across those systems belong to the same constituent or household. The same individual may appear across multiple platforms under different identifiers, different names, or different records, and AI can get confused unless that ambiguity has already been resolved.

That reduces structural ambiguity: different systems use different schemas, identifiers, and relationships to represent similar information. It also reduces identity ambiguity: uncertainty about whether records across systems represent the same constituent, household, organization, or relationship

Combine the standardized metric and semantic definitions with the Common Data Model, and that’s when things start to sing! Connected data gives us a broader view of the constituent relationship. A common data foundation provides a consistent structure for representing and analyzing those definitions and relationships.

AI can then work from that foundation rather than inferring meaning from disconnected systems.

ROI Solutions | Common Data Model Overview

AI Makes the Foundation More Important, Not Less

There is a certain irony to all of this. We’re entering an era of increasingly sophisticated AI that can interact with enormous amounts of organizational data through natural language. Yet its usefulness may depend on whether organizations have done some decidedly less exciting work underneath.

That’s the larger lesson I take from Anthropic’s experience. The headline-worthy result is 95-98% accuracy, sure. But look underneath those numbers, because there is a lot of deliberate work to reduce ambiguity and give Claude the context it needs to understand the data it’s being asked to analyze so it doesn’t just give you answers, but the right answers.

For nonprofits exploring AI, that distinction matters. The question isn’t simply what AI can do with our data. We also need to ask whether we’ve built the foundation that lets AI understand what our data mean, how our systems relate to one another, and who our data actually represents. Our business definitions differ from the for-profit world, and the data foundation we build for analytics must be designed specifically for our work.

Everyone Wants Branches; We Have to Talk About the Trunk

At ROI Solutions, we have been talking a lot recently about the data platform tree – think of the roots as the messy source system data; tangled, always changing, overlapping, splitting, and growing. The Common Data Model and semantic definitions are the trunk. All of the fun dashboards, queries, and natural language conversations are the branches.  

Everyone is talking about the branches. We’re investing in the trunk.

While much of the industry is understandably focused only on copilots, agents, and conversational analytics, we’re continuing to invest in the Common Data Model and data foundation within Unite Analytics. 

The exciting things we want AI to do depend on the work underneath them.

Because trustworthy AI doesn’t begin with AI. It begins with trustworthy data. Would you like to learn more? Let’s Talk!

Key Takeaways

  • AI analytics are transforming data interactions, but ambiguity in definitions complicates understanding.
  • Shared language is essential; without it, different interpretations can lead to inaccurate insights.
  • Anthropic highlights that strong data foundations enhance AI accuracy; clarity in definitions prevents ambiguity.
  • Unite Analytics’ Common Data Model provides a consistent structure, but semantic clarity remains crucial.
  • Trustworthy AI relies on solid data foundations, emphasizing that effective technology builds on clear definitions and governance.
Get the latest ROI Solutions 
News & Insights