Perfect query, rubbish evidence: why data quality matters more in the age of AI



For years, social and digital analysts have spent an enormous amount of time perfecting the way we search for and collate information. We have debated boolean strings, keyword combinations, exclusions, taxonomies and query design. More recently, that conversation has shifted toward semantic search, natural language queries and AI-powered retrieval as the technology we use day-to-day has become better at finding, connecting and summarising information.
But crafting the search is only one piece of the puzzle –and is useless if the dataset we interrogate is garbage. When the underlying dataset is incomplete, outdated, distorted or simply irrelevant to the audience we are trying to understand, even the most sophisticated search won't produce decision-ready intelligence.
Start with the audience, not the dataset
Before we optimise data retrieval, we have to define the evidence universe that is relevant to the question we are trying to answer.Everything we do as researchers and analysts needs to start with the audience we want to understand. When it comes to internet and social data, that means looking for evidence in the places that actually matter to that audience.
“Social” intelligence has expanded far beyond the major social platforms as audience behaviour has fragmented across news, Substacks, forums, creator content, private and niche communities, podcasts and countless other corners of the internet.
If I am trying to understand what is shaping the reputation of a company amongst both investors and government stakeholders, for example, I need to look for that information in very different places. Only then can I capture evidence that genuinely reflects the voices and opinions of those audiences – and the people who influence them.
It changes the starting point of the analysis. You must first define who you are trying to understand, then determine where they spend their time and who they listen to. Only after these steps can we begin to think about the best way to capture the conversation (e.g, via boolean, semantic or natural language).
AI is only as good as the evidence we give it
Data quality is even more consequential when we position AI more centrally within the analytical workflow. AI can process enormous quantities of information, find patterns and make connections at a speed that we as human analysts will simply never replicate. But fast-tracking retrieval and pattern matching doesn't remove the fundamental requirement for solid, rigorous inputs.
If an AI system is searching a knowledge base containing inaccurate, stale or low-quality information, the output may then just be a highly coherent analysis of poor evidence. It can sound convincing and be technically “correct” in its synthesis of the available information, while still leading you to the wrong conclusion.
As organisations increasingly use agents and LLMs to retrieve and synthesise information, the data and source architecture underneath them needs to be continuously interrogated. We need to make sure we understand:
- What sources can the system access?
- More crucially, what can't it see?
- How frequently is the data refreshed?
- Is original material being prioritised, or are we analysing endless repetitions of the same information?
Not every signal means what it appears to mean
There is a third challenge that is particularly importantfor social and digital analysts: data still requires interpretation.
Online conversation has never been a perfectly clean dataset, but the environment is becoming considerably messier. AI-generated content can reproduce and amplify the same idea across hundreds or thousands of pieces of content in a very short period. Bots and coordinated automated activity can create apparent momentum around a topic. LLMs can surface information because it is structured and easily retrievable rather than because it is particularly authoritative, influential or even written by a real person.
A “spike” is therefore not automatically a trend. Large volumes are not automatically important. And repeated messages and narratives do not necessarily represent genuine human consensus.
AI can identify patterns in language, but it cannot reliably determine the meaning, authenticity and significance of every signal without context and human judgement. As we move through the analysis, we still need to ask:
- Who actually started this?
- Are these real people?
- Is the same exact content being repeated?
- Is this conversation occurring among the audience that matters?
- Does this represent a genuine change in opinion,or simply an amplification mechanism?
Those questions become even more important when AI is doingthe first layer of analysis for us.
There is also a risk that is much more subtle – because AI makes it incredibly easy to retrieve evidence, it can make it unbelievably easy to find data points for almost any argument we want to make. Instead of beginning our research with an open question and allowing the evidence to lead us toward an answer, you may find yourself taking the easy route and starting with the conclusion – then and ask AI to assemble the supporting proof.
That isn't intelligence. It is confirmation bias at machinespeed.
What does this mean for our industry?
I think it means the analyst's job is moving upstream.
Better data retrieval through semantic and natural language search gives analysts the opportunity to spend less time manually finding and cleaning information and more time deciding which information matters, whether it can be trusted and what it actually means.
Ultimately, that changes where the value of analysis sits.
The technology we use to retrieve information will continue to change: boolean will evolve (or perhaps even disappear one day), semantic search will improve and agents will become increasingly capable. But better retrieval does not remove the need for robust and high-quality datasets. Because however sophisticated the query becomes, rubbish evidence still produces rubbish intelligence.
