How much can you trust geolocation in social intelligence?



Geolocation has always been one of those areas in social intelligence where the practitioners had to settle for something far less than ideal and had to rely on proxies and inference. The social listening tools rely on and infer geo location on data that’s usually extracted from the bios of the authors. And that is if such information is publicly listed in the author’s profile and this still only works on a limited number of platforms, mainly Twitter/X and later BlueSky and Threads. Platforms like Reddit and Instagram fall under Unknown Geo category due to API restrictions and insufficient meta data fed into the tools.
It is no surprise this approach leaves vast swaths of content, sometimes above 80%, receive the “Unknown” label. Additionally, recent evidence showed geotags and location can be manipulated via VPNs for ill gains too, further complicating the issue and casting doubts on the accuracy of the data the platforms themselves provide. Last, but not least, even the little geo meta data we get from the social media platforms might disappear too as the platforms continue to curb their access to content and data and users demand more privacy.
Modern tech solutions still fail to address this age-old problem
Lately some tool vendors have been trying to enrich the geo location meta data by looking at visual clues in Instagram and TikTok content and location tags in the posts. However, this approach also has some drawbacks and can pull false positives (Americans eating McD in Paris in front of the Eiffel tower for example).
@justbeth_ France it’s a love hate baby girl. It’s not the same without the plastic cup and straw so sue me #mcdonalds #cannes #coke ♬ Ironic - Alanis Morissette

For a few lucky social listening practitioners in the industry, language Boolean operators and filters can suffice. However, for example UK/English, Portuguese (Brazil/Potugal) and of course Spanish remained challenging and frustrating.
When conversations stop respecting borders
What’s more, Instagram and TikTok users started chasing global clout and virality, often posting in English, just using hashtags without any other text, or even just a video or image with text/hashtag overlayed on top of it and no other audio clues about language.


What’s more, we are increasingly living in a “flat word” and recent Publicis Groupe APAC and TikTok research also examined how trends travel and transform across markets, further complicating attempts to treat online conversations as neatly geographically bounded. On top of it all, GenAI translation capabilities and in-platform automatic translation features have made it even easier for people to converse seamlessly on social media in their native languages.

This raises the question how far the social listening tools can get you when platforms outright lack geo location data, the geo location data is incomplete or comprises a small percentage of the overall data set, or sometimes the tagging is slightly off (not exactly wrong but up for dispute, see the example above).
The built-in navigation tools
It must be noted that some tools offer features to help you mitigate some of these issues. Depending on the Boolean operators and filters available, one can narrow down the subreddits covered in the search and include or exclude country and city specific subreddits like r/UK for example.
Additionally, a few tools on the market allow to run searches on the bios of authors. This will not solve the Reddit issue but can help extract relevant content from Instagram and Tiktok for example. All you need to do is run a search for authors with bios that contain the name of the country, state (if applicable), major cities, and the flag of the respective country. This might produce a decent sample of content. If you go this route, my advice is to run separately a broad language search and compare the results with the Country-specific data set to see how the latter set stacks against the broader set and trends and narratives.
Looking at the data through the prism of the consumer path journey
Once these options have been exhausted, there is one more thing a practitioner can do. The first step is to decide how willing you are to forgo relevant content and how important quant metrics like volume of conversation are to you. If you are not willing to give up relevant content and quant metrics are not a priority, these next steps help you alleviate the geo location issue.
Cast a broad search with Language as the primary geo location proxy and try to think as a consumer – what is this person looking for, where can this person find the answers online and what will the person encounter on social media. In their seminal book “Absolute value - What Really Influences Customers in the Age of (Nearly) Perfect Information” Itamar Simonson and Manuel Rosen point out consumers flock and migrate to the best available content and sources of information on the internet.
Examine the research subject and break it down through the prism of the “flat world” paradigm. Are the pleasure and pain points of the consumer universal or unique and tied to the geographic location of the consumer? For example, the side effects of the GLP-1 drugs do not vary across borders and people from different countries will go to subreddits like r/GLP1, r/Ozempic or r/Mounjaro to discuss them. On the other hand, price and access to GLP-1 drugs varies significantly and the discussions around these topics will take place on the respective country's subreddits and local forums.
One thing to do diligently is verify manually the authors of content that has gone viral or seems like an ideal verbatim to substantiate or refute your work hypothesis. In such cases it is highly advisable to check what kind of accounts the author is following, the other content the author has posted, participation in other subreddits and threads on forums, etc.
What lies ahead
Unfortunately, restrictions on the social media platforms’ data have become routine in our industry and consumers’ desire for more privacy will probably further exacerbate the issue. Thus, it will increasingly fall on the discretion of the analyst how to tackle the issue. There is a fledgling progress in GenAI inference about geo location and demographic data, but this approach raises some legal and ethical concerns. Thus, it will be the analyst’s task to gauge how important geographic certainty is for the research question and it will be up to the analyst's acumen, cultural knowledge, and familiarity with different layers of context to discern which pieces of digital evidence can measure up and answer the questions with enough level of confidence.
