What video listening tools still can't see
.png)


Although Instagram and YouTube have been around for a while and image analytics/first frame analytics have progressed in the past couple of years, it was TikTok that really accelerated social video listening technology.
The platform’s rapid rise in popularity, the medium (video and audio), the sheer volume and nature of content posted, as well as the restricted access to this content continue to cause FOMO and nightmares among social intelligence practitioners. In quick succession, several video listening vendors entered the market to answer the call.
While the technology has come a long way, I still wasn’t convinced it had reached the level of sophistication required to generate meaningful insights without human intervention. So, I conducted an experiment.
Identifying unspoken signals
I recently worked on a sneaker project involving large TikTok and Instagram data sets. Runners have always been active on forums and subreddits and the Natural Language Processing (NLP) algorithms have generally been able to pick up signals related to price, quality, durability, and other brand/product attributes from text-based content. However, TikTok and Instagram – where runners tend to post their daily routines and photos and videos of races – posed a different challenge.
Despite the excellent image/first video frame and audio transcript analytics available, the algorithms failed to idenitify any signals about brand\product attributes, and we had to resort to manual interpretation and analysis. For example, watching this video I thought “Elderly lady walking 4 miles a day, pushing a stroller with her dog (bonus points for cuteness!) - these ON running shoes must be really comfortable.” Yet there were no meta tags for quality or comfort in the tool.
@ahnestkitchen Everything Umma packs for her daily 4-mile walk featuring Mr. Chicago. #biodancepartner AD @Biodance Store US #biodancejellymist #jellymist #sprayonmask ♬ What Floor? - idokay
This made me wonder how full stack video listening tools would fare with this same data set: What kind of outputs could they produce? And could they reach the conclusions my team and I were drawing from these videos?
Thankfully, YouScan, ViralMoment, ConvoTrack, Social Voice, Nebula Social, Aggero, and Forage responded to my call to help answer these questions. However, I want to make it clear that this is not a tool vendor ranking. Although it involved data and an experiment, the design and analysis didn’t follow scientific standards and rigour.
Can AI recognise the extraordinary?
I submitted two more videos for analysis. One was of Wayne Rooney scoring a bicycle kick against Manchester City. As a friend of mine who inquired about the experiment commented: “One word: that goal, I don’t even need to watch it: iconic.”
I was curious to understand if tools can recognize the athletic feat and significance of such an extraordinary moment. And if so, how?
The other was a recent video of a Duke - UConn game. (For the record I am a Duke fan and stayed till the early hours of the morning to see the game).
@spartans.senter The freshman #braylonmullins #gamewinner #edit #finalfour #uconn @braylonmullins24 ♬ E85 - Don Toliver
My reason for choosing it was similar to Wayne Rooney’s goal but the clock running out of time added an extra dimension. This video also had plenty of pointers to allow the tools to expand and enrich the context of the video footage and enrich the outputs: the live commentary, team logos, Final Four sign/logo, even the music overlay at the end.
All the video listening vendors provided impressive outputs, a testimony to how far video analytics has progressed in the past couple of years. Across the board they successfully identifed logos, objects, activities, people, animals, etc. and transcribed and analysed the video audio transcript.
Some of the tools even captured (or claimed to capture) the clock ticking down, linking it to the events on the court in the Duke video. A couple enriched their analysis by trawling the internet for additional information and their analysis carried deep GenAI traces. This worked well for the Wayne Rooney and the Duke videos as both events have generated a significant number of comments and reactions online. The approach failed with the ON sneakers video as it had no digital footprint.
At this stage, I could confidently state that social listening tools, with the help of GenAI, can churn out troves of visual analytics metadata points. However, it still hasn’t convinced me that the tools provide anything more than data. It was humans that assigned and extracted value and insights from this data.
The blind spots only humans can see
My thinking is grounded on the Economic Subjective Theory of Value: value is not intrinsic to a product; rather, it’s assigned by the consumer based on the item's utility/ its ability to satisfy a specific desire. This exercise confirmed humans’ innate ability to process new information and appreciate beauty, deviancy or the extraordinary. It’s this that’s essential for insights generation and hard for technology to replicate.
Deploying GenAI agents, pretraining and directing the algorithms to look for data points and clues might produce similar insights, but they’re still dependent on the analyst’s knowledge, expectations and values. Even if there’s no human involvement and algorithms work "autonomously", the agent relies on historic human generated data and information to build context and draw connections and conclusions. They’re stuck on what we, and other analysts, already know and have documented. Plus, there are huge challenges around scalability and cost, both financial and environmental.
This leaves a blind spot when it comes to the unknown-unknowns, the “aha” moments. By design, LLMs tend to reproduce and reinforce the most statistically prevalent ideas. However, as another good friend of mine told me, “There is beauty and value in deviancy. All kinds of cool things happen there.” Insights and emerging trends are usually found in the fringes.
And then there are the awesome moments. Humans need not have seen a bicycle kick or a buzzer beater to appreciate and understand their beauty, regardless if they saw one in the local park or at Wembley or Madison Square Garden.
Can the algorithm recognise these acts and signals, predict how and why a human would react? On their own, no. This is where we as researchers shine: interpreting changing attitudes and behaviour and drawing new connections and observations.
