In recent months, video large language models (Video-LLMs) have been hailed as capable of tracking a single character through a long video, reporting changes in their outfit or actions. However, a recent study reveals that these systems are not truly 'watching' the named person; instead, they rely on shallow cues such as gender to answer. This finding has deep implications for businesses that depend on artificial intelligence for surveillance, automated customer service, or behavior analysis.
The research shows that when the character's name in the question is changed, the models barely alter their responses (between 4% and 31% of the time). This demonstrates no real identity tracking. Moreover, if the name is swapped for one of the same gender, the change rate is even lower: the models only react when the gender differs. In open-ended tests (without multiple choice), accuracy drops drastically: open-source models fall from 37-38% to 12-19%, and none of the 151 answers were fully correct.
What does this mean for the business world? Imagine a video surveillance system that must identify a specific suspect among several people, or a virtual assistant that must remember a particular customer's preferences over a long interaction. If the model relies on coarse cues like clothing color or gender, it will make costly mistakes. Companies investing in AI solutions must be aware of these limitations before deploying critical systems.
At Q2BSTUDIO, we tackle these challenges from a technical and business perspective. We do not believe in generic AI solutions that fail in real-world contexts. That is why we develop custom software applications that integrate language and vision models with specific business logic. For example, we combine AI with customized workflows that verify character identity using additional metadata (such as timestamps or sensor data), reducing reliance on superficial visual cues.
Our approach includes using cloud infrastructure on AWS or Azure to scale video processing without performance loss, and implementing cybersecurity layers to protect the sensitive data these systems handle. Additionally, we integrate Business Intelligence dashboards with Power BI so that executives can audit model behavior and detect biases or tracking failures.
One of our most innovative lines is specialized AI agents. Instead of a single Video-LLM, we deploy multiple agents that collaborate: one analyzes motion, another facial identity (using biometric recognition), and a third cross-references the information with user databases. This modular architecture corrects the errors of monolithic models and provides robust tracking even in long videos with multiple characters.
The study also reveals that adding subtitles, using the most informative frames, or doubling the number of images does not improve character tracking. This indicates the bottleneck is not the amount of visual data, but how the model links the video to the named person. To solve this, at Q2BSTUDIO we design hybrid systems that combine computer vision with symbolic logic: for example, we assign a unique identifier to each character based on their initial appearance and maintain it throughout the sequence using data association algorithms, regardless of clothing changes.
The lesson for businesses is clear: do not blindly trust Video-LLM benchmarks. High scores can hide fundamental weaknesses. For critical applications such as security, customer service, or retail video analytics, a customized approach combining AI, cloud, cybersecurity, and BI is necessary. At Q2BSTUDIO, we have been developing custom software for years that overcomes these limitations, helping our clients achieve reliable and actionable results.
If your company needs a video person tracking system that actually works, we invite you to contact us. We will analyze your case, design a tailored architecture, and deploy it in the cloud with full security. Do not let a superficial model put your operations at risk. Artificial intelligence must be as precise as your business demands.





