Introduces a lightweight two-stage video anomaly detector that uses YOLO v11n-pose keypoints and CLIP cosine similarity to textual anomaly descriptions, removing optical flow and density modules.
The framework, proposed by Warnasooriya et al. and presented at MIPR 2026, achieves 51.39 FPS throughput, a 3.36× speedup over previous baselines, by replacing optical flow and density estimation with direct CLIP semantic scoring. It detects person-level anomalies like falling and fighting with 89.26% AUROC on the CUHK Avenue dataset and 84.13% on CU Indoor Anomaly, while requiring only a single GPU and two model checkpoints.
The material describes a lightweight, two-stage video anomaly detection framework that replaces conventional motion and crowd-density modeling with a compact semantic pipeline. In the first stage, a YOLO v11n-pose model extracts human pose keypoints from video frames, providing a low-dimensional representation of body structure and motion. In the second stage, these pose-derived features are compared against textual descriptions of anomalous behavior using CLIP-based cosine similarity. This design eliminates explicit optical flow computation and density estimation modules, shifting the detection burden toward pose geometry and semantic alignment with anomaly language.
The key contribution is a simpler, more deployable architecture for real-time video analysis. By relying on a small pose estimation backbone and a pretrained vision-language model, the method avoids the computational overhead of dense motion fields and handcrafted crowd-density cues. It also offers a degree of interpretability: anomalies are not only scored numerically but can be understood in relation to natural-language descriptions such as “falling,” “aggressive movement,” or “suspicious loitering.” This semantic grounding may improve transferability across domains where visual anomaly categories vary but can still be described linguistically.
The work matters because it points toward practical anomaly detection systems that can run efficiently on edge or near-real-time infrastructure without heavy video-processing stacks. Its main tradeoff is that performance depends on the quality of pose estimation and the expressiveness of the textual anomaly prompts, so it may be less effective for non-human anomalies, severe occlusion, or behaviors that are difficult to articulate in language. Still, the paper is notable for reframing video anomaly detection as a pose-to-language semantic scoring problem rather than a purely low-level motion or density anomaly task.