arXiv:2609.20263v1 Announce Type: cross Abstract: The proliferation of generative video models has shifted the practical detection threat from fully fabricated clips to partially manipulated footages. Although modern detectors achieve strong accuracy using foundation backbones of 400M+ parameters, their resource footprint precludes edge deployment. In this paper, we present a lightweight full-fra

Topological visualization of AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments
Brave API

AI or Real is a lightweight detector for partially manipulated AI-generated videos designed for edge deployment without requiring face-detection preprocessing. It utilizes knowledge distillation to compress a large DINOv2-Base teacher model into a frozen MobileNetV3-Small student, achieving an AUC of 0.672 and 50.6% detection rate at the difficult 6.2% fake-frame ratio. The system addresses specific partial-manipulation challenges through within-video temporal hard negatives to prevent false positives on legitimate scene cuts and calibration-aware sampling to correct miscalibration on the dominant pure-real class.

Key technical and performance specifications include: Architecture: A residual feature adapter reconditions ImageNet-pretrained features toward diffusion-artifact statistics, while attention-diversity regularization prevents spatial attention heads from collapsing. Efficiency: The model operates at 3.65 ms per 16-frame clip on an RTX A4000 GPU with a 150.4 MB checkpoint, making it compatible with edge-device memory and latency budgets. Performance Gap: The student model closes 58% of the accuracy gap to the uncompressed DINOv2-Base teacher (AUC 0.766), which is not deployable on edge hardware due to its ~80 ms latency and 330 MB size. Evaluation: Tested on a 55,393-sample spliced test set across fake-frame ratios from 6.2% to 31.2%, demonstrating robustness against temporal signal dilution and semantic consistency in modern generative models.

Generated 16d ago
Open-Weights Reasoning

AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments addresses the growing mismatch between the capability of video-forgery detectors and the environments where they need to run. As generative video models become more capable, the practical threat is increasingly localized or partial manipulation—such as object insertion, face swapping, background editing, or temporal inconsistencies—rather than entirely fabricated clips. The paper frames detection as an efficiency problem: state-of-the-art detectors often achieve strong accuracy by relying on foundation-model-style backbones with 400M+ parameters, but their memory, latency, and compute requirements make them poorly suited for edge devices, mobile capture pipelines, or real-time monitoring systems.

The proposed work introduces a lightweight full-frame detection approach aimed at identifying partially altered videos while remaining deployable under resource constraints. Its central contribution is not merely higher benchmark accuracy in a cloud-scale setting, but a shift toward practical, compact video forensics: preserving sensitivity to subtle manipulation cues without depending on very large foundation video encoders. In doing so, the paper targets a deployment-oriented design space where detection must be fast, memory-efficient, and capable of processing video streams directly on constrained hardware.

This matters because many real-world authentication workflows cannot assume unlimited cloud inference. On-device or edge-side screening could enable real-time triage of manipulated footage in social media, journalism, surveillance, mobile capture, and privacy-sensitive applications where raw video should not be uploaded for analysis. By focusing on partially altered content in resource-constrained settings, the work is relevant to the next stage of deepfake detection: moving from laboratory-accurate models toward deployable systems that can operate where video is created, consumed, and shared.

Generated 16d ago
Sources