arXiv:2609.30629v1 Announce Type: cross Abstract: Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-trained split interfaces, while stronger channel-aware codecs can impose substantial onboard co

Topological visualization of FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception
Brave API

FRESHLATENT (arXiv:2609.30629v1) is a lightweight, channel-aware latent adapter designed to enable robust split vision-language model (VLM) perception on resource-constrained UAVs operating under wireless corruption. It trains a power-normalized encoder-decoder through simulated wireless distortion while keeping the surrounding VLM frozen, effectively bridging the gap between clean-trained interfaces and real-world channel conditions.

Key performance metrics on an NVIDIA Jetson AGX Xavier demonstrate that FRESHLATENT uses 37–40× fewer encoder parameters, 7.7–9.9× lower latency, and 8.8–10.0× lower energy than heavier feature-JSCC codecs. At 0 dB SNR with tight communication budgets, it recovers 63.5–69.1% of the robustness gain achieved by significantly larger codecs, improving generalized intersection over union (gIoU) and class-specific IoU (cIoU) by approximately 20.8 points over clean split compression.

Generated 6d ago
Open-Weights Reasoning

Summary

FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception addresses a practical bottleneck in edge–cloud vision-language model (VLM) deployment for UAVs. In split VLM perception, the UAV typically runs a vision encoder and transmits intermediate latent features to a remote server that performs higher-level language-conditioned reasoning. The paper highlights a key deployment mismatch: these split interfaces are often trained on clean latents, but real wireless links introduce compression, noise, packet loss, or other corruption that degrades downstream perception. Simply using stronger channel-aware codecs can improve fidelity, but doing so may introduce unacceptable onboard computational and energy overhead, which is especially problematic for resource-constrained mission-critical UAVs.

The central contribution is a channel-aware latent adaptation approach that makes the split interface itself more robust to transmission impairments and resource constraints. Rather than relying solely on heavier communication codecs, FRESHLATENT adapts the transmitted latent representation—or a lightweight adaptation layer around it—based on channel or resource-related signals, so that the server-side VLM receives features that better match its expected input distribution. This reframes robustness as a property of the latent interface rather than an external communication problem, aiming to preserve VLM perception accuracy while limiting onboard compute and communication cost.

This work matters because embodied VLM perception is moving toward realistic deployment settings where bandwidth, latency, energy, and hardware constraints are first-order design variables. For UAVs operating in mission-critical scenarios, reliable vision-language understanding cannot assume ideal feature transmission. By targeting the mismatch between clean-trained split interfaces and corrupted real-world transmissions, FRESHLATENT offers a practical direction for building more resilient edge–cloud VLM systems that can operate effectively under constrained wireless and onboard-resource conditions.

Generated 6d ago
Sources