Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests into pixel-level visual grounding, offering a natural perception interface for embodied agents. However, existing benchmarks largely focus on generic visual scenes and overlook the domain and resource constraints encountered in flood-response platforms. We present FloodReasonBench, a benchm

Topological visualization of FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge
Brave API

FloodReasonBench is a benchmark designed to evaluate Vision-Language Models (VLMs) on reasoning segmentation specifically for embodied flood response scenarios. It addresses the gap between generic visual grounding datasets and domain-specific needs by introducing FloodResponseSeg, a dataset focused on pixel-level visual grounding for people, buildings, and vehicles in flood environments.

The benchmark evaluates VLMs under edge computing constraints, characterizing performance through lightweight visual encoding, hierarchical split inference, and compressed intermediate representations. Unlike general benchmarks, FloodReasonBench demonstrates that flood-adapted models exhibit more consistent accuracy across different hardware partitions compared to generic pre-adapted models, which show significant performance variation depending on the specific edge device partition.

Generated Sep 2, 2026
Open-Weights Reasoning

FloodReasonBench addresses a gap in vision-language model (VLM) evaluation by targeting reasoning segmentation in a domain where perception must be tightly coupled to operational decision-making: embodied flood response. Reasoning segmentation asks a VLM not merely to recognize objects, but to translate mission-relevant language requests into pixel-level visual grounding. The paper argues that existing segmentation and visual grounding benchmarks are largely built around generic scenes and therefore underrepresent the challenges of flood-response platforms, where agents must interpret domain-specific language, operate in hazardous and visually ambiguous environments, and often run under edge-compute and latency constraints.

The key contribution is a benchmark for evaluating how well VLMs can ground flood-response language into actionable visual outputs, such as masks for relevant hazards, infrastructure, or operational targets. By framing segmentation as a reasoning and perception interface for embodied agents, the work shifts the evaluation focus from generic object localization toward task-oriented visual grounding in a safety-critical setting. The benchmark is also notable for explicitly accounting for the resource constraints of edge deployment, making it more relevant to real-world embodied systems than benchmarks designed primarily for high-capacity server-side models.

This matters because reliable flood-response agents need interpretable, low-latency perception that can be directed by natural-language mission goals. FloodReasonBench provides a structured testbed for measuring whether current VLMs can support that interface, exposing limitations in domain grounding, robustness, and resource-aware inference. As a result, it offers a useful foundation for developing and comparing VLM-based perception pipelines that are practical for embodied flood response rather than only accurate in generic visual benchmarks.

Generated Sep 2, 2026
Sources