Collaborative inference pools distributed resources to run compute-intensive Vision Transformers (ViTs) in satellite edge computing. Model partitioning enables such collaboration by assigning consecutive layer groups to different nodes, but the large volume of intermediate activation data incurs substantial transfer overhead that can erase its benefit. Token compression reduces downstream computat

Topological visualization of Iapetus: Content-Aware Hierarchical Scheduling for Collaborative ViT Inference in LEO Satellite Networks
Brave API

Iapetus is a content-aware hierarchical scheduler designed for collaborative Vision Transformer (ViT) inference in Low-Earth-Orbit (LEO) satellite networks. It addresses the high communication overhead of ViT intermediate activations by jointly optimizing token compression and layer offloading into complete trajectories, rather than treating them as independent decisions.

The system employs a Lyapunov-guided hierarchical approach that screens constellation-wide candidates to bound search space, then refines them using content-aware quality prediction and joint planning. This unified objective balances per-task latency, energy consumption, and quality loss against accumulated workload and battery pressures.

Implemented on an NVIDIA Jetson AGX Orin testbed, Iapetus demonstrates significant performance gains, accomplishing 91.6% of tasks at 5 tasks/s. This outperforms the strongest baseline, MARATD3, by 26.1 percentage points, while reducing mean latency by 53.0% and battery draw by 70.8%.

Generated 28d ago
Cerebras Thinking

Iapetus: Content-Aware Hierarchical Scheduling for Collaborative ViT Inference in LEO Satellite Networks

The material presents Iapetus, a framework for running compute-intensive Vision Transformer inference collaboratively across LEO satellite edge nodes. It targets a key bottleneck in model-partitioned satellite inference: while splitting consecutive ViT layer groups across multiple satellites can reduce per-node computation, the intermediate activation tensors exchanged between stages are large and can consume enough satellite-to-satellite or satellite-to-ground bandwidth to offset the computational gains. Iapetus addresses this by combining content-aware token compression with hierarchical scheduling. Instead of transmitting full activation feature maps, the system selectively retains or compresses tokens/patches based on their downstream importance, thereby reducing both bandwidth usage and subsequent compute. The scheduler then coordinates model placement, compression choices, and activation routing across a hierarchy of satellite nodes under time-varying visibility, link capacity, and compute constraints.

The central contribution is the joint treatment of token compression and resource scheduling as a content-aware, network-aware optimization problem. The key insight is that activation traffic in ViTs is not uniformly valuable: some visual tokens carry most of the information needed by later layers, while others can be discarded or represented more coarsely. This means that the effective partitioning strategy and communication plan should depend on the input image, the ViT stage, and the current state of the LEO network rather than being fixed statically. Hierarchical scheduling is especially relevant because LEO constellations are not a flat, stable cluster; they involve multiple communication tiers, intermittent links, and rapidly changing topology, so layer placement and data transfer must be coordinated across nodes and time windows to preserve the latency and energy benefits of collaboration.

This work matters because it points to a practical way to run perception-heavy models in space-based or satellite-edge environments where no single satellite can provide enough compute, memory, or energy. If effective, such systems could support near-real-time visual analytics from LEO constellations—such as Earth observation, target tracking, or anomaly detection—without downlinking raw imagery or relying on a single powerful platform. More broadly, the paper highlights content-adaptive compression as a critical enabling mechanism for collaborative deep-learning inference in bandwidth-constrained, heterogeneous, and mobile edge networks.

Generated 28d ago
Open-Weights Reasoning

Summary

The paper tackles collaborative Vision Transformer (ViT) inference in LEO satellite edge networks, where a ViT is split into consecutive layer groups and executed across multiple satellites or edge nodes. It identifies a central practical bottleneck: model partitioning can distribute compute, but ViTs generate large intermediate activation/token tensors between stages, and transferring those tensors over constrained, dynamic satellite links can dominate latency and energy cost. Iapetus addresses this by coupling token compression with hierarchical scheduling, reducing the amount of intermediate data that must be moved while preserving enough semantic information for downstream layers to maintain inference quality.

Its main contribution is a content-aware, multi-level scheduling framework that adapts collaborative inference to both input complexity and satellite-network conditions. The “content-aware” component exploits redundancy in ViT token representations to compress, prune, or selectively retain tokens before they are transmitted between nodes, thereby reducing both communication volume and downstream computation. The hierarchical scheduler then coordinates decisions across nodes, clusters, and inference stages—such as where to place layer groups, how aggressively to compress activations, and how to allocate limited satellite resources—under coupled accuracy, latency, bandwidth, and energy constraints. The core insight is that satellite collaborative inference should not treat inter-node transfer as a fixed overhead; instead, it should jointly shape token-level model execution and task placement to match the dynamic constraints of LEO networks.

This matters because LEO constellations can offer distributed edge compute, but their intermittent connectivity, limited backhaul, and heterogeneous resources make naive model parallelism impractical for transformer-based vision workloads. By reducing intermediate activation traffic and aligning computation with network and resource availability, Iapetus makes high-accuracy ViT inference more feasible in space-based edge settings, where onboard compute alone may be insufficient and terrestrial cloud inference may be too slow or energy-intensive. The work is therefore relevant to satellite AI, distributed edge inference, and latency/accuracy-aware scheduling for large transformer models in bandwidth-constrained environments.

Generated 28d ago
Sources