Brave API

This curated collection on AI Reasoning & Multi-Agent Systems aggregates 15 frontier papers and surveys from arXiv and sources like NBER, spanning reinforcement learning (RL), graph neural networks (GNNs), large language model (LLM) optimization, and agentic AI in economic contexts. It emphasizes scalable multi-agent coordination (e.g., process rewards, monotonic improvement guarantees via MonoScale, and principal-agent framings requiring mechanism design), autonomous decision-making (e.g., end-to-end belief-policy optimization in shared human-AI autonomy and IRL-DAL for safe trajectory planning with diffusion models), and reasoning advancements (e.g., GNNs executing graph algorithms exactly or disentangling multispecific antibody functions). Additional works explore efficient architectures like separable neural nets for agile RL, tensorized orthonormalization (TEON) for LLM pre-training, and small LMs generating dynamic game content.

Key themes interconnect through a focus on emergent behaviors and scalability in agentic systems. Multi-agent papers link coordination challenges to economic principles, such as principal-agent dynamics and market-like interactions in virtual economies, echoed in NBER and web surveys on AI agents transforming economic research, financial services (with model risk management), and Industry 4.0 manufacturing. Reasoning components (GNNs, RL architectures) provide foundational tools for exactness and adaptability, preventing issues like catastrophic forgetting, while tying into broader agent economies where AI interactions yield unintended dynamics. These threads converge on designing robust, verifiable systems that scale without performance cliffs.

These topics matter profoundly for deploying production-grade AI in high-stakes domains like autonomous driving, finance, and distributed manufacturing. By addressing coordination failures, safety via energy-guided planning, and economic incentives, the collection advances toward reliable multi-agent RL and agentic workflows—critical as AI agents proliferate in simulated economies and real-world applications. This work not only mitigates risks like non-monotonic scaling but also unlocks novel insights into emergent intelligence, positioning AI as a transformative tool for scientific discovery and industrial automation.

Generated Feb 21, 2026
Cerebras Thinking

This collection explores the frontier of AI reasoning and multi-agent systems (MAS), shifting the research focus from isolated model performance to complex, interactive, and autonomous ecosystems. A significant portion of the work investigates agentic architectures through an economic and strategic lens, arguing that multi-agent systems exhibit principal-agent dynamics that require careful mechanism design to ensure alignment and prevent issues like "lazy" agents or catastrophic performance drops. This includes comprehensive surveys and working papers on virtual agent economies, financial service applications, and industrial automation (Industry 4.0), highlighting how AI agents collaborate, compete, and evolve in market-like environments. The collection also bridges the gap between theoretical reasoning and practical execution, covering advancements in Reinforcement Learning (RL)—such as monotonic scaling guarantees (MonoScale), agile adaptation, and shared autonomy—alongside specific applications in mathematical problem-solving, legal reasoning, and dynamic game content generation using small language models.

A recurring technical theme is the pursuit of robustness, efficiency, and verification within these sophisticated systems. The research connects high-level reasoning with low-level safety mechanisms, distinguishing between weak and strong verification for trustworthiness and utilizing Graph Neural Networks (GNNs) for tasks ranging from exact algorithm execution to antibody characterization. Efficiency is addressed through innovations like TEON for optimized LLM pre-training and sink-aware pruning for diffusion language models, while safety and alignment are tackled via methods like MARS (margin-aware reward modeling) and energy-guided diffusion for safe trajectory planning. Furthermore, the collection emphasizes human-AI collaboration, defining protocols for counterfactual harm and user-specified requirements in high-stakes decision-making, alongside automated tools like FAMOSE for feature discovery. These topics are critical as they represent the necessary evolution from static large language models to dynamic, reliable, and economically viable autonomous agents capable of operating safely in the real world.

Generated Feb 22, 2026
Open-Weights Reasoning

AI Reasoning & Multi-Agent Systems: This curated research collection focuses on frontier work in AI reasoning, multi-agent systems, reinforcement learning, and autonomous decision-making. The collection includes 15 research cards from sources such as arXiv and NBER, covering various aspects of these topics.

Key Themes: One key theme in the collection is scaling multi-agent systems and improving coordination and performance. Several papers investigate the use of process-based rewards, monotonic improvement guarantees, and even treating multi-agent systems as principal-agent problems to address these challenges. Another theme is the application of advanced machine learning techniques, such as reinforce learning, inverse reinforcement learning, and graph neural networks, to develop autonomous agents that can perform complex tasks, execute graph algorithms exactly, characterize functional properties of multispecific antibodies, and generate high-quality dynamic game content. Lastly, the collection explores the application of these technologies in various industries like finance and the emergent economic behaviors in virtual environments populated by AI agents.

Why it Matters: AI reasoning and multi-agent systems play a crucial role in creating advanced autonomous agents and systems that can collaborate and make informed decisions. This collection highlights the importance of developing more efficient and effective methods for scaling multi-agent systems and advancing machine learning techniques. These advancements can lead to improvements in various fields, including finance, manufacturing, and gaming, and can open new avenues for research in Artificial General Intelligence (AGI) and beyond. By studying the latest research in this area, we can gain insights into emerging trends and innovations in AI and continue to push the boundaries of what's possible.

Generated Feb 21, 2026
Research Materials (100)
Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs
Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, resource ordering, and placement must be selected according to the observed pool
A Technique for Load Shifting Low-latency Applications in Multi-Region Renewables Harvesting via SMT Core Pooling
Load shifting across geographic regions to chase intermittent renewable energy availability is commonly used in reducing cloud infrastructure carbon footprint. However, it often omits low-latency applications due to high latency variances of wide area networks (WAN) that interconnect regions. This paper addresses accommodating low-latency applications into load shifting by minimizing their shiftin
Iapetus: Content-Aware Hierarchical Scheduling for Collaborative ViT Inference in LEO Satellite Networks
Collaborative inference pools distributed resources to run compute-intensive Vision Transformers (ViTs) in satellite edge computing. Model partitioning enables such collaboration by assigning consecutive layer groups to different nodes, but the large volume of intermediate activation data incurs substantial transfer overhead that can erase its benefit. Token compression reduces downstream computat
Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs
As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensi
Barnacle: Adaptive Multi-Leader Scheduling for DAG-Based Consensus
In DAG-based consensus, all validators propose blocks concurrently, and designated leader blocks drive transaction commit. Having multiple leader slots per round cuts queuing latency, yet production deployments run a single leader because of head-of-line blocking: a slow leader stalls the pipeline for at least one leader timeout, and for several waves when its slot must wait for the fallback indir
Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
Distributing an AI computation across the GPUs of a multi-GPU server is one of the central problems in systems-for-AI. We present Einsummable, a prototype system that accepts a PyTorch-like description of an AI computation and automatically distributes it across a multi-GPU server, with no device assignments, sharding annotations, or communication operations written by the programmer. Einsummable
JuPyLive: Seamless Migration of Jupyter Notebook Resources from Laptop to HPC
This work introduces JuPyLive, a migration mechanism that enables seamless transition of Jupyter notebooks between local resources of user's workstation and remote resources of high-performance computing~(HPC) environments, while preserving the user experience. JuPyLive eliminates the underlying complexities of migration process, enabling users to freely choose among available local and remote res
Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic Programming
Constant optimization refines the numerical coefficients of candidate expressions in tree-based genetic programming for symbolic regression. But its per-generation cost has led modern GPU-accelerated frameworks to omit it or restrict it to lightweight forms. We present a GPU-resident, batched Levenberg--Marquardt solver that optimizes constants across a structurally heterogeneous population of exp
FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train Embedding
Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows across input indices are not fully reused and intermediate results are repeatedly materialized off-chip between sequential TT-core contractions. We present FlowTT, a flow-aware GPU execution framework th
BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training
Long-context reasoning for large language models (LLMs) is becoming increasingly important, but training over long sequences remains challenging due to massive memory and communication requirements. Sequence parallelism has emerged as an essential technique for addressing bottlenecks in long sequence LLM training. However, we observe that existing sequence parallelism methods are batch-agnostic an
Skywing: A Platform for Decentralized Mathematical Computing in Unreliable Environments
Emerging edge, autonomous, and cyber-physical systems increasingly require mathematical computation across heterogeneous devices connected by unreliable communication networks. Traditional high-performance computing and distributed data-processing frameworks provide powerful abstractions for managed environments but are less suited to decentralized settings where centralized coordination, reliable
Fully Fluctuating Sleepy Consensus from Minimal Assumptions
Bitcoin's proof-of-work (PoW)-based protocol is remarkable for how little it asks of its participants. Not only can miners take breaks from work whenever they please, but it is almost unique in offering a path of contrition: corrupt miners can reclaim honest status simply by resuming mining on the longest chain. The protocol only requires that honest miners hold the majority of computational power
eAVID: Asynchronous Verifiable Information Dispersal with Post-Dissemination Pruning
Asynchronous verifiable information dispersal (AVID) lets a sender spread a message across $N=3F+1$ nodes such that it remains recoverable despite up to $F$ Byzantine failures. Because dispersal must complete on $N-F$ responses, standard AVID protocols fix a $(F{+}1,\, N)$ erasure code and pay a $3\times$ storage blowup, whereas a synchronous system achieves the optimal $3/2\times$. This cost is p
MM-BEV: Enhancing Timeliness by Computing Where and When it Matters
Multimodal bird's-eye-view (BEV) perception combines LiDAR depth accuracy with dense camera semantics, but its high computational cost and imperfect sensing conditions make real-time deployment challenging. Existing methods largely compress individual detectors and overlook three opportunities: structured sparsity within camera and LiDAR inputs, timing misalignment between modalities, and the fact
FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge
Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests into pixel-level visual grounding, offering a natural perception interface for embodied agents. However, existing benchmarks largely focus on generic visual scenes and overlook the domain and resource constraints encountered in flood-response platforms. We present FloodReasonBench, a benchm
LOCAL: Enabling Learning On-device Contiguously for Agent LLMs
On-device LLM agents interact repeatedly with users on local hardware, producing private traces that are valuable for adaptation but should not be sent to a remote trainer. Ideally, such agents would learn contiguously---adapting from every interaction without pausing or suspending user-facing inference---yet existing inference runtimes assume stable weights and existing RL systems assume separate
P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving
Long-context LLM applications such as retrieval-augmented generation (RAG) and agentic systems often process tens of thousands of input tokens to produce short outputs, making end-to-end request latency an important serving objective. We show that the maximum number of batched tokens (MBT), which controls the token scheduling budget in vLLM, has a scheduling-pressure-dependent effect on latency. L
TERRA: A Hierarchical Parallel Training and Memory Orchestration Framework for High-Resolution AI-based Earth Modeling
Training high-resolution AI-based Earth forecasting models is memory-intensive. Window-based Swin Transformers reduce the quadratic cost of global attention, but existing distributed systems such as AERIS primarily target pixel-level models and do not jointly support convolutional sampling modules and shifted-window execution. Long-lead rollout finetuning further increases activation memory. To ad
From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench,
Collective Communication for Distributed LLM Systems: Planning, Runtime Adaptation, and Computation Coordination
Distributed large language model (LLM) systems increasingly rely on collective communication primitives such as AllReduce (AR), ReduceScatter (RS), AllGather (AG), and AlltoAll (A2A). In modern LLM training and serving clusters, heterogeneous GPU interconnects, multi-NIC networking, mixed parallelism strategies, low-latency inference requests, and high-throughput training pipelines have motivated
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads
Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage, KV-cache, and activation terms parameterized by step count, tool invocations, and context expansion. We evaluate this decomposition empirically within a strictly scoped measurement study: a LangGraph-based CUDA-kernel-synthesis agent (AgentK), a 4-bit quantization family (Q4 K M), a single NVIDIA H10
PAS-QFL: Personalized Ansatz Selection for Quantum Federated Learning under Client Data Heterogeneity
Quantum federated learning (QFL) lets multiple quantum clients collaboratively train quantum neural networks (QNNs) without sharing private local data. However, existing QFL methods commonly assume that all clients use the same ansatz, overlooking how heterogeneous client data affects ansatz suitability. Under class-imbalanced non-IID data, different clients may favor different ansatz structures,
When Does Distributed AI Inference Need More Wide-Area Bandwidth? A Co-Design Evaluation of Optical, Packet, and Software Levers
Wide-area bandwidth per unit of GPU compute falls every hardware generation: in compute-intensity-ratio terms (CIR, bytes per FLOP), the gap between on-package memory and the conventional WAN is four to five orders of magnitude, widening at roughly 12-19% per year. Position papers - including our own - argued this makes elastic optical wide-area capacity necessary for cross-site AI inference. Revi
Evaluating Agentic Code Repair Capabilities in Distributed Systems
LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system debugging, however, remains an under-explored regime: bugs span processes, nodes, and protocol interactions, with root causes rarely recoverable from source alone and brute-force exploration intractable across non-deterministic int
Enabling Hybrid HPCQC Workflows with a Heterogeneous Software Stack
In this work, we demonstrate hybrid High Performance Computing-Quantum Computing (HPCQC) workflows on a production petascale system. The demonstration combines three components: the SuperMUC-NG supercomputer at the Leibniz Supercomputing Centre (LRZ), a 20-qubit superconducting quantum processor provided by IQM Quantum Computers (IQM), and Munich Quantum Valley (MQV)'s Munich Quantum Software Stac
Porting and Benchmarking Chapel on Emerging RISC-V Hardware: an HPC Viability Study
The Chapel programming language recently added support for the RISC-V architecture. Here we discuss what changes were needed for Chapel to work on RISC-V as well as lessons learned from the porting process. We use some of Chapel's extensive benchmark suite to gain further insight into the suitability of the RISC-V architecture for future HPC use. We compare performance on SiFive P550 and Unmatched
Validating LLM-Modernized Scientific Software Through Differential Fault Injection
Large language model (LLM) agents are increasingly used to modernize the legacy Fortran underlying production scientific software, but validation of these transformations emphasizes nominal executions and may not test whether a modernization preserves the original code's response to faults, perturbations, and reduced precision. We present a differential fault-injection validation method: a harness
Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training
Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using task feedback, but current on-policy RL runtimes execute rollout, reference scoring, and actor training in strict serial phases. While effective for text-only RL, this phase-granular execution is waste
Could Model Partitioning Make Federated Learning More Sustainable?
As federated learning (FL) extends from distributed machine learning between low-power devices to cross-silo scenarios involving edge servers and data centres, its carbon footprint has become a growing concern. Addressing this, methods for sustainable FL align training with low-carbon energy availability or low grid demand and reduce the energy consumption of clients powered by high-carbon sources
Large-scale workflow placement in serverless computing using integer nonlinear programming
Serverless edge computing has become a powerful cloud framework that enables the execution of large workflows without the need for the user to manage the underlying servers and edge devices. In this work, we address the challenge of deploying these workflows on a large number of different existing servers and edge devices such that monetary costs for the users and workflow evaluation times are min
Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning
Electrocardiogram (ECG) recordings are sensitive biomedical data, limiting the ability of hospitals and wearable devices to share raw signals for centralized model training. Federated learning addresses this practical privacy constraint by enabling collaborative model training while keeping raw biosignal data at their respective sources. However, federated ECG classification remains challenging du
Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions
Large language models (LLMs) have become core components of cloud-based intelligent services in academia and industry, yet their training and deployment are hindered by high computational costs, data centralization, and privacy concerns. Federated learning (FL) offers a decentralized training paradigm that enables clients to collaboratively train a learning model without sharing raw data, making i
A Barrier-Free Synchronization Algorithm for Multi-Engine AI Accelerators
Multi-engine AI accelerators such as AWS Trainium comprise specialized compute engines that execute in parallel, and the compiler must synchronize the data dependencies between them. For straight-line code this is simple: each dependency reduces to waiting for a threshold count of instruction completions, which the compiler computes statically. Loops admit no such static threshold; a simple soluti
Balancing Workload Performance and Slurm Stress: Four Nextflow Deployment Strategies
Wide Nextflow fan-outs on shared Slurm clusters can submit tens of thousands of short tasks. Deployment settings route them through individual jobs, arrays, or nested schedulers inside enclosing allocations. These settings determine workflow turnaround and RPC volume, a shared cost that can degrade scheduler responsiveness. Existing comparisons evaluate whole workflow systems, while per-task queue
Adaptive Snapshots Require Visible Reads
Snapshots are widely used to record the state of a running execution. Snapshots have been extensively studied in the literature, with the goal of improving performance and extending functionality. In this work, we consider $adaptive$ snapshots over a set of $m$ components. Adaptive snapshots provide a Click() operation that logically creates a new snapshot and an Observe$(i)$ operation that return
Triangle-Free Coloring in LOCAL via Resilient Lovász Local Lemma
The Lovász Local Lemma (LLL) is a probabilistic tool that has been shown to be of central importance in the study of distributed algorithms. For example, the constructive LLL is known to be complete for the class of locally-checkable labeling problems with $o(\log n)$ randomized complexities in the LOCAL model. One classic application of the LLL is in coloring graphs with some sparse structure, su
Fast Tendermint: Speeding Up a Foundational Consensus Protocol
Tendermint is among the most widely studied and deployed Byzantine fault-tolerant (BFT) consensus protocols, owing in part to its native leader-rotation mechanism that subsumes complex view changes. Like most partially-synchronous BFT protocols, Tendermint tolerates $f < n/3$ Byzantine processes and decides in three communication steps. Motivated by the push for lower-latency blockchains, a recent
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving
Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LLMs). Autoscaling is the key mechanism for cluster resource management, yet a basic system design question is open for serving LLMs: what should be the unit of scaling? Existing approaches primarily treat the entire model
A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family
Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. A kernel can pass that test and still be silently wrong. It can return an ordinary number where the true answer is a NaN or an infinity, differ from run to run,
Efficient Randomized LL/SC that Preserves History Independence
We study the fundamental problem of implementing $m$ linearizable LL/SC objects with constant expected step complexity in a system of $n$ processes, using bounded base objects commonly available in hardware. Assuming that each process may have at most $τ$ outstanding LL operations, the best known deterministic algorithm requires $Ω(n^2τ+ m)$ base objects (CAS and registers) [Blelloch and Wei, DISC
vToken: Token-Level Virtualization for Reclaimable KV Caches
Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction algorithms operate at a token granularity finer than block-level management. This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memo
LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service
As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models without sacrificing reliability has become a central concern. Existing semantic caching methods largely rely on empirical similarity thresholds; while such thresholds improve hit rates, they tend to introduce silent misclassifications near decision boundaries.
Meshlib: In-Process Policy Enforcement for Sidecar-less Service Meshes
Service meshes facilitate service-to-service communication and enforce security policies in microservice architectures. However, they often depend on per-pod sidecar proxies, which introduce significant latency and resource overhead due to redundant application-layer parsing on every request. Eliminating sidecars without compromising security guarantees remains a central challenge. To address this
Validation-Centric AI-Assisted GPU Porting of a 250,000+ Line Legacy Weather Simulation Code
Recent advances in large language models have made CLI-based AI agents a practical tool for accelerating GPU porting of large legacy scientific applications. Such applications, however, are not merely old code bases; they are scientific assets whose credibility has been accumulated through long-term development, comparison with observations, and use in domain studies. GPU porting must therefore pr
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below $n^* \approx 156$--$168$ tokens, HBM weight streaming dominates---cost attaches to $activated replicas$, not tokens
InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers
The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is built, making direct measurement costly, slow, and sometimes infeasible. We pres
A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings
Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limited, and clinical decision-making requires integrating heterogeneous modalities. We introduce a cloud--edge collaborative architecture that addresses these constraints: lightweight, domain-specific mode
Offering Microsecond-Scale Cross-VM Core Elasticity on Colocated Lightweight Virtual Machines
Serverless platforms commonly colocate many diverse workloads, each in a fast-booting, memory-lean virtual machine (VM), to improve deployment density. Overprovisioning each VM for its peak protects tail latency during traffic bursts but hurts density; maintaining high density while effectively protecting tail latency requires the infrastructure to be able to shift physical cores, at a microsecond
RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning
Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, while token routing determines sparse expert work on expert-parallel ranks. Optimizing either alone can shift the bottleneck to the other. In MoE RL, rollout-time routing replay exposes every sample's se
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route decision remains on device. We formalize the ready-cohort boundary using fixed-partition share F, exact offline share P
Evaluating OpenMP Offloading for Intra-node Multi-GPU Programming across NVIDIA, AMD, and Intel Architectures: A 3D Heat Transfer Case Study
Currently, most supercomputers are equipped with GPUs from manufacturers such as NVIDIA, AMD, or Intel, which provide substantial parallelism and high throughput. It is common for a single compute node (intra-node) to host multiple GPUs, typically four or more. Therefore, effectively leveraging all these GPUs within a single compute node is essential for applications in scientific and engineering
User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling
Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources
Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines
Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis. Running all models on the GPU creates a serial bottleneck that limits real-time throughput as pipeline stages grow. Modern edge SoC
Enabling Differentiated QoS Degradation for Replicated Databases under Failures
Elasticity is commonly presented as the default response to capacity loss after failures, since replacement replicas can compensate for failed nodes and restore pre-incident service levels. Replacement capacity entails both delay and additional resource commitment, as replicas must be provisioned and synchronized before they can serve traffic. Under fixed budgets or constrained operating condition
Descriptive Dispatch of Computational Work
Agents powered by AI/ML are becoming ingrained in orchestration. Dispatch of work is the task of receiving a request, transforming it for a workload manager, and successfully submitting it. Running scientific workflows across multi-cluster environments introduces substantial challenges of dynamic job transformation, dispatch, and submission to heterogeneous clusters. These tasks are well-suited to
Scheduling Mixed RL Rollouts Beyond Prefix Locality
Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how heterogeneous rollout sessions compete for KV-cache capacity. When reinforcement learning with verifia
QSimAdv: A Late-Bound, Vendor-Agnostic Architecture for High-Performance Quantum-Circuit Simulation
Portability in high-performance quantum-circuit simulation need not begin at the kernel. We present QSimAdv, which makes late binding, rather than a common kernel, the basis of vendor independence. Representation, operator lowering, and data placement are bound only when their required inputs become available. Before full-state allocation, circuit, noise, and output inspection can route eligible g
An Event-Driven Cloud-Native Wearable Analytics Framework for Real-Time Clinical Workloads
Continuous physiological monitoring using consumer-grade wearables offers a transformative opportunity for clinical care and research, yet integration remains hindered by device heterogeneity, proprietary data formats, and strict regulatory requirements. We present an event-driven, cloud-native system designed to ingest, normalize, and analyze high-frequency vital signs from wearables at scale and
Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data
Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations, and intended use. However, these practices remain focused on static artifacts (the datasets and trained models themselves) while overlooking the workflow executions that produce, transform, and evaluate them. Such execu
ClusterBench: A Framework for Cluster-Wide Continuous Benchmarking and Regression Testing
Data centers need tooling that validates an entire installation rather than individual nodes, at acceptance and at regular intervals thereafter. This requires dispatching identical benchmarks to every node in a single submission, and therefore cluster-aware scheduling. This paper presents ClusterBench, a framework for cluster-wide continuous benchmarking. It ships with a benchmark collection targe
SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training
In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin. Existing diagnosis often relies on in-process monitors that cannot report after the trainer blocks or terminates, or on post-mortem logs that preserve only synchronized symptoms; offline health tests lose the workload and operating conditions that trigg
Cross-Species Transfer Learning for Electrophysiology-to-Transcriptomics Mapping in Cortical GABAergic Interneurons
Replicates and extends Gouwens et al.'s electrophysiology-to-transcriptomics framework using Allen Institute Patch-seq data from mouse/human cortex, focusing on GABAergic interneuron subclasses.
RCTs & Human Uplift Studies: Methodological Challenges and Practical Solutions for Frontier AI Evaluation
Examines human uplift studies (RCTs measuring AI effects on human performance) for frontier AI, highlighting underexamined interactions with AI properties in high-stakes decisions.
Leech Lattice Vector Quantization for Efficient LLM Compression
Explores Leech lattice-based vector quantization for LLMs, enabling joint parameter encoding beyond scalar limits without explicit codebooks via optimal sphere packing.
LLMGreenRec: LLM-Based Multi-Agent Recommender System for Sustainable E-Commerce
Introduces LLMGreenRec, a multi-agent LLM framework for e-commerce recommenders that promotes sustainable products while minimizing digital carbon footprints and capturing eco-friendly user intents.
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
Characterizes mechanisms by which VLMs predict artistic style and evaluates their alignment with art historians' criteria through interdisciplinary collaboration.
LiTo: Surface Light Field Tokenization
Introduces a 3D latent representation that jointly models object geometry and view-dependent appearance by encoding random subsamples of surface light fields from RGB-depth images into compact latent vectors.
COMIC: Agentic Sketch Comedy Generation
Proposes a fully automated AI system using agent populations mimicking studio roles to generate SNL-style comedic videos via iterative competition, evaluation, and improvement. Key contribution: LLM critics aligned with real viewer preferences through preference analysis.
Instruction set for the representation of graphs
Presents IsalGraph, a method encoding any finite simple graph as a compact string over a 9-character alphabet using a virtual machine with a CDLL of nodes and traversal pointers, where every string decodes to a valid graph.
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
Introduces V2M-Zero, a zero-pair video-to-music generator that aligns music temporally with video events by matching shared change timing and magnitude, ignoring semantic differences.
Neural Field Thermal Tomography: A Differentiable Physics Framework for Non-Destructive Evaluation
Presents Neural Field Thermal Tomography (NeFTY), a differentiable physics framework parameterizing 3D diffusivity as a continuous neural field for quantitative reconstruction of material properties from transient surface temperatures.
OrchMAS: Orchestrated Reasoning with Multi Collaborative Heterogeneous Scientific Expert Structured Agents
Demonstrates OrchMAS multi-agent system with reinforcement learning achieves consistent strong performance across diverse reasoning and scientific benchmarks, with public code available.
Beyond Factual Correctness: Mitigating Preference-Inconsistent Explanations in Explainable Recommendation
Introduces PURE, a select-then-generate framework for preference-consistent explanations in LLM recommenders, addressing inconsistencies missed by standard metrics.
On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions
Demonstrates Transformers approximate maxout networks, inheriting ReLU-like universal approximation with comparable complexity.
Safe and Robust Domains of Attraction for Discrete-Time Systems: A Set-Based Characterization and Certifiable Neural Network Estimation
Develops a framework for estimating safe, robust domains of attraction in uncertain, constrained nonlinear discrete-time systems.
Proactive Guiding Strategy for Item-side Fairness in Interactive Recommendation
Proposes proactive fairness in recommenders by guiding user preferences toward long-tail items, avoiding preference misalignment from direct insertion.
Compact Prompting in Instruction-tuned LLMs for Joint Argumentative Component Detection
Discusses argumentative component detection (ACD) in argument mining as a challenging task, with existing methods simplifying to labeling or pipelines.
From Complex Dynamics to DynFormer: Rethinking Transformers for PDEs
Critiques Transformer-based neural operators for uniformly treating spatial points in PDE solving, ignoring scale separation and incurring high costs.
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
Proves Adam's superiority over SGD via second-moment normalization under bounded variance using martingale analysis, explaining empirical convergence gaps.
Odin: Multi-Signal Graph Intelligence for Autonomous Discovery in Knowledge Graphs
Presents Odin, a production graph engine using COMPASS score (PageRank + NPLL) for autonomous pattern discovery in knowledge graphs.
Multi-Scale Adaptive Neighborhood Awareness Transformer For Graph Fraud Detection
Highlights GNN limitations in graph fraud detection due to homogeneity assumptions and poor global modeling, proposing solutions to these challenges.
Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
Introduces Procedure-Aware Evaluation (PAE) for LLM-based agents, assessing procedures via structured observations across Utility, Efficiency, Interaction Quality, and Procedural Integrity beyond mere task completion.
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
Addresses gradient instability in black-box attacks on LVLMs due to ViT sensitivity, improving transfer-based methods like M-Attack.
A.R.I.S.: Automated Recycling Identification System for E-Waste Classification Using Deep Learning
Presents A.R.I.S., a YOLOx-based portable sorter for real-time e-waste material classification to boost recycling efficiency.
Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting
Critiques scaling in time series foundation models for inefficiency despite performance gains, advocating alternatives.
MARS: Margin-Aware Reward-Modeling with Self-Refinement
Proposes difficulty-aware data augmentation for reward models in RLHF/RLAIF to improve alignment without costly human labels.
Multi-Round Human-AI Collaboration with User-Specified Requirements
Defines counterfactual harm and complementarity principles for conversational AI to reliably aid high-stakes human decisions via user-defined rules.
Sink-Aware Pruning for Diffusion Language Models
DLMs suffer high inference costs from iterative denoising, and unlike AR LLMs, their attention-sink positions show high variance across generation, invalidating inherited pruning heuristics.
Mine and Refine: Optimizing Graded Relevance in E-commerce Search Retrieval
Introduces 'Mine and Refine' contrastive training for semantic embeddings handling graded relevance in e-commerce search with long-tail queries.
[2601.12538] Agentic Reasoning for Large Language Models
Differentiates in-context vs. post-training reasoning in agentic frameworks across domains like science and robotics.
(PDF) Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
Introduces MATTRL, injecting structured textual experience into multi-agent deliberation at inference time.
MAR: Multi-Agent Reflexion Improves Reasoning Abilities in LLMs
Reports MAR outperforms ReAct (32% EM) and Reflexion+ReAct (44% EM) on HotPotQA with 47% EM, attributing modest gains to EM metric limitations.
Benchmarking Multi-Agent AI: Insights & Practical Use | Galileo
Benchmark supports diverse agent architectures for multi-agent system comparisons.
LLM-Based Multi-Agent Systems for Mathematical Problem Solving: A Comprehensive Literature Review[v1] | Preprints.org
Details math benchmarks (MATH500 etc.) with hierarchical multi-agent setups using CoT prompting and RL fine-tuning.
Towards a Science of Scaling Agent Systems
Defines multi-agent scaling via agents, coordination, models, and tasks, evaluated on benchmarks like Finance-Agent.
From single-agent to multi-agent: a comprehensive review of LLM-based legal agents
Reviews legal AI enhancements like syllogism prompts, logic benchmarks, retrieve-then-read, and emotional interaction.
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
Analyzes lazy agents in multi-agent LLM frameworks and promotes online RL for balanced contributions.
MARS: Reinforcing Multi-Agent Reasoning of LLMs through Self-Play in Strategic Games
Emphasizes developing LLMs for effective cooperation and competition in multi-agent systems toward advanced intelligence.
AgentAsk: Multi-Agent Systems Need to Ask
Introduces AgentAsk, a three-stage pipeline (distillation, supervision, E-GRPO optimization) for agent querying that improves accuracy, latency, and cost on math, reasoning, and coding benchmarks.
An Economy of AI Agents
Examines emergent economic behaviors when AI agents interact in market-like environments.