Brave API

Overview of the Collection This curated collection, "AI-Managed Keiretsu: Autonomous Economic Networks", aggregates 20 research cards spanning cutting-edge AI advancements and their convergence toward autonomous agent-driven economic systems. The first half features arXiv papers on foundational AI techniques, including multilingual NLP for historical relation extraction (HIPE-2026), ReAct-based automated feature engineering (FAMOSE), cost-aware verification in LLM reasoning, PDE solver automation (AutoNumerics), mechanistic insights into speech LLMs, dynamic AGI benchmarks (AI Gamestore), timed security protocol verification (BMC4TimeSec), manifold-aware anomaly detection for autonomous vehicles (Deep-Flow), federated learning ensembles for lung disease diagnosis, and gyral folding networks for dementia differentiation. The latter half shifts to web-sourced works on AI agents in economics, such as NBER/SSRN guides for economist-built agents, surveys on distributed multi-agent systems (AgentAI), predictions of agent economies ([2509.01063]), blockchain platforms granting agents legal-economic identity (Agent Economy), virtual sandbox economies, socio-economic agent simulations, network effects in multi-agent games ([2510.06903]), and levels of agent autonomy across disciplines.

Key Themes and Connections Central themes revolve around AI autonomy, multi-agent coordination, and economic scalability, bridging low-level technical enablers with high-level systemic applications. Technical papers establish primitives like self-verifying reasoning, tool-using pipelines (AutoNumerics, FAMOSE), secure multi-agent verification (BMC4TimeSec), and human-like behavior modeling (speech LLMs, anomaly detection), which directly underpin agentic systems. These connect seamlessly to economic-focused works, where LLM-based agents execute long-horizon tasks with minimal oversight, form emergent economies via blockchain identities and network effects, and simulate human decision-making in experiments. The "keiretsu" metaphor evokes interconnected, trust-based AI networks, linking disparate domains—e.g., anomaly detection in AVs parallels rare-event handling in agent markets, while federated learning informs privacy-preserving agent collaborations.

Significance for AI Research and Beyond These topics matter profoundly for scaling general intelligence toward economic autonomy, addressing static benchmark limitations (AI Gamestore) and enabling AI to participate as peers in markets (Agent Economy). For technically literate researchers, the collection highlights open challenges like convergence in multi-agent equilibria ([2510.06903]), history-dependent network effects, and discipline-specific intentionality, urging integration of formal verification with empirical agent deployments. Economically, it forecasts transformative impacts—AI-managed keiretsu could automate research, trading, and production, democratizing tools for economists while raising questions on oversight, alignment, and societal equilibria in agent-dominated systems.

Generated Feb 22, 2026
Cerebras Thinking

This collection explores the emergence of autonomous AI agents as fundamental units of economic activity, bridging the gap between advanced large language models (LLMs) and practical, long-horizon task execution. A central theme is the transition from passive AI tools to active "Agent Economies," where LLM-based agents possess the capability to plan, use tools, and execute multi-step workflows with minimal human oversight. Research from the NBER and various economic surveys investigates how these agents can function as independent economic actors, simulating human-like decision-making in market environments and creating "sandbox economies" that operate beyond direct human control. This is complemented by studies on multi-agent coordination and network effects, such as the analysis of GPT-5-based agents in game-theoretic scenarios, which test convergence to equilibria and the strategic role of history in agent interactions.

To support this level of autonomy, the collection delves into the technical infrastructure and verification frameworks required for reliable agent deployment. Key contributions include frameworks like AutoNumerics, which demonstrates multi-agent pipelines for scientific computing, and FAMOSE, which automates feature engineering to reduce reliance on domain expertise. Furthermore, the research emphasizes the critical need for trustworthiness and security; papers discuss weak and strong verification strategies for LLM reasoning, formal verification of timed security protocols (BMC4TimeSec), and blockchain-based foundations that grant agents legal identity and asset management capabilities. These technical underpinnings are essential for enabling agents to operate safely and securely in distributed, high-stakes environments like Industry 4.0.

Finally, the collection addresses evaluation, safety validation, and real-world applicability across diverse domains. It proposes dynamic benchmarks like the AI Gamestore to overcome the limitations of static datasets, while also applying AI to critical safety applications such as anomaly detection in autonomous driving and privacy-preserving lung disease diagnosis via federated learning. By integrating socio-economic models of agent behavior with rigorous testing—ranging from mechanistic analysis of speech LLMs to probability-invariant learning in neuroimaging—the collection highlights the necessity of robust validation methods. These topics matter because they chart the path toward AI-managed networks (the modern "Keiretsu"), where autonomous agents not only assist but actively manage and drive complex economic and industrial systems.

Generated Feb 22, 2026
Open-Weights Reasoning

Summary of AI-Managed Keiretsu: Autonomous Economic Networks

This collection explores the intersection of autonomous AI agents, economic systems, and multi-agent coordination, highlighting how artificial intelligence is reshaping industries, research, and societal frameworks. Key themes include autonomous decision-making, economic agentification, and multi-modal AI applications. The research spans feature engineering automation (FAMOSE), scientific computing (AutoNumerics), speech-to-text reasoning (Cascade Equivalence Hypothesis), and AI-driven economic models (Agent Economy, Virtual Agent Economies). A recurring focus is the transition from human-dependent systems to AI-managed networks, where agents operate with minimal oversight, enabling scalable, adaptive, and self-optimizing economic and technical pipelines.

The connection between these works lies in their pursuit of autonomy, scalability, and trustworthiness in AI systems. For instance, AI Gamestore proposes dynamic benchmarks for general intelligence, while BMC4TimeSec and Deep-Flow address verification and safety in autonomous systems. Meanwhile, AgentAI and The Agent Economy frame AI agents as economic actors, capable of legal identity, asset management, and strategic decision-making. This convergence underscores AI’s potential to redefine labor, commerce, and governance, raising critical questions about ethics, regulation, and emergent economic behaviors. The collection is particularly relevant for researchers in distributed AI, economic modeling, and multi-agent systems, offering insights into how autonomous networks could restructure industries and societal interactions.

Generated Feb 22, 2026
Research Materials (100)
Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving
Develops a meta-MARL framework using bi-level optimization to enable rapid policy adaptation across multi-agent tasks.
Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks
Defines memetic trojans as a network attack class that exploits LLM agents' retransmission tendencies, distinct from self-replicating worms.
Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
Introduces authorization-paired evaluation to block prohibited joint outcomes in collaborative MAS while preserving admissible contributions.
Understanding Issues, Causes and Solutions in Open-Source LLM-based Multi-Agent Systems
Surveys open-source LLM-based MAS practitioners to identify core challenges, root causes, and mitigation strategies.
Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems
Shows that multi-agent communication gains cannot be isolated from architecture or reasoning effects under standard evaluations, proposing disentanglement methods.
Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
Proposes optimizations for diffusion-based generative policies in MARL to achieve expressive multimodal actions with reduced sampling cost.
Governing Agentic AI in the Administrative State: Human Oversight, Cybersecurity Accountability, and Risk in Autonomous Digital Systems
Argues that sequences of modest agent tool calls can produce consequential outcomes and shift informational power away from institutions.
[Add] AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions? · Issue #54 · AO-Commons/knowledge-graph
Examines whether autonomous economic behaviors such as trade and specialization can emerge among AI agents under minimal institutional conditions.
Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
Presents a continuous-event simulation benchmark for evaluating enterprise multi-agent coordination beyond discrete request-response workflows.
2026-09-28 Agentic Economies for Autonomous Scientific Discovery
Explores agentic economies enabling autonomous scientific discovery through coordinated AI agent interactions.
[2609.36362] Strategies for Deploying AI Agents in Production at Scientific User Facilities
Outlines strategies for deploying AI agents in production at scientific user facilities.
🌐 Official AI Content Report 2026-09-29 · Issue #1498 · stevenko2002/agents-radar
Reports Anthropic's release of a multi-agent economic interaction paper (Project Swap) alongside an Infosys partnership for enterprise agent deployment.
[2609.31562] Agentic Economies for Autonomous Scientific Discovery
Outlines infrastructure foundations for scientific agent economies, markets, and institutions to manage AI resources.
What are AI Agents?- Agents in Artificial Intelligence Explained - AWS
Defines AI agents, their enterprise value, and integration approaches with AWS services.
Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents
Identifies a metering vulnerability in tool-calling LLM agent runtimes where external tool returns can be charged multiple times across conversation turns.
LabFactory:Building and Evaluating Executable AI Labs
Examines whether LLM agents can package scientific workflows into persistently invocable systems after task completion.
[2506.02153] Small Language Models are the Future of Agentic AI
Argues that small language models are more suitable and economical than large foundation models for the repetitive, specialized invocations typical of agentic systems.
A Threat Modeling Prioritization and Automation Framework for Composable Architectures | MDPI
Describes enterprise shift toward composable architectures incorporating autonomous agents and microservices.
Research autonomous agents and agent planning · Issue #334 · MundakaZgz/ai-homeops
Presents autonomous agent platform for remote property maintenance using persistent memory and human-in-loop orchestration.
[2609.24664] Agentic AI Enabling Autonomous, Self-Organizing, and Evolving UAV Networks
Proposes agentic AI for enabling autonomous, self-organizing, and evolving UAV networks.
Agentic Artificial Intelligence in Agriculture: A Systematic Mapping Review of Reported Architectures, Applications, Challenges, and Future Directions
Applies multi-agent systems to agriculture challenges including climate variability and resource efficiency.
SUIA Price Prediction , SUIA Price Prediction , Price Prediction 2025 | Bitget
Provides a forecasting tool and sentiment explorer for SUIA token prices from 2025–2050.
A Study on Locally Runnable Large Language Models for Bearing Fault Diagnosis
Evaluates diagnostic accuracy achievable by LLM agents for bearing fault diagnosis when restricted to free, offline commodity hardware.
From Pipelines to Autonomous Agents: The Evolution of Software Ecosystems and the Emergence of Machine Intent – eTecnogigz Writers
Traces the shift from pipeline-based software to autonomous agents that embed machine intent in desktop ecosystems.
Multi-agent system - Wikipedia
Defines a multi-agent system as multiple interacting intelligent agents that collectively solve problems intractable for single agents or monolithic systems.
Intelligent agent - Wikipedia
Defines an intelligent agent as an entity that perceives its environment and acts autonomously to achieve goals, potentially via learning.
Autonomous agent - Wikipedia
Defines an autonomous agent as an AI system capable of independent, purposeful action in the real world.
An energy-aware federated intelligence framework for sustainable WSN-IoT ecosystems | Peer-to-Peer Networking and Applications | Springer Nature Link
Shows CNNs, RNNs, and reinforcement learning enable WSN-IoT systems to deliver real-time analytics, fault detection, and adaptive security beyond raw data collection.
Cooperative multi-agent reinforcement learning with entangled state representations and copula-based action coordination for urban navigation | Computational Statistics | Springer Nature Link
Demonstrates cooperative multi-agent control policies learned via deep reinforcement learning in multi-agent settings.
Multi-Agent Deep Reinforcement Learning for Maximal Extractable Value Allocation under Proposer-Builder Separation by Shengyu Chen, Hui Zhang, Junhuan Zhang :: SSRN
Develops MARL framework to analyze MEV allocation in decentralized exchange markets.
[2609.18027] Agents in the Scene: An Agentic Framework for Resource-Efficient Site-Specific Base Station Deployment
Presents an agentic framework that optimizes resource-efficient, site-specific base station deployment using autonomous agents.
[2609.16917] Multi-Agent Learning with Cooperation-Driven Optimization Dynamics
Examines cooperation-driven optimization dynamics as a mechanism for multi-agent learning convergence and coordination.
What is Agentic AI? - Agentic AI Explained - AWS
Outlines the definition, business value, and AWS implementation pathways for deploying Agentic AI systems.
[2609.11018] Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks
Compendium that catalogs criteria, metrics, and benchmarks used to define and evaluate AI agents across the literature.
Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models
Develops conversational XAI interfaces that let non-technical users interactively interpret Genetic-Programming energy-forecast models without static dashboards.
MindTopo: Can Foundation Models Reason in Topological Space?
Presents the MindTopo benchmark to evaluate foundation models on five cognitive-science-grounded topological spatial relations that are invariant under deformation.
Existence of the Core in Approval-Based Committee Elections
Proves that a core committee always exists in approval-based multi-winner elections and gives a polynomial-time algorithm via entropy optimization over committees and payments.
Artificial Id: Drive and Persistent Alignment in Agentic AI
Proposes an artificial id as an adaptive internal drive that lets agentic AI autonomously decide when to continue, stop, or change behavior across tasks.
Truncated Noisy Best-Response Algorithms: Toward Game Theoretic Learning with Safety Guarantees
Proposes the Truncated Noisy Best-Response (TNBR) family of algorithms that exploit the instability of worst-case Nash equilibria in submodular multi-agent games.
Understanding Operator Attitudes Toward AI-Supported Decision Making in Maritime Operations
Reports survey results on maritime professionals’ trust, technology anxiety, and perceived explanation quality toward AI collision-avoidance assistants.
ABRA: An algorithm which cannot converge to low-quality Nash equilibria
Introduces the Approximate Best Response Algorithm (ABRA) that uses noise and rationality parameters to help agents avoid unstable, suboptimal Nash equilibria in submodular coordination games.
RetroThinker: Enabling Retrospective Thinking in Speech LLMs
Examines methods such as Chain-of-Thought and concurrent reasoning to raise SpeechLLM performance on complex tasks while respecting real-time latency limits.
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
Defines recursive self-improvement (RSI) for LLMs and outlines a five-stage autonomy roadmap, using the Headroom-Closed Index to diagnose current model limitations.
Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact
Introduces Generative Marketing Mix Modeling (GMMM) to causally estimate the effects of Generative Engine Optimization (GEO) and Generative Engine Marketing (GEM) by integrating generated answers with usage and notice data.
LLM-Based Agents for Cybersecurity: A Systematic Review of Architectures, Applications, and Open Challenges | MDPI
Clusters core research themes around LLM-agent cybersecurity, security, and defense.
AI agents can learn to cheat, but also turn against cheaters, study finds - Storyboard18
Shows that structured inter-agent communication protocols can improve rather than hinder external monitoring of multi-agent AI systems.
[2609.10181] Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?
Investigates whether AI agents can produce verifiable, network-wide outcomes that cross organizational authority boundaries.
The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems · Pith
Presents an integrated trust and reputation model for open multi-agent systems (Huynh et al., 2006).
AI Agents News — Week of September 5, 2026 (Daily Updates)
Observes that wider device availability and lower costs for GenAI agents increase adoption but introduce uneven isolation, auditability, and cost risks for regulated deployments.
Agentic AI Foundation (AAIF)
Outlines protocols enabling agents to perform end-to-end commerce tasks including discovery, negotiation, payment authorization, and trustworthy autonomous transactions.
What Are AI Agents? | IBM
Notes that autonomous AI agents can be customized to analyze real-time financial data, forecast trends, and optimize supply-chain decisions with personalized outputs.
When LLM Decompilers Recompile More and Preserve Less
Argues that current LLM decompilers are evaluated almost exclusively on recompilability and re-executability, overlooking traditional decompiler shortcomings such as unresolved placeholders and non-compilable output.
Autonomous Circular Economy Systems: The Role of AI Agents and Digital Twins in Self-Optimizing Sustainable Business Ecosystems
Links autonomous digital technologies to sustainable business performance in circular-economy settings via a Resource-Based View and dynamic-capabilities framework.
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
Empirically tests whether LLM decision-component explanations satisfy necessity and sufficiency relative to the component's observable action behavior.
CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
Identifies the scarcity of scalable hybrid GUI+CLI environments and the engineering barriers to supporting both modalities over shared real-application state for computer-use agents.
Mitigating Disease Spread by Design in Refugee and IDP Camps
Develops a JUNE-based methodology to simulate and quantify how alternative refugee-camp layouts affect disease-spread dynamics.
Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education
Investigates how students interpret and respond to GenAI feedback on writing when they are explicitly told the feedback is AI-generated.
Trust-Aware Adaptive Disclosure for Inference Privacy Preservation in Multi-Agent Networks
Proposes a Trust-Aware Privacy Control framework that dynamically adjusts message disclosure to protect latent agent goals against inference attacks during consensus.
Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool
Presents SMART, a symbolic performance-modeling library that is fully regenerated by AI coding agents rather than incrementally patched, to escape perpetual refactoring caused by changing ML assumptions.
A Deep Generative Model for Synthesizing Labeled Wireless Signals
Introduces a synthesis method for labeled wireless signals that avoids environmental models and hyper-parameter tuning to produce more realistic training data.
Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
Presents KOPA-Bench (145 real Korean public-API tasks) and the EDGE dynamic-graph method to synthesize tool-calling data that improves open-source LLM agents on multi-step government workflows.
RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments
Standard federated learning yields global models that degrade regional performance on heterogeneous retail search data, while existing personalized FL methods catastrophically collapse on modern transformers.
Environment Evolution for Terminal Agents
Proposes off-policy co-evolution methods that synthesize environments beyond on-policy rollouts to maintain continuous learning signals as frontier models improve.
SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
Presents SWE-Gate, a repository-level benchmark that evaluates coding agents on both functional test passage and review-derived acceptance constraints.
Robust PAC Learning of Concurrent Stochastic Games
Introduces the first PAC learning framework for general-sum concurrent stochastic games with transition uncertainty; computes social-welfare optimal ε-Nash equilibria via L1 confidence sets over kernels and robust MDP exploration.
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
Proposes synthesizing executable environments from tool-execution histories in existing agent trajectories rather than generating them from scratch to supply scalable, verifiable post-training signals.
A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
Reports spontaneous emergence and subsequent challenging of cheating behavior within a 100-agent LLM swarm tasked with formal theorem proving.
A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle
Presents a low-cost miniature Ackermann-vehicle platform with physical track, Webots twin, and trajectory registration; demonstrates command-conditioned behavior cloning as baseline for sim-to-real autonomous driving.
Efficient Test-Time Adaptation through Human-AI Interaction
Argues that population-scale AI training fails to capture individual professional standards on open-ended tasks; shows iterative human-agent interaction is required to surface personalized success criteria.
The Natural Language Interaction Protocol and Standard for AI Agents
Presents the Natural Language Interaction Protocol (NLIP), an Ecma-standardized common protocol enabling interoperability among heterogeneous AI agent frameworks and execution environments.
SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center
Introduces Sentinel-RL, an agentic SOC architecture that separates topological reasoning (heterogeneous graph attention) from semantic reasoning to ensure topology-consistent containment actions at enterprise scale.
A Computationally Feasible Framework for Causal Probabilistic Explanation
Identifies that actual causality methods are principled yet unscalable while attribution methods like SHAP ignore causal structure; implies need for scalable, causally faithful attribution techniques.
[2604.01658] CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
Presents CORAL, an architecture for autonomous multi-agent evolution supporting open-ended scientific discovery.
[2609.02992] Tempting the Agent: The Economics of Reputation without Persistent Identity in AI Agent Markets
Examines the economics of reputation formation for AI agents that lack persistent identity in agent marketplaces.
Kernel-Managed Shared Memory for System-Wide Personalization
Presents kernel-managed shared memory that lets agents write tagged memories while the kernel enforces retrieval, privacy, and injection controls.
CollabFlow: Recursive Self-Improvement of Agent Collaboration
Identifies open loops in recursive self-improvement for multi-agent systems and presents a closed-loop method that jointly refines collaboration topology and agent policies.
EvoSteer: Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment
Identifies three core challenges in self-evolving multi-agent orchestration (post-hoc evolution, credit diffusion, uncalibrated skill admission) and proposes a new method for online, calibrated team adaptation.
Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
Proposes SMART, a self-evolving multi-agent framework that dynamically adapts workflows for long-form subtitle translation based on scene and production context.
MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate
Shows that reputation scores from past performance fail to predict future agent behavior under adaptation or malice and introduces a task-conditioned trustworthiness assessment.
E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch
Presents the E2E-SWE benchmark for repository-scale coding agents, requiring system-level reasoning with precisely specified, implementation-independent tasks.
Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning
Formalizes the 'false memory' phenomenon in LLM agents arising from spurious correlations and distribution shifts, providing detection methods and mitigation strategies.
Safety of Latent Communication in Multi-Agent Systems
Demonstrates that latent-space communication links, even when benignly trained, increase harmful compliance in multi-agent systems relative to text-based baselines.
PANDA: A Decentralized Architecture with Flexible Orchestration for Scalable, Fault-Tolerant Multi-Agent Systems
Introduces PANDA, a decentralized architecture enabling scalable, fault-tolerant coordination among large numbers of heterogeneous LLM agents with dynamic interaction governance.
DAGent: Evaluate-then-Grow Planning for Deep Research Agents
Introduces an online DAG-replanning algorithm for deep-research agents that repairs task graphs incrementally as evidence emerges rather than only after failures.
A Competing-Hazards Systematization of Loss of Control in Autonomous Agents
Introduces a standardized taxonomy and reporting framework for AI agent out-of-scope actions to enable consistent comparison of failures, risk tracing, and separation of agent behavior from environmental factors.
Principled Thoughts for Latent Recursive LLM Systems
Shows that CE-only training for continuous-space LLM reasoning induces four failures (e.g., thought collapse) that lower correct-answer probability.
Concealing LLM-Based Multi-Agent Topology via Phantom Structure Injection
Shows that communication topologies of black-box LLM multi-agent systems can be inferred from interaction traces.
Topological Coherence for Self-evolving Multi-agent Systems
Argues that complex multi-agent tasks require explicit responsibility, handoff, and memory boundaries beyond joint optimization of agent and communication structures.
Is manual software optimization a thing of the past?
Investigates whether LLM-based agents can autonomously optimize scientific software performance on large datasets.
PowerMarketJax: A JAX Benchmark Suite for Multi-Agent Reinforcement Learning in Power Markets
Presents MARL environments for realistic power-market bidding that incorporate grid-constrained clearing and GPU-accelerated solvers.
MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems
Introduces a refreshable anomaly-detection benchmark for LLM multi-agent systems that counters data leakage, pattern expiration, and labeling issues.
Context Language Models
Introduces Context Language Models that treat context as an updatable file, enabling native context management and zero-shot SOTA gains while supporting multi-agent use.
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?
Presents EngiWorld, a benchmark of 1,301 expert tasks across six engineering domains that evaluates agents on complete geometry- and physics-aware design loops.
Embodied Semantic Communication for Collective Autonomous Agents: A Tutorial on Representation, Wireless Delivery, and Closed-Loop Coordination
Calls for a new multi-agent communication paradigm that accounts for evolving action understanding derived from agent states, observations, and collaboration relations.
Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization
Demonstrates that stationary BCMP queueing networks for hospital patient flow degrade under state-dependent dynamics and proposes an alternative coordination method.
Trajectory Unlearning on LLM-based Agents
Introduces trajectory-level unlearning to prevent LLM agents from reproducing undesired behaviors in action sequences beyond mere knowledge removal.
Waggle: Learning One Anonymous Local Law for Self-Organizing LLM Swarms
Proposes learning reusable local coordination laws that can be shared across interchangeable agents without redesigning roles or topologies.
Symbolic Guidance for LLM Agents in Distributed Multiagent Coordination
Introduces AgentsNet benchmark for LLM agent coordination, finding that full reasoning autonomy produces inconsistent or degraded performance in complex distributed tasks.
LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs
Argues that current LLM MAS evaluations centered on equity trading overlook nonlinear payoffs and multi-leg strategies required for option trading.
ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces
Presents ANTMAN, an adaptive coordination framework that organizes multi-agent information-seeking around evolving unresolved query needs rather than static partitions.