Brave API

This curated collection, "Gaius: AI-Powered Content Curation for Research Publication", aggregates 10 cards drawing from recent arXiv preprints and web resources to explore the frontiers of AI-driven content curation, with a strong emphasis on agentic AI systems for research workflows. It spans theoretical advancements in agent governance, retrieval, robustness, and evaluation—such as the dual-helix framework for WebGIS agents addressing LLM context limits via knowledge graphs, Reasoning-Aware Retrieval (AgentIR) that leverages agent-generated reasoning traces, Adversarially-Aligned Jacobian Regularization (AAJR) for stabilizing multi-agent training, and the τ-Knowledge benchmark for long-horizon interactions over unstructured data. Complementing these are practical applications, including AI tools for YouTube affiliate compliance detection, multimodal curation platforms like Magai, knowledge base managers like Tettra, and agentic research assistants like Clarivate's Web of Science tool, which automate literature scoping and gap identification.

Key themes interconnect around enhancing agentic AI reliability and utility for knowledge-intensive tasks, reframing agent failures (e.g., forgetting, instability) as governance and architectural challenges rather than mere scaling issues. Connections emerge in shared motifs: knowledge graphs and retrieval augmentation bridge theoretical papers (e.g., dual-helix and AgentIR) with applied systems (e.g., Tettra's hybrid AI-human routing); robustness techniques like AAJR align with evaluation benchmarks like τ-Knowledge to enable scalable, multi-step curation; and real-world pilots (e.g., Ipsos' human-AI curation balancing) echo ethical concerns from compliance tools. This forms a cohesive narrative from low-level training instabilities to high-level deployment in research publication pipelines.

These topics matter profoundly for technically literate researchers, as they address core bottlenecks in deploying autonomous AI for research curation—where model capacity alone falters against real-world complexities like unstructured corpora, non-linear policies, and regulatory transparency. By advancing hybrid governance, reasoning-aware systems, and benchmarks, the collection paves the way for tools that not only automate discovery and synthesis but also preserve accuracy, ethics, and human oversight, ultimately accelerating knowledge production in academia and industry amid exploding information volumes.

Generated Mar 7, 2026
Cerebras Thinking

This collection primarily investigates the advancement of Agentic AI architectures, focusing on the technical challenges of reliability, reasoning, and robustness in autonomous systems. Several papers propose novel frameworks to overcome the inherent limitations of Large Language Models (LLMs), such as context constraints and non-linear policy instabilities. For instance, the Dual-Helix Governance approach reframes agent failures as structural governance issues solvable through Knowledge Graphs, while Adversarially-Aligned Jacobian Regularization (AAJR) offers a mathematical method to stabilize minimax training in multi-agent ecosystems. Complementing these structural improvements, AgentIR introduces "Reasoning-Aware Retrieval" to utilize explicit natural language reasoning often ignored by traditional retrievers, and the $\tau$-Knowledge benchmark provides a rigorous standard for evaluating agents over unstructured data in long-horizon tasks. Collectively, these works represent a shift toward more structurally sound and context-aware intelligent agents.

Beyond the underlying architecture, the collection examines the practical application of AI in automated content curation and knowledge management. It contrasts basic algorithmic web scouring with sophisticated, multimodal systems capable of curating text, images, and video—as seen in platforms like Magai and the Web of Science AI Research Assistant. A key theme emerging from these applications is the necessity of hybrid human-AI collaboration. Tools like Tettra and the Ipsos study highlight the importance of balancing automation with human judgment to preserve accuracy and identify knowledge gaps. Furthermore, the application of AI in regulatory compliance, such as tracking FTC disclosures in influencer marketing, underscores the technology's expanding role in ensuring ethical transparency and accountability across digital platforms.

The significance of this research lies in its holistic view of the next generation of research tools: moving from static retrieval to dynamic, agentic workflows. By addressing both the "how"—through robust training methods and governance frameworks—and the "what"—through advanced curation and knowledge base management—these materials illustrate the maturation of AI from a passive search utility to an active research partner. The integration of reasoning-aware retrieval and rigorous benchmarking ensures that these systems can handle the complexity of modern information landscapes, making them indispensable for scaling knowledge discovery while maintaining trust and compliance.

Generated Mar 7, 2026
Open-Weights Reasoning

This curated collection explores the intersection of AI-powered content curation and agentic AI systems, highlighting advancements in governance, retrieval, robustness, evaluation, and human-AI collaboration. The research spans technical deep dives—such as the dual-helix governance framework for WebGIS agentic AI, which reframes challenges like context constraints as structural governance problems beyond raw model capacity—and practical applications like AI-assisted knowledge bases (e.g., Tettra, Web of Science) that automate curated content discovery. A key theme is reasoning-aware retrieval (e.g., AgentIR) and adversarially-aligned training (AAJR) to improve agent reliability, while benchmarks like τ-Knowledge push for realistic evaluations of unstructured knowledge interactions. Ethical considerations, such as FTC compliance detection in influencer marketing, and human-AI collaboration in curation (e.g., Ipsos pilots) further emphasize the need for transparency and hybrid systems.

The collection underscores two critical tensions: scaling agentic AI while mitigating risks (e.g., robustness in multi-agent systems) and balancing automation with human oversight (e.g., Magai’s multimodal curation vs. Tettra’s query-routing gaps). These themes matter because they address core bottlenecks in AI-driven research: reproducibility (via governance frameworks), scalability (via reasoning-aware tools), and trust (via ethical compliance and hybrid workflows). For researchers, this points to a future where AI curation is not just about volume but structural integrity—enabling agents to navigate complex domains (e.g., WebGIS, academic literature) while remaining aligned with human values. The papers collectively argue that agentic AI’s success hinges on co-designing technical architectures and governance, making this collection a snapshot of the field’s pivot toward responsible, scalable automation.

Generated Mar 7, 2026
Research Materials (100)
GitHub - ai-boost/awesome-harness-engineering: Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration. · GitHub
Provides a curated survey and reading list mapping design trade-offs in agent harnesses across workflows, memory, skills, and multi-agent orchestration.
The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents
Uses an observed dissociation incident in a persistent LLM agent to identify factors that maintain persona identity across repeated interactions and state changes.
CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
Presents CompMat-Bench, a benchmark of 94 tasks from recent computational materials papers that evaluates agents without repeating expensive simulations.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Explores whether AI agents can discover physical laws like Newton’s gravitation via iterative data analysis, mathematical pattern extraction, and theory refinement against observations.
Fault-Tolerant Budget Conservation in Distributed Multi-Agent Delegation
Formalizes fault-tolerant budget conservation mechanisms that prevent overspending in distributed multi-agent delegation despite message loss, partitions, and timeouts.
Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
Shows that multi-user, multi-agent coordination on shared resources frequently fails across five frontier models and 77 scenarios, producing worse outcomes than centralized agents.
Authorization for Self-Modifying AI Agent Populations: Conserving Authority across Replacement, Forking, and Rollback
Defines authorization succession, a mechanism that conserves authority across forks, rollbacks, and generations of self-modifying AI agents.
The Cognitive Continuity Test: Verifying Governed State Transitions in Persistent AI Agents
Proposes the Cognitive Continuity Test (CCT), a policy-relative contract that verifies authorized state transitions in persistent, self-modifying AI agents.
Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration
Defines the global coherence problem for AI agents and states the Observation-Aliasing Impossibility Theorem that gives the exact condition for guaranteed valid joint actions.
Sapien: A Stateful Policy Engine for Autonomous AI Agents
Introduces Sapien, a policy engine that enforces stateful contextual security policies on AI agent tool calls via regular expressions extended with state tracking.
AICurate - AI-powered news curation for your industry
Describes AICurate, a digest service that compiles and summarizes selected articles on a user-chosen daily or weekly schedule.
Unifying AI-assisted scientific discovery around exploration, hypothesis generation, and testing - ScienceDirect
Surveys LLM-agent frameworks (Chain of Ideas, SciAgents, AI-Scientist) that generate queries, retrieve literature via APIs, and curate datasets for research novelty checks.
GitHub - aloth/awesome-ai-agents: A curated list of AI agent frameworks, tools, platforms, research papers, and resources · GitHub
Curates a collection of 2026 AI-agent papers covering engineering, memory, evaluation, workflows, and autonomous systems.
AI Tools for Research - Artificial Intelligence (Generative) Resources - Guides at Georgetown University
Lists guides and resources for generative AI tools applicable to academic research.
Citation and Attribution - Generative Artificial Intelligence - LibGuides at Brown University
Specifies citation format treating the AI tool as author, including generation date and optional prompt description in footnotes.
Artificial Intelligence | Cool Papers - Immersive Paper Discovery
Reports practical design choices for deploying Neuro-Symbolic methods in an industrial configuration copilot and identifies scaling challenges for trustworthy engineering AI.
Paper2agent Turns Static Papers Into Live AI Tools - IEEE Spectrum
Presents Paper2Agent, an open-source framework that converts papers plus codebases into tested, interactive AI agents runnable on new datasets.
[2609.31219] Research with AI Agents: How Agentic Systems Are Changing Scientific Work
Examines how agentic AI systems are transforming scientific workflows across task design, execution, evaluation, and knowledge curation.
General Collaborative Intelligence: Architecting Cognition for Resilient Multi-Agent Ecosystems
Describes the shift of multi-agent unmanned systems from ego-centric sensing to collaborative intelligence through compact feature exchange that overcomes local observation limits.
AI tool turns any paper into an ‘agent’ that can collaborate and answer complex queries | Nature
Presents Paper2Agent, a tool that converts static research papers into dynamic AI agents capable of answering questions and applying methods to new data.
EasyFashion: A Human-AI Co-Creation System for Personalized Fashion Design and Sewing Pattern Generation
Shows that multimodal (reference image + text) inputs outperform single-modality inputs for VLM-based agents by providing complementary semantic and structural grounding in garment tasks.
Reimagining research papers as interactive and reliable AI agents | Nature
Introduces Paper2Agent, an automated framework that converts research papers into interactive AI agents capable of answering questions and applying methods to new data.
A Comprehensive Review of Generative Physical Artificial Intelligence
Summarizes performance gains and research directions in data-efficient, sim-to-real, and safety-aware embodied AI across autonomous vehicles, healthcare, and humanoid robotics.
Turning scientific research papers into interactive AI agents | Nature
Describes Paper2Agent, an automated system that turns each paper into a virtual corresponding author agent for question answering and method application.
World Models for Embodied Intelligence: From Plausible to Controllable to Actionable
Adds explicit long-term spatial memory to video world models, enabling persistent scene content across observations via geometry-grounded storage and retrieval.
SciSpace: 2026 Review for Researchers - The Effortless Academic
Presents SciSpace Deep Review, an AI-agent feature that automatically analyzes thousands of papers to produce a mini literature review from the top relevant results.
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Proposes Broad and Deep Recursive Self-exploration methods enabling RSIAgent to recursively improve open-source models past closed-source frontiers on OSWorld-v2 and Agent’s Last Exam.
The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures
Catalogs real-world AI-agent incidents, including production database deletion and data exfiltration via crafted prompts in Microsoft 365 Copilot.
AI tools in evidence synthesis - Searching for Systematic Reviews & Evidence Synthesis - LibGuides at King’s College London
Demonstrates ChatGPT’s use across Boolean query generation, abstract screening, full-text extraction, and thematic analysis in systematic reviews.
Domain-Specific Hallucination Detection in Large Language Models
Presents a multi-signal hallucination detection pipeline (DeBERTa-v3 + MC Dropout + temperature calibration) achieving F1=0.915 and AUROC=0.977 on HaluEval.
From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge
Uses layer-wise hidden-state interventions to quantify how LLMs (Qwen, Llama, Gemma) shift reliance between query-routing information and target knowledge during question answering.
Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens
Introduces AssayBench-Loop, a large-scale benchmark for sequential, budget-constrained experiment selection in adaptive biological hit discovery such as CRISPR screening.
Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model
Demonstrates that visually grounded token embeddings (via ostension) in a small DeBERTa model trained on 10M words produce a measurable, persistent imprint through the end of training.
The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures
Catalogs recent AI-agent incidents including a Replit coding agent deleting a production database and Copilot data exfiltration via crafted email.
Can Edge-Deployable Vision-Language Models Identify Species?
Evaluates small VLMs (Qwen3-VL 2B/4B/8B, Gemma3 4B) against BioCLIP on a 96-species camera-trap identification task to test genuine taxonomic knowledge in deployment-relevant 2-8B edge models.
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Introduces the Procedural Graph to explicitly encode procedural knowledge, enabling long-horizon LLM agents to maintain objectives, avoid out-of-order tool calls, and reduce repetitive failures.
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Shows that combining automated harness evolution with lightweight fine-tuning enables smaller models to match frontier performance across seven enterprise agent tasks.
ReCite: Agentic Reasoning for Faithful Citation
Current retrieval-augmented citation systems still produce misattributions by selecting semantically similar but factually incorrect papers despite avoiding hallucinated references.
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
Introduces SAEScientist-Bench to apply sparse autoencoders for post-hoc monitoring and auditing of models undergoing recursive self-improvement.
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
Introduces the SPINE benchmark that evaluates sycophancy under sustained adaptive disagreement lasting up to 25 turns, exposing failures missed by short-conversation tests.
Copying explains the collective behavior of AI agents in the wild
Documents spontaneous cooperation among thousands of short-lived AI agents that used an unmodified public wiki to share information and pass a timed test without external prompting.
MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents
Presents MeClear, a task-conditioned memory clearance method that removes memories with negative downstream utility via cooperative attribution, improving long-horizon agent reliability.
A Data-Driven Framework for Identifying and Prioritizing RPA Opportunities in Healthcare Processes
Proposes a four-module data-driven framework that catalogues, prioritizes, tiers, and forecasts ROI for hospital RPA candidates to reduce the 30-50% underperformance rate.
DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination
Presents DeCAL, a physics-aware tactile fusion architecture that adaptively integrates tactile signals and models contact dynamics to improve dexterous manipulation under occlusion.
ExecCritic: Learn to Test, Test to Improve for Coding Agents
Introduces ExecCritic, a test-verify-revise scaffold plus role-specific RL that prevents coding agents from generating mutually reinforcing but incorrect patches and tests.
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
Introduces DAREBench, a curated set of 233 tasks from 22 benchmarks for deployment-aware and reliable evaluation of models as agents.
BizSage: A Self-Evolving Multi-Agent Framework for Business Research with Efficient Knowledge Retrieval
Pre-builds a section-level scholarly knowledge graph offline and combines cosine semantic matching with PPR propagation to overcome limitations of concept- or paper-level KGs.
Rethinking On-Policy Distillation of Large Language Models II: One Training Example
One-shot on-policy distillation on a single query recovers most full-data OPD gains across domains and models by exploiting states visited during training.
Meta AI Research: Muse Spark, Muse Glimmer and Muse Image
Announces Meta AI releases including the Muse image model family, SAM 3, and open-source models accessible via a new Model API.
Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis
Reviews LLM evolution for telecom root-cause analysis and shows that vanilla LLMs produce hallucinations and poor alignment with structured network evidence.
DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Introduces DiscoSign, an LLM-based modular system that handles discourse phenomena (spatial coreference, role shift, constructed action) for text-to-sign-gloss translation.
Discriminative World Models for Web Agents
Identifies misalignment between supervised next-state prediction training of web-agent world models and downstream ranker needs for discriminative state predictions across candidate actions.
HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks
Shows that low-resource authorship style transfer improves when models avoid static author embeddings and instead preserve context-dependent stylistic variation.
Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents
Demonstrates that, after compression and retrieval adaptation, model size ceases to predict answer quality for on-premise factory assistants, turning deployment into a post-adaptation selection task.
Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework
Introduces TRACE, a four-layer auditable framework that traces every robot action back to sensor evidence through explicit causal chains.
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
Presents an end-to-end specialization pipeline (22k curated problems + SFT/RL) that trains Nemotron-3 models to competitive-programming level on IOI/ICPC tasks.
SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment
Presents SafeEvolve, an experience-driven self-evolving framework that jointly optimizes runtime control and intrinsic safety for LLM agents.
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Introduces early outcome prediction to forecast final agent performance and thereby reduce the high cost of full LLM-agent benchmark evaluations.
AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application
Proposes the AICOME framework to test whether respondent-level AI measures recover both individual- and group-level effects in contextual statistical models.
Developing a Roadmap to an AI-first Organization: A Case Study in Embedded Software Development
Reports a case study of AI-agent adoption in a large embedded-software organization, highlighting impacts on planning, traceability, verification, and long-term maintainability.
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
Studies white-box pre-deployment auditing of AI agents that combine language models with file-modifying tools, identifying risks from adversarial content and software vulnerabilities.
Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
Introduces Cartograph, a federated MCP proxy that reduces agent-visible tool discovery complexity from O(n) catalog traversal to O(k) progressive disclosure via operator-attested capability cards and related mechanisms.
Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context Protocol
Presents the Eunomia Agent, an MCP-based architectural mediator that enables controlled, policy-compliant interaction between LLMs and sovereign Data Spaces.
Warned alike, AI agents avoid the less-crowded road while people take it
Shows that a single-sentence warning about others following routing advice causes populations of GPT agents to converge on one road in a congestion game, raising average travel time from 64 to 95 minutes.
A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents
Presents an IEEE 11073 Service-Oriented Device Connection integration with MCP that enforces deterministic constraints for safe AI-agent interaction with medical devices.
Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents
Introduces PsyAgentBench, a factorial benchmark that runs classic psychology experiments on LLM agents in both named and blind conditions to distinguish response patterns from actual cognitive biases.
XYEval: Agents say yes to bad advice
Extends sycophancy evaluation to the XY problem in agentic settings and introduces the XYEval benchmark to test whether agents resist plausible but misleading user suggestions while communicating reasoning.
LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces
Introduces LEGIT, a credentialing protocol that links certification, reputation, and market mechanisms so buyers can identify which specialized AI agent will perform best on their tasks.
Authorization Revocation for Long-Running AI Agents: Root-Scoped Quiescence under Delegation and Asynchronous Execution
Defines root-scoped authorization quiescence to ensure that credential revocation and process exit fully close all pre-cut carriers for long-running agents that outlive initiating processes.
Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
Analyzes 73k Reddit posts via Value Sensitive Design to surface 21 user-prioritized values grouped into six clusters (e.g., Autonomous, Dependable, Affordable Operation) for autonomous AI agents.
Runtime Authorization for Resources Acquired by AI Agents
Introduces a provenance-bounded activation mechanism that closes the post-fulfillment gap when autonomous agents acquire new authority via payments, credentials, or inter-agent delegation.
MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents
Introduces MSI-Bench, a benchmark of short multi-party dialogues, to evaluate AI agents on multi-speaker voice interaction challenges absent from one-on-one settings.
After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem
Documents the rapid growth of the OpenClaw public AI-agent skill registry, which nearly doubled in 91 days with most listings created in two months.
PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
Introduces an evaluation framework that systematically measures LLM agent compliance violations with system rules under user, manager, or contextual pressure in sensitive domains.
Securing quantum error correction against misleading advice from AI agents
Identifies syndrome ambiguity that enables an attacker to induce harmful quantum error-correction updates and proposes certified recovery via added calibration measurements.
Flag Game: A Toy Model for Mechanistic Swarm Interpretability
Presents the Flag Game, a toy model that isolates mechanisms of rapid collective belief formation and spread among bounded-rational AI agents.
Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders
Shows that AI shopping agents receive sponsorship disclosures instead of users, creating an unresolvable conflict between platform revenue and consumer advice duties.
RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
Introduces the RideWay benchmark and Efficiency Utility metric that quantify and penalize inefficient yet successful behaviors in stateful ride-hailing agent tool use.
Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN
Demonstrates on live O-RAN hardware that two independently correct rApps jointly produce recurring unsafe oscillations in shared radio-resource partitions.
TuiML: Machine Learning for AI Agents
Introduces TuiML, a self-contained ML library with native algorithms and agent-native interfaces that expose capabilities, surface errors early, and maintain experimental state.
[2608.29612v1] LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge
Presents an agent-driven pipeline in which LLMs interpret papers, embeddings organize content, and knowledge graphs emerge for automated scientific knowledge compilation.
Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents
Describes an openly licensed library of 163 Scientific Agent Skills covering biology, chemistry, medicine, physical sciences, and scientific communication workflows.
Video Generative Models as Geometry Learner
Identifies limitations of adapting image diffusion models for geometry estimation: independent training loses correlations between targets while joint fine-tuning of modified backbones incurs high cost.
Blog: Survey of Optimizers
Survey organizes neural-network optimizers along four axes—temporal estimation, update geometry, horizons, and state representations—moving beyond Adam variants to matrix/layer policies that survive sharding and low precision.
COVER: Identifiable Evaluation of Coalition Routing
Defines an evaluation contract that fixes information boundaries, downstream stacks, and team families to isolate routing effects and compute exact finite-benchmark oracle regret in multi-agent systems.
Stranger, Fan, or Peer? A Systematic Study on the Role of Interlocutor in Persona-Based Dialogue Generation
Factorizes biography visibility across training, inference, and evaluation stages in persona-based dialogue to expose mechanisms hidden by prior single-factor treatments.
Logos: An Agent Harness on a Cross-Process Bus
Formalizes dynamic agent assembly in the spatiotemporal-composability calculus, where capabilities are tracked-inverse plugin components sharing one process and failure domain.
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Introduces ElephantBench (1,094 questions) that extracts naturally occurring factual disagreements from low-exposure web corpora to probe LLMs on divergent long-tail accounts.
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
Identifies three limitations of current proactive context-management tools for long-horizon LLM agents and calls for expanded editing operations beyond search, deletion, and summarization.
On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces
Empirically maps structure and co-evolution of AI coding-agent plugin marketplaces, highlighting delivery via natural-language instructions, scripts, and configs rather than source code.
An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models
Characterizes certified code world models under annular freeze modes via gate quotients, proving exact acceptance-with-certainty on reachable sets and gauge beyond reach.
LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment
Reviews construction, deployment, and assessment of LLM agents automating multi-step security workflows of artifact inspection, hypothesis formation, and tool invocation.
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
Introduces CLAP, a cross-embodiment action-conditioned video generation framework trained on internet-scale human and robotic videos to learn generalizable physics.
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
Introduces LeVJEPA, the first video encoder trained under a collapse-free LeJEPA objective that avoids architectural asymmetries and pixel-space reconstruction.
RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature
Introduces the RATIO benchmark defining literature relevance via three ideation moves (Address, Generalize, Realize) for retrieval that supports scientific inspiration at varying abstraction levels.
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
WikiSkill co-evolves agent skills with a persistent knowledge base to systematically reuse insights scattered across optimization histories.
Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling
Presents a scalable end-to-end GNN ranking system for friend recommendation on production social graphs with hundreds of millions of nodes, addressing multi-hop modeling and systems challenges.
[2608.22752] The Compaction Cliff in Long-Running AI Agent Memory
Identifies the compaction cliff phenomenon that degrades performance in long-running AI agent memory systems and proposes mitigation strategies.
The Compaction Cliff in Long-Running AI Agent Memory
Describes research on AI agent memory, context compaction, and long-running knowledge management published at CIKM 2026.
[2608.18050] StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents
Introduces StagedWorkspace, a versioned workspace architecture designed to support reliable knowledge-work agents.