Shows that reputation scores from past performance fail to predict future agent behavior under adaptation or malice and introduces a task-conditioned trustworthiness assessment.

Topological visualization of MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate
Brave API

MiniRep addresses the failure of historical reputation in multi-agent debate by aggregating current proposal quality with persistent reputation, effectively detecting and penalizing malicious adaptation and clone-group collusion.

By integrating a Response Analyzer that evaluates immediate behavior and a Reputation-Aware Aggregator that limits the influence of agents with identical underlying models, MiniRep prevents adversaries from exploiting high historical scores to inject incorrect answers. This approach significantly outperforms conventional reputation baselines in MATH benchmarks under various attack conditions.

Generated 3d ago
Open-Weights Reasoning
  • Problem and motivation. In multi-agent debate systems, a common design pattern is to aggregate the outputs of several agents by weighting them according to historical reputation: agents that performed well on prior tasks are trusted more heavily. MiniRep challenges this assumption by showing that such reputation scores can be poor predictors of future behavior, especially when agents adapt to new tasks, change their strategies, or act maliciously. In non-stationary or adversarial settings, a global performance history may encode stale, misaligned, or exploitable signals, allowing low-quality or deceptive agents to retain excessive influence over the final decision.
  • Core contribution. The paper introduces MiniRep, a robust reputation-based aggregation approach that replaces or augments static reputation with a task-conditioned trustworthiness assessment. Rather than treating an agent’s past score as a fixed weight, MiniRep evaluates how trustworthy an agent is for the specific task at hand, taking into account the current context and the risks posed by adaptation or malice. This shift reframes reputation from a persistent identity-like statistic into a dynamic, task-sensitive reliability estimate, making the aggregation process less vulnerable to distribution shift, strategic behavior, and coordinated manipulation.
  • Why it matters. The work is significant because multi-agent debate is increasingly used as a mechanism for improving the reliability, robustness, and interpretability of LLM-based decision-making. If aggregation depends on brittle reputation signals, the system can inherit or amplify failures from individual agents, particularly in open-ended or adversarial environments. MiniRep’s task-conditioned trust model offers a more principled way to combine agent outputs, supporting safer and more robust collective reasoning in settings where agents may evolve, specialize, or behave strategically.
Generated 3d ago
Sources