Shows that reputation scores from past performance fail to predict future agent behavior under adaptation or malice and introduces a task-conditioned trustworthiness assessment.
MiniRep addresses the failure of historical reputation in multi-agent debate by aggregating current proposal quality with persistent reputation, effectively detecting and penalizing malicious adaptation and clone-group collusion.
By integrating a Response Analyzer that evaluates immediate behavior and a Reputation-Aware Aggregator that limits the influence of agents with identical underlying models, MiniRep prevents adversaries from exploiting high historical scores to inject incorrect answers. This approach significantly outperforms conventional reputation baselines in MATH benchmarks under various attack conditions.