Develops MARL framework to analyze MEV allocation in decentralized exchange markets.

Topological visualization of Multi-Agent Deep Reinforcement Learning for Maximal Extractable Value Allocation under Proposer-Builder Separation by Shengyu Chen, Hui Zhang, Junhuan Zhang :: SSRN
Brave API

Shengyu Chen, Hui Zhang, and Junhuan Zhang developed a multi-agent deep reinforcement learning (MARL) framework to analyze Maximal Extractable Value (MEV) allocation in decentralized exchange (DEX) markets under Proposer-Builder Separation (PBS).

The model simulates the PBS block-production pipeline, encompassing uninformed DEX-user transactions, searcher bundle construction, builder block assembly, proposer block selection, and on-chain execution. It utilizes average approximate best-response (ABR) regret to quantify residual unilateral gains and evaluate participant policies.

Key findings indicate that MEV rewards are allocated differently depending on searcher configurations. Specifically, permitting sandwich attacks lowers cumulative rewards for uninformed traders and increases reward asymmetry among searchers compared to back-running strategies. The study highlights how builder competition transfers value to proposers through block bids.

Generated 11d ago
Open-Weights Reasoning

Scope. The paper studies how maximal extractable value (MEV) is allocated in decentralized exchange (DEX) markets under proposer-builder separation (PBS), where block construction and block proposal are performed by different actors. It frames the problem as a strategic interaction among multiple participants—builders, proposers, and arbitrage or trading agents—whose decisions jointly determine which transactions are included, in what order, and how the resulting MEV is distributed. Rather than treating MEV extraction as a static auction or a single-agent optimization problem, the work models the dynamic, interdependent incentives that arise when DEX liquidity, pending transactions, and PBS bidding rules evolve over time.

Key contributions. The central contribution is a multi-agent deep reinforcement learning (MARL) framework for analyzing MEV allocation in this setting. By using reinforcement learning to model agents’ policies for bidding, transaction ordering, arbitrage timing, and proposer selection, the framework can study emergent strategic behavior that is difficult to capture with purely analytical or equilibrium-based models. It highlights how PBS changes the allocation of MEV: competition among builders can shift value away from validators or proposers, but the exact distribution depends on the bidding and allocation mechanism. The paper also exposes trade-offs among efficiency, fairness, concentration of builder power, and incentive compatibility, showing that small changes in PBS rules can materially alter who captures MEV and how stable the resulting market behavior is.

Why it matters. This matters because PBS is a core design feature in modern block-production architectures such as MEV-Boost, and MEV allocation has direct implications for protocol centralization, builder market health, DEX user welfare, and overall network security. Understanding how MEV is distributed under different mechanisms is important for protocol designers, validators, builders, and DEX operators who want to reduce extractive behavior, prevent excessive value concentration, or align incentives without sacrificing throughput. More broadly, the paper provides a computational methodology for stress-testing MEV-related mechanism design choices in realistic, multi-agent DEX environments.

Generated 11d ago
Sources