arXiv:2610.00714v1 Announce Type: new Abstract: Memory tiering has been used to expand memory capacity, particularly in datacenters, by combining fast DRAM with slower tiers, including CXL-attached memory. Its effectiveness depends on keeping useful pages in the fast tier, but existing heuristic policies can lag behind changing hot sets in phased or bursty workloads. To explore these limitations,
MANTA (arXiv:2610.00714v1) is a Machine Learning Augmented Tiering Advisor that improves datacenter memory efficiency by predicting future page usefulness to keep hot data in fast DRAM, outperforming heuristic-based systems like ARMS. Developed by Johannes Freischuetz, Shivaram Venkataraman, and colleagues, MANTA integrates a lightweight learned model into the ARMS framework to better handle phased or bursty workloads where existing policies lag behind changing hot sets.
Key performance gains include geometric-mean speedups over ARMS of 1.12× on emulated CXL with Linux 6.2 and 1.69× across six Optane workloads at 4 GB of fast memory. The system uses a scalable offline optimizer (ChOMP) and a trace-driven simulator to identify performance opportunities, achieving up to 5.6× speedup on individual Optane workloads and maintaining effectiveness across different Linux kernels and memory capacities.
MANTA: Machine Learning Augmented Tiering Advisor
The paper examines memory tiering as a way to expand effective memory capacity in datacenter systems by combining fast DRAM with slower memory tiers, including CXL-attached memory. Its central concern is that the performance of such systems depends heavily on keeping “useful” pages in the fast tier, but existing heuristic tiering policies can struggle when workloads change rapidly. In particular, phased or bursty workloads can shift hot sets quickly, causing static or reactive heuristics to place pages suboptimally and increase accesses to slower memory.
The key contribution is MANTA, a machine-learning-augmented tiering advisor intended to improve page-placement decisions across memory tiers. Rather than relying solely on fixed heuristics, the approach uses learned signals to better predict or infer which pages are likely to be valuable in fast memory under changing workload behavior. This positions MANTA as a more adaptive alternative for tiering management, especially in environments where access patterns are nonstationary and difficult to capture with simple frequency- or recency-based rules.
This work matters because CXL-attached memory and other slower memory tiers are becoming increasingly relevant for scaling capacity without proportionally increasing DRAM cost. However, the benefit of tiering is only realized if the system can keep latency-sensitive pages in the fast tier. A learning-based advisor that tracks shifting hot sets could improve effective bandwidth, reduce tail latency, and make heterogeneous memory systems more practical for datacenter workloads.