arXiv:2609.37399v1 Announce Type: cross Abstract: Multi-engine deep learning (DL) accelerators are becoming increasingly prevalent as they address the heterogeneity and growing complexity of modern DL workloads. To efficiently process diverse DL workloads, these accelerators must incorporate combinations of engines with complementary capabilities to match the distinct computational characteristic

Topological visualization of MEDEM: Multi-Engine DL Accelerator Design Methodology
Brave API

MEDEM (Multi-Engine DL Accelerator Design Methodology), published on arXiv:2609.37399v1 on September 29, 2026, is a systematic two-stage methodology for designing flexible multi-engine deep learning accelerators. It decouples engine co-design from combination selection to efficiently navigate exponentially large design spaces and minimize aggregate execution costs.

The methodology employs two core strategies: the Hierarchical Co-designer (HCo), which prioritizes high-impact hardware-mapping axes for exhaustive exploration, and Balanced Average and Specialization Search (BASS), which selects engine multisets that balance common-case replication with specialized heterogeneity. Comprehensive evaluations using 51 DL workloads demonstrate that MEDEM identifies accelerators outperforming state-of-the-art designs, achieving geometric-mean improvements of up to 4.84× in energy-delay product (EDP) and 1.59× in throughput.

Generated 4d ago
Open-Weights Reasoning

MEDEM: Multi-Engine DL Accelerator Design Methodology addresses the architectural challenge of designing deep learning accelerators that integrate multiple specialized compute engines rather than relying on a single general-purpose datapath. As modern DL workloads become increasingly heterogeneous—spanning convolutional, attention-based, memory-bound, and irregular kernels—no single engine type can simultaneously optimize latency, throughput, energy efficiency, and utilization. The paper frames this as a design-space problem: an accelerator must select and combine engines with complementary computational characteristics so that different workload components can be matched to the engines best suited to their arithmetic intensity, data-movement patterns, and precision requirements.

Methodologically, MEDEM contributes a structured approach for composing such multi-engine architectures, moving accelerator design away from ad hoc engine selection toward a more systematic evaluation of workload demands, engine capabilities, and system-level integration. The central insight is that efficiency comes not only from building a fast core, but from aligning a heterogeneous set of engines with the distinct phases and operators of DL workloads, while accounting for the trade-offs introduced by multi-engine coordination. This makes the methodology useful for architects who need to balance performance, energy efficiency, flexibility, and implementation complexity in next-generation AI hardware.

The work matters because multi-engine DL accelerators are becoming a practical path to serving diverse workloads on shared silicon, especially in edge, datacenter, and heterogeneous SoC environments where a single fixed engine would often be underutilized or inefficient. By providing a design methodology, MEDEM helps reduce the search space and clarifies the trade-offs involved in choosing engine mixes, potentially leading to more predictable performance, better cost/energy outcomes, and more adaptable accelerator designs for evolving DL workloads.

Generated 4d ago
Sources