Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows across input indices are not fully reused and intermediate results are repeatedly materialized off-chip between sequential TT-core contractions. We present FlowTT, a flow-aware GPU execution framework th

Topological visualization of FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train Embedding
Brave API

FlowTT is a flow-aware GPU execution framework designed to optimize Tensor-Train (TT) embedding lookup in recommendation models by exploiting computation flow reuse. It addresses the inefficiency of standard TT-based lookups, where partially shared computation paths across input indices are not fully reused and intermediate results are repeatedly materialized off-chip.

The framework reformulates TT gather as a set of prefix-shared irregular computation flows, grouping common operations along overlapping paths to keep them in on-chip memory. Key innovations include: Flow-aligned on-chip execution: Reorganizes computation into a single pipeline to minimize redundant operations and off-chip data movement. Prefix-based index grouping: Groups inputs sharing high-order TT core indices to eliminate redundant core operations. Dynamic load balancing: Uses persistent threads and chunk-based work stealing to maintain high GPU utilization under skewed workloads. Inter-block sharing: Stores reusable intermediates in L2 cache checkpoint buffers to allow other blocks to reload and reuse them.

In evaluations on Meta’s synthetic benchmarks, FlowTT reduces inference latency by up to 42.2% and training latency by 49.2% relative to EcoRec, while achieving the lowest peak memory usage. This demonstrates that exposing and reusing prefix-shared computation is critical for efficient TT-based embedding execution.

Generated 29d ago
Cerebras Thinking

This material addresses a systems bottleneck in using Tensor-Train (TT) decomposition for large embedding tables in recommendation models. TT decomposition can dramatically reduce the memory footprint of high-dimensional embeddings by representing a full lookup table as a sequence of smaller tensor cores, but the resulting lookup operation is not automatically efficient on GPUs. For a batch of sparse, irregular input indices, the required computation is a chain of tensor contractions across TT cores, and many of these contractions share partial computation flows across indices. Existing implementations often fail to exploit this sharing and instead materialize intermediate tensors to off-chip memory between sequential core updates, creating redundant work and heavy memory traffic.

FlowTT is presented as a flow-aware GPU execution framework designed specifically for irregular TT-based embedding lookup. Its central idea is to treat the sequence of TT-core contractions as a dataflow graph in which partial results can be reused across related input indices, rather than recomputing or persisting them independently. By identifying shared computation flows, keeping intermediate results resident on-chip where possible, and fusing or rescheduling the sequential core contractions, FlowTT reduces both redundant arithmetic and global memory access. The result is a more efficient execution model for compressed embedding lookups that preserves the compression benefits of TT decomposition while addressing the runtime inefficiencies that have limited its practical adoption.

The work matters because embedding lookup is a dominant cost in large-scale recommendation systems, where models routinely contain sparse, high-cardinality embedding tables. TT decomposition is attractive because it can compress these tables without requiring full dense materialization, but its value depends on an execution engine that can efficiently evaluate the resulting tensor contractions under irregular access patterns. By focusing on computation-flow reuse and on-chip data movement, FlowTT targets the practical gap between embedding compression and low-latency GPU execution. If effective, this line of work makes TT-based embeddings more viable for high-throughput training and serving, where memory bandwidth and kernel latency are first-order constraints.

Generated 29d ago
Open-Weights Reasoning

Tensor-Train decomposition is a promising way to compress very large embedding tables in recommendation systems, but the paper identifies a key systems-level bottleneck: actual TT-based lookup is often dominated by inefficient execution rather than by the decomposition itself. In a standard TT lookup, an input index is processed by sequentially contracting a chain of small TT cores, and for batched workloads many of these contraction paths share partial computation. However, existing GPU implementations typically treat each index’s contraction chain too independently, so partially shared subcomputations are recomputed, and intermediate tensors are repeatedly written to and read from off-chip memory between sequential core contractions. This creates unnecessary memory traffic and underutilizes on-chip storage.

FlowTT addresses this by treating TT embedding lookup as a flow-aware execution problem. The framework models the batch of lookup operations as a computation flow, identifies shared subflows across input indices, and schedules those shared computations so their intermediate results can be kept in on-chip memory and reused instead of being materialized repeatedly in global memory. In effect, FlowTT fuses and reorganizes the sequence of TT-core contractions around data reuse, reducing redundant arithmetic and off-chip traffic while better matching the irregular, partially overlapping structure of batched TT lookups to GPU execution constraints.

The contribution matters because large embedding tables are a central component of modern recommendation models, and TT compression can only be practical if lookup is fast enough for serving or large-scale training. By targeting the execution pattern rather than only the mathematical representation, FlowTT provides a systems-level path to making TT-based embeddings more deployable. More broadly, the work highlights an important insight for tensor-decomposition inference: when decomposed operations are executed in batches, exploiting shared computation flows and keeping intermediate state on-chip can be as important as improving the underlying numerical factorization.

Generated 29d ago
Sources