arXiv:2609.21281v1 Announce Type: cross Abstract: Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the personalization-scale paradox

Topological visualization of Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale
Brave API

Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale (arXiv:2609.21281) is a deployed production system that resolves the "personalization–scale paradox" by separating modeling depth from inventory breadth. The architecture utilizes a billion-scale GPU pathway for deep, interaction-heavy pre-ranking and a tens-of-billions-scale CPU pathway for broad, lightweight semantic coverage.

In production A/B tests against a legacy CPU-only system, the hybrid architecture achieved a 4.51% improvement in DCG@20 and a 2.01% improvement in Good Search Result Rate (GSRR). The two pathways contribute structurally distinct candidates, allowing for independent evolution and cost-effective scaling where GPU capacity is roughly four times more expensive than CPU for matched workloads.

Generated 13d ago
Open-Weights Reasoning
  • Problem framing: The paper studies embedding-based retrieval for personalized search over user-generated content at the trillion-document scale. Its central observation is what it calls the personalization–scale paradox: production systems must simultaneously support rich, intent-aware personalization for individual queries and maintain broad recall across an enormous inventory, all under strict latency and compute budgets. This tension is especially acute for dense retrieval, where expressive user/query embeddings can improve relevance but make exhaustive or high-quality search over massive corpora prohibitively expensive.
  • Likely contribution: The work appears to propose or analyze a hybrid GPU–CPU retrieval architecture that resolves this conflict by allocating different retrieval responsibilities to the hardware layer best suited to each task. In such a design, GPU resources can be reserved for the most personalized, latency-sensitive, or high-value retrieval paths—where deep embeddings, user context, or complex scoring matter most—while CPU-based retrieval handles broader coverage, larger candidate pools, or lower-cost stages. The key insight is that a single homogeneous retrieval stack is often suboptimal at this scale; a coordinated hybrid approach can balance recall, personalization quality, latency, and infrastructure cost more effectively.
  • Why it matters: For large-scale search and recommendation systems, this is an important systems-level contribution because it moves beyond isolated model improvements and addresses the practical constraints of production deployment. If personalized dense retrieval can be made viable at trillion-document scale without sacrificing coverage or exceeding latency targets, the result is a more capable search stack for e-commerce, content platforms, and other user-generated ecosystems. The paper is therefore relevant not only to retrieval researchers, but also to engineers designing large-scale ANN infrastructure, personalization pipelines, and cost-efficient serving systems.
Generated 13d ago
Sources