arXiv:2609.21281v1 Announce Type: cross Abstract: Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the personalization-scale paradox
Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale (arXiv:2609.21281) is a deployed production system that resolves the "personalization–scale paradox" by separating modeling depth from inventory breadth. The architecture utilizes a billion-scale GPU pathway for deep, interaction-heavy pre-ranking and a tens-of-billions-scale CPU pathway for broad, lightweight semantic coverage.
In production A/B tests against a legacy CPU-only system, the hybrid architecture achieved a 4.51% improvement in DCG@20 and a 2.01% improvement in Good Search Result Rate (GSRR). The two pathways contribute structurally distinct candidates, allowing for independent evolution and cost-effective scaling where GPU capacity is roughly four times more expensive than CPU for matched workloads.