Demonstrates that, after compression and retrieval adaptation, model size ceases to predict answer quality for on-premise factory assistants, turning deployment into a post-adaptation selection task.
Model size ceases to be a reliable predictor of answer quality after structural compression and retrieval-grounded adaptation, as general capability declines linearly with parameter count while retrieval-augmented quality does not.
Deployment becomes a post-adaptation selection problem where one sub-network is committed per device based on judged answer quality and measured on-device throughput, constrained by a configurable general-capability floor and memory budget.
A weight-shared supernetwork trained with sandwich-style in-place distillation keeps this selection process inexpensive, optimizing for size, speed, or quality without sacrificing necessary capability or throughput.
This material examines how to choose the right model for on-premise, retrieval-augmented factory agents once models have been compressed and adapted to a local retrieval stack. The setting is practically constrained: factory assistants must operate with limited compute, low latency, local data, and domain-specific knowledge, often relying on retrieval over internal documents rather than purely parametric knowledge. The paper’s focus is not simply on whether a smaller model can replace a larger one, but on how deployment decisions should be made after the model has undergone compression and retrieval-specific adaptation.
Its key insight is that model size becomes a poor predictor of answer quality in this regime. After compression and retrieval adaptation, candidate sub-networks no longer rank in a clean monotonic way by parameter count; a smaller or more aggressively compressed model may perform competitively, while a larger one may not. The contribution is therefore a measurement-driven selection framework: instead of defaulting to the largest available model or relying on generic pre-adaptation benchmarks, practitioners should evaluate post-adaptation candidates on task-relevant answer quality, retrieval behavior, and deployment constraints, then select the best sub-network from that measured set.
This matters because it reframes on-premise RAG deployment from a “pick the biggest model that fits” problem into a post-adaptation model-selection task. For factory and industrial assistants, where reliability, cost, latency, and data locality are critical, the result supports a more efficient deployment practice: compress, adapt to the local retrieval pipeline, then measure and select the sub-network that best fits the operational workload. The broader implication is that practical RAG systems may benefit more from careful sub-network selection and workload-specific evaluation than from naive scaling.