On-device LLM agents interact repeatedly with users on local hardware, producing private traces that are valuable for adaptation but should not be sent to a remote trainer. Ideally, such agents would learn contiguously---adapting from every interaction without pausing or suspending user-facing inference---yet existing inference runtimes assume stable weights and existing RL systems assume separate

Topological visualization of LOCAL: Enabling Learning On-device Contiguously for Agent LLMs
Brave API

On-device LLM agents adapt continually by distilling successful and failed task episodes into reusable insights stored in an external memory, rather than updating model weights. This approach allows agents to improve in the field while keeping all private data on the local device and avoiding the need to send traces to a remote trainer.

To manage the resource constraints of edge hardware, systems like CURATOR govern this experience memory using a budgeted value score that balances retention, cross-agent sharing, and trust. This method shifts the continual-learning problem from parameter updates to efficient memory access, ensuring agents remain responsive without exceeding fixed RAM or energy limits.

Generated Sep 2, 2026
Open-Weights Reasoning

The material introduces LOCAL, a system for enabling contiguous on-device learning in LLM-based agents. It targets a growing deployment pattern in which LLM agents run locally on user hardware and interact repeatedly with users, generating private interaction traces that are highly useful for personalization and adaptation. The central problem is that these traces are valuable but should not necessarily be transmitted to a remote trainer, while at the same time the agent should be able to improve from them without pausing user-facing inference. The paper frames this as a systems gap: conventional inference runtimes are designed around stable weights and predictable serving behavior, whereas existing reinforcement-learning or continual-learning pipelines often assume a separation between training and inference phases.

LOCAL’s key contribution is to make on-device adaptation a first-class runtime concern rather than an offline post-processing step. It couples inference and learning tightly enough that the agent can absorb local interaction experience, derive updates, and continue serving users without suspending the agent loop or exposing raw traces externally. In doing so, it addresses the practical constraints of local hardware—limited compute, memory pressure, latency requirements, and the need to keep the model usable while its weights or policies are changing. The broader insight is that for on-device agent LLMs, learning and serving are not separable workloads; they must be co-designed so that adaptation can happen continuously, safely, and privately.

This matters because many high-value LLM agent applications are inherently local and interactive: personal assistants, on-device copilots, private productivity agents, and edge-based automation systems. If such agents cannot learn from their local usage without sending data away or interrupting service, their ability to personalize and improve over time is severely limited. LOCAL therefore bridges a significant gap between LLM serving systems, reinforcement learning, and on-device continual learning, pointing toward a more practical architecture for privacy-preserving, always-available agent intelligence.

Generated Sep 2, 2026
Sources