Shows that CE-only training for continuous-space LLM reasoning induces four failures (e.g., thought collapse) that lower correct-answer probability.

Topological visualization of Principled Thoughts for Latent Recursive LLM Systems
Brave API

Principled Thoughts for Latent Recursive LLM Systems identifies that Cross-Entropy (CE)-only training fails to constrain latent thoughts, leading to four specific failures: thought collapse (distinct questions sharing identical representations), retention of irrelevant information (encoding input rather than the producer’s plan), lack of producer uncertainty encoding, and inability to distinguish difficult examples from collapsed states.

To address these, the paper introduces REST (REpresentation-Supervised Thoughts), which adds differentiable loss terms for four axioms of valid thought representation: causality, minimality, separability, and stability.

Empirical results show that REST improves accuracy by up to 7.5 percentage points and increases convergence on final answers by 30% compared to CE-only training, while reducing token usage by encouraging more efficient reasoning paths.

Generated 4d ago
Open-Weights Reasoning

The paper analyzes latent recursive LLM systems, where a model reasons by iteratively updating continuous latent thoughts rather than generating an explicit token-level chain of thought. Its central claim is that training these systems with a cross-entropy-only objective—supervising the final answer while leaving the latent reasoning dynamics to be shaped indirectly—can produce systematic pathologies. The authors identify four failure modes, including thought collapse, in which latent thoughts become degenerate, repetitive, or too weakly differentiated to carry useful intermediate information. These pathologies matter because they do not merely make reasoning harder to inspect; they reduce the probability that the system emits the correct final answer by weakening the causal and informational link between latent computation and prediction.

The contribution is a principled diagnostic of why “latent reasoning” can fail even when the output distribution appears well trained. By treating latent thoughts as intermediate computations that should preserve, transform, and summarize task-relevant information, the paper reframes CE-only training as an incomplete objective: it can optimize surface answer likelihood without ensuring that hidden states remain informative, diverse, and aligned with the reasoning process. This is especially important for continuous-space LLMs, where the lack of discrete intermediate tokens removes many of the natural regularization and interpretability signals that token-based chain-of-thought provides.

The work matters because it challenges a common assumption in latent or recursive LLM research: that maximizing final-answer cross-entropy is sufficient to induce healthy internal reasoning. It suggests that future systems need explicit diagnostics—possibly involving latent-state diversity, uncertainty, information flow, and correct-answer probability—rather than relying solely on answer-level losses. More broadly, it provides a cautionary framework for designing objectives that support reliable continuous reasoning instead of allowing latent thoughts to collapse into unhelpful, low-signal representations.

Generated 4d ago
Sources