Demonstrates that visually grounded token embeddings (via ostension) in a small DeBERTa model trained on 10M words produce a measurable, persistent imprint through the end of training.

Topological visualization of Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model
Brave API

Yes, the study confirms that visually grounded token embeddings produce a measurable, persistent imprint in a small DeBERTa model trained on approximately 10M words. This visual initialization leaves a lasting effect on the model's embeddings until the end of training, particularly evident in object-property knowledge (tested via the COMPS benchmark) and a tailored Visual-Property Swap (VP-Swap) benchmark.

However, the study also highlights significant limitations in how this effect is perceived: Benchmark Blindness: Standard BabyLM benchmarks, which primarily probe abstract grammatical competence, fail to detect the benefits of visual grounding for most words. Function Words: While abstract and function words retain strong visual seeds that aid the training objective (reducing masked-LM loss), these improvements are not registered by existing evaluation metrics. * Relational Persistence: The visual imprint survives as a relational structure among embeddings rather than as retrievable anchor points, meaning the geometry of related concepts is preserved even if individual embedding positions shift significantly.

Ultimately, the research demonstrates that while ostension provides a causal, persistent advantage for concrete lexical semantics, current evaluation frameworks are misaligned with the specific types of knowledge that visual grounding enriches.

Generated 22d ago
Open-Weights Reasoning

This paper investigates whether ostensive definition—the act of pointing to a visual referent to teach a word—can meaningfully shape the internal representations of a small language model. It uses a compact DeBERTa-style model trained on roughly 10M words, with token embeddings grounded through visual ostension rather than relying solely on standard next-token prediction. The central question is whether this kind of perceptual grounding leaves a detectable trace in the model after training, or whether ordinary language training simply overwrites it.

Its key contribution is an empirical demonstration that visually grounded embeddings do not vanish: they produce a measurable and persistent imprint in the model’s representations by the end of training. In other words, even a modest amount of ostensive grounding can bias the model’s semantic geometry and leave a stable signature in its final state. At the same time, the work appears to delineate the limits of this approach: ostension can anchor referential meaning, but it is not a complete substitute for the broader linguistic, compositional, and contextual knowledge acquired from language data.

This matters because it speaks to a long-standing question in both cognitive science and modern NLP: how much semantic structure can be bootstrapped from perception alone, and how much must come from language itself. For practitioners, the result suggests that targeted visual grounding may be a useful, low-cost way to stabilize or steer small language models, especially when efficient training or interpretable semantic anchors are desired. More broadly, it provides a concrete test of the “Augustinian” intuition that words can be learned through demonstration, while also showing where that intuition breaks down.

Generated 22d ago
Sources