Introduces DiscoSign, an LLM-based modular system that handles discourse phenomena (spatial coreference, role shift, constructed action) for text-to-sign-gloss translation.

Topological visualization of DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Brave API

DiscoSign is a modular Large Language Model (LLM) framework that introduces discourse-aware text-to-sign language gloss translation, addressing limitations in traditional sentence-level systems. It explicitly models three key phenomena: spatial coreference resolution (maintaining consistent entity locations in 3D space), Question-Answer Clauses (QACs) for discourse-pragmatic structures, and concept-gloss consistency (stable mappings between English concepts and ASL signs).

The system operates by processing sentences sequentially while maintaining explicit state registries for spatial indices, concept mappings, and QAC decisions. It employs programmatic constraint enforcement and a deterministic post-processor to correct violations, ensuring cross-sentence coherence.

Key outcomes include: Novel Evaluation Metrics: Introduces metrics like Spatial Coreference Accuracy (SCA) to assess discourse coherence, which traditional metrics (BLEU, COMET) fail to capture. Performance: Significantly improves spatial consistency and entity tracking compared to sentence-only baselines while maintaining competitive single-sentence gloss quality. * Architecture: Utilizes a pipeline where an LLM generates glosses based on hard constraints from accumulated discourse state, verified by a post-processor.

DiscoSign establishes the first systematic framework for discourse-level text-to-gloss translation, demonstrating that explicit modeling of discourse phenomena is necessary for accurate sign language comprehension.

Generated Sep 3, 2026
Open-Weights Reasoning

DiscoSign addresses a persistent weakness in text-to-sign language gloss translation: the tendency of systems to produce locally fluent gloss sequences that fail to represent discourse-level structure. The paper introduces an LLM-based modular system designed to handle phenomena that are central to sign language but difficult to model with word-level translation alone, including spatial coreference, role shift, and constructed action. Rather than treating gloss generation as a simple lexical or syntactic mapping from written text, DiscoSign frames it as a discourse-aware planning problem in which referents, perspectives, and event structure must be tracked across the utterance.

A key contribution is the modular architecture, which allows different discourse mechanisms to be handled as explicit components rather than left to the LLM to infer implicitly. Spatial coreference is especially important because sign languages frequently assign entities to locations in signing space and return to them later; role shift requires the signer to change perspective, often by shifting facial grammar or body orientation; constructed action involves signing an action to represent a described event. By making these mechanisms visible in the generation pipeline, DiscoSign can produce glosses that are more faithful to how sign language encodes discourse, not merely how written text decomposes into sign glosses.

The work matters because it pushes text-to-sign gloss translation beyond surface fluency toward semantically and pragmatically grounded output. For downstream applications such as sign language avatars, accessibility tools, and sign language NLP research, gloss sequences that preserve coreference, perspective, and event representation are far more useful than those that only approximate lexical content. DiscoSign’s approach suggests that LLMs can be made more effective for sign language generation when paired with explicit discourse modeling, offering a path toward systems that better capture the structure of signed communication.

Generated Sep 3, 2026
Sources