Proposes behavioral correctness assumptions framework with a taxonomy of preserving vs. altering alterations to meta-evaluate reference-based NLG evaluators under controlled conditions.
The search context does not contain information regarding a specific framework titled "Behavioral Correctness Assumptions" or a taxonomy of "preserving vs. altering alterations" for meta-evaluating reference-based NLG evaluators.
However, recent research highlights related concepts in NLG evaluation:
If you are referring to a specific paper not included in the provided search results, please provide additional details or the author's name for a more targeted response.
This material addresses a key weakness in the meta-evaluation of reference-based automatic NLG metrics: reliance on aggregate scores, such as correlation with human judgments, can obscure whether an evaluator actually behaves in a principled way. The paper proposes a behavioral correctness assumptions framework for assessing reference-based automatic evaluation methods under controlled transformations. Rather than asking only how well a metric’s scores rank overall, it asks whether the metric responds appropriately to specific, interpretable changes in system outputs and references.
The central contribution is a taxonomy of controlled alterations, divided into preserving alterations and altering alterations. Preserving alterations are changes that should not materially affect the evaluator’s score because they leave the relevant quality or meaning intact, while altering alterations are changes that should move the score in a predictable direction, such as worsening or improving the output relative to the reference. This structure allows evaluators to be tested against explicit behavioral expectations—such as invariance to irrelevant surface changes, sensitivity to meaningful degradations, and monotonic response to quality shifts—rather than being judged solely by a single global correlation statistic.
The work matters because it shifts the evaluation of automatic NLG metrics from “which score correlates best?” to “does the metric behave correctly in the ways that matter for downstream use?” This diagnostic perspective can reveal failure modes that aggregate metrics mask, including overpenalization of superficial variation, insensitivity to real errors, excessive reference dependence, or non-monotonic scoring. For researchers and practitioners selecting or combining metrics for tasks such as summarization, machine translation, or dialogue generation, the framework offers a more fine-grained and theoretically grounded way to assess whether an automatic evaluator is trustworthy.