Performs a controlled evaluation of 13 LLMs spanning Qwen and GPT families on ontology learning using a fixed OntoLearner RAG pipeline to measure scale effects.
The provided search context does not contain a study with the specific title "When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning." However, the OntoLearner framework and related research provide insights into the effects of LLM scale on ontology learning tasks.
Scaling provides diminishing returns and task-dependent benefits: Term Typing: Scaling from Qwen3-0.6B to Qwen3-14B produces dramatic gains (e.g., Finance: 22.2% → 71.8%), showing that larger models improve representation and instruction following for this task. Taxonomy Discovery: Scaling yields negligible gains (e.g., Biology & Life Science: 2.0% → 10.8%), indicating that model size alone cannot solve non-hierarchical reasoning challenges. * Model Family vs. Scale: Model architecture and training alignment matter more than parameter count. For instance, Gemma3-1B outperformed much larger Qwen models on specific relation extraction tasks (e.g., Biology & Life Science: 20.9% vs. 15.4%), demonstrating that task-aware selection outperforms defaulting to the largest available model.
Key Findings on Scale and Complexity: Failure Modes: Failures in ontology learning scale with ontological complexity rather than model size. No single retrieval architecture handles all ontology learning tasks. Reasoning vs. Discipline: Open-ended chain-of-thought reasoning (often found in larger, reasoning-optimized models) can conflict with the strict schema compliance required for structured ontology extraction, sometimes leading to zero F1 scores in strict evaluation settings. * Domain Transfer: Models excelling in high-resource domains (e.g., Finance) do not necessarily transfer performance to specialized domains (e.g., Medicine), regardless of scale.
This paper examines whether larger language models improve performance on ontology learning tasks, using a controlled benchmark in which 13 LLMs from the Qwen and GPT families are evaluated within a fixed OntoLearner RAG pipeline. By holding the surrounding system constant—prompting, retrieval-augmented generation, and ontology extraction workflow—the study isolates model scale as the primary variable. This design allows the authors to assess whether increases in model size translate into measurable gains in ontology construction, such as concept extraction, relation identification, and hierarchical or taxonomic reasoning, rather than conflating scale effects with changes in prompt engineering or pipeline architecture.
A key contribution is the empirical clarification of when “bigger” actually helps. The study suggests that scale effects are not uniform across models or ontology tasks: larger models may provide advantages in complex relational reasoning, disambiguation, or structurally coherent output generation, but their benefits can be limited when retrieval already supplies strong domain context or when the task is relatively straightforward. The comparison across Qwen and GPT families also highlights that model family, training objectives, and architectural differences can matter as much as raw parameter count, offering a more nuanced view than a simple size-versus-accuracy framing.
The work matters because ontology learning is a high-stakes application area where model selection directly affects cost, latency, reproducibility, and downstream knowledge-graph quality. By providing a controlled study of LLM scale, the paper gives practitioners evidence-based guidance for choosing smaller, cheaper models when they are sufficient, and for reserving larger models for tasks where their reasoning advantages are likely to pay off. More broadly, it contributes to the growing body of work questioning whether scale alone is the best predictor of performance in specialized NLP pipelines, and it helps bridge LLM evaluation with practical ontology engineering.