Demonstrates LLMs can generate both instance-level solutions and general algorithms for inventory, queueing, and assortment problems, with level-2 algorithm synthesis outperforming hand-crafted baselines on several tasks.
LLMs can design near-optimal operations research (OR) algorithms, with the strongest models matching or outperforming the best existing methods on nearly all evaluated instances of inventory control, queueing network control, and assortment optimization.
This paper investigates whether large language models can move beyond solving individual decision problems and instead synthesize reusable operations research (OR) algorithms. It evaluates LLMs on classic OR domains—inventory management, queueing, and assortment optimization—where decisions are typically made using analytic models, heuristics, or simulation-based search. The central distinction is between instance-level solution generation, where an LLM produces a concrete decision for a given problem instance, and algorithm-level synthesis, where it produces a general procedure that can be applied to new instances. The paper’s key insight is that LLMs can be useful not only as optimizers or planners, but as designers of decision policies and near-optimal algorithms.
A notable contribution is the empirical finding that LLM-generated “level-2” algorithms can outperform hand-crafted baselines on several benchmark tasks. This suggests that LLMs can capture useful structural regularities in OR problems—such as threshold policies, service-rate choices, or product-mix rules—without requiring explicit symbolic derivation from the modeler. For a technically literate audience, the result is significant because it reframes LLMs as a potential source of algorithmic design: rather than merely fitting a black-box policy to data, they may be able to propose interpretable, generalizable decision rules that approximate the performance of more carefully engineered OR methods.
The broader importance of the work lies in its implications for automating parts of quantitative decision-making. If LLMs can reliably generate near-optimal algorithms for structured operations problems, they could reduce the cost of algorithm design, accelerate prototyping in supply chain, service operations, and revenue management, and make sophisticated OR tools more accessible. At the same time, the paper implicitly raises important evaluation questions: how to verify optimality or robustness, how to handle distribution shift, and how to distinguish genuinely transferable algorithms from benchmark-specific heuristics. In that sense, the material is both a demonstration of capability and a prompt for more rigorous standards in LLM-assisted algorithm design.