Presents KOPA-Bench (145 real Korean public-API tasks) and the EDGE dynamic-graph method to synthesize tool-calling data that improves open-source LLM agents on multi-step government workflows.
KOPA-Bench is a benchmark consisting of 145 multi-step tasks across six domains of live Korean public APIs, designed to evaluate LLM agents on chaining dependent calls and handling high-cardinality responses. The benchmark highlights that open-source models frequently fail by skipping prerequisite lookups or prematurely answering from partial results.
To address these gaps, the paper introduces EDGE (Execution-grounded Dynamic Graph for tool-calling data synthEsis), a data-synthesis pipeline that builds a dependency graph of tool outputs and inputs, verifying links against live API execution to prune failures. EDGE synthesizes executable trajectories by labeling output-to-input handoffs, effectively converting high-cardinality responses into valid sequential paths.
Training small open-source models (4B and 9B) with GRPO on the EDGE dataset yielded substantial improvements, with the 9B model nearly matching the performance of an untuned 27B model. These gains extended to out-of-distribution benchmarks like BFCL, demonstrating that training on verified, live-API trajectories significantly enhances multi-step tool-calling capabilities.
This paper targets a concrete gap in agentic tool-use research: reliable multi-step tool calling over real, non-English public APIs, especially Korean government and public-service APIs. It introduces KOPA-Bench, a benchmark of 145 real Korean public-API tasks that require agents to plan and execute sequences of API calls rather than answer single-call questions. The tasks are drawn from actual open public APIs and reflect the kinds of constraints found in civic and government workflows, including parameter dependencies, multi-hop retrieval, validation steps, and service-specific calling patterns. This makes the benchmark more realistic than many existing tool-calling evaluations, which often rely on synthetic tools, English-only APIs, or isolated function-calling examples.
A central contribution is EDGE, a dynamic-graph-based data-synthesis recipe for generating tool-calling training data. Rather than hand-writing large numbers of multi-step trajectories, EDGE models the structure of API interactions as an evolving graph, allowing the method to sample and compose sequences of calls that respect API dependencies and task logic. The resulting synthetic data is used to improve open-source LLM agents on multi-step government workflows, demonstrating that structured, graph-aware synthesis can produce coherent and useful tool-use trajectories even when high-quality real-world examples are scarce. The key insight is that tool-calling data should be generated from the relational structure of APIs and workflows, not merely from isolated function schemas or surface-level prompts.
The work matters because it provides both an evaluation resource and a practical data-generation method for building open-source agents that can operate in real public-sector environments. Korean government APIs are a particularly important test case because they are real, heterogeneous, and often underrepresented in LLM training and benchmarking data. By pairing a realistic benchmark with a reproducible synthesis recipe, the paper offers a template for extending agentic tool-use to other localized, high-stakes, and multi-step API domains where reliability, correctness, and domain coverage are more important than simple single-turn function invocation.