Constructing a Japanese Financial Tool-Use Benchmark for Large Language Models
Abstract
Applying large language models (LLMs) to many tasks in financial practice requires referring to external information that cannot be obtained from the model's internal knowledge alone, such as current market data and interest rates, customer-specific information, product documents, and internal rules. Tool use is therefore an important capability for such tasks, but reaching the necessary tool responses is not the same as correctly interpreting and integrating the retrieved information into the final answer to the user. In this paper, we construct a tool-use benchmark of 108 scenarios in nine categories built around Japanese financial tasks, and evaluate performance by separating tool-execution histories from final answers. The evaluation combines a content audit of each answer against the tool responses the model actually observed with rubric-based scoring. At the same time, we record two diagnostic indicators: whether the model exactly reproduced the tool calls assumed at design time (canonical-call reproduction) and whether it reached the tool responses containing the necessary evidence (evidence-tool reach). We evaluated a set of models including open-weight and closed models, and analyzed the validity of the evaluation and the observed trends. In the results, the correlation between evidence-tool reach and the final-answer score remained around 0.6, and there were many cases in which a model fully reproduced the expected tool calls (canonical calls) yet received a low score. Scores tended to increase with model size, and failures in numerical processing and cross-checking decreased with scale. In judgments about information freshness and timing, differences between models were observed: failures remained in some frontier models, while they were almost eliminated in the most recent frontier model. These results indicate the need to evaluate directly not only the tool calls themselves but also the use of the retrieved information and the final answer produced for the task.