CostFit: Cost-Aware Agent Design for Hybrid Database and API Question Answering
Abstract
Enterprise questions often require an AI agent to combine information from relational databases and REST APIs. In this setting, a good agent must do more than return the correct answer: it should also avoid unnecessary reasoning turns, tool calls, tokens, and latency. We present CostFit, a measurement-first methodology for designing cost-aware agents for hybrid database and API question answering. Rather than selecting an agent architecture first and optimizing it afterward, CostFit measures five properties of a labeled workload schema cost, endpoint usage, trajectory depth, step independence, and intent concentration and uses explicit break-even rules to determine which architectural components are justified. On an enterprise marketing workload spanning an 18-table relational graph and a REST control plane, CostFit removes two of four initially planned components and selects single-pass DAG planning. In matched-pair experiments, the selected agent reduces tool calls by 26% and limits trajectories to at most four turns, although only the tool-call reduction remains statistically significant after Holm–Bonferroni correction and per-query token cost does not significantly decrease. We further find that the baseline model already follows the intended planning behavior on half of eligible queries, demonstrating the importance of measuring architectural conformance before attributing improvements to an agent scaffold. Applying CostFit unchanged to four public benchmarks produces different architectural recommendations across workloads, including two cases where it rejects the architecture used in our deployment. These results support a simple principle: measure the workload first, then build only the agent components that the workload justifies.