Evaluating Small and Large Language Models in a Self-Correcting Text-to-SQL Agent for Enterprise ERP Reporting
DOI:
https://doi.org/10.65091/icicset.v3i1.104Abstract
Retrieving data from enterprise relational databases
requires SQL. Many business users do not have SQL skills,
which drives the need for natural language to SQL systems.
Most leading systems depends on large, cloud-hosted LLMs,
which sits poorly with commercially sensitive ERP data and
their surrounding retrieval and validation pipelines are rarely
evaluated separately from the generation model itself. This article
investigates whether a small locally deployed LLM can compete
as a generation model in the same retrieval augmented, rule
validated agent pipeline as larger models, against which the
retrieval and validation pipelines also evaluated with the data.
We built an agent for an on-premise Oracle-based Synergy ERP
schema, using BGE-M3 embeddings over Qdrant vector database
to index schema and example retrieval, and a deterministic
validator with bounded self-correction before execution. We use
and compare four interchangeable LLMs: locally run Qwen2.5-
Coder-14B (4bit quantization) and cloud run Mistral Large,
Minimax M3 and Qwen3.8-Max based on semantic equivalence,
execution accuracy, component-level F1, row-count matching,
and generation success rate. Qwen3.8-Max performed best on
every metrics, while MiniMax M3 outperformed the larger
Mistral Large. The much smaller Qwen2.5-Coder-14B matched
or exceeded Mistral Large on several metrics. Schema retrieval
was reliable enough with Hit@5 = 0.93. An ablation comparing
before and after validation and SQL correction shows the largest
accuracy gains for weaker models, with limited or slightly
negative effect for the strongest model. The locally hosted model
was also the fastest (2.671 s/query) and cheapest ($0.0011/query)
among the four models. These findings suggest model scale
does not fully determine performance within a fixed pipeline,
supporting compact local models as a practical, low-cost, lowlatency
alternative for enterprise deployment.