Edge-Native Semantic Firewall for Autonomous LLM Agents: A Structured Chain-of-Thought Verification Framework

Authors

  • Sushant Poudel Nepal Engineering College
  • Rakhee Pandey Nepal Engineering College
  • Aashika Pandey Nepal Engineering College

DOI:

https://doi.org/10.65091/icicset.v3i1.75

Abstract

An autonomous agent that executes
actions rather than proposing them sits outside the
reach of role-based access control, which
authenticates an identity but not intent. Routing
every proposal to a cloud-hosted frontier model
closes that gap but adds round-trip latency and, in
regulated settings, is often prohibited outright. We
describe an edge-native semantic firewall: a 3.8Bparameter
Phi-3-mini model, 4-bit quantized under
a 4.2 GiB VRAM ceiling, acting as the evaluator in
a Generator-Evaluator pipeline on one consumer
laptop. Its mechanism is a structured Chain-of-
Thought JSON schema that requires the evaluator to
name the governing rule and justify the match before
emitting a decision, turning an opaque verdict into
an auditable trace. We evaluate it on a 600-scenario
corpus across three policy rules, including 155
adversarial scenarios spanning eight promptinjection
techniques, 60 compound and 30 boundary
cases; all 1,800 generations ran locally at
temperature 0. The results invert a natural
assumption: constraining output format without
requiring the reasoning step produced the least safe
evaluator of the three. The JSON-only arm approved
46.2% of proposals the policy would block or route
to human review, worse than unconstrained freeform
at 17.2%. The full schema cut that to 23.5%
and raised decision accuracy from 52.3% to 66.3%.
On the irreversible class, accuracy runs 62.5% under
JSON-only, 72.1% under free-form, and 90.8%
under the proposed schema, with hard-denial
approvals falling from 71 to 6. The gain is real and
not sufficient: the best configuration still approved 6
of 208 hard-denial actions, and was more permissive
than free-form on ambiguous proposals that should
have reached a human. Peak VRAM was 3.95 GiB
and median latency 2.55 s. An edge model of this
size can serve as one layer of a defence-in-depth
stack; the evidence does not support treating it as a
sole control.

Downloads

Published

2026-10-02

How to Cite

[1]
S. Poudel, R. Pandey, and A. Pandey, “Edge-Native Semantic Firewall for Autonomous LLM Agents: A Structured Chain-of-Thought Verification Framework”, ICICSET2025, vol. 3, no. 1, Oct. 2026.