GPTProto

When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs

HuggingFace Daily Papers(社区热门论文)·Sep 28, 2026, 8:00 AM·Meta / Llama

View PDF HTML (experimental)

Abstract:Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matched random directions matches calibrated target gain. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches, and promising point estimates on Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty. SAKIKO establishes the necessity of outcome-resolved adjudication before claiming internal repair. Code: this https URL.
Comments: Preprint
Subjects: Computation and Language (cs.CL); Software Engineering (cs.SE)
Cite as: arXiv:2609.36138 [cs.CL]
  (or arXiv:2609.36138v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.36138

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruizhe Li [view email]
[v1] Mon, 28 Sep 2026 19:12:55 UTC (1,660 KB)