[Submitted on 2 Sep 2026]
Abstract:LLMs are trained to create natural language. However, assorted strands of evidence signify that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding inner example computation. We current the term ``linguistic illegibility'' to broadly mention to scenarios in which an LLM's externalized or mechanistically-probed tongue artifacts neglect to portray how the example really thinks. We contend that the specter of linguistic illegibility is unavoidable for LLMs whose inner computations are not immediately expressed via language, but fairly math complete activation spaces (with lossy translations between activation spaces and natural tongue happening at the bookends). If linguistic illegibility is continually possible, afterward safety mechanisms that depend on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, lawful self-critique, activation probing for linguistically-defined characteristic vectors) can never be entirely sound; the example sandbox volition continually need isolation techniques whose guarantees do not depend on study a model's linguistic province at all. We contend that observing a model's outputs using taint tracking is a promising method for an productive sandbox: despite of how a example linguistically self-reports, a taint tracking guideline can define, a priori, assorted pieces of scheme province that should never be influenced by model-produced data. We additionally conversation multiple additional sandboxing mechanisms (e.g., sturdy virtualization, third-party auditing of sandboxing configurations) which collectively provision a crucial flat below linguistic monitoring, and would have mitigated latest sandbox exploits by frontier models.
Submission history
From: James Mickens [view email]
[v1] Wed, 2 Sep 2026 17:37:22 UTC (33 KB)