The Implications of Linguistic Illegibility for LLM Security

Hacker News by 2 min read 34x views
The Implications of Linguistic Illegibility for LLM Security

Share Post

[Submitted on 2 Sep 2026]

View PDF HTML (experimental)

Abstract:LLMs are trained to create natural language. However, assorted strands of evidence signify that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding inner example computation. We current the term ``linguistic illegibility'' to broadly mention to scenarios in which an LLM's externalized or mechanistically-probed tongue artifacts neglect to portray how the example really thinks. We contend that the specter of linguistic illegibility is unavoidable for LLMs whose inner computations are not immediately expressed via language, but fairly math complete activation spaces (with lossy translations between activation spaces and natural tongue happening at the bookends). If linguistic illegibility is continually possible, afterward safety mechanisms that depend on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, lawful self-critique, activation probing for linguistically-defined characteristic vectors) can never be entirely sound; the example sandbox volition continually need isolation techniques whose guarantees do not depend on study a model's linguistic province at all. We contend that observing a model's outputs using taint tracking is a promising method for an productive sandbox: despite of how a example linguistically self-reports, a taint tracking guideline can define, a priori, assorted pieces of scheme province that should never be influenced by model-produced data. We additionally conversation multiple additional sandboxing mechanisms (e.g., sturdy virtualization, third-party auditing of sandboxing configurations) which collectively provision a crucial flat below linguistic monitoring, and would have mitigated latest sandbox exploits by frontier models.

Submission history

From: James Mickens [view email]
[v1] Wed, 2 Sep 2026 17:37:22 UTC (33 KB)

Other Article Hacker News
Close Right Ads
Close Left Ads