Authors:
,
,
,
,
,
,
,
,
,
Abstract
Vision-Language Models can procedure content as rendered images, but accuracy degrades alongside compression; LensVLM addresses this by scanning compressed images and selectively expanding applicable parts through learned tools, maintaining elevated accuracy equal at elevated compression ratios.
Vision Language Models (VLMs) recommendation the breathtaking possible of handling content as rendered images, bypassing the need for tokenizing the content into lengthy token sequences. Since VLM depiction encoders map fixed-size images to a fixed figure of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the imagination encoder's productive resolution, making them indistinguishable. To location this, we propose LensVLM, an conclusion example and post-training formula that enables VLMs to scan compressed images, afterward selectively develop lone the applicable images to their uncompressed form via learned tools. Building on Qwen3.5-9B-Base, LensVLM maintains accuracy comparable to the full-text high border at 4.3x productive compression and outperforms retrieval-based, text- and visual-compression baselines up to 10.1x productive compression throughout seven text QA benchmarks. LensVLM additionally generalizes to multimodal document and code understanding tasks, alongside the accuracy acquire complete baselines expanding as compression increases. Our inspection validates this approach: training makes visual compression sturdy to rendering choices, and as compression grows the example increasingly relies on expanded satisfied fairly than unreliable ocular reading. The inspection additionally yields applicable tool-choice guidance: content enlargement is preferable for rendered text, during high-resolution depiction enlargement suits native documents whose layout cues transport task-relevant information.
Get this document in your agent:
hf document peruse 2605.07019
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this document 2
![]()
apple/LensVLM-9B
Image-Text-to-Text •
9B •
Updated concerning 24 hours ago
•
233
•
67
![]()
suryatmodulus/LensVLM-9B
Image-Text-to-Text •
9B •
Updated concerning 2 hours ago
Datasets citing this document 0
No dataset linking this paper
Cite arxiv.org/abs/2605.07019 in a dataset README.md to nexus it from this page.
Spaces citing this document 1
Collections including this document 0
No Collection including this paper
Add this document to a collection to nexus it from this page.