NeurIPS 2026

Reading the Unreadable

Text-Aware Image Super-Resolution Needs Reasoning

Jaeseong Lee  ·  Jinwoo Kim  ·  Jinho Jeong  ·  Seon Joo Kim

Yonsei University

Text-aware image super-resolution (TAISR) restores a high-resolution (HR) image from a low-resolution (LR) input with unknown degradations, while preserving the text in it.

Existing methods treat this as a visual restoration problem, recovering text from local visual evidence alone.

But humans do not rely on local evidence alone. When text is nearly illegible, they can often still read it by reasoning over the surrounding context, logical patterns, and prior knowledge.

This raises a fundamental question:

Does TAISR need reasoning?

1. Benchmark

To answer this question, we construct the ReasonText benchmark, the first TAISR benchmark that separates locally readable text from reasoning-required text.

We label each text instance by how a human can read it in the LR image:

Level 1 (locally readable): a human can read the text given only the cropped LR text image.

Level 2 (reasoning-required): a human can only read it from the LR image by reasoning over the surrounding context.

2. Diagnosis

A: GAN-based  ·  B: Multi-step diffusion  ·  C: Few-step diffusion  ·  D: AR-based  ·  E: Text-aware SR

GroupModel Fidelity Perceptual Text Restoration
ReferenceNo Reference
PSNR↑SSIM↑ LPIPS↓DISTS↓MANIQA↑MUSIQ↑CLIP-IQA↑ Overall↑Level 1↑Level 2↑
HR (oracle)80.00001.00000.00000.00000.426166.23670.540394.681396.294290.6911
AReal-ESRGAN24.04530.76940.21340.17190.446368.53970.461140.316752.96479.0268
SwinIR24.77620.78590.19890.16260.428866.94830.435745.676059.407111.7066
BDiffBIR24.10970.70390.25910.18800.612771.84270.641754.973669.612318.7588
SUPIR24.29710.70980.26220.19470.460863.59980.592255.298468.529122.5670
FaithDiff24.20380.73630.20610.16230.509571.65400.582152.009765.165319.4640
DiT4SR22.65390.69330.24770.18300.497370.85020.607459.155573.147124.5416
CResShift24.91700.76450.22190.18390.411063.30330.490247.706061.003414.8096
SinSR25.16440.76680.21160.17470.455166.19940.557443.118155.986311.2835
OSEDiff23.95960.75330.21470.16660.480371.15590.567244.620458.437910.4372
DPURE22.29700.67060.26520.18930.523370.00360.653134.389045.72416.3470
ETeReDiff23.51880.72330.23620.17960.563870.51810.618351.238365.108316.9252
UniT22.76720.65030.28140.20720.515968.73820.631561.388676.453824.1185

For each metric, the best score among the twelve SR models is shown in bold (the HR oracle is excluded).

We evaluate twelve recent SR models on ReasonText, including real-world SR models for natural images (A–D) and TAISR models designed for text (E).

All models struggle to restore reasoning-required text, regardless of architecture, generative backbone, or training paradigm, and even TAISR models designed specifically for text are no exception.

These results answer our question:

TAISR does need reasoning.

3. Method

The remaining question is how to supply this reasoning to SR models. Many recent SR models already take a text caption as an additional input, and our method builds on two observations:

Observation 1. These text-conditioned SR models restore reasoning-required text substantially better when the caption contains the ground-truth transcripts.

Observation 2. Modern multimodal large language models (MLLMs) can read much of the reasoning-required text in an LR image by reasoning over its context.

RTC method: Caption by Reasoning then Restore by Conditioning
Overall pipeline of Reasoning Transfer via Captioning (RTC), illustrated with an example.

Inspired by these two observations, we propose Reasoning Transfer via Captioning (RTC), which works in two stages:

Stage 1 (Caption by Reasoning): an MLLM reads the LR image with a reasoning prompt and produces a two-sentence caption: a scene description, followed by the text it infers by reasoning.

Stage 2 (Restore by Conditioning): this caption replaces the default caption of a text-conditioned SR model, which restores the image.

Both models stay frozen, so RTC can be attached to any text-conditioned SR model.

GroupModel Fidelity Perceptual Text Restoration
ReferenceNo Reference
PSNR↑SSIM↑ LPIPS↓DISTS↓MANIQA↑MUSIQ↑CLIP-IQA↑ Overall↑Level 1↑Level 2↑
HR (oracle)80.00001.00000.00000.00000.426166.23670.540394.681396.294290.6911
BSUPIR24.29710.70980.26220.19470.460863.59980.592255.298468.529122.5670
+ RTC23.95340.70710.24800.18460.520266.86680.641267.316378.848338.7870
FaithDiff24.20380.73630.20610.16230.509571.65400.582152.009765.165319.4640
+ RTC24.27240.73870.20200.16020.497571.45770.579761.023174.002328.9140
DiT4SR22.65390.69330.24770.18300.497370.85020.607459.155573.147124.5416
+ RTC22.68160.69360.24790.18300.500970.67450.611073.568884.777745.8392
COSEDiff23.95960.75330.21470.16660.480371.15590.567244.620458.437910.4372
+ RTC23.89030.75040.21710.16820.488671.22810.576749.411364.253112.6939
ETeReDiff23.51880.72330.23620.17960.563870.51810.618351.238365.108316.9252
+ RTC23.38840.72120.23770.17950.577770.98520.623756.638270.125423.2722
UniT22.76720.65030.28140.20720.515968.73820.631561.388676.453824.1185
+ RTC22.84250.65550.27510.20410.516968.81300.633870.970482.554242.3131

For each + RTC value, blue/red marks an improvement/drop over the baseline.

RTC consistently improves all six text-conditioned SR models, especially on reasoning-required text, while preserving general image quality.

Remarkably, with RTC, the real-world SR model DiT4SR outperforms UniT, a TAISR model that builds on the same backbone and adds text-aware modules. This suggests that, once reasoning is supplied, current text-aware modules may bring no additional benefit.

It's time to design TAISR that reasons!

Citation

@inproceedings{lee2026reading,
  title     = {Reading the Unreadable: Text-Aware Image Super-Resolution Needs Reasoning},
  author    = {Lee, Jaeseong and Kim, Jinwoo and Jeong, Jinho and Kim, Seon Joo},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}