Mean gradient alignment versus mean transfer for seven understanding capabilities, averaging over 19 I2I sources.
Does visual generation help visual understanding?
To isolate the effect of confounding factors, we construct paired I2I and I2T tasks that require solving the same task from the same input, differing only in whether the answer is produced as an image or as text. This allows us to directly test whether supervision from an I2I objective transfers to the corresponding I2T task. This setup eliminates differences in annotation quality, image source, and other factors.
In Jigsaw, patches of an image are shuffled, and the model must recover their spatial arrangement. In Zoom-In, views of the same image at different zoom levels are shuffled, and the model must recover their correct order.
For each input, the I2I objective produces the correctly ordered image, while the I2T objective predicts the same ordering as a permutation in text. The objectives require the same visual operation on the same input and differ only in output modality.
Which training recipe transfers best?
We therefore use I2I → I2T as the default recipe in subsequent experiments.
An initial I2I training stage that updates parameters shared with the I2T objective provides a useful initialization for subsequent I2T learning.
OmniTaskonomy: A unified taxonomy of visual capabilities
Existing benchmarks typically organize visual tasks by task formulation or output modality, leaving visual generation and understanding tasks separated. This makes it difficult to compare and study tasks that rely on similar visual capabilities but produce different outputs. We introduce OmniTaskonomy, which organizes I2I tasks and understanding tasks in a shared hierarchy. For a visual understanding sample, we consider the primary visual capability required to solve it; for an I2I task, we consider the visual capability directly supervised by its training objective.
Which generation tasks help which capabilities?
| I2I supervision task | |||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Recognition | Reconstruction | Reorganization | |||||||||||||||||
| I2T capability | Object editing | Attribute editing | Colorization | Z-depth | Euclidean depth | Surface normals | Principal curvature | Occlusion edges | 3D keypoints | Reshading | Inpainting | Semantic segmentation | 2D edges | 2D keypoints | 2D segmentation | 2.5D segmentation | Object pointing | Jigsaw | Localization |
| Category recognition | |||||||||||||||||||
| Appearance understanding | |||||||||||||||||||
| Visual similarity | |||||||||||||||||||
| State recognition | |||||||||||||||||||
| Activity understanding | |||||||||||||||||||
| Anomaly detection | |||||||||||||||||||
| Situation understanding | |||||||||||||||||||
| OCR and text recognition | |||||||||||||||||||
| Lighting understanding | |||||||||||||||||||
| Depth understanding | |||||||||||||||||||
| Metric 3D relation | |||||||||||||||||||
| Orientation understanding | |||||||||||||||||||
| 2D spatial relation | |||||||||||||||||||
| Multi-view reasoning | |||||||||||||||||||
| Localization | |||||||||||||||||||
| Connectivity | |||||||||||||||||||
| Counting | |||||||||||||||||||
| 2D Ordering | |||||||||||||||||||
| Visual correspondence | |||||||||||||||||||
Visual generation supervision yields significant gains for specific visual understanding capabilities, through both related tasks and transfer across tasks.
Shared visual operations predict several of the strongest gains.
Localization and object pointing produce the largest improvements in counting (+2.5 and +2.0 pp), consistent with all three tasks requiring individual objects to be identified and spatially localized.
Similarly, Z-depth, Euclidean depth, and surface normals improve metric 3D relation by 3.6, 3.8, and 3.4 pp, respectively, connecting dense geometric reconstruction to relational 3D judgments. Jigsaw improves 2D ordering by 6.8 pp, consistent with both tasks requiring the relative spatial arrangement of image regions.
Useful transfer is not confined to closely matched capabilities.
Inpainting improves both counting (+1.5 pp) and 2D ordering (+7.2 pp), despite neither target explicitly requiring missing-region reconstruction. Solving inpainting tasks may encourage the model to infer object quantity and global spatial structure from incomplete local evidence. 2.5D segmentation likewise improves category recognition (+1.2 pp), potentially because capturing object boundaries can shape cues that support category recognition.
What explains visual generation-to-understanding transfer?
Alignment and transfer for all 133 source–target pairs.
average alignment and average transfer are strongly positively correlated across the seven capabilities (r=0.795). Capabilities whose gradients are more compatible with generation objectives, on average, therefore, also tend to receive larger gains from I2I training.
Alignment and transfer are positively correlated across these 133 pairs (r=0.529).
Gradient alignment is concentrated in early pre-attention normalization layers and is positively associated with downstream transfer across both understanding capabilities and individual source-target pairs.
Citation
% BibTeX pending.

