CVPR 2026

Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning

Jiachen Tu*,1Guanghui Qin*,2Theodore Zhengde Zhao2Jeya Maria Jose Valanarasu2Sheng Zhang2Tristan Naumann2Fan Lam1Sheng Wang3Hoifung Poon2
*Equal contribution  ·  1University of Illinois Urbana-Champaign  ·  2Microsoft  ·  3University of Washington

Abstract

Effective medical image analysis requires representations that capture both global anatomical structure and fine-grained tissue texture. Current self-supervised approaches exhibit limited capacity to address both requirements simultaneously. Invariance-based methods learn through augmentation consistency but face challenges in medical imaging where common augmentations may discard diagnostically relevant intensity patterns. Masked image modeling approaches employ high masking ratios to enforce holistic reasoning, yet inherently limit exposure to fine-grained texture. Recent work in general-domain vision demonstrates that generative and semantic objectives can mutually benefit each other, yet this paradigm remains unexplored for 3D medical imaging. We introduce Masked-Diffusion Autoencoders (MDAE), a self-supervised framework that imposes concurrent spatial masking and diffusion corruption, encouraging the model to learn complementary objectives: masked region reconstruction for structural coherence and visible region denoising for textural characteristics. This dual corruption enables the network to learn structure-texture representations within a unified time-conditioned objective. Evaluated on brain MRI across tumor classification, molecular marker detection, and dense segmentation benchmarks, MDAE consistently outperforms state-of-the-art baselines, with improvements most pronounced in cross-modal generalization tasks.

Method

MDAE Framework: Dual corruption pipeline combining spatial masking and diffusion noise for 3D medical volume representation learning

MDAE Framework. A 3D medical volume is corrupted with voxel Gaussian noise at a sampled diffusion timestep t, then block-masked at a variable masking ratio. The encoder processes the corrupted, masked input, and the decoder—conditioned on t via FiLM layers—reconstructs the clean volume. The loss combines masked region reconstruction (structural learning) and visible region denoising (texture learning).

Comparison of corruption strategies: MAE, DSM, and MDAE

Corruption strategy comparison. Masked Autoencoding (MAE) removes spatial patches; Denoising Score Matching (DSM) adds global noise; MDAE combines both, enabling complementary structure and texture learning.

Citation

@inproceedings{tu2026mdae,
  title     = {Masked-Diffusion Autoencoders for 3D Medical Vision
               Representation Learning},
  author    = {Tu, Jiachen and Qin, Guanghui and Zhao, Theodore Zhengde and
               Valanarasu, Jeya Maria Jose and Zhang, Sheng and
               Naumann, Tristan and Lam, Fan and Wang, Sheng and
               Poon, Hoifung},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer
               Vision and Pattern Recognition (CVPR)},
  year      = {2026}
}