Abstract
Magnetic resonance imaging (MRI) inherently suffers from noise, which limits downstream medical analyses. In MRI, noise-free images are unobtainable; therefore, existing denoising approaches formulate surrogate training objectives, compromising between preserving detail and concealing noise, causing domain shifts or incomplete denoising. To enable denoiser training directly on unmodified, noisy images, we exploit repeated acquisitions. This naturally constitutes a physical Noise2Noise (pN2N) setting. For unrepeated data, we introduce a diffusion-based re-noiser that synthesizes noisy image pairs, extending pN2N to Renoise2Noise (ReN2N). Furthermore, we demonstrate that ReN2N improves generalization to unseen datasets. Additionally, we propose to combine pN2N or ReN2N with optional guidance from co-acquired contrast, yielding four versions of our novel denoising framework: YADO (You Accurately Denoise real Observations). Across 14 test conditions, YADO consistently outperforms 17 state-of-the-art baselines, matching the quality of physically noise-suppressed images obtained via brute-force averaging of independent acquisitions. YADO thus establishes practical denoising for real-world acquisition settings.
Also presented in the ECCV MedFM-Bench workshop: Poster.
Background
What does MRI "Denoising" actually mean?
While acquired images are often assumed or defined to be clean, they are not: If we re-acquire an image, it will look different every time. This is the whole point of denoising.
Every MRI is noisy — there is no "clean" MRI in the real world.
Thus, we have to treat the acquisition of an MRI Xt as a random process. Here we model image formation as:
Hence, we assume X′ = E[X] as the (surrogate, inaccessible) clean image *We thus incorporate all systematic effects like the Rician bias or undersampling artifacts into the noise.. This definition is inspired by, and thus theoretically grounded in, the only way to approximate the clean image X′, which is brute-force averaging:
Denoising is therefore the task of recovering X′ given a noisy Xin, i.e. finding such that:
Results
Matching NEx=7 (TA=45 min) quality from routine MRI
Visual comparison of denoising performance of 4 YADO variants from routine Rhineland (single-acquisition, NEx=1) T1w scans to physically noise-suppressed (NEx=7, >45 min scan time) reference images.
Use slider to compare, click the zoom in (click the zoom badge to reset)
Benchmark
Quantitative Results
SSIM (%) and CNR against 17 state-of-the-art baselines for T1w denoising, evaluated on models trained on OASIS-3 (ca. 1.0 mm) and HCP-A (0.8 mm).
| Test dataset | Venue | OASIS-3 trained (ca. 1.0 mm) | HCP-A trained (0.8 mm) | Rank (all) |
|||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SSIM (%) *Based on repeated data to evaluate removal of true, physical noise. | CNR | SSIM (%) *Based on repeated data to evaluate removal of true, physical noise. | CNR | ||||||||||
| OAS3 | CHDI | Kirby | HCP-E | -YA | RS | PVS | HCP-E | -YA | RS | PVS | |||
| uYADO (pN2N) | ECCV '26 (proposed) |
89.39 | 94.25 | 95.69 | 89.65 | 94.06 | 93.83 | 2.63 | 87.05 | 85.85 | 90.44 | 2.15 | 4.55 |
| uYADO (ReN2N) | 89.41 | 94.33 | 95.99 | 90.60 | 94.50 | 94.37 | 2.74 | 88.25 | 86.66 | 91.23 | 2.32 | 2.09 | |
| gYADO (pN2N) | 90.03 | 94.11 | (94.63) | 89.96 | 94.11 | 94.31 | 2.85 | 86.87 | 86.05 | 89.93 | 2.25 | 4.17 | |
| gYADO (ReN2N) | 89.77 | 94.59 | (95.63) | 90.91 | 94.59 | 94.59 | 2.85 | 88.74 | 86.98 | 91.69 | 2.29 | 1.46 | |
| Noisier2N | CVPR '20 | 88.59 | 93.53 | 95.58 | 89.47 | 93.91 | 93.42 | 2.28 | 84.63 | 82.65 | 88.45 | 2.10 | 9.26 |
| N2Score | NeurIPS '21 | 89.07 | 93.77 | 95.21 | 87.89 | 93.61 | 92.75 | 1.06 | 79.97 | 65.71 | 90.24 | 1.71 | 12.24 |
| N2Void | CVPR '19 | 88.60 | 93.78 | 95.72 | 88.19 | 90.86 | 92.63 | 2.24 | 86.99 | 85.14 | 90.11 | 2.20 | 8.58 |
| S-N2Void | ISBI '20 | 88.69 | 93.67 | 95.59 | 88.43 | 91.38 | 92.62 | 2.17 | 87.42 | 85.89 | 90.56 | 2.22 | 7.81 |
| Neigh2Neigh | CVPR '21 | 87.94 | 93.03 | 94.59 | 89.08 | 93.61 | 92.68 | 2.31 | 87.36 | 85.40 | 89.63 | 1.83 | 9.75 |
| RED-WGAN | MedIA '19 | 86.35 | 91.38 | 93.42 | 87.21 | 89.48 | 90.49 | 2.18 | 84.70 | 83.53 | 88.34 | 1.88 | 15.35 |
| CNN3D | MedIA '19 | 87.79 | 93.00 | 95.11 | 88.86 | 92.62 | 92.62 | 2.25 | 86.28 | 84.51 | 89.31 | 1.95 | 11.62 |
| DDM² | ICLR '23 | 51.05 | 43.04 | 39.29 | 35.15 | 39.49 | 48.27 | 0.16 | 23.64 | 30.25 | 39.82 | 0.02 | 22.00 |
| DDM² Reg. *Applying single-step regression sampling as previously introduced with YODA. | ICLR '23 | 87.28 | 91.14 | 92.73 | 84.10 | 89.80 | 91.28 | 1.76 | 72.24 | 75.28 | 84.41 | 1.01 | 17.70 |
| BME-X | Nat. BE '25 | 83.10 | 87.81 | 86.64 | 82.87 | 88.38 | 88.05 | 1.55 | 82.71 | 82.17 | 88.14 | 1.56 | 18.68 |
| N2Contrast | IPMI '23 | 87.32 | 93.12 | 95.34 | 87.98 | 90.46 | 92.16 | 2.42 | 85.51 | 83.39 | 88.97 | 2.12 | 11.82 |
| YODA | TMI '26 | 88.29 | 87.95 | (70.80) | 83.48 | 86.87 | 89.18 | 2.17 | 71.19 | 80.11 | 84.17 | 1.69 | 17.30 |
| DIP | CVPR '18 | 87.87 | 93.27 | 95.31 | 89.04 | 92.23 | 92.88 | 2.33 | 86.77 | 84.86 | 89.58 | 1.97 | 9.84 |
| Pixel2Pixel | TPAMI '25 | 85.76 | 90.81 | 92.16 | 87.20 | 91.39 | 90.23 | 1.41 | 86.20 | 83.87 | 87.94 | 1.21 | 16.46 |
| ANTs NLM | JMRI '10 | 87.89 | 93.11 | 94.97 | 88.95 | 93.21 | 92.62 | 2.36 | 87.22 | 85.42 | 89.73 | 1.91 | 9.49 |
| BM4D | TIP '12 | 88.28 | 93.40 | 95.33 | 89.37 | 93.71 | 92.93 | 2.13 | 87.81 | 86.09 | 90.27 | 1.83 | 7.96 |
| BME-X (as-is) | Nat. BE '25 | 52.72 | 55.39 | 55.48 | 46.47 | 68.41 | 53.13 | 1.53 | 35.33 | 68.85 | 64.45 | 1.53 | 20.52 |
| Identity | proposed*While the operation is — of course — trivial, we are not aware of any prior works comparing trained denoisers against the Identity function. | 86.91 | 92.66 | 95.03 | 87.09 | 89.50 | 91.60 | 2.33 | 84.61 | 82.68 | 88.47 | 2.10 | 14.35 |
T2w/FLAIR: SSIM (%) and CNR (RS-trained, 0.8 mm), plus WMH lesion-segmentation F1 agreement on FLAIR (via SHIVA). We only evaluated representative, well-performing methods from T1w denoising.
| Test dataset | Venue | RS-trained T2w (0.8 mm) | RS-trained FLAIR (0.8 mm) | Rank (all) |
|||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| SSIM (%) *Based on repeated data to evaluate removal of true, physical noise. | CNR | SSIM (%) *Based on repeated data to evaluate removal of true, physical noise. | CNR | F1 | |||||||
| RS | HCP-E | -YA | PVS | RS | still | nodd | WMH | WMH | |||
| uYADO (pN2N) | ECCV '26 (proposed) |
88.80 | 90.21 | 89.50 | 2.69 | 88.36 | 84.89 | 73.84 | 3.14 | 61.58 | 5.60 |
| uYADO (ReN2N) | 88.58 | 91.10 | 89.76 | 2.59 | 88.25 | 84.94 | 73.88 | 3.08 | 61.02 | 4.60 | |
| gYADO (pN2N) | 89.65 | 90.38 | 89.24 | 3.12 | 89.60 | (69.94) | (62.75) | 3.40 | 64.23 | 3.55 | |
| gYADO (ReN2N) | 88.98 | 91.43 | 90.28 | 2.91 | 89.28 | (59.29) | (52.94) | 3.33 | 62.42 | 3.63 | |
| Noisier2N | CVPR '20 | 87.00 | 90.23 | 88.58 | 2.21 | 85.78 | 83.21 | 70.28 | 3.34 | 62.82 | 7.65 |
| S-N2V | CVPR '19 | 87.69 | 90.83 | 89.26 | 2.38 | 86.60 | 83.58 | 71.45 | 3.36 | 62.75 | 5.52 |
| DIP | CVPR '18 | 87.44 | 90.95 | 89.79 | 2.29 | 87.04 | 83.90 | 71.12 | 3.28 | 62.68 | 5.30 |
| ANTs NLM | JMRI '10 | 86.76 | 90.27 | 89.05 | 2.32 | 87.20 | 83.52 | 72.80 | 3.37 | 62.89 | 6.00 |
| BM4D | TIP '12 | 87.28 | 90.93 | 89.68 | 2.52 | 87.76 | 84.51 | 73.87 | 3.21 | 62.83 | 4.74 |
| Identity | proposed*While the operation is — of course — trivial, we are not aware of any prior works comparing trained denoisers against the Identity function. | 86.99 | 90.14 | 88.52 | 2.22 | 85.73 | 83.18 | 70.27 | 3.37 | 62.45 | 8.41 |
YADO for FLAIR & T2w denoising
We also trained YADO versions for FLAIR and T2w denoising, thus, covering all main structural contrasts.
OOD Performance
Beyond routine, in-distribution scans, we tested YADO on several (severely) OOD images:
Huntington's disease alters brain structure causing severe atrophy and results in involuntary motion which can induce severe motion artifacts during MRI. See more results in the paper.
Applying 3T HCP (0.8 mm MPRAGE) to ultra-high-field MP2RAGE at 0.5 mm.
Although trained on (predominantly) healthy participants and with non-enhanced MPRAGE, YADO can handle tumor cases and contrast-enhanced T1w images.
NEx=4
Human-trained YADO also generalizes to monkey brain, i.e. denoising is not tied to human-only priors.
Cite
BibTeX
@inproceedings{rassmann2026yado,
title = {Rethinking Real-World MRI Denoising: Learning from Physical Noise},
author = {Rassmann, Sebastian and K\"ugler, David and Brunheim, Sascha and Ehses, Philipp and Reuter, Martin},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}