Frozen CLIP Priors for Robust Self-Supervised Poisson Inverse Problems
Department of Computer Science, High Dimensional Signal Processing Group, Universidad Industrial de Santander, Colombia
- Session
- Poster Session 3
- When
- CEST
- Where
- ExHall · board #50
Abstract
Self-supervised learning for imaging inverse problems is increasingly important in photon-limited settings, where acquiring clean ground truth is impractical and reconstruction must remain stable under dataset and acquisition shifts. This challenge is amplified under Poisson noise, whose signal-dependent statistics interact with sampling operators (e.g., CFA mosaicing). Meanwhile, foundation vision encoders trained at web scale offer distortion-invariant, content-related representations that generalize well across domains, suggesting a promising route to build priors that transfer beyond the training distribution without expensive fine-tuning.
This paper proposes an ADMM-inspired unrolled plug-and-play solver for Poisson inverse problems that decouples a closed-form data-consistency update from a parameter-efficient prior. The prior is implemented as a lightweight decoder operating on frozen CLIP RN50 dense multi-scale features, adapting foundation representations with less trainable parameters. For self-supervision, the method integrates GR2R measurement-domain re-corruption with an Equivariant Imaging regularizer via virtual acquisitions. Experiments on Poisson CFA demosaicing and deblurring show competitive quality, improved robustness under shifts, and self-supervised performance approaching supervised training.
Method
ADMM-unrolled plug-and-playPhoton-limited acquisition. A is the known forward operator — CFA mosaicing or blur — and γ sets the shot-noise level.
Data consistency
A closed-form update for both operators — one division per pixel under CFA sampling, and the same in the Fourier domain under blur. No inner solver, no learned parameters.
Frozen CLIP prior
The proximal step is a denoiser built on frozen CLIP RN50 dense multi-scale features. Only a lightweight decoder is trained — 10.99 M parameters — initialised from Transfer CLIP weights; the encoder never moves.
Self-supervision
GR2R re-corruption in the measurement domain (α = 0.2), unbiased for Poisson data, combined with an Equivariant Imaging regulariser over virtual acquisitions (τ = 0.1). No clean image is ever required.
The solver unrolls T = 2 ADMM iterations for demosaicing and T = 3 for deblurring, with parameters shared across iterations, so a single decoder serves every layer. Freezing the encoder is what preserves CLIP's distortion-invariant, content-related features: the ablation shows that learning those same features from scratch collapses performance, confirming the gains come from the pretrained representation rather than from the unrolling architecture alone.
Results
PSNR [dB] / SSIM| BSDS500 | DIV2K | ||||
|---|---|---|---|---|---|
| Method | γ | PSNR | SSIM | PSNR | SSIM |
| DPIR | 0.01 | 25.96 | 0.7039 | 25.95 | 0.7093 |
| Transfer CLIP | 0.01 | 27.88 | 0.7833 | 27.50 | 0.7841 |
| GSPnP | 0.01 | 29.25 | 0.8061 | 29.36 | 0.8237 |
| RAM | 0.01 | 30.46 | 0.8622 | 30.48 | 0.8729 |
| Ours (Self) | 0.01 | 30.53 | 0.8609 | 30.30 | 0.8641 |
| Ours (Sup) | 0.01 | 30.75 | 0.8703 | 30.56 | 0.8735 |
| DPIR | 0.05 | 24.71 | 0.6358 | 24.55 | 0.6397 |
| Transfer CLIP | 0.05 | 21.91 | 0.5716 | 21.67 | 0.6122 |
| GSPnP | 0.05 | 26.50 | 0.6945 | 26.58 | 0.7193 |
| RAM | 0.05 | 26.17 | 0.7178 | 26.33 | 0.7343 |
| Ours (Self) | 0.05 | 26.98 | 0.7579 | 26.60 | 0.7435 |
| Ours (Sup) | 0.05 | 27.04 | 0.7522 | 26.91 | 0.7669 |
| BSDS500 | DIV2K | ||||
|---|---|---|---|---|---|
| Method | γ | PSNR | SSIM | PSNR | SSIM |
| DPIR | 0.01 | 28.81 | 0.7949 | 28.25 | 0.7894 |
| Transfer CLIP | 0.01 | 24.63 | 0.7389 | 24.33 | 0.7298 |
| GSPnP | 0.01 | 29.37 | 0.8149 | 28.90 | 0.8143 |
| RAM | 0.01 | 28.67 | 0.8028 | 28.40 | 0.8078 |
| Ours (Self) | 0.01 | 29.51 | 0.8257 | 28.97 | 0.8238 |
| Ours (Sup) | 0.01 | 29.80 | 0.8380 | 29.27 | 0.8362 |
| DPIR | 0.05 | 26.72 | 0.7113 | 26.18 | 0.7085 |
| Transfer CLIP | 0.05 | 19.87 | 0.5339 | 19.89 | 0.5478 |
| GSPnP | 0.05 | 27.03 | 0.7227 | 26.51 | 0.7223 |
| RAM | 0.05 | 25.80 | 0.6792 | 25.04 | 0.6679 |
| Ours (Self) | 0.05 | 27.45 | 0.7532 | 27.00 | 0.7564 |
| Ours (Sup) | 0.05 | 27.58 | 0.7613 | 27.18 | 0.7666 |
The method stays competitive across both photon regimes and under the BSDS500 → DIV2K dataset shift. The margins over the baselines widen where the photon count is lowest — at γ = 0.05 the self-supervised model leads every baseline on both tasks — and the self-supervised objective recovers most of the supervised performance, never falling more than 0.31 dB behind its own supervised counterpart across all eight settings.
Qualitative comparison
BSDS500 · γ = 0.01
Real sensor data
SID · zero-shot · supplementary
| Method | PSNR [dB] | SSIM |
|---|---|---|
| Raw mosaic input | 14.95 | 0.1731 |
| Ours (Self) | 26.37 | 0.6482 |
| Ours (Sup) | 29.66 | 0.8205 |
Synthetic Poisson simulations only go so far. This experiment moves to real photon-limited data from the SID dataset (Chen et al., Learning to See in the Dark) under a zero-shot adaptation protocol: a single real low-light image, one 512×512 patch held out for testing, and the remaining patches used only to optimise the self-supervised objective. No clean target and no paired supervision are used at any point.
Because real sensors exhibit channel-dependent photon statistics, the Poisson scaling is estimated independently for each RGB channel from the raw measurements, giving γrgb = (0.018, 0.017, 0.026). Reconstruction is performed in the linear RGB domain, then a standard affine colour transform is applied for sRGB visualisation and metric evaluation. Adapting from that one image alone lifts the raw mosaic from 14.95 dB to 26.37 dB — +11.42 dB with nothing clean to learn from — which is the evidence that the GR2R + EI training signal survives contact with real acquisition statistics rather than only the simulated ones.
Efficiency
3×256×256 · one RTX 4070| Method | Params [M] | TFLOPs | Time [s] |
|---|---|---|---|
| DPIR | 32.64 | 11.49 | 0.5193 |
| Transfer CLIP | 10.99 | 9.00 | 2.4683 |
| GSPnP | 17.01 | 9.12 | 0.9247 |
| RAM | 34.13 | 0.32 | 0.8026 |
| Ours | 10.99 | 0.09 | 0.0242 |
Keeping the physics in closed form is what makes the solver cheap: the same 10.99 M-parameter decoder as Transfer CLIP runs at a hundredth of the TFLOPs and roughly a hundred times faster per image, because the data-consistency step needs no inner optimisation.
Citation
@inproceedings{diazdelgado2026frozenclip,
title = {Frozen {CLIP} Priors for Robust Self-Supervised
Poisson Inverse Problems},
author = {Diaz-Delgado, Laura C. and Martinez, Emmanuel
and Arguello, Henry},
booktitle = {Proceedings of the European Conference on
Computer Vision (ECCV)},
year = {2026},
eprint = {2608.20524},
archivePrefix = {arXiv},
primaryClass = {eess.IV}
}