Laura C. Diaz-Delgado ECCV 2026 · Malmö
ECCV 2026 European Conference on Computer Vision · Main conference

Frozen CLIP Priors for Robust Self-Supervised Poisson Inverse Problems

Laura C. Diaz-Delgado, Emmanuel Martinez, Henry Arguello

Department of Computer Science, High Dimensional Signal Processing Group, Universidad Industrial de Santander, Colombia

Session
Poster Session 3
When
CEST
Where
ExHall · board #50
Overview of the proposed method: an ADMM-unrolled network whose layers alternate an x-update for data consistency, a z-update implementing the CLIP prior as a frozen encoder with a trainable decoder, and a v-update, mapping the noisy measurement y to the reconstruction x-hat.
Overview of the proposed method. The reconstruction network with T iterations alternates between a data-consistency update xt and a CLIP-based prior update zt. The prior consists of a frozen CLIP encoder coupled with a learnable decoder. After T iterations, the updates yield the reconstruction .

Abstract

Self-supervised learning for imaging inverse problems is increasingly important in photon-limited settings, where acquiring clean ground truth is impractical and reconstruction must remain stable under dataset and acquisition shifts. This challenge is amplified under Poisson noise, whose signal-dependent statistics interact with sampling operators (e.g., CFA mosaicing). Meanwhile, foundation vision encoders trained at web scale offer distortion-invariant, content-related representations that generalize well across domains, suggesting a promising route to build priors that transfer beyond the training distribution without expensive fine-tuning.

This paper proposes an ADMM-inspired unrolled plug-and-play solver for Poisson inverse problems that decouples a closed-form data-consistency update from a parameter-efficient prior. The prior is implemented as a lightweight decoder operating on frozen CLIP RN50 dense multi-scale features, adapting foundation representations with less trainable parameters. For self-supervision, the method integrates GR2R measurement-domain re-corruption with an Equivariant Imaging regularizer via virtual acquisitions. Experiments on Poisson CFA demosaicing and deblurring show competitive quality, improved robustness under shifts, and self-supervised performance approaching supervised training.

Method

ADMM-unrolled plug-and-play
y=γ Poisson(Ax/γ)

Photon-limited acquisition. A is the known forward operator — CFA mosaicing or blur — and γ sets the shot-noise level.

1

Data consistency

A closed-form update for both operators — one division per pixel under CFA sampling, and the same in the Fourier domain under blur. No inner solver, no learned parameters.

2

Frozen CLIP prior

The proximal step is a denoiser built on frozen CLIP RN50 dense multi-scale features. Only a lightweight decoder is trained — 10.99 M parameters — initialised from Transfer CLIP weights; the encoder never moves.

3

Self-supervision

GR2R re-corruption in the measurement domain (α = 0.2), unbiased for Poisson data, combined with an Equivariant Imaging regulariser over virtual acquisitions (τ = 0.1). No clean image is ever required.

Data consistency — closed form for both operators
(AA+ρI) xt+1 = Ay+ρ (ztut)
CLIP prior — frozen encoder, trainable decoder
zt+1= 𝒟θ (z˜)= 𝒢θ (CLIP (z˜))
Self-supervised objective — no clean target
self= GR2R+ τ EI

The solver unrolls T = 2 ADMM iterations for demosaicing and T = 3 for deblurring, with parameters shared across iterations, so a single decoder serves every layer. Freezing the encoder is what preserves CLIP's distortion-invariant, content-related features: the ablation shows that learning those same features from scratch collapses performance, confirming the gains come from the pretrained representation rather than from the unrolling architecture alone.

Results

PSNR [dB] / SSIM
30.53 dB
Self-supervised demosaicing, BSDS500 at γ = 0.01
0.31 dB
Worst-case gap to its own supervised counterpart
100×
Fewer TFLOPs than Transfer CLIP
24 ms
Per image, 3×256×256 on one RTX 4070
Poisson demosaicing. Trained on BSDS500; DIV2K columns measure dataset shift.
BSDS500DIV2K
MethodγPSNRSSIMPSNRSSIM
DPIR0.0125.960.703925.950.7093
Transfer CLIP0.0127.880.783327.500.7841
GSPnP0.0129.250.806129.360.8237
RAM0.0130.460.862230.480.8729
Ours (Self)0.0130.530.860930.300.8641
Ours (Sup)0.0130.750.870330.560.8735
DPIR0.0524.710.635824.550.6397
Transfer CLIP0.0521.910.571621.670.6122
GSPnP0.0526.500.694526.580.7193
RAM0.0526.170.717826.330.7343
Ours (Self)0.0526.980.757926.600.7435
Ours (Sup)0.0527.040.752226.910.7669
Poisson deblurring. 9×9 Gaussian kernel, same photon regimes.
BSDS500DIV2K
MethodγPSNRSSIMPSNRSSIM
DPIR0.0128.810.794928.250.7894
Transfer CLIP0.0124.630.738924.330.7298
GSPnP0.0129.370.814928.900.8143
RAM0.0128.670.802828.400.8078
Ours (Self)0.0129.510.825728.970.8238
Ours (Sup)0.0129.800.838029.270.8362
DPIR0.0526.720.711326.180.7085
Transfer CLIP0.0519.870.533919.890.5478
GSPnP0.0527.030.722726.510.7223
RAM0.0525.800.679225.040.6679
Ours (Self)0.0527.450.753227.000.7564
Ours (Sup)0.0527.580.761327.180.7666

The method stays competitive across both photon regimes and under the BSDS500 → DIV2K dataset shift. The margins over the baselines widen where the photon count is lowest — at γ = 0.05 the self-supervised model leads every baseline on both tasks — and the self-supervised objective recovers most of the supervised performance, never falling more than 0.31 dB behind its own supervised counterpart across all eight settings.

Qualitative comparison

BSDS500 · γ = 0.01
Three rows of Poisson demosaicing results. Columns: mosaiced noisy measurement, DPIR, Transfer CLIP, GSPnP, RAM, Ours (Self), Ours (Sup), and the reference, with PSNR and SSIM overlaid on each reconstruction.
Poisson demosaicing. Columns show the mosaiced noisy Measurement, reconstructions by DPIR, Transfer CLIP, GSPnP and RAM, followed by the proposed method (Ours (Self)) and its supervised counterpart (Ours (Sup)), and the Reference. PSNR/SSIM are overlaid for direct visual comparison.
Three rows of Poisson deblurring results with the same column layout: blurred noisy measurement, DPIR, Transfer CLIP, GSPnP, RAM, Ours (Self), Ours (Sup), and the reference, with PSNR and SSIM overlaid.
Poisson deblurring. The blurred measurements lose high-frequency detail and exhibit noise-dependent grain, making texture recovery and edge localisation challenging. Per-image PSNR/SSIM are overlaid.

Real sensor data

SID · zero-shot · supplementary
Four panels on a real low-light SID capture: the raw mosaic input at 14.95 dB, Ours (self) at 26.37 dB, Ours (sup) at 29.66 dB, and the reference.
Zero-shot self-supervised reconstruction on real photon-limited data from the SID dataset. A 512×512 patch is reserved for testing, while the remaining patches from the same low-light image are used to train with self. The Poisson scaling parameter is estimated per RGB channel, and an affine colour transform is applied for sRGB visualisation and metric evaluation.
One real low-light capture, no clean target anywhere in training.
MethodPSNR [dB]SSIM
Raw mosaic input14.950.1731
Ours (Self)26.370.6482
Ours (Sup)29.660.8205

Synthetic Poisson simulations only go so far. This experiment moves to real photon-limited data from the SID dataset (Chen et al., Learning to See in the Dark) under a zero-shot adaptation protocol: a single real low-light image, one 512×512 patch held out for testing, and the remaining patches used only to optimise the self-supervised objective. No clean target and no paired supervision are used at any point.

Because real sensors exhibit channel-dependent photon statistics, the Poisson scaling is estimated independently for each RGB channel from the raw measurements, giving γrgb = (0.018, 0.017, 0.026). Reconstruction is performed in the linear RGB domain, then a standard affine colour transform is applied for sRGB visualisation and metric evaluation. Adapting from that one image alone lifts the raw mosaic from 14.95 dB to 26.37 dB — +11.42 dB with nothing clean to learn from — which is the evidence that the GR2R + EI training signal survives contact with real acquisition statistics rather than only the simulated ones.

Efficiency

3×256×256 · one RTX 4070
Runtime and complexity. Trainable parameters only; the CLIP encoder is frozen and shared across every unrolled layer. TFLOPs count the denoiser over the maximum iteration count.
MethodParams [M]TFLOPsTime [s]
DPIR32.6411.490.5193
Transfer CLIP10.999.002.4683
GSPnP17.019.120.9247
RAM34.130.320.8026
Ours10.990.090.0242

Keeping the physics in closed form is what makes the solver cheap: the same 10.99 M-parameter decoder as Transfer CLIP runs at a hundredth of the TFLOPs and roughly a hundred times faster per image, because the data-consistency step needs no inner optimisation.

Citation

@inproceedings{diazdelgado2026frozenclip,
  title     = {Frozen {CLIP} Priors for Robust Self-Supervised
               Poisson Inverse Problems},
  author    = {Diaz-Delgado, Laura C. and Martinez, Emmanuel
               and Arguello, Henry},
  booktitle = {Proceedings of the European Conference on
               Computer Vision (ECCV)},
  year      = {2026},
  eprint    = {2608.20524},
  archivePrefix = {arXiv},
  primaryClass  = {eess.IV}
}