Diffusion Controller is a framework that treats the text-to-image denoising process as a continuous control problem and supplements a frozen pretrained base model with a lightweight "steering damper" network. Evaluated on a Stable Diffusion v1.4 backbone using supervised fine-tuning (SFT), reward-weighted loss (RWL) and PPO, the method improved human preference alignment and the fully fine-tuned white-box variant reached a 90% win rate over the baseline.
The problem addressed
Modern text-to-image models such as Nano Banana, Stable Diffusion and Flux have made high-fidelity, photorealistic image synthesis from text widely available. Steering these large models to meet precise user intent, downstream requirements, or strict visual constraints is still challenging: a requested attribute might be omitted or, if forced, may cause distortions that degrade overall image quality.
Historically, guidance techniques fall into two categories: inference-time adjustments (for example classifier-free diffusion guidance) that alter generation on the fly, and heavier fine-tuning approaches that change model behavior through adapters like LoRA, reward-weighted regression, or policy gradients. Because these approaches have been developed and used separately, the field lacked a single principled mathematical framework to analyze and unify control strategies, leaving practitioners to balance preference alignment and image fidelity by trial and error.
Diffusion Controller: core idea
Diffusion Controller reframes denoising as a smoothly controlled trajectory. Instead of rebuilding or modifying the core of a large pretrained model (analogous to swapping out a motorcycle engine), the method attaches a compact steering damper network while keeping the base model frozen.
During generation the steering damper observes intermediate denoising states and injects small, carefully computed corrections that bias the trajectory toward objectives defined by a user or reward function (for example specific stylistic or contextual targets). The framework uses mathematical optimization with feedback to nudge the generation toward new preferences without sacrificing the base model’s visual clarity or its stability. For the earlier example, the damper can encourage the model to include sunglasses on a lizard while preventing distortions of scale or texture.
From theory to practice: two fine-tuning methods
To make the control formulation practical, the authors implemented two reward-based fine-tuning strategies:
- Policy gradient with PPO (the "steady optimizer"): an iterative method that applies incremental changes and includes a clipping rule to avoid large, destabilizing updates during training.
- Reward-weighted loss (RWL, the "shortcut"): a direct optimization that heavily rewards high-quality generations and provides a mathematical guarantee that the model will learn to produce the desired outputs reliably.
The steering damper and access-restricted models
A key advantage of the steering damper is that it enables precise control even when the base model is a closed or partly accessible system (gray-box or black-box). The damper combines the pretrained model’s outputs with a small corrective signal, observing intermediate reverse means as the image clears and injecting microscopic steering adjustments. This lets engineers customize behavior without modifying the underlying model weights.
Experiments and concrete results
The framework was evaluated on Stable Diffusion v1.4 across three fine-tuning regimes: supervised fine-tuning (SFT), reward-weighted loss (RWL) and PPO. Performance was measured with the standardized Human Preference Score HPS-v2, which assesses how well generated images match prompts and aesthetic preferences.
Four network variants were implemented:
- Diffusion Controller: a gray-box steering damper that uses the intermediate reverse mean and a proposed side adapter stream as inputs.
- Diffusion Controller-Naive: a simpler design without the intermediate reverse mean or side adapter stream.
- Diffusion Controller-J: a white-box approach jointly training the steering damper and base model.
- Diffusion Controller-S: another white-box setup training the steering damper and base model separately.
Key findings:
- Diffusion Controller outperformed corresponding baselines in both white-box and gray-box settings, yielding a better quality-efficiency trade-off.
- In the SFT and RWL tracks the gray-box Diffusion Controller achieved higher HPS-v2 win rates than LoRA, a widely used parameter-efficient white-box method, despite manipulating significantly fewer internal layers.
- Human evaluation panels rated Diffusion Controller highest for subjective image quality and prompt matching on complex, multi-attribute prompts.
- The fully unlocked white-box fine-tuned Diffusion Controller reached a 90% win rate over the baseline.
Runtime flexibility
At inference time a single guidance strength parameter controls the intensity of the steering. Adjusting that scalar lets users gradually increase or decrease prompt alignment without breaking baseline stability or introducing the visual distortions that older guidance methods can cause.
Future directions
Because the control layer is separated from the base model, Diffusion Controller opens opportunities beyond prompt matching: personalized generation tools, stronger safety mechanisms to reduce harmful outputs, and extensions of the steering damper concept to control complex future video models.
Acknowledgements
The authors thank collaborators at Google Research, Google DeepMind and academic partners for their contributions.
Conclusion
Diffusion Controller offers a mathematically grounded, lightweight control layer for image generation that improves prompt alignment while preserving image quality and model stability. It works both on access-restricted models and in fully open training setups, providing a unified alternative to the previous patchwork of guidance and fine-tuning techniques.



