Seyedmorteza Sadat

dblp:359/6070 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0003-4668-5703ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 91% Probabilistic and Bayesian machine learning · 7% Efficient and distributed learning · 2%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 87% Image and video processing · 13%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model › guided diffusion
classifier-free guidance
3.442025
Token Perturbation Guidance for Diffusion Models · NeurIPS 2025
No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion Models · ICLR 2025
Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models · ICLR 2025
Machine learning › Generative modeling
diffusion model
3.442025
Token Perturbation Guidance for Diffusion Models · NeurIPS 2025
No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion Models · ICLR 2025
Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models · ICLR 2025
Machine learning › Generative modeling › diffusion model
conditional sampling
1.622025
No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion Models · ICLR 2025
CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed Sampling · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning
sampling
0.912025
Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models · ICLR 2025
Machine learning › Generative modeling › diffusion model › guided diffusion
training-free guidance
0.912025
Token Perturbation Guidance for Diffusion Models · NeurIPS 2025
Visual content generation and editing › image generation
diffusion-based image generation
0.912025
HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling · SIGGRAPH Asia 2025
Visual content generation and editing › image generation
high-resolution image synthesis
0.912025
HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling · SIGGRAPH Asia 2025
Visual content generation and editing
image generation
0.912025
HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling · SIGGRAPH Asia 2025
Visual content generation and editing › image generation
training-free generation
0.912025
HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling · SIGGRAPH Asia 2025
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.812024
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models · NeurIPS 2024
Machine learning › Generative modeling
image generation
0.522025
Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models · ICLR 2025
CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed Sampling · ICLR 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312025
Token Perturbation Guidance for Diffusion Models · NeurIPS 2025
Machine learning › Generative modeling › generative model
unconditional generation
0.312025
No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion Models · ICLR 2025
Image and video processing › image enhancement
detail enhancement
0.312025
HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling · SIGGRAPH Asia 2025
Image and video processing
image enhancement
0.312025
HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling · SIGGRAPH Asia 2025
Machine learning › Efficient and distributed learning
model compression
0.212024
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

wavelet transform · 0.9token perturbation · 0.9timestep guidance · 0.9norm-preserving shuffling · 0.9momentum · 0.9independent condition guidance · 0.9gradient ascent · 0.9diffusion model · 0.9adaptive projected guidance · 0.9DDIM inversion · 0.9gaussian noise annealing · 0.8discrete wavelet transform · 0.8condition-annealed sampling · 0.8
YearPublicationVenuePosition
2025 Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models
abstract
Classifier-free guidance (CFG) is crucial for improving both generation quality and alignment between the input condition and final output in diffusion models. While a high guidance scale is generally required to enhance these aspects, it also causes oversaturation and unrealistic artifacts. In this paper, we revisit the CFG update rule and introduce modifications to address this issue. We first decompose the update term in CFG into parallel and orthogonal components with respect to the conditional model prediction and observe that the parallel component primarily causes oversaturation, while the orthogonal component enhances image quality. Accordingly, we propose down-weighting the parallel component to achieve high-quality generations without oversaturation. Additionally, we draw a connection between CFG and gradient ascent and introduce a new rescaling and momentum method for the CFG update rule based on this insight. Our approach, termed adaptive projected guidance (APG), retains the quality-boosting advantages of CFG while enabling the use of higher guidance scales without oversaturation. APG is easy to implement and introduces practically no additional computational overhead to the sampling process. Through extensive experiments, we demonstrate that APG is compatible with various conditional diffusion models and samplers, leading to improved FID, recall, and saturation scores while maintaining precision comparable to CFG, making our method a superior plug-and-play alternative to standard classifier-free guidance.
Seyedmorteza Sadat, Otmar Hilliges, Romann M. Weber
ICLR1
2025 No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion Models
abstract
Classifier-free guidance (CFG) has become the standard method for enhancing the quality of conditional diffusion models. However, employing CFG requires either training an unconditional model alongside the main diffusion model or modifying the training procedure by periodically inserting a null condition. There is also no clear extension of CFG to unconditional models. In this paper, we revisit the core principles of CFG and introduce a new method, independent condition guidance (ICG), which provides the benefits of CFG without the need for any special training procedures. Our approach streamlines the training process of conditional diffusion models and can also be applied during inference on any pre-trained conditional model. Additionally, by leveraging the time-step information encoded in all diffusion networks, we propose an extension of CFG, called time-step guidance (TSG), which can be applied to *any* diffusion model, including unconditional ones. Our guidance techniques are easy to implement and have the same sampling cost as CFG. Through extensive experiments, we demonstrate that ICG matches the performance of standard CFG across various conditional diffusion models. Moreover, we show that TSG improves generation quality in a manner similar to CFG, without relying on any conditional information.
Seyedmorteza Sadat, Manuel Kansy, Otmar Hilliges, Romann M. Weber
ICLR1
2025 Token Perturbation Guidance for Diffusion Models
abstract
Classifier-free guidance (CFG) has become an essential component of modern diffusion models to enhance both generation quality and alignment with input conditions. However, CFG requires specific training procedures and is limited to conditional generation. To address these limitations, we propose Token Perturbation Guidance (TPG), a novel method that applies perturbation matrices directly to intermediate token representations within the diffusion network. TPG employs a norm-preserving shuffling operation to provide effective and stable guidance signals that improve generation quality without architectural changes. As a result, TPG is training-free and agnostic to input conditions, making it readily applicable to both conditional and unconditional generation. We also analyze the guidance term provided by TPG and show that its effect on sampling more closely resembles CFG compared to existing training-free guidance techniques. We extensively evaluate TPG on SDXL and Stable Diffusion 2.1, demonstrating nearly a 2x improvement in FID for unconditional generation over the SDXL baseline and showing that TPG closely matches CFG in prompt alignment. Thus, TPG represents a general, condition-agnostic guidance method that extends CFG-like benefits to a broader class of diffusion models.
Javad Rajabi, Soroush Mehraban, Seyedmorteza Sadat, Babak Taati
NeurIPS3
2025 HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling
abstract
Diffusion models have emerged as the leading approach for image synthesis, demonstrating exceptional photorealism and diversity. However, training diffusion models at high resolutions remains computationally prohibitive, and existing zero-shot generation techniques for synthesizing images beyond training resolutions often produce artifacts, including object duplication and spatial incoherence. In this paper, we introduce HiWave, a training-free, zero-shot approach that substantially enhances visual fidelity and structural coherence in ultra-high-resolution image synthesis using pretrained diffusion models. Our method employs a two-stage pipeline: generating a base image from the pretrained model followed by a patch-wise DDIM inversion step and a novel wavelet-based detail enhancer module. Specifically, we first utilize inversion methods to derive initial noise vectors that preserve global coherence from the base image. Subsequently, during sampling, our wavelet-domain detail enhancer retains low-frequency components from the base image to ensure structural consistency, while selectively guiding high-frequency components to enrich fine details and textures. Extensive evaluations using Stable Diffusion XL demonstrate that HiWave effectively mitigates common visual artifacts seen in prior methods, achieving superior perceptual quality. A user study confirmed HiWave’s performance, where it was preferred over the state-of-the-art alternative in more than 80% of comparisons, highlighting its effectiveness for high-quality, ultra-high-resolution image synthesis without requiring retraining or architectural modifications.
Tobias Vontobel, Seyedmorteza Sadat, Farnood Salehi, Romann M. Weber
SIGGRAPH Asia2
2024 CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed Sampling
abstract
While conditional diffusion models are known to have good coverage of the data distribution, they still face limitations in output diversity, particularly when sampled with a high classifier-free guidance scale for optimal image quality or when trained on small datasets. We attribute this problem to the role of the conditioning signal in inference and offer an improved sampling strategy for diffusion models that can increase generation diversity, especially at high guidance scales, with minimal loss of sample quality. Our sampling strategy anneals the conditioning signal by adding scheduled, monotonically decreasing Gaussian noise to the conditioning vector during inference to balance diversity and condition alignment. Our Condition-Annealed Diffusion Sampler (CADS) can be used with any pretrained model and sampling algorithm, and we show that it boosts the diversity of diffusion models in various conditional generation tasks. Further, using an existing pretrained diffusion model, CADS achieves a new state-of-the-art FID of 1.70 and 2.31 for class-conditional ImageNet generation at 256$\times$256 and 512$\times$512 respectively.
Seyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges, Romann M. Weber
ICLR1
2024 Factorized Motion Diffusion for Precise and Character-Agnostic Motion Inbetweening
abstract
Animation is a challenging and time-consuming process where animators must manipulate hundreds of controls over space and time to create compelling motions. Recent advances in motion diffusion models have shown impressive results for general motion generation and hold the potential to reduce the number of controls manipulated by animators to achieve high quality results. However, these models are limited by their inability to match sparse constraints precisely, preventing frame-level joint control required by artists. Additionally, recent models are trained for specific characters, preventing reuse, and are incompatible for characters with only a small datasets available. To tackle these shortcomings, we propose a novel factorization of motion between a character-agnostic Bézier Motion Model (BMM), which can be trained on a large motion dataset, followed by a character-specific posing model, trainable on a much smaller pose dataset, that enables reuse across many characters. BMM provides accuracy for meeting sparse joint-level constraints by working in a reduced space of Bézier curves that better aligns the condition signal with the prediction space of our model. Additionally, the Bézier curves offer animators an intuitive interface compatible with existing authoring software. Through quantitative and qualitative comparisons, we show the effectiveness of our factorization and parametric subspace, enabling user control with higher fidelity.
Justin Studer, Dhruv Agrawal, Dominik Borer, Seyedmorteza Sadat, Robert W. Sumner, Martin Guay, Jakob Buhmann
MIG4
2024 LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
abstract
Advances in latent diffusion models (LDMs) have revolutionized high-resolution image generation, but the design space of the autoencoder that is central to these systems remains underexplored. In this paper, we introduce LiteVAE, a new autoencoder design for LDMs, which leverages the 2D discrete wavelet transform to enhance scalability and computational efficiency over standard variational autoencoders (VAEs) with no sacrifice in output quality. We investigate the training methodologies and the decoder architecture of LiteVAE and propose several enhancements that improve the training dynamics and reconstruction quality. Our base LiteVAE model matches the quality of the established VAEs in current LDMs with a six-fold reduction in encoder parameters, leading to faster training and lower GPU memory requirements, while our larger model outperforms VAEs of comparable complexity across all evaluated metrics (rFID, LPIPS, PSNR, and SSIM).
Seyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges, Romann M. Weber
NeurIPS1