Yusuf Dalva

dblp:301/8337 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-8402-8291ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Generative modeling · 57% 3D vision · 18% Language models and text generation · 10%
Computer graphics and multimedia
5 papers
Visual content generation and editing · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.622025
LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers · NeurIPS 2025
NoiseCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions in Diffusion Models · CVPR 2024
Visual content generation and editing
image editing
1.522025
FluxSpace: Disentangled Semantic Editing in Rectified Flow Models · CVPR 2025
StyleRes: Transforming the Residuals for Real Image Editing with StyleGAN · CVPR 2023
Visual content generation and editing
image-to-image translation
1.222023
Image-to-Image Translation With Disentangled Latent Vectors for Face Editing · IEEE Trans. Pattern Anal. Mach. Intell. 2023
VecGAN: Image-to-Image Translation with Interpretable Latent Directions · ECCV (16) 2022
Machine learning › Generative modeling › flow matching
rectified flow model
0.912025
FluxSpace: Disentangled Semantic Editing in Rectified Flow Models · CVPR 2025
Natural language and speech › Language models and text generation › knowledge editing
representation editing
0.912025
FluxSpace: Disentangled Semantic Editing in Rectified Flow Models · CVPR 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers · NeurIPS 2025
Computer vision › 3D vision
3d human reconstruction
0.812024
Refining 3D Human Texture Estimation From a Single Image · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Representation and self-supervised learning › latent space
latent space manipulation
0.812024
NoiseCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions in Diffusion Models · CVPR 2024
Computer vision › 3D vision
UV mapping
0.812024
Refining 3D Human Texture Estimation From a Single Image · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Generative modeling › generative adversarial network
GAN inversion
0.712023
StyleRes: Transforming the Residuals for Real Image Editing with StyleGAN · CVPR 2023
Machine learning › Generative modeling › diffusion model › diffusion inversion
image inversion
0.712023
StyleRes: Transforming the Residuals for Real Image Editing with StyleGAN · CVPR 2023
Visual content generation and editing
face editing
0.712023
Image-to-Image Translation With Disentangled Latent Vectors for Face Editing · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Visual content generation and editing › image editing
GAN-based image editing
0.712023
StyleRes: Transforming the Residuals for Real Image Editing with StyleGAN · CVPR 2023
Machine learning › Trustworthy machine learning
interpretability
0.612022
VecGAN: Image-to-Image Translation with Interpretable Latent Directions · ECCV (16) 2022
Visual content generation and editing › image editing › semantic image editing
attribute editing
0.212023
StyleRes: Transforming the Residuals for Real Image Editing with StyleGAN · CVPR 2023

Methods — techniques the papers use, named apart from their topics

rectified flow transformer · 3.5latent mask blending · 1.7LoRA · 1.7cycle-consistency loss · 1.4residual feature learning · 1.3disentanglement loss · 1.3StyleGAN · 1.3uncertainty-based reconstruction loss · 0.8deformable convolution · 0.8contrastive learning · 0.8orthogonality constraint · 0.7latent space factorization · 0.7cycle consistency loss · 0.7generative adversarial network · 0.6
YearPublicationVenuePosition
2025 FluxSpace: Disentangled Semantic Editing in Rectified Flow Models
abstract
Rectified flow models have emerged as a dominant approach in image generation, showcasing impressive capabilities in high-quality image synthesis. However, despite their effectiveness in visual generation, rectified flow models often struggle with disentangled editing of images. This limitation prevents the ability to perform precise, attribute-specific modifications without affecting unrelated aspects of the image. In this paper, we introduce FluxSpace, a domain-agnostic image editing method leveraging a representation space with the ability to control the semantics of images generated by rectified flow transformers, such as Flux. By leveraging the representations learned by the transformer blocks within the rectified flow models, we propose a set of semantically interpretable representations that enable a wide range of image editing tasks, from fine-grained image editing to artistic creation. This work offers a scalable and effective image editing approach, along with its disentanglement capabilities.
Yusuf Dalva, Kavana Venkatesh, Pinar Yanardag Delul
CVPR1
2025 LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers
abstract
We introduce LoRAShop, the first framework for multi-concept image generation and editing with LoRA models. LoRAShop builds on a key observation about the feature interaction patterns inside Flux-style diffusion transformers: concept-specific transformer features activate spatially coherent regions early in the denoising process. We harness this observation to derive a disentangled latent mask for each concept in a prior forward pass and blend the corresponding LoRA weights only within regions bounding the concepts to be personalized. The resulting edits seamlessly integrate multiple subjects or styles into the original scene while preserving global context, lighting, and fine details. Our experiments demonstrate that LoRAShop delivers better identity preservation compared to baselines. By eliminating retraining and external constraints, LoRAShop turns personalized diffusion models into a practical `photoshop-with-LoRAs' tool and opens new avenues for compositional visual storytelling and rapid creative iteration.
Yusuf Dalva, Hidir Yesiltepe, Pinar Yanardag Delul
NeurIPS1
2024 NoiseCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions in Diffusion Models
abstract
Generative models have been very popular in the recent years for their image generation capabilities. GAN-based models are highly regarded for their disentangled latent space, which is a key feature contributing to their success in controlled image editing. On the other hand, diffusion models have emerged as powerful tools for generating high-quality images. However, the latent space of diffusion models is not as thoroughly explored or understood. Existing methods that aim to explore the latent space of diffusion models usually relies on text prompts to pinpoint specific semantics. However, this approach may be restrictive in ar-eas such as art, fashion, or specialized fields like medicine, where suitable text prompts might not be available or easy to conceive thus limiting the scope of existing work. In this paper, we propose an unsupervised method to discover la-tent semantics in text-to-image diffusion models without relying on text prompts. Our method takes a small set of un la-beled images from specific domains, such as faces or cats, and a pre-trained diffusion model, and discovers diverse se-mantics in unsupervised fashion using a contrastive learning objective. Moreover, the learned directions can be ap-plied simultaneously, either within the same domain (such as various types of facial edits) or across different domains (such as applying cat and face edits within the same image) without interfering with each other. Our extensive experi-ments show that our method achieves highly disentangled edits, outperforming existing approaches in both diffusion-based and GAN-based latent space editing methods.
Yusuf Dalva, Pinar Yanardag Delul
CVPR1
2024 Refining 3D Human Texture Estimation From a Single Image
abstract
Estimating 3D human texture from a single image is essential in graphics and vision. It requires learning a mapping function from input images of humans with diverse poses into the parametric (uv) space and reasonably hallucinating invisible parts. To achieve a high-quality 3D human texture estimation, we propose a framework that adaptively samples the input by a deformable convolution where offsets are learned via a deep neural network. Additionally, we describe a novel cycle consistency loss that improves view generalization. We further propose to train our framework with an uncertainty-based pixel-level image reconstruction loss, which enhances color fidelity. We compare our method against the state-of-the-art approaches and show significant qualitative and quantitative improvements.
Said Fahri Altindis, Adil Meric, Yusuf Dalva, Ugur Güdükbay, Aysegul Dundar
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Benchmarking the Robustness of Instance Segmentation Models
abstract
This article presents a comprehensive evaluation of instance segmentation models with respect to real-world image corruptions as well as out-of-domain image collections, e.g., images captured by a different set-up than the training dataset. The out-of-domain image evaluation shows the generalization capability of models, an essential aspect of real-world applications, and an extensively studied topic of domain adaptation. These presented robustness and generalization evaluations are important when designing instance segmentation models for real-world applications and picking an off-the-shelf pretrained model to directly use for the task at hand. Specifically, this benchmark study includes state-of-the-art network architectures, network backbones, normalization layers, models trained starting from scratch versus pretrained networks, and the effect of multitask training on robustness and generalization. Through this study, we gain several insights. For example, we find that group normalization (GN) enhances the robustness of networks across corruptions where the image contents stay the same but corruptions are added on top. On the other hand, batch normalization (BN) improves the generalization of the models across different datasets where statistics of image features change. We also find that single-stage detectors do not generalize well to larger image resolutions than their training size. On the other hand, multistage detectors can easily be used on images of different sizes. We hope that our comprehensive study will motivate the development of more robust and reliable instance segmentation models.
Yusuf Dalva, Hamza Pehlivan, Said Fahri Altindis, Aysegul Dundar
IEEE Trans. Neural Networks Learn. Syst.1
2023 StyleRes: Transforming the Residuals for Real Image Editing with StyleGAN
abstract
We present a novel image inversion framework and a training pipeline to achieve high-fidelity image inversion with high-quality attribute editing. Inverting real images into StyleGAN's latent space is an extensively studied problem, yet the trade-off between the image reconstruction fidelity and image editing quality remains an open challenge. The low-rate latent spaces are limited in their expressiveness power for high-fidelity reconstruction. On the other hand, high-rate latent spaces result in degradation in editing quality. In this work, to achieve high-fidelity inversion, we learn residual features in higher latent codes that lower latent codes were not able to encode. This enables preserving image details in reconstruction. To achieve high-quality editing, we learn how to transform the residual features for adapting to manipulations in latent codes. We train the framework to extract residual features and transform them via a novel architecture pipeline and cycle consistency losses. We run extensive experiments and compare our method with state-of-the-art inversion methods. Qualitative metrics and visual comparisons show significant improvements. Code: https://github.com/hamzapehlivanIStyleRes.
Hamza Pehlivan, Yusuf Dalva, Aysegul Dundar
CVPR2
2023 Image-to-Image Translation With Disentangled Latent Vectors for Face Editing
abstract
We propose an image-to-image translation framework for facial attribute editing with disentangled interpretable latent directions. Facial attribute editing task faces the challenges of targeted attribute editing with controllable strength and disentanglement in the representations of attributes to preserve the other attributes during edits. For this goal, inspired by the latent space factorization works of fixed pretrained GANs, we design the attribute editing by latent space factorization, and for each attribute, we learn a linear direction that is orthogonal to the others. We train these directions with orthogonality constraints and disentanglement losses. To project images to semantically organized latent spaces, we set an encoder-decoder architecture with attention-based skip connections. We extensively compare with previous image translation algorithms and editing with pretrained GAN works. Our extensive experiments show that our method significantly improves over the state-of-the-arts.
Yusuf Dalva, Hamza Pehlivan, Öykü Irmak Hatipoglu, Cansu Moran, Aysegul Dundar
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 VecGAN: Image-to-Image Translation with Interpretable Latent Directions
Yusuf Dalva, Said Fahri Altindis, Aysegul Dundar
ECCV (16)1