VLDB 2026 Research / reviewers in the wild / expert
Han Yan 0003
dblp:63/49-3
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Learning to Disentangle the Colors, Textures, and Shapes of Fashion Items: A Unified FrameworkabstractToday,fashion design can be readily performed by most people due to the rapid development of design tools. However, not everyone possesses the professional skills to produce an aesthetically pleasing design. In order to assist an inexperienced user during the design process, this research explored a new fashion-related disentanglement task, with the goal of creating novel fashion items with controllable attributes. The key idea is to develop a unified framework, called CTS-GAN, by disentangling the colors, textures, and shapes of fashion items simultaneously using a generative adversarial network (GAN). Specifically, we first introduced a fashion attribute encoder to decompose input fashion items into three latent spaces, i.e., color, texture, and shape. A fashion item pattern-making module (FIPM)-based generator was then proposed to control the corresponding parameters of color and texture in FIPMs independently and combine them with the shape features in order to accomplish the final generation of new fashion items. Furthermore, three independent pathways were introduced to extract the representations of color, texture, and shape in fashion items to optimize our CTS-GAN in an unsupervised manner. Extensive experimental results demonstrate the effectiveness of our CTS-GAN and suggest that it can generate diverse, novel fashion images by taking full advantage of the controllability of the colors, textures, and shapes of different fashion items. Han Yan 0003, Haijun Zhang 0002, Zhao Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Learning to Synthesize Compatible Fashion Items Using Semantic Alignment and Collocation Classification: An Outfit Generation FrameworkabstractThe field of fashion compatibility learning has attracted great attention from both the academic and industrial communities in recent years. Many studies have been carried out for fashion compatibility prediction, collocated outfit recommendation, artificial intelligence (AI)-enabled compatible fashion design, and related topics. In particular, AI-enabled compatible fashion design can be used to synthesize compatible fashion items or outfits to improve the design experience for designers or the efficacy of recommendations for customers. However, previous generative models for collocated fashion synthesis have generally focused on the image-to-image translation between fashion items of upper and lower clothing. In this article, we propose a novel outfit generation framework, i.e., OutfitGAN, with the aim of synthesizing a set of complementary items to compose an entire outfit, given one extant fashion item and reference masks of target synthesized items. OutfitGAN includes a semantic alignment module (SAM), which is responsible for characterizing the mapping correspondence between the existing fashion items and the synthesized ones, to improve the quality of the synthesized images, and a collocation classification module (CCM), which is used to improve the compatibility of a synthesized outfit. To evaluate the performance of our proposed models, we built a large-scale dataset consisting of 20 000 fashion outfits. Extensive experimental results on this dataset show that our OutfitGAN can synthesize photo-realistic outfits and outperform the state-of-the-art methods in terms of similarity, authenticity, and compatibility measurements. Dongliang Zhou, Haijun Zhang 0002, Kai Yang 0018, Han Yan 0003, Xiaofei Xu 0001, Zhao Zhang 0001, Shuicheng Yan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | FashionDiff: A Controllable Diffusion Model Using Pairwise Fashion Elements for Intelligent DesignabstractThe process of fashion design involves creative expression through various methods, including sketch drawing, brush painting, and choices of textures and colors, all of which are employed to characterize the originality and uniqueness of the designed fashion items. Despite recent advances in intelligence-driven fashion design, the complexity of the diverse elements of a fashion item, such as its texture, color and shape, which are associated with the semantic information conveyed, continues to present challenges in terms of generating high-quality fashion images as well as achieving a controllable editing process. To address this issue, we propose a unified framework, FashionDiff, that leverages the diverse elements in fashion items to generate new items. Initially, we collected a large number of fashion images with multiple categories and created pairwise data in terms of sketch and additional data, such as brush areas, textures, or colors. To eliminate semantic discrepancies between these pairwise datasets, we introduce a feature modulation fusion (FMFusion) process, which enables interactive communication among different images, allowing them to be fused into latent spaces characterized by different resolutions. In order to produce high-quality editable fashion images, we develop a generator based on a state-of-the-art diffusion model called FD-ControlNet, which integrates latent spaces into different layers of the generator to generate ready-to-wear fashion items. Qualitative and quantitative experimental results demonstrate the effectiveness of our proposed method, and suggest that our model can offer flexible control over the generated images in terms of sketches, brush areas, textures, and colors. Han Yan 0003, Haijun Zhang 0002, Xiangyu Mu, Jicong Fan 0001, Zhao Zhang 0001 |
ACM Multimedia | 1 |
| 2023 | InspirNET: An Unsupervised Generative Adversarial Network with Controllable Fine-grained Texture Disentanglement for Fashion GenerationabstractTexture constitutes the color and fabric of fashion items. Its choice in fashion items can directly express the personality and emotional state of a wearer. Despite the rapid development of intelligence-driven fashion design, it remains challenging to achieve independent control over texture without affecting other attributes, due to the highly intertwined nature of texture space. To accomplish fine-grained texture disentanglement, we propose InspirNET, an unsupervised disentangled generative adversarial framework, that manipulates textures in a fine-grained latent space so as to produce new textures effectively, aiming to broaden the range of fashion options available to common users with distinct textures as well as boosting designers' potential for fashion innovation and inspiration. Specifically, we first introduce an auto-fashion attribute encoder to map the input fashion item into texture and structure spaces. To achieve unsupervised fine-grained texture disentanglement, our model proposes a K-textures disentanglement module that decomposes the texture space into several orthogonal vectors, each of which is empowered to control an independent texture element. In particular, by employing an orthogonal eigenvector to interpolate with another, a multitude of new textures can be generated easily. Qualitative and quantitative experiments demonstrate that our InspirNET can effectively utilize decomposed orthogonal vectors to generate a wide range of fashion items with diverse textures. Our model exhibits superior performance over state-of-the-art methods in terms of maintaining the authenticity of texture transfer. Han Yan 0003, Haijun Zhang 0002, Jicong Fan 0001, Zhao Zhang 0001 |
ACM Multimedia | 1 |
| 2023 | Texture Brush for Fashion Inspiration Transfer: A Generative Adversarial Network With Heatmap-Guided Semantic DisentanglementabstractAutomatically accomplishing intelligent fashion design with certain ‘inspiration’ images can greatly facilitate a designer’s design process, as well as allow users to interactively participate in the process. In this research, we propose a generative adversarial network with heatmap-guided semantic disentanglement (HSD-GAN) to perform an ‘intelligent’ design with ‘inspiration’ transfer. Our model aims to learn how to integrate the feature representations, from the styles of both source fashion items and target fashion items, in an unsupervised manner. Specifically, a semantic disentanglement attention-based encoder is proposed to capture the most discriminative regions of different input fashion items and disentangle the features into two key factors: attribute and texture. A generator is then developed to synthesize mixed-style fashion items by utilizing the two factors. In addition, a heatmap-based patch loss is introduced to evaluate the visual-semantic matching degree between the texture of the generated fashion items and the input texture information. Extensive experimental results show that our proposed HSD-GAN consistently achieves superior performance, compared to other state-of-the-art methods. Han Yan 0003, Haijun Zhang 0002, Jianyang Shi, Jianghong Ma |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Toward Intelligent Design: An AI-Based Fashion Designer Using Generative Adversarial Networks Aided by Sketch and Rendering GeneratorsabstractThe traditional fashion industry is heavily dependent on designers whose talent and vision have a significant impact on their innovative designs. Through taking advantage of recent advances in image-to-image translation by generative adversarial networks (GANs), marked improvement in designers’ efficiency is now possible. Considering both randomness and controllability in the design process, this article presents a novel artificial intelligence (AI)-based framework for fashion design. Under this framework, a sketch-generation module which is based on latent space is firstly introduced for designing various sketches. Secondly, a rendering-generation module is proposed to learn mapping between textures and sketches to complete the task of fashion design. In order to achieve effectiveness in synthesizing semantic-aware textures on sketches, a multi-conditional feature interaction module is developed in the rendering-generation model. Moreover, two different training schemes are introduced to optimize both the sketch-generation module and the rendering-generation module. In order to evaluate the performance of our proposed models, we built a large-scale dataset which consists of 115,584 pairs of fashion item images. Experimental results demonstrate the effectiveness of our proposed method, and indicate that our model can facilitate designers’ design process by taking full advantage of the controllability of different conditions (e.g., sketch and texture) and the randomness of latent space. Han Yan 0003, Haijun Zhang 0002, Dongliang Zhou, Xiaofei Xu 0001, Zhao Zhang 0001, Shuicheng Yan |
IEEE Trans. Multim. | 1 |
| 2023 | Toward Intelligent Fashion Design: A Texture and Shape Disentangled Generative Adversarial NetworkabstractTexture and shape in fashion, constituting essential elements of garments, characterize the body and surface of the fabric and outline the silhouette of clothing, respectively. The selection of texture and shape plays a critical role in the design process, as they largely determine the success of a new design for fashion items. In this research, we propose a texture and shape disentangled generative adversarial network (TSD-GAN) to perform “intelligent” design with the transformation of texture and shape in fashion items. Our TSD-GAN aims to learn how to disentangle the features of texture and shape of different fashion items in an unsupervised manner. Specifically, a fashion attribute encoder is developed to decompose the input fashion items into independent representations of texture and shape. Then, to learn the coarse or fine styles hidden in the features of texture and shape, a texture mapping network and a shape mapping network are proposed to disentangle the features into different hierarchical representations. The different hierarchical representations of texture and shape are then fed into a multi-factor-based generator to generate mixed-style fashion items. In addition, a multi-discriminator framework is developed to distinguish the authenticity and texture similarity between the generated images and the real images. Experimental results on different fashion categories demonstrate that our proposed TSD-GAN may be useful for assisting designers to accomplish the design process by transforming the texture and shape of fashion items. Han Yan 0003, Haijun Zhang 0002, Jianyang Shi, Jianghong Ma, Xiaofei Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | KPN-MFI: A Kernel Prediction Network with Multi-frame Interaction for Video Inverse Tone MappingabstractUp to now, the image-based inverse tone mapping (iTM) models have been widely investigated, while there is little research on video-based iTM methods. It would be interesting to make use of these existing image-based models in the video iTM task. However, directly transferring the imagebased iTM models to video data without modeling spatial-temporal information remains nontrivial and challenging. Considering both the intra-frame quality and the inter-frame consistency of a video, this article presents a new video iTM method based on a kernel prediction network (KPN), which takes advantage of multi-frame interaction (MFI) module to capture temporal-spatial information for video data. Specifically, a basic encoder-decoder KPN, essentially designed for image iTM, is trained to guarantee the mapping quality within each frame. More importantly, the MFI module is incorporated to capture temporal-spatial context information and preserve the inter-frame consistency by exploiting the correction between adjacent frames. Notably, we can readily extend any existing image iTM models to video iTM ones by involving the proposed MFI module. Furthermore, we propose an inter-frame brightness consistency loss function based on the Gaussian pyramid to reduce the video temporal inconsistency. Extensive experiments demonstrate that our model outperforms state-ofthe-art image and video-based methods. The code is available at https://github.com/caogaofeng/KPNMFI. Gaofeng Cao, Fei Zhou 0001, Han Yan 0003, Anjie Wang, Leidong Fan |
IJCAI | 3 |