VLDB 2026 Research / reviewers in the wild / expert
Dongliang Zhou
dblp:178/4432
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech RecognitionabstractVisual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have significantly enhanced the performance of lip reading models. Despite these advancements, existing datasets predominantly feature stable video recordings with limited variability in lip movements. This limitation results in models that are highly sensitive to variations encountered in real-world scenarios. To address this issue, we propose a novel framework, LipGen, which aims to improve model robustness by leveraging speech-driven synthetic visual data, thereby mitigating the constraints of current datasets. Additionally, we introduce an auxiliary task that incorporates viseme classification alongside attention mechanisms. This approach facilitates the efficient integration of temporal information, directing the model’s focus toward the relevant segments of speech, thereby enhancing discriminative capabilities. Our method demonstrates superior performance compared to the current state-of-the-art on the lip reading in the wild (LRW) dataset and exhibits even more pronounced advantages under challenging conditions. Dongliang Zhou, Liang Xie 0012, Jianlong Wu, Erwei Yin |
ICASSP | 2 |
| 2025 | PanoGen++: Domain-adapted text-guided panoramic environment generation for vision-and-language navigation
Dongliang Zhou, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
Neural Networks | 2 |
| 2025 | AVE Speech: A Comprehensive Multimodal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic SignalsabstractThe global aging population faces considerable challenges, particularly in communication, due to the prevalence of hearing and speech impairments. To address these, we introduce the AVE speech, a comprehensive multimodal dataset for speech recognition tasks. The dataset includes a 100-sentence Mandarin corpus with audio signals, lip-region video recordings, and six-channel electromyography data, collected from 100 participants. Each subject read the entire corpus ten times, with each sentence averaging approximately two seconds in duration, resulting in over 55 hours of multimodal speech data per modality. Experiments demonstrate that combining these modalities significantly improves recognition performance, particularly in cross-subject and high-noise environments. To our knowledge, this is the first publicly available sentence-level dataset integrating these three modalities for large-scale Mandarin speech recognition. We expect this dataset to drive advancements in both acoustic and nonacoustic speech recognition research, enhancing cross-modal learning and human–machine interaction. Dongliang Zhou, Yakun Zhang 0002, Jinghan Wu, Liang Xie 0012, Erwei Yin |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2024 | Towards Intelligent Design: A Self-Driven Framework for Collocated Clothing Synthesis Leveraging Fashion Styles and TexturesabstractCollocated clothing synthesis (CCS) has emerged as a pivotal topic in fashion technology, primarily concerned with the generation of a clothing item that harmoniously matches a given item. However, previous investigations have relied on using paired outfits, such as a pair of matching upper and lower clothing, to train a generative model for achieving this task. This reliance on the expertise of fashion professionals in the construction of such paired outfits has engendered a laborious and time-intensive process. In this paper, we introduce a new self-driven framework, named style- and texture-guided generative network (ST-Net), to synthesize collocated clothing without the necessity for paired outfits, leveraging self-supervised learning. ST-Net is designed to extrapolate fashion compatibility rules from the style and texture attributes of clothing, using a generative adversarial network. To facilitate the training and evaluation of our model, we have constructed a large-scale dataset specifically tailored for unsupervised CCS. Extensive experiments substantiate that our proposed method outperforms the state-of-the-art baselines in terms of both visual authenticity and fashion compatibility. Minglong Dong, Dongliang Zhou, Jianghong Ma, Haijun Zhang 0002 |
ICASSP | 2 |
| 2024 | BC-GAN: A Generative Adversarial Network for Synthesizing a Batch of Collocated ClothingabstractCollocated clothing synthesis using generative networks has become an emerging topic in the field of fashion intelligence, as it has significant potential economic value to increase revenue in the fashion industry. In previous studies, several works have attempted to synthesize visually-collocated clothing based on a given clothing item using generative adversarial networks (GANs) with promising results. These works, however, can only accomplish the synthesis of one collocated clothing item each time. Nevertheless, users may require different clothing items to meet their multiple choices due to their personal tastes and different dressing scenarios. To address this limitation, we introduce a novel batch clothing generation framework, named BC-GAN, which is able to synthesize multiple visually-collocated clothing images simultaneously. In particular, to further improve the fashion compatibility of synthetic results, BC-GAN proposes a new fashion compatibility discriminator in a contrastive learning perspective by fully exploiting the collocation relationship among all clothing items. Our model was examined in a large-scale dataset with compatible outfits constructed by ourselves. Extensive experiment results confirmed the effectiveness of our proposed BC-GAN in comparison to state-of-the-art methods in terms of diversity, visual authenticity, and fashion compatibility. Dongliang Zhou, Haijun Zhang 0002, Jianghong Ma, Jianyang Shi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Learning to Synthesize Compatible Fashion Items Using Semantic Alignment and Collocation Classification: An Outfit Generation FrameworkabstractThe field of fashion compatibility learning has attracted great attention from both the academic and industrial communities in recent years. Many studies have been carried out for fashion compatibility prediction, collocated outfit recommendation, artificial intelligence (AI)-enabled compatible fashion design, and related topics. In particular, AI-enabled compatible fashion design can be used to synthesize compatible fashion items or outfits to improve the design experience for designers or the efficacy of recommendations for customers. However, previous generative models for collocated fashion synthesis have generally focused on the image-to-image translation between fashion items of upper and lower clothing. In this article, we propose a novel outfit generation framework, i.e., OutfitGAN, with the aim of synthesizing a set of complementary items to compose an entire outfit, given one extant fashion item and reference masks of target synthesized items. OutfitGAN includes a semantic alignment module (SAM), which is responsible for characterizing the mapping correspondence between the existing fashion items and the synthesized ones, to improve the quality of the synthesized images, and a collocation classification module (CCM), which is used to improve the compatibility of a synthesized outfit. To evaluate the performance of our proposed models, we built a large-scale dataset consisting of 20 000 fashion outfits. Extensive experimental results on this dataset show that our OutfitGAN can synthesize photo-realistic outfits and outperform the state-of-the-art methods in terms of similarity, authenticity, and compatibility measurements. Dongliang Zhou, Haijun Zhang 0002, Kai Yang 0018, Han Yan 0003, Xiaofei Xu 0001, Zhao Zhang 0001, Shuicheng Yan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Toward Intelligent Interactive Design: A Generation Framework Based on Cross-domain Fashion ElementsabstractTraditional fashion design typically requires the expertise of designers, which limits the involvement of ordinary users during the design process. While it would be desirable for users to participate in the preliminary design phase, their lack of basic design knowledge may render them too inexperienced to produce satisfactory designs. To improve the design efficiency for common users, we present a novel interactive fashion design framework based on generative adversarial network (GAN). This framework can assist users in designing fashion items by drawing only rough scribbles and providing simple fashion styles. Specifically, we propose a new cross-domain feature fusion encoder network that maps design image features from different domains into a series of style vectors which are then fed into a generator. We demonstrate that the learned style vectors can decouple the representations of cross-domain design elements and control the design results through scribbles and style images. Furthermore, we propose a method for rewriting our model with scribbles and style images, to allow designers to train our model more easily. To examine the effectiveness of our proposed model, we constructed a large-scale dataset containing 90,000 pairs of fashion item images. Experimental results show that our proposed method outperforms state-of-the-art methods and can effectively control cross-domain image features, suggesting the potential of our model for providing users with an intelligence-driven interactive design tool. Jianyang Shi, Haijun Zhang 0002, Dongliang Zhou, Zhao Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | FCBoost-Net: A Generative Network for Synthesizing Multiple Collocated Outfits via Fashion Compatibility BoostingabstractOutfit generation is a challenging task in the field of fashion technology, in which the aim is to create a collocated set of fashion items that complement a given set of items. Previous studies in this area have been limited to generating a unique set of fashion items based on a given set of items, without providing additional options to users. This lack of a diverse range of choices necessitates the development of a more versatile framework. However, when the task of generating collocated and diversified outfits is approached with multimodal image-to-image translation methods, it poses a challenging problem in terms of non-aligned image translation, which is hard to address with existing methods. In this research, we present FCBoost-Net, a new framework for outfit generation that leverages the power of pre-trained generative models to produce multiple collocated and diversified outfits. Initially, FCBoost-Net randomly synthesizes multiple sets of fashion items, and the compatibility of the synthesized sets is then improved in several rounds using a novel fashion compatibility booster. This approach was inspired by boosting algorithms and allows the performance to be gradually improved in multiple steps. Empirical evidence indicates that the proposed strategy can improve the fashion compatibility of randomly synthesized fashion items as well as maintain their diversity. Extensive experiments confirm the effectiveness of our proposed framework with respect to visual authenticity, diversity, and fashion compatibility. Dongliang Zhou, Haijun Zhang 0002, Jianghong Ma, Jicong Fan 0001, Zhao Zhang 0001 |
ACM Multimedia | 1 |
| 2023 | IASA: An IoU-aware tracker with adaptive sample assignment
Kai Yang 0018, Haijun Zhang 0002, Dongliang Zhou, Li Dong 0011, Jianghong Ma |
Neural Networks | 3 |
| 2023 | Toward Intelligent Design: An AI-Based Fashion Designer Using Generative Adversarial Networks Aided by Sketch and Rendering GeneratorsabstractThe traditional fashion industry is heavily dependent on designers whose talent and vision have a significant impact on their innovative designs. Through taking advantage of recent advances in image-to-image translation by generative adversarial networks (GANs), marked improvement in designers’ efficiency is now possible. Considering both randomness and controllability in the design process, this article presents a novel artificial intelligence (AI)-based framework for fashion design. Under this framework, a sketch-generation module which is based on latent space is firstly introduced for designing various sketches. Secondly, a rendering-generation module is proposed to learn mapping between textures and sketches to complete the task of fashion design. In order to achieve effectiveness in synthesizing semantic-aware textures on sketches, a multi-conditional feature interaction module is developed in the rendering-generation model. Moreover, two different training schemes are introduced to optimize both the sketch-generation module and the rendering-generation module. In order to evaluate the performance of our proposed models, we built a large-scale dataset which consists of 115,584 pairs of fashion item images. Experimental results demonstrate the effectiveness of our proposed method, and indicate that our model can facilitate designers’ design process by taking full advantage of the controllability of different conditions (e.g., sketch and texture) and the randomness of latent space. Han Yan 0003, Haijun Zhang 0002, Dongliang Zhou, Xiaofei Xu 0001, Zhao Zhang 0001, Shuicheng Yan |
IEEE Trans. Multim. | 4 |
| 2023 | COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion StylesabstractHow to recommend outfits has gained considerable attention in both academia and industry in recent years. Many studies have been carried out regarding fashion compatibility learning, to determine whether the fashion items in an outfit are compatible or not. These methods mainly focus on evaluating the compatibility of existing outfits and rarely consider applying such knowledge to ‘design’ new fashion items. We propose the new task of generating complementary and compatible fashion items based on an arbitrary number of given fashion items. In particular, given some fashion items that can make up an outfit, the aim of this paper is to synthesize photo-realistic images of other, complementary, fashion items that are compatible with the given ones. To achieve this, we propose an outfit generation framework, referred to as COutfitGAN, which includes a pyramid style extractor, an outfit generator, a UNet-based real/fake discriminator, and a collocation discriminator. To train and evaluate this framework, we collected a large-scale fashion outfit dataset with over 200 K outfits and 800 K fashion items from the Internet. Extensive experiments show that COutfitGAN outperforms other baselines in terms of similarity, authenticity, and compatibility measurements. Dongliang Zhou, Haijun Zhang 0002, Qun Li 0011, Jianghong Ma, Xiaofei Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | PaaRPN: Probabilistic anchor assignment with region proposal network for visual tracking
Kai Yang 0018, Haijun Zhang 0002, Dongliang Zhou, Li Dong 0011 |
Inf. Sci. | 3 |
| 2021 | Clothing generation by multi-modal embedding: A compatibility matrix-regularized GAN model
Haijun Zhang 0002, Dongliang Zhou |
Image Vis. Comput. | 3 |
| 2021 | TGAN: A simple model update strategy for visual tracking via template-guidance attention network
Kai Yang 0018, Haijun Zhang 0002, Dongliang Zhou |
Neural Networks | 3 |
| 2016 | OCEAN: Fast Discovery of High Utility Occupancy Itemsets
Bilong Shen, Zhaoduo Wen, Dongliang Zhou |
PAKDD (1) | 4 |