Xingxing Zou

dblp:248/9715 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards intelligent online cross-selling
Kaicheng Pang, Xingxing Zou, Zowie Broach, Wai Keung Wong
Expert Syst. Appl.2
2026 Revisiting kernel complexity in Gaussian process regression: An empirical study on text-to-visual embedding mapping in fashion design
Dongmei Mo, Daniel Augusto R. M. A. de Souza, Xingxing Zou, Wai Keung Wong
Neurocomputing3
2026 LayerDiffusion: High-Fidelity Layered Virtual Try-On via Semantic-Aware Diffusion
abstract
Image-based Virtual Try-On (VTON) aims to generate virtual try-on results by transferring input garment images onto target person images. Current algorithms have achieved high generalization and quality on public datasets including VITON and DressCode. However, current algorithms primarily process images at the pixel level while neglecting the semantic information of garments. Moreover, layered outerwear, as an essential aspect of try-on, has been largely overlooked. Layering increases the difficulty of precise garment control due to the need to consider inner garments. Furthermore, the task-specific nature of current methods prevents them from capturing layering characteristics in try-on. To address these limitations, we propose LayerDiffusion, a novel framework that enhances both generalization capability and semantic controllability for virtual try-on generation. Our approach makes three key contributions. First, we construct LayerDataset, a specialized dataset focusing on diverse outerwear garments with balanced gender representation and multi-view captures. Second, we integrate a pretrained image encoder to capture rich semantic garment information, combined with multi-conditional masking and adapter fine-tuning to enable flexible layering control. Third, we introduce a joint training strategy on both try-on and try-off tasks, which substantially improves model generalization by learning bidirectional garment-body correlations. Extensive experiments demonstrate that LayerDiffusion significantly outperforms existing state-of-the-art methods on public benchmarks (VITON-HD and DressCode) across multiple metrics including LPIPS, SSIM, FID, KID, and CLIP-based scores. Our method excels in both single-garment and layered try-on scenarios while naturally supporting try-off capabilities, offering superior semantic preservation and detail fidelity compared to existing approaches.
Fangjian Liao, Xingxing Zou, Wai Keung Wong
IEEE Trans. Circuits Syst. Video Technol.2
2026 WordCon: Word-Level Typography Control in Visual Text Rendering
abstract
Visual text rendering represents a fundamental capability of large-scale text-to-image (T2I) models, yet achieving precise word-level controllability remains a significant challenge in this domain. While existing approaches primarily focus on text content accuracy, they often fail to provide fine-grained control over typographic attributes at the word level. To address this limitation, we introduce a comprehensive solution comprising three key components: (1) a novel word-level controlled scene text dataset and benchmark, (2) the Text-Image Alignment (TIA) framework that leverages cross-modal correspondence between textual queries and local image regions through grounding models, and (3) WordCon, a hybrid parameter-efficient fine-tuning (PEFT) method that employs selective parameter reparameterization to enhance both computational efficiency and model portability. The proposed framework incorporates additive supervision mechanisms: a masked loss at the latent level to focus on text regions, and a joint-attention loss at the feature level to promote disentanglement between different words. Extensive experimental evaluations demonstrate that our approach outperforms state-of-the-art methods in both qualitative and quantitative metrics. The proposed method exhibits remarkable versatility, enabling seamless integration with diverse pipelines, including artistic text rendering and image-conditioned text generation. Our datasets, source code, and models will be available for academic research.
Wenda Shi, Yiren Song, Zihan Rao, Dengming Zhang, Xingxing Zou
IEEE Trans. Circuits Syst. Video Technol.6
2025 FonTS: Text Rendering with Typography and Style Controls
abstract
Visual text rendering are widespread in various real-world applications, requiring careful font selection and typographic choices. Recent progress in diffusion transformer (DiT)-based text-to-image (T2I) models show promise in automating these processes. However, these methods still encounter challenges like inconsistent fonts, style variation, and limited fine-grained control, particularly at the word-level. This paper proposes a two-stage DiT-based pipeline to address these problems by enhancing controllability over typography and style in text rendering. We introduce typography control fine-tuning (TC-FT), an parameter-efficient fine-tuning method (on $5\%$ key parameters) with enclosing typography control tokens (ETC-tokens), which enables precise word-level application of typographic features. To further address style inconsistency in text rendering, we propose a text-agnostic style control adapter (SCA) that prevents content leakage while enhancing style consistency. To implement TC-FT and SCA effectively, we incorporated HTML-render into the data synthesis pipeline and proposed the first word-level controllable dataset. Through comprehensive experiments, we demonstrate the effectiveness of our approach in achieving superior word-level typographic control, font consistency, and style consistency in text rendering tasks. The datasets and models will be available for academic use.
Wenda Shi, Yiren Song, Dengming Zhang, Xingxing Zou
ICCV5
2025 A two-stage classification-regression method for prediction of flexural strength of fiber reinforced polymer strengthened reinforced concrete beams
Xingxing Zou, Lesley H. Sneed
Eng. Appl. Artif. Intell.2
2025 LoopNet for fine-grained fashion attributes editing
Xingxing Zou, Shumin Zhu, Wai Keung Wong
Expert Syst. Appl.1
2025 Any Fashion Attribute Editing: Dataset and Pretrained Models
abstract
Fashion attribute editing is essential for combining the expertise of fashion designers with the potential of generative artificial intelligence. In this work, we focus on 'any' fashion attribute editing: 1) the ability to edit 78 fine-grained design attributes commonly observed in daily life; 2) the capability to modify desired attributes while keeping the rest components still; and 3) the flexibility to continuously edit on the edited image. To this end, we present the Any Fashion Attribute Editing (AFED) dataset, which includes 830 K high-quality fashion images from sketch and product domains, filling the gap for a large-scale, openly accessible fine-grained dataset. We also propose Twin-Net, a twin encoder-decoder GAN inversion method that offers diverse and precise information for high-fidelity image reconstruction. This inversion model, trained on the new dataset, serves as a robust foundation for attribute editing. Additionally, we introduce PairsPCA to identify semantic directions in latent space, enabling accurate editing without manual supervision. Comprehensive experiments, including comparisons with ten state-of-the-art image inversion methods and four editing algorithms, demonstrate the effectiveness of our Twin-Net and editing algorithm. All data and models are available at https://github.com/ArtmeScienceLab/AnyFashionAttributeEditing.
Shumin Zhu, Xingxing Zou, Wenhan Yang, Wai Keung Wong
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Generative AI in Fashion: Overview
abstract
Generative Artificial Intelligence (GenAI) has recently gained immense popularity by offering various applications for generating high-quality and aesthetically pleasing content of image, 3D, and video data format. The innovative GenAI solutions have shifted paradigms across various design-related industries, particularly fashion. In this paper, we explore the incorporation of GenAI into fashion-related tasks and applications. Our examination encompasses a thorough review of more than 470 research papers and an in-depth analysis of over 300 applications, focusing on their contributions to the field. These contributions are identified as 13 tasks within four categories: multi-modal fashion understanding, and fashion synthesis of image, 3D, and dynamic (video and animatable 3D) formats We delve into these methods, recognizing their potential to propel future endeavours toward achieving state-of-the-art (SOTA) performance. Furthermore, we present a comprehensive overview of 53 publicly available datasets suitable for training and benchmarking fashion-centric models, accompanied by the relevant evaluation metrics. Finally, we review real-world applications, unveiling existing challenges and future directions. With comprehensive investigation and in-depth analysis, this paper is targeted to serve as a useful resource for understanding the current landscape of GenAI in fashion, paving the way for future innovations in this dynamic field. Papers discussed in this paper, along with public code and datasets links are available at: https://github.com/wendashi/Cool-GenAI-Fashion-Papers/ .
Wenda Shi, Wai Keung Wong, Xingxing Zou
ACM Trans. Intell. Syst. Technol.3
2024 Uni-DlLoRA: Style Fine-Tuning for Fashion Image Translation
abstract
Image-to-image (i2i) translation has achieved notable success, yet remains challenging in scenarios like real-to-illustrative style transfer of fashion. Existing methods focus on enhancing the generative model with diversity while lacking ID-preserved domain translation. This paper introduces a novel model named Uni-DlLoRA to release this constraint. The proposed model combines the original images within a pretrained diffusion-based model using the proposed Uni-adapter extractors, while adopting the proposed Dual-LoRA module to provide distinct style guidance. This approach optimizes generative capabilities and reduces the number of additional parameters required. In addition, a new multimodal dataset featuring higher-quality images with captions built upon an existing real-to-illustration dataset is proposed. Experimentation validates the effectiveness of our proposed method.
Fangjian Liao, Xingxing Zou, Wai Keung Wong
ACM Multimedia2
2024 Learning Visual Body-shape-Aware Embeddings for Fashion Compatibility
abstract
Body shape is a crucial factor in outfit recommendation. Previous studies that directly used body measurement data to investigate the relationship between body shape and outfit have achieved limited performance due to oversimplified body shape representations. This paper proposes a Visual Body-shape-Aware Network (ViBA-Net) to improve the fashion compatibility model’s awareness of human body shape through visual-level information. Specifically, ViBA-Net consists of three modules: a body-shape embedding module, which extracts visual and anthropometric features of body shape from a newly introduced large-scale body shape dataset; an outfit embedding module, which learns the outfit representation based on visual features extracted from a try-on image and textual features extracted from fashion attributes; and a joint embedding module, which jointly models the relationship between the representations of body shape and outfit. ViBA-Net is designed to generate attribute-level explanations for the evaluation results based on the computed attention weights. The effectiveness of ViBA-Net is evaluated on two mainstream datasets through qualitative and quantitative analysis. Data and code are released1.
Kaicheng Pang, Xingxing Zou, Wai Keung Wong
WACV2
2024 Attentional pixel-wise deformation for pose-based human image generation
Fangjian Liao, Xingxing Zou, Wai Keung Wong
Expert Syst. Appl.2
2024 Learning Structured Relation Embeddings for Fine-Grained Fashion Attribute Recognition
abstract
Fashion attribute recognition is a not-new topic, but rather a core task in understanding fashion from the perspective of computer vision. This article proposes a structured relation-aware network (sRA-Net), which exploits multiple hidden relations in fashion images to enrich and achieve accurate attribute representations to boost the performance of fashion attribute recognition. Specifically, it deconstructs the features of a clothing fashion item into three levels, including low-level attribute-related image region information, mid-level attribute dependency information, and high-level clothing look information. To learn these multi-relational embeddings, we present three relation-aware attention mechanisms. The attribute attention mechanism describes the relationship among different attribute vectors through self-attention and uses the attention map to update the attribute embedding. Then, the spatial attention mechanism associates the attribute with the image features and enhances the attribute embedding by leveraging the attribute-related image region. Finally, the channel attention mechanism selects attribute-related image feature channels to obtain a more fine-grained attribute embedding. Furthermore, we introduce structure-aware embedding to constrain attribute recognition in images from a global perspective by identifying the inner structure of the clothing. Without bells and whistles, sRA-Net outperforms all state-of-the-art attribute recognition methods on two mainstream fashion attribute datasets, namely the DeepFashion-C dataset and iFashion-Attribute dataset, with over 1%-3% improvement.
Shumin Zhu, Xingxing Zou, Jianjun Qian, Wai Keung Wong
IEEE Trans. Multim.2
2023 Personalized Fashion Recommendation via Deep Personality Learning
Dongmei Mo, Xingxing Zou, Wai Keung Wong
BMVC2
2023 CLOTH4D: A Dataset for Clothed Human Reconstruction
abstract
Clothed human reconstruction is the cornerstone for creating the virtual world. To a great extent, the quality of recovered avatars decides whether the Metaverse is a passing fad. In this work, we introduce CLOTH4D, a clothed human dataset containing 1,000 subjects with varied appearances, 1,000 3D outfits, and over 100,000 clothed meshes with paired unclothed humans, to fill the gap in large-scale and high-quality 4D clothing data. It enjoys appealing characteristics: 1) Accurate and detailed clothing textured meshes-all clothing items are manually created and then simulated in professional software, strictly following the general standard in fashion design. 2) Separated textured clothing and under-clothing body meshes, closer to the physical world than single-layer raw scans. 3) Clothed human motion sequences simulated given a set of 289 actions, covering fundamental and complicated dynamics. Upon CLOTH4D, we novelly designed a series of temporally-aware metries to evaluate the temporal stability of the generated 3D human meshes, which has been over-looked previously. Moreover, by assessing and retraining current state-of-the-art clothed human reconstruction methods, we reveal insights, present improved performance, and propose potential future research directions, confirming our dataset's advancement. The dataset is available at.
Xingxing Zou, Xintong Han, Wai Keung Wong
CVPR1
2023 Towards private stylists via personalized compatibility learning
Dongmei Mo, Xingxing Zou, Kaicheng Pang, Wai Keung Wong
Expert Syst. Appl.2
2022 Dress Well via Fashion Cognitive Learning
Kaicheng Pang, Xingxing Zou, Wai Keung Wong
BMVC2
2022 How Good Is Aesthetic Ability of a Fashion Model?
abstract
We introduce A100 (Aesthetic 100) to assess the aesthetic ability of the fashion compatibility models. To date, it is the first work to address the AI model's aesthetic ability with detailed characterization based on the professional fashion domain knowledge. A100 has several desirable characteristics: 1. Completeness. It covers all types of standards in the fashion aesthetic system through two tests, namely LAT (Liberalism Aesthetic Test) and AAT (Academicism Aesthetic Test); 2. Reliability. It is training data agnostic and consistent with major indicators. It provides a fair and objective judgment for model comparison. 3. Explainability. Better than all previous indicators, the A100 further identifies essential characteristics of fashion aesthetics, thus showing the model's performance on more fine-grained dimensions, such as Color, Balance, Material, etc. Experimental results prove the advance of the A100 in the aforementioned aspects. All data can be found at https://github.com/AemikaChow/AiDLab-fAshIon-Data.
Xingxing Zou, Kaicheng Pang, Wen Zhang 0010, Wai Keung Wong
CVPR1
2022 Neural stylist: Towards online styling service
Dongmei Mo, Xingxing Zou, Wai Keung Wong
Expert Syst. Appl.2