EDBT 2026 Demo / reviewers in the wild / expert
Yanghong Zhou
dblp:172/2788
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-0372-996XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Low-Light Scene Restoration via Illumination TransitionabstractSynthesizing normal-light novel views from low-light multiview images is an important yet challenging task, given the low visibility and high ISO noise present in the input images. Existing low-light enhancement methods often struggle to effectively preprocess such low-light inputs, as they fail to consider correlations among multiple views. Although other state-of-the-art methods have introduced illumination-related components offering alternative solutions to the problem, they often result in drawbacks such as color distortions and artifacts, and they provide limited denoising effectiveness. In this paper, we propose a novel Robust Low-light Scene Restoration framework (RoSe), which enables effective synthesis of novel views in normal lighting conditions from low-light multiview image inputs, by formulating the task as an illuminance transition estimation problem in 3D space, conceptualizing it as a specialized rendering task. This multiview-consistent illuminance transition field establishes a robust connection between low-light and normal-light conditions. By further exploiting the inherent low-rank property of illumination to constrain the transition representation, we achieve more effective denoising without complex 2D techniques or explicit noise modeling. To implement RoSe, we design a concise dual-branch architecture and introduce a low-rank denoising module. Experiments demonstrate that RoSe significantly outperforms state-of-the-art models in both rendering quality and multiview consistency on standard benchmarks. The codes and data are available at https://pegasus2004.github.io/RoSe. Feng Zhang 0052, Xiatian Zhu, Yanghong Zhou, P. Y. Mok 0001 |
ICCV | 5 |
| 2025 | Feature Field Fusion for few-shot novel view synthesisabstractReconstructing neural radiance fields from limited or sparse views has given very promising potential for this field of research. Previous methods usually constrain the reconstruction process with additional priors, e.g. semantic-based or patch-based regularization. Nevertheless, such regularization is given to the synthesis of unseen views, which may not effectively assist the field of learning, in particular when the training views are sparse. Instead, we propose a feature Field Fusion (FFusion) NeRF in this paper that can learn structure and more details from features extracted from pre-trained neural networks for the sparse training views, and use as extra guide for the training of the RGB field. With such extra feature guides, FFusion predicts more accurate color and density when synthesizing novel views. Experimental results have shown that FFusion can effectively improve the quality of the synthesized novel views with only limited or sparse inputs. • A novel NeRF method using spatial feature maps to improve sparse view synthesis. • Concurrently learning both feature & color fields with pixel-level loss supervisions. • A fusion scheme regularizing color prediction with learned features as condition guide. • Extensive tests validate the performance and model-agnostic advantages of the method. Junting Li, Yanghong Zhou, Jintu Fan, Dahua Shou 0001, Sa Xu, P. Y. Mok 0001 |
Image Vis. Comput. | 2 |
| 2025 | A cross-feature interaction network for 3D human pose estimationabstractThe task of estimating 3D human poses from single monocular images is challenging because, unlike video sequences, single images can hardly provide any temporal information for the prediction. Most existing methods attempt to predict 3D poses by modeling the spatial dependencies inherent in the anatomical structure of the human skeleton, yet these methods fail to capture the complex local and global relationships that exist among various joints. To solve this problem, we propose a novel Cross-Feature Interaction Network to effectively model spatial correlations between body joints. Specifically, we exploit graph convolutional networks (GCNs) to learn the local features between neighboring joints and the self-attention structure to learn the global features among all joints. We then design a cross-feature interaction (CFI) module to facilitate cross-feature communications among the three different features, namely the local features, global features, and initial 2D pose features, aggregating them to form enhanced spatial representations of human pose. Furthermore, a novel graph-enhanced module (GraMLP) with parallel GCN and multi-layer perceptron is introduced to inject the skeletal knowledge of the human body into the final representation of 3D pose. Extensive experiments on two datasets (Human3.6M (Ionescu et al., 2013) and MPI-INF-3DHP (Mehta et al., 2017)) show the superior performance of our method in comparison to existing state-of-the-art (SOTA) models. The code and data are shared at https://github.com/JihuaPeng/CFI-3DHPE • A novel CFI Network for enhanced learning of 3D pose representations. • A specific multi-head cross-attention to model dependences across features. • A graph-enhanced GraMPL module with parallel MLP and GCN for feature aggregation. • Outperforming existing models for 3D pose estimation based on single-image inputs. Jihua Peng, Yanghong Zhou, P. Y. Mok 0001 |
Pattern Recognit. Lett. | 2 |
| 2025 | CoDE-GAN: Content Decoupled and Enhanced GAN for Sketch-guided Flexible Fashion EditingabstractRapid advancements in generative models, including Generative Adversarial Networks (GANs) and diffusion models, have made possible automated image editing through the use of text descriptions, semantic segmentation, and/or reference style images. Nevertheless, in terms of fashion image editing, it often requires more flexible, and typically iterative, modifications to the image content that existing methods struggle to achieve. This article proposes a new model called Content Decoupled and Enhanced GAN (CoDE-GAN), which is formulated and trained for the task of image editing, drawing on methods from image reconstruction, more specifically, image inpainting with sketch guidance. Through this proxy task, the trained model can be used for flexible image editing, generating new images with consistent colors and required textures based on sketch inputs. In this new model, a content decoupling block is introduced including specially designed dual encoders, which pre-process inputs and transform into separated structure and texture representations. Moreover, a content enhancing module is designed and applied to the decoder, improving the color consistency and refining the texture of the generated images. The proposed CoDE-GAN can achieve coarse-to-fine results in one single stage. Extensive experiments on three datasets, covering human, garment-only, and scene images, show that CoDE-GAN outperforms other state-of-the-art methods in terms of both generated image quality and editing flexibility. The code and dataset are available at https://github.com/Taited/CoDE-GAN . Zhengwentai Sun, Yanghong Zhou, P. Y. Mok 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | EHFusion: an efficient heterogeneous fusion model for group-based 3D human pose estimationabstractAbstract Stimulated by its important applications in animation, gaming, virtual reality, augmented reality, and healthcare, 3D human pose estimation has received considerable attention in recent years. To improve the accuracy of 3D human pose estimation, most approaches have converted this challenging task into a local pose estimation problem by dividing the body joints of the human body into different groups based on the human body topology. The body joint features of different groups are then fused to predict the overall pose of the whole body, which requires a joint feature fusion scheme. Nevertheless, the joint feature fusion schemes adopted in existing methods involve the learning of extensive parameters and hence are computationally very expensive. This paper reports a new topology-based grouped method ‘EHFusion’ for 3D human pose estimation, which involves a heterogeneous feature fusion (HFF) module that integrates grouped pose features. The HFF module reduces the computational complexity of the model while achieving promising accuracy. Moreover, we introduce motion amplitude information and a camera intrinsic embedding module to provide better global information and 2D-to-3D conversion knowledge, thereby improving the overall robustness and accuracy of the method. In contrast to previous methods, the proposed new network can be trained end-to-end in one single stage. Experimental results not only demonstrate the advantageous trade-offs between estimation accuracy and computational complexity achieved by our method but also showcase the competitive performance in comparison with various existing state-of-the-art methods (e.g., transformer-based) when evaluated on two public datasets, Human3.6M and HumanEva. The data and code are available at doi:10.5281/zenodo.11113132 Jihua Peng, Yanghong Zhou, P. Y. Mok 0001 |
Vis. Comput. | 2 |
| 2024 | KTPFormer: Kinematics and Trajectory Prior Knowledge-Enhanced Transformer for 3D Human Pose EstimationabstractThis paper presents a novel Kinematics and Trajectory Prior Knowledge-Enhanced Transformer (KTPFormer), which overcomes the weakness in existing transformer-based methods for 3D human pose estimation that the derivation of Q, K, V vectors in their self-attention mechanisms are all based on simple linear mapping. We propose two prior attention modules, namely Kinematics Prior Attention (KPA) and Trajectory Prior Attention (TPA) to take advantage of the known anatomical structure of the human body and motion trajectory information, to facilitate effective learning of global dependencies and features in the multi-head self-attention. KPA models kinematic relationships in the human body by constructing a topology of kinematics, while TPA builds a trajectory topology to learn the information of joint motion trajectory across frames. Yielding Q, K, V vectors with prior knowledge, the two modules enable KTPFormer to model both spatial and temporal correlations simultaneously. Extensive experiments on three benchmarks (Human3.6M, MPI-INF-3DHP and HumanEva) show that KTPFormer achieves superior performance in comparison to state-of-the-art methods. More importantly, our KPA and TPA modules have lightweight plug-and-play designs and can be integrated into various transformer-based networks (i.e., diffusion-based) to improve the performance with only a very small increase in the computational overhead. The code is available at: https://github.com/JihuaPeng/KTPFormer. Jihua Peng, Yanghong Zhou, P. Y. Mok 0001 |
CVPR | 2 |
| 2024 | Zero-Shot Sketch Based Image Retrieval via Modality Capacity Guidance
Yanghong Zhou, P. Y. Mok 0001 |
IJCAI | 1 |
| 2024 | Knowledge enhanced multi-task learning for simultaneous optimization of human parsing and pose estimation
Yanghong Zhou, P. Y. Mok 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Enhancing human parsing with region-level learningabstractAbstract Human parsing is very important in a diverse range of industrial applications. Despite the considerable progress that has been achieved, the performance of existing methods is still less than satisfactory, since these methods learn the shared features of various parsing labels at the image level. This limits the representativeness of the learnt features, especially when the distribution of parsing labels is imbalanced or the scale of different labels is substantially different. To address this limitation, a Region‐level Parsing Refiner (RPR) is proposed to enhance parsing performance by the introduction of region‐level parsing learning. Region‐level parsing focuses specifically on small regions of the body, for example, the head. The proposed RPR is an adaptive module that can be integrated with different existing human parsing models to improve their performance. Extensive experiments are conducted on two benchmark datasets, and the results demonstrated the effectiveness of our RPR model in terms of improving the overall parsing performance as well as parsing rare labels. This method was successfully applied to a commercial application for the extraction of human body measurements and has been used in various online shopping platforms for clothing size recommendations. The code and dataset are released at this link https://github.com/applezhouyp/PRP . Yanghong Zhou, P. Y. Mok 0001 |
IET Comput. Vis. | 1 |
| 2023 | SGDiff: A Style Guided Diffusion Model for Fashion SynthesisabstractThis paper reports on the development of a novel style guided diffusion model (SGDiff) which overcomes certain weaknesses inherent in existing models for image synthesis. The proposed SGDiff combines image modality with a pretrained text-to-image diffusion model to facilitate creative fashion image synthesis. It addresses the limitations of text-to-image diffusion models by incorporating supplementary style guidance, substantially reducing training costs, and overcoming the difficulties of controlling synthesized styles with text-only inputs. This paper also introduces a new dataset -- SG-Fashion, specifically designed for fashion image synthesis applications, offering high-resolution images and an extensive range of garment categories. By means of comprehensive ablation study, we examine the application of classifier-free guidance to a variety of conditions and validate the effectiveness of the proposed model for generating fashion images of the desired categories, product attributes, and styles. The contributions of this paper include a novel classifier-free guidance method for multi-modal feature fusion, a comprehensive dataset for fashion image synthesis application, a thorough investigation on conditioned text-to-image synthesis, and valuable insights for future research in the text-to-image synthesis domain. The code and dataset are available at: https://github.com/taited/SGDiff. Zhengwentai Sun, Yanghong Zhou, Honghong He, P. Y. Mok 0001 |
ACM Multimedia | 2 |
| 2023 | Unbiased feature position alignment for human pose estimation
Chen Wang 0136, Yanghong Zhou, Feng Zhang 0052, P. Y. Mok 0001 |
Neurocomputing | 2 |
| 2023 | A Pose-Aware Global Representation Network for Human ParsingabstractMany recognition tasks including image/video classification, segmentation and object detection can be improved by the integration of global information. Although global information may be better represented in some recognition tasks than the others, it is worth exploring how global information from related tasks can be effectively used to improve the performance of a target task. The task of pose estimation predicts the locations of human joints, thus providing global information about the human body. In this paper, we propose a pose-aware global representation network model (PAGRnet) that exploits global information from pose estimation to enhance feature learning in human parsing. In our PAGRnet model, a novel learning module with three integrated parts is used to learn global information. The first part generates a global joint representation, while the second part learns the relationship between the pixels and joints. By integrating the global joint representation with the pixel-joint relationship, the resulting pose-aware global representation is augmented for the parsing task. Our experimental results show competitive performance of our method on the LIP, the Pascal-person-part and the ATR datasets, with reduced computation costs in comparison to other proposals of global information fusion. We also demonstrate the advantages of our feature fusion model over concatenation, pixel-wise and channel-wise relation models. Yanghong Zhou, P. Y. Mok 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Modeling Field-Level Factor Interactions for Fashion RecommendationabstractPersonalized fashion recommendation aims to explore patterns from historical interactions between users and fashion items and thereby predict the future ones. It is challenging due to the sparsity of the interaction data and the diversity of user preference in fashion. To tackle the challenge, this paper investigates multiple factor fields in fashion domain, such as colour, style, brand, and tries to specify the implicit user-item interaction into field level. Specifically, an attentional factor field interaction graph (AFFIG) approach is proposed which models both the user-factor interactions and cross-field factors interactions for predicting the recommendation probability at specific field. In addition, an attention mechanism is equipped to aggregate the cross-field factor interactions for each field. Extensive experiments have been conducted on three E-Commerce fashion datasets and the results demonstrate the effectiveness of the proposed method for fashion recommendation. The influence of various factor fields on recommendation in fashion domain is also discussed through experiments. Yujuan Ding, P. Y. Mok 0001, Xun Yang 0001, Yanghong Zhou |
ICME | 4 |
| 2019 | Fashion recommendations through cross-media information retrieval
Wei Zhou 0028, P. Y. Mok 0001, Yanghong Zhou, Yangping Zhou, Jialie Shen 0001, Qiang Qu 0001, K. P. Chau |
J. Vis. Commun. Image Represent. | 3 |