P. Y. Mok 0001

dblp:50/226 · also Pik-Yin Mok, Tracy Pik Yin Mok · DBLP profile ↗
← Back
42ranked-venue papers
5as first author
27since 2021 · last 2025
0000-0002-0635-5318ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Robust Low-Light Scene Restoration via Illumination Transition
abstract
Synthesizing normal-light novel views from low-light multiview images is an important yet challenging task, given the low visibility and high ISO noise present in the input images. Existing low-light enhancement methods often struggle to effectively preprocess such low-light inputs, as they fail to consider correlations among multiple views. Although other state-of-the-art methods have introduced illumination-related components offering alternative solutions to the problem, they often result in drawbacks such as color distortions and artifacts, and they provide limited denoising effectiveness. In this paper, we propose a novel Robust Low-light Scene Restoration framework (RoSe), which enables effective synthesis of novel views in normal lighting conditions from low-light multiview image inputs, by formulating the task as an illuminance transition estimation problem in 3D space, conceptualizing it as a specialized rendering task. This multiview-consistent illuminance transition field establishes a robust connection between low-light and normal-light conditions. By further exploiting the inherent low-rank property of illumination to constrain the transition representation, we achieve more effective denoising without complex 2D techniques or explicit noise modeling. To implement RoSe, we design a concise dual-branch architecture and introduce a low-rank denoising module. Experiments demonstrate that RoSe significantly outperforms state-of-the-art models in both rendering quality and multiview consistency on standard benchmarks. The codes and data are available at https://pegasus2004.github.io/RoSe.
Feng Zhang 0052, Xiatian Zhu, Yanghong Zhou, P. Y. Mok 0001
ICCV6
2025 Blockchain-driven innovation in fashion supply chain contractual party evaluations as an emerging collaboration model
abstract
202509 bcch
Minhao Qiao, Xuanchang Chen, Yangping Zhou, P. Y. Mok 0001
Blockchain Res. Appl.4
2025 TiPGAN: High-quality tileable textures synthesis with intrinsic priors for cloth digitization applications
abstract
Seamless textures play an important role in 3D modeling, animation, video games, and Augmented Reality/Virtual Reality, enhancing the realism and aesthetics of the digital environments. Despite its significance, generating seamless textures is not trivial, requiring the edges of the synthesized texture image to represent a continuous pattern when tiled. Although traditional methods and deep learning models have made good progress in texture synthesis, they often fail in ensuring the seamless property of the synthesized textures. In this paper, we report on TiPGAN, a Generative Adversarial Network (GAN) model, which we developed to generate seamless textures. Leveraging the inherent intrinsics of seamless textures as priors , our model introduces two novel modules: a Patch Swapping Module, for maintaining texture continuity through diagonal patch swapping, and a Patch Tiling Module, for ensuring seamless repetition across tiles. To overcome the limitations of existing image quality metrics in evaluating tileability, we introduce a new metric, termed Relative Total Variation (RTV), for assessing the smoothness and continuity of the synthesized textures. Our experimental results demonstrate that TiPGAN outperforms existing methods in generating high-quality, seamless textures, as validated by both conventional image quality metrics and our newly proposed RTV metric. This research represents a significant advancement in texture generation, offering valuable applications in graphic design, virtual reality, and digital art. Our code and dataset are available at https://github.com/VickyHEHonghong/TiPGAN . • TiPGAN leverages intrinsic priors to improve seamless texture synthesis. • RTV metric is introduced to quantitatively evaluates the texture tileability. • TiPGAN outperforms existing methods as verified by image quality metrics and RTV. • TiPGAN contributes to cloth digitization, providing value to digital art.
Honghong He, Zhengwentai Sun, Jintu Fan, P. Y. Mok 0001
Comput. Aided Des.4
2025 Feature Field Fusion for few-shot novel view synthesis
abstract
Reconstructing neural radiance fields from limited or sparse views has given very promising potential for this field of research. Previous methods usually constrain the reconstruction process with additional priors, e.g. semantic-based or patch-based regularization. Nevertheless, such regularization is given to the synthesis of unseen views, which may not effectively assist the field of learning, in particular when the training views are sparse. Instead, we propose a feature Field Fusion (FFusion) NeRF in this paper that can learn structure and more details from features extracted from pre-trained neural networks for the sparse training views, and use as extra guide for the training of the RGB field. With such extra feature guides, FFusion predicts more accurate color and density when synthesizing novel views. Experimental results have shown that FFusion can effectively improve the quality of the synthesized novel views with only limited or sparse inputs. • A novel NeRF method using spatial feature maps to improve sparse view synthesis. • Concurrently learning both feature & color fields with pixel-level loss supervisions. • A fusion scheme regularizing color prediction with learned features as condition guide. • Extensive tests validate the performance and model-agnostic advantages of the method.
Junting Li, Yanghong Zhou, Jintu Fan, Dahua Shou 0001, Sa Xu, P. Y. Mok 0001
Image Vis. Comput.6
2025 A cross-feature interaction network for 3D human pose estimation
abstract
The task of estimating 3D human poses from single monocular images is challenging because, unlike video sequences, single images can hardly provide any temporal information for the prediction. Most existing methods attempt to predict 3D poses by modeling the spatial dependencies inherent in the anatomical structure of the human skeleton, yet these methods fail to capture the complex local and global relationships that exist among various joints. To solve this problem, we propose a novel Cross-Feature Interaction Network to effectively model spatial correlations between body joints. Specifically, we exploit graph convolutional networks (GCNs) to learn the local features between neighboring joints and the self-attention structure to learn the global features among all joints. We then design a cross-feature interaction (CFI) module to facilitate cross-feature communications among the three different features, namely the local features, global features, and initial 2D pose features, aggregating them to form enhanced spatial representations of human pose. Furthermore, a novel graph-enhanced module (GraMLP) with parallel GCN and multi-layer perceptron is introduced to inject the skeletal knowledge of the human body into the final representation of 3D pose. Extensive experiments on two datasets (Human3.6M (Ionescu et al., 2013) and MPI-INF-3DHP (Mehta et al., 2017)) show the superior performance of our method in comparison to existing state-of-the-art (SOTA) models. The code and data are shared at https://github.com/JihuaPeng/CFI-3DHPE • A novel CFI Network for enhanced learning of 3D pose representations. • A specific multi-head cross-attention to model dependences across features. • A graph-enhanced GraMPL module with parallel MLP and GCN for feature aggregation. • Outperforming existing models for 3D pose estimation based on single-image inputs.
Jihua Peng, Yanghong Zhou, P. Y. Mok 0001
Pattern Recognit. Lett.3
2025 CoDE-GAN: Content Decoupled and Enhanced GAN for Sketch-guided Flexible Fashion Editing
abstract
Rapid advancements in generative models, including Generative Adversarial Networks (GANs) and diffusion models, have made possible automated image editing through the use of text descriptions, semantic segmentation, and/or reference style images. Nevertheless, in terms of fashion image editing, it often requires more flexible, and typically iterative, modifications to the image content that existing methods struggle to achieve. This article proposes a new model called Content Decoupled and Enhanced GAN (CoDE-GAN), which is formulated and trained for the task of image editing, drawing on methods from image reconstruction, more specifically, image inpainting with sketch guidance. Through this proxy task, the trained model can be used for flexible image editing, generating new images with consistent colors and required textures based on sketch inputs. In this new model, a content decoupling block is introduced including specially designed dual encoders, which pre-process inputs and transform into separated structure and texture representations. Moreover, a content enhancing module is designed and applied to the decoder, improving the color consistency and refining the texture of the generated images. The proposed CoDE-GAN can achieve coarse-to-fine results in one single stage. Extensive experiments on three datasets, covering human, garment-only, and scene images, show that CoDE-GAN outperforms other state-of-the-art methods in terms of both generated image quality and editing flexibility. The code and dataset are available at https://github.com/Taited/CoDE-GAN .
Zhengwentai Sun, Yanghong Zhou, P. Y. Mok 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Computer-Aided Colorization State-of-the-Science: A Survey
abstract
This article reviews published research in the field of computer-aided colorization technology. We argue that within this context, the colorization task can be considered to originate from computer graphics, advance by introducing computer vision, and progress towards the fusion of vision and graphics. Hence, we propose a specific taxonomy and organize the research work chronologically. We extend the existing reconstruction-based colorization evaluation techniques on the basis that aesthetic assessment should be introduced to ensure the computer-coloredimages closely satisfy human visual-related requirements. We then perform an aesthetic assessment using the proposed metric and existing evaluations, comparing the colorization performance of seven representative unconditional colorization models. Finally, we identify unresolved issues and propose fruitful areas for future research and development.
Yu Cao 0019, Xin Duan, Xiangqiao Meng, P. Y. Mok 0001, Ping Li 0016, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2025 EHFusion: an efficient heterogeneous fusion model for group-based 3D human pose estimation
abstract
Abstract Stimulated by its important applications in animation, gaming, virtual reality, augmented reality, and healthcare, 3D human pose estimation has received considerable attention in recent years. To improve the accuracy of 3D human pose estimation, most approaches have converted this challenging task into a local pose estimation problem by dividing the body joints of the human body into different groups based on the human body topology. The body joint features of different groups are then fused to predict the overall pose of the whole body, which requires a joint feature fusion scheme. Nevertheless, the joint feature fusion schemes adopted in existing methods involve the learning of extensive parameters and hence are computationally very expensive. This paper reports a new topology-based grouped method ‘EHFusion’ for 3D human pose estimation, which involves a heterogeneous feature fusion (HFF) module that integrates grouped pose features. The HFF module reduces the computational complexity of the model while achieving promising accuracy. Moreover, we introduce motion amplitude information and a camera intrinsic embedding module to provide better global information and 2D-to-3D conversion knowledge, thereby improving the overall robustness and accuracy of the method. In contrast to previous methods, the proposed new network can be trained end-to-end in one single stage. Experimental results not only demonstrate the advantageous trade-offs between estimation accuracy and computational complexity achieved by our method but also showcase the competitive performance in comparison with various existing state-of-the-art methods (e.g., transformer-based) when evaluated on two public datasets, Human3.6M and HumanEva. The data and code are available at doi:10.5281/zenodo.11113132
Jihua Peng, Yanghong Zhou, P. Y. Mok 0001
Vis. Comput.3
2025 Recycling/upcycling graphic design: automatic design elements extraction and vectorization
abstract
Abstract A graphic design image can be considered as an arrangement of design elements filled with specific colors according to certain layout rules. This study focuses on developing an automatic pipeline to identify and extract core design elements from any unknown input design images, and it also converts the extracted design elements to a neat vector format with a new optimization algorithm. More specifically, an unsupervised learning technique is applied for core design element extraction to address the problem of no available dataset on labeled designs. Next, a novel design element vectorization method is proposed based on color features. Extensive experiments have demonstrated the effectiveness of the proposed method on design image processing. The proposed method can process, namely extract and vectorize, core design elements of an unknown image within around 29 s. The output vector graphics of design elements have very compact format, compared to other commercial software, and are optimized for reuse and re-editing. Source code and demonstration of the proposed pipeline can be found at https://github.com/ZoeQU/Recycling-Upcycling-Graphic-Design-Automatic-Design-Elements-Extraction-and-Vectorization .
K. P. Chau, P. Y. Mok 0001
Vis. Comput.3
2024 SewPCT: Sewing Pattern Reconstruction from Point Cloud with Transformer
Hao Tian 0014, Yu Cao 0019, P. Y. Mok 0001
CGI (2)3
2024 KTPFormer: Kinematics and Trajectory Prior Knowledge-Enhanced Transformer for 3D Human Pose Estimation
abstract
This paper presents a novel Kinematics and Trajectory Prior Knowledge-Enhanced Transformer (KTPFormer), which overcomes the weakness in existing transformer-based methods for 3D human pose estimation that the derivation of Q, K, V vectors in their self-attention mechanisms are all based on simple linear mapping. We propose two prior attention modules, namely Kinematics Prior Attention (KPA) and Trajectory Prior Attention (TPA) to take advantage of the known anatomical structure of the human body and motion trajectory information, to facilitate effective learning of global dependencies and features in the multi-head self-attention. KPA models kinematic relationships in the human body by constructing a topology of kinematics, while TPA builds a trajectory topology to learn the information of joint motion trajectory across frames. Yielding Q, K, V vectors with prior knowledge, the two modules enable KTPFormer to model both spatial and temporal correlations simultaneously. Extensive experiments on three benchmarks (Human3.6M, MPI-INF-3DHP and HumanEva) show that KTPFormer achieves superior performance in comparison to state-of-the-art methods. More importantly, our KPA and TPA modules have lightweight plug-and-play designs and can be integrated into various transformer-based networks (i.e., diffusion-based) to improve the performance with only a very small increase in the computational overhead. The code is available at: https://github.com/JihuaPeng/KTPFormer.
Jihua Peng, Yanghong Zhou, P. Y. Mok 0001
CVPR3
2024 Hypergraph-Enhanced Contrastively Regularized Transformer for Multi-Behavior E-commerce Product Recommendation
abstract
Multi-behavior sequential recommendation systems, which predict users' subsequent interactions based on historical behavior sequences, play a vital role in e-commerce operations. Despite the promising results of existing state-of-the-art methods, which employ either advanced attention mechanisms or graph-based neural networks, they still face a few open challenges, including the diversity of user actions, the intricate relationships within time-evolving interactions, and the complications introduced by limited or noisy data. To address these limitations, we propose a Similarity Graph-enhanced Multi-Scale Transformer (SG-MST) for multi-behavior e-commerce product recommendation. SG-MST integrates a Similarity Augmented Multi-Behavior Hypergraph that captures complex behavior-aware dependencies among items and strengthens connections through item context similarities, producing more informative latent representations. Additionally, we introduce a Contrastively Regularized Multi-Scale Transformer, which leverages contrastive learning to capture enriched sequential patterns across both long-term and short-term dynamically sampled and augmented temporal scales, thereby improving the robustness of user behavior prediction. Extensive evaluation on three real-world benchmark datasets demonstrates that the proposed SG-MST outperforms other state-of-the-art methods.
Shuiying Liao, P. Y. Mok 0001
ICDM2
2024 Zero-Shot Sketch Based Image Retrieval via Modality Capacity Guidance
Yanghong Zhou, P. Y. Mok 0001
IJCAI3
2024 Reproducibility Companion Paper: Recommendation of Mix-and-Match Clothing by Modeling Indirect Personal Compatibility
abstract
ICMR '24: International Conference on Multimedia Retrieval, Phuket, Thailand, June 10-14, 2024
Shuiying Liao, Yujuan Ding, P. Y. Mok 0001, Qiushi Huang, Jialun Cao
ICMR3
2024 Knowledge enhanced multi-task learning for simultaneous optimization of human parsing and pose estimation
Yanghong Zhou, P. Y. Mok 0001
Eng. Appl. Artif. Intell.2
2024 Enhancing human parsing with region-level learning
abstract
Abstract Human parsing is very important in a diverse range of industrial applications. Despite the considerable progress that has been achieved, the performance of existing methods is still less than satisfactory, since these methods learn the shared features of various parsing labels at the image level. This limits the representativeness of the learnt features, especially when the distribution of parsing labels is imbalanced or the scale of different labels is substantially different. To address this limitation, a Region‐level Parsing Refiner (RPR) is proposed to enhance parsing performance by the introduction of region‐level parsing learning. Region‐level parsing focuses specifically on small regions of the body, for example, the head. The proposed RPR is an adaptive module that can be integrated with different existing human parsing models to improve their performance. Extensive experiments are conducted on two benchmark datasets, and the results demonstrated the effectiveness of our RPR model in terms of improving the overall parsing performance as well as parsing rare labels. This method was successfully applied to a commercial application for the extraction of human body measurements and has been used in various online shopping platforms for clothing size recommendations. The code and dataset are released at this link https://github.com/applezhouyp/PRP .
Yanghong Zhou, P. Y. Mok 0001
IET Comput. Vis.2
2024 Personalizing human avatars based on realistic 3D facial reconstruction
abstract
Abstract Personalized 3D human avatars have aroused a great deal of interest because it is attractive to most people, particularly generation Z, to have the digital twins in their own appearance to live, work, interact, and shop in the metaverse. Nevertheless, personalized avatars are rarely used in practice because of the computational cost and hardware restrictions in the creation process. This has resulted in avatars of diverse topologies being used on different platforms/systems for various applications, which further hinders the utilization of personalized avatars. This paper reports on a new method for personalizing human avatars, which includes the reconstruction of personalized face models from single images and transferring the reconstructed 3D facial shape and appearance to avatars with varying topologies. This newly developed method is compared with state-of-the-art face reconstruction and personalized avatar reconstruction methods. Based upon the results obtained, it was concluded that the new method created more realistic and true-to-life avatars. This method has been applied in an augmented reality (AR) mobile application, enabling users to engage in virtual try-on experience of fashion. The code will be released once the paper is published.
Yueming Ding, Honghong He, P. Y. Mok 0001
Multim. Tools Appl.3
2024 Improved 3D human face reconstruction from 2D images using blended hard edges
abstract
Abstract This study reports an effective and robust edge-based scheme for the reconstruction of 3D human faces from input of single images, addressing drawbacks of existing methods in case of large face pose angles or noisy input images. Accurate 3D face reconstruction from 2D images is important, as it can enable a wide range of applications, such as face recognition, animations, games and AR/VR systems. Edge features extracted from 2D images contain wealthy and robust 3D geometric information, which were used together with landmarks for face reconstruction purpose. However, the accurate reconstruction of 3D faces from contour features is a challenging task, since traditional edge or contour detection algorithms introduce a great deal of noise, which would adversely affect the reconstruction. This paper reports on the use of a hard-blended face contour feature from a neural network and a Canny edge extractor for face reconstruction. The quantitative results indicate that our method achieves a notable improvement in face reconstruction with a Euclidean distance error of 1.64 mm and a normal vector distance error of 1.27 mm when compared to the ground truth, outperforming both traditional and other deep learning-based methods. These metrics show particularly significant advancements, especially in face shape reconstruction under large pose angles. The method also achieved higher accuracy and robustness on in-the-wild images under conditions of blurring, makeup, occlusion and poor illumination.
Yueming Ding, P. Y. Mok 0001
Neural Comput. Appl.2
2024 AnimeDiffusion: Anime Diffusion Colorization
abstract
Being essential in animation creation, colorizing anime line drawings is usually a tedious and time-consuming manual task. Reference-based line drawing colorization provides an intuitive way to automatically colorize target line drawings using reference images. The prevailing approaches are based on generative adversarial networks (GANs), yet these methods still cannot generate high-quality results comparable to manually-colored ones. In this article, a new AnimeDiffusion approach is proposed via hybrid diffusions for the automatic colorization of anime face line drawings. This is the first attempt to utilize the diffusion model for reference-based colorization, which demands a high level of control over the image synthesis process. To do so, a hybrid end-to-end training strategy is designed, including phase 1 for training diffusion model with classifier-free guidance and phase 2 for efficiently updating color tone with a target reference colored image. The model learns denoising and structure-capturing ability in phase 1, and in phase 2, the model learns more accurate color information. Utilizing our hybrid training strategy, the network convergence speed is accelerated, and the colorization performance is improved. Our AnimeDiffusion generates colorization results with semantic correspondence and color consistency. In addition, the model has a certain generalization performance for line drawings of different line styles. To train and evaluate colorization methods, an anime face line drawing colorization benchmark dataset, containing 31,696 training data and 579 testing data, is introduced and shared. Extensive experiments and user studies have demonstrated that our proposed AnimeDiffusion outperforms state-of-the-art GAN-based methods and another diffusion-based model, both quantitatively and qualitatively.
Yu Cao 0019, Xiangqiao Meng, P. Y. Mok 0001, Tong-Yee Lee, Xueting Liu 0001, Ping Li 0016
IEEE Trans. Vis. Comput. Graph.3
2023 Attention-Aware Anime Line Drawing Colorization
abstract
Automatic colorization of anime line drawing has attracted much attention in recent years since it can substantially benefit the animation industry. User-hint based methods are the mainstream approach for line drawing colorization, while reference-based methods offer a more intuitive approach. Nevertheless, although reference-based methods can improve feature aggregation of the reference image and the line drawing, the colorization results are not compelling in terms of color consistency or semantic correspondence. In this paper, we introduce an attention-based model for anime line drawing colorization, in which a channel-wise and spatial-wise Convolutional Attention module is used to improve the ability of the encoder for feature extraction and key area perception, and a Stop-Gradient Attention module with cross-attention and self-attention is used to tackle the cross-domain long-range dependency problem. Extensive experiments show that our method outperforms other SOTA methods, with more accurate line structure and semantic color information.
Yu Cao 0019, Hao Tian 0014, P. Y. Mok 0001
ICME3
2023 Recommendation of Mix-and-Match Clothing by Modeling Indirect Personal Compatibility
abstract
Fashion recommendation considers both product similarity and compatibility, and has drawn increasing research interest. It is a challenging task because it often needs to use information from different sources, such as visual content or textual descriptions for the prediction of user preferences. In terms of complementary recommendation, existing approaches were dedicated to modeling either product compatibility or users’ personalization in a direct and decoupled manner, yet overlooked additional relations hidden within historical user-product interactions. In this paper, we propose a Normalized indirect Personal Compatibility modeling scheme based on Bayesian Personalized Ranking (NiPC-BPR) for mix-and-match clothing recommendations. We exploit direct and indirect personalization and compatibility relations from the user and product interactions, and effectively integrate various multi-modal data. Extensive experimental results on two benchmark datasets show that our method outperforms other methods by large margins.
Shuiying Liao, Yujuan Ding, P. Y. Mok 0001
ICMR3
2023 Modeling Multi-Relational Connectivity for Personalized Fashion Matching
abstract
Personalized fashion matching task aims to predict the compatible fashion items given available ones for specific users through the effective modeling of the third-order interaction patterns among the user and item pairs. To achieve this, previous methods separately model two key components, user-item and item-item relationships, which ignore the inherent correlations between them and lead to undesirable performance. With a new perspective, this paper proposes to formulate the personalized item matching as the multi-relational connectivity and apply a single-component translation operation to model the targeted third-order interactions. With user-item-item interactions naturally constructing a multi-relational graph, we further device two graph learning modules to enhance the translation-based matching approach from two perspectives,C ontext and Path. The proposed method, named CP-TransMatch, has been tested with extensive experiments on three benchmark fashion datasets and proven effective. It sets the new SOTA for the personalized fashion matching task.
Yujuan Ding, P. Y. Mok 0001, Yi Bin, Xun Yang 0001, Zhiyong Cheng 0001
ACM Multimedia2
2023 SGDiff: A Style Guided Diffusion Model for Fashion Synthesis
abstract
This paper reports on the development of a novel style guided diffusion model (SGDiff) which overcomes certain weaknesses inherent in existing models for image synthesis. The proposed SGDiff combines image modality with a pretrained text-to-image diffusion model to facilitate creative fashion image synthesis. It addresses the limitations of text-to-image diffusion models by incorporating supplementary style guidance, substantially reducing training costs, and overcoming the difficulties of controlling synthesized styles with text-only inputs. This paper also introduces a new dataset -- SG-Fashion, specifically designed for fashion image synthesis applications, offering high-resolution images and an extensive range of garment categories. By means of comprehensive ablation study, we examine the application of classifier-free guidance to a variety of conditions and validate the effectiveness of the proposed model for generating fashion images of the desired categories, product attributes, and styles. The contributions of this paper include a novel classifier-free guidance method for multi-modal feature fusion, a comprehensive dataset for fashion image synthesis application, a thorough investigation on conditioned text-to-image synthesis, and valuable insights for future research in the text-to-image synthesis domain. The code and dataset are available at: https://github.com/taited/SGDiff.
Zhengwentai Sun, Yanghong Zhou, Honghong He, P. Y. Mok 0001
ACM Multimedia4
2023 Unbiased feature position alignment for human pose estimation
Chen Wang 0136, Yanghong Zhou, Feng Zhang 0052, P. Y. Mok 0001
Neurocomputing4
2023 Personalized fashion outfit generation with user coordination preference learning
Yujuan Ding, P. Y. Mok 0001, Yunshan Ma 0001, Yi Bin
Inf. Process. Manag.2
2023 A Pose-Aware Global Representation Network for Human Parsing
abstract
Many recognition tasks including image/video classification, segmentation and object detection can be improved by the integration of global information. Although global information may be better represented in some recognition tasks than the others, it is worth exploring how global information from related tasks can be effectively used to improve the performance of a target task. The task of pose estimation predicts the locations of human joints, thus providing global information about the human body. In this paper, we propose a pose-aware global representation network model (PAGRnet) that exploits global information from pose estimation to enhance feature learning in human parsing. In our PAGRnet model, a novel learning module with three integrated parts is used to learn global information. The first part generates a global joint representation, while the second part learns the relationship between the pixels and joints. By integrating the global joint representation with the pixel-joint relationship, the resulting pose-aware global representation is augmented for the parsing task. Our experimental results show competitive performance of our method on the LIP, the Pascal-person-part and the ATR datasets, with reduced computation costs in comparison to other proposals of global information fusion. We also demonstrate the advantages of our feature fusion model over concatenation, pixel-wise and channel-wise relation models.
Yanghong Zhou, P. Y. Mok 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Modeling Field-Level Factor Interactions for Fashion Recommendation
abstract
Personalized fashion recommendation aims to explore patterns from historical interactions between users and fashion items and thereby predict the future ones. It is challenging due to the sparsity of the interaction data and the diversity of user preference in fashion. To tackle the challenge, this paper investigates multiple factor fields in fashion domain, such as colour, style, brand, and tries to specify the implicit user-item interaction into field level. Specifically, an attentional factor field interaction graph (AFFIG) approach is proposed which models both the user-factor interactions and cross-field factors interactions for predicting the recommendation probability at specific field. In addition, an attention mechanism is equipped to aggregate the cross-field factor interactions for each field. Extensive experiments have been conducted on three E-Commerce fashion datasets and the results demonstrate the effectiveness of the proposed method for fashion recommendation. The influence of various factor fields on recommendation in fashion domain is also discussed through experiments.
Yujuan Ding, P. Y. Mok 0001, Xun Yang 0001, Yanghong Zhou
ICME2
2019 Fashion recommendations through cross-media information retrieval
Wei Zhou 0028, P. Y. Mok 0001, Yanghong Zhou, Yangping Zhou, Jialie Shen 0001, Qiang Qu 0001, K. P. Chau
J. Vis. Commun. Image Represent.2
2017 A new design concept: 3D to 2D textile pattern design for garments
Shufang Lu, P. Y. Mok 0001, Xiaogang Jin 0001
Comput. Aided Des.2
2014 From design methodology to evolutionary design: An interactive creation of marble-like textile patterns
Shufang Lu, P. Y. Mok 0001, Xiaogang Jin 0001
Eng. Appl. Artif. Intell.2
2014 Shape Deformation Using Skeleton Correspondences for Realistic Posed Fashion Flat Creation
abstract
We propose a 2D shape deformation method to fit technical drawings of garments (“flats”) to body figure drawings with a diversity of fashion poses. We first dress a flat onto a body figure in a standard standing pose using Radial Basis Function (RBF) mapping. For different types of clothing, we suggest two levels of treatment to determine handles automatically. We deform the flats using the selected handles to create realistic fashion sketches that fit the garments onto fashion figures in different poses. Shape deformation is performed to minimize the distortion of all the triangles of the garment mesh and preserve garment properties in the deformation. Finally, the garment details, such as style lines and seams, are deformed accordingly for realistic deformation results. Experimental results have shown that our method can deform various garment flats to fit fashion figures in different poses.
Xianmei Wan, P. Y. Mok 0001, Xiaogang Jin 0001
IEEE Trans Autom. Sci. Eng.2
2013 An IGA-based design support system for realistic and practical fashion designs
P. Y. Mok 0001, Jie Xu 0055, X. X. Wang, J. T. Fan, Y. L. Kwok, John H. Xin
Comput. Aided Des.1
2013 An efficient human model customization method based on orthogonal-view monocular photos
Shuaiyin Zhu, P. Y. Mok 0001, Y. L. Kwok
Comput. Aided Des.2
2012 Computer aided clothing pattern design with 3D editing and pattern alteration
Yuwei Meng, P. Y. Mok 0001, Xiaogang Jin 0001
Comput. Aided Des.2
2012 A robust adaptive clustering analysis method for automatic identification of clusters
P. Y. Mok 0001, Haiqiao Huang, Yi-lin Kwok, J. S. Au
Pattern Recognit.1
2011 Interactive Sketch Design Recognition System using Evolutionary Techniques
P. Y. Mok 0001, X. X. Wang, J. T. Fan, Y. L. Kwok, John H. Xin
ICINCO (1)1
2010 Interactive virtual try-on clothing design systems
Yuwei Meng, P. Y. Mok 0001, Xiaogang Jin 0001
Comput. Aided Des.2
2009 A Parameter Free Approach for Clustering Analysis
Haiqiao Huang, P. Y. Mok 0001, Yi-lin Kwok, Sau-Chuen Au
CAIP2
2009 A decision support system for the production control of a semiconductor packaging assembly line
P. Y. Mok 0001
Expert Syst. Appl.1
2009 Solving the two-dimensional irregular objects allocation problems by using a two-stage packing approach
Wai Keung Wong, X. X. Wang, P. Y. Mok 0001, Sunney Yung-Sun Leung, C. K. Kwong 0001
Expert Syst. Appl.3
2009 A fashion mix-and-match expert system for fashion retailers using fuzzy screening approach
Wai Keung Wong, X. H. Zeng, W. M. R. Au, P. Y. Mok 0001, Sunney Yung-Sun Leung
Expert Syst. Appl.4
2004 An ICA design of intraday stock prediction models with automatic variable selection
abstract
Independent component analysis (ICA) provides a mechanism of decomposing non-Gaussian data signals into statistically independent components. In this paper, ICA is used to extract the underlying news factors from intraday stock data. A, prediction algorithm is developed to improve stock index predictions using such extracted "news". Both linear regression model and nonlinear artificial neural network model are proposed to predict stock indexes of Open, Close, High and Low using the ICA extracted "news". These models are compared with models using only raw intraday data as "news". It is demonstrated that ICA helps in extracting market underlying affecting "news", and thus improves the stock prediction accuracy. It shows that the proposed ICA prediction algorithm is a simple to use and versatile algorithm that automatically extracts the most relevant news for different stock index predictions.
P. Y. Mok 0001, Kai-Pui Lam, H. S. Ng
IJCNN1