Liang Chang 0001

dblp:72/6746-1 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-9450-5960ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Representation Learning with Mutual Influence of Modalities for Node Classification in Multi-Modal Heterogeneous Networks
abstract
Nowadays, numerous online platforms can be described as multi-modal heterogeneous networks (MMHNs), such as Douban's movie networks and Amazon's product review networks. Accurately categorizing nodes within these networks is crucial for analyzing the corresponding entities, which requires effective representation learning on nodes. However, existing multi-modal fusion methods often adopt either early fusion strategies which may lose the unique characteristics of individual modalities, or late fusion approaches overlooking the cross-modal guidance in GNN-based information propagation. In this paper, we propose a novel model for node classification in MMHNs, named Heterogeneous Graph Neural Network with Inter-Modal Attention (HGNN-IMA). It learns node representations by capturing the mutual influence of multiple modalities during the information propagation process, within the framework of heterogeneous graph transformer. Specifically, a nested inter-modal attention mechanism is integrated into the inter-node attention to achieve adaptive multi-modal fusion, and modality alignment is also taken into account to encourage the propagation among nodes with consistent similarities across all modalities. Moreover, an attention loss is augmented to mitigate the impact of missing modalities. Extensive experiments validate the superiority of the model in the node classification task, providing an innovative view to handle multi-modal data, especially when accompanied with network structures. The full version including Appendix is available at http://arxiv.org/abs/2505.07895.
Jiafan Li, Jiaqi Zhu 0001, Liang Chang 0001, Yilin Li 0003, Miaomiao Li 0007, Yang Wang 0102, Yi Yang 0060, Hongan Wang
IJCAI3
2024 Scattering-based hybrid network for facial attribute classification
Fan Zhang 0062, Liang Chang 0001, Fuqing Duan
Frontiers Comput. Sci.3
2023 Degradation Conditioned GAN for Degradation Generalization of Face Restoration Models
abstract
Face restoration models are usually trained on synthetic degraded data to output an image that matches the clean version of itself. Most previous methods use a single model to deal with all the degradation levels, resulting in a domain generalization problem. We explore the value of degradation information and propose a Degradation Conditioned GAN (DeCGAN). The architecture consists of modulated convolution, bias, and fusion modules, inspired by deblurring, denoising and super-resolution. The whole network can be modulated by the degradation levels to achieve delicate and precise restoration effects. Experiments are conducted on conventional and modulated face restoration tasks. DeCGAN can achieve more faithful restoration and better metrics (FID, LPIPS, etc.) than previous methods do. Moreover, our model performs well on real-world low-quality face images.
Qi Song 0003, Wu Shi, Guojing Ge, Liang Chang 0001
ICIP4
2023 Facial attribute classification by deep mining inter-attribute correlations
abstract
Abstract Face attribute classification (FAC) has received considerable attention due to its excellent application value in bio‐metric verification and face retrieval. Current FAC methods suffer two typical challenges: complex inter‐attribute correlations and imbalanced learning. Aims at the challenges, presents an end‐to‐end FAC framework with integrated use of multiple strategies, which consists of a convolutional neural network (CNN) and a graph convolutional network (GCN). The GCN is used to model the semantic correlations among attributes and capture inter‐dependency among them. The correlation information learnt via the GCN is used to guide the learning of the inter‐dependent classification features of the FAC network. An adaptive thresholding strategy and a boosting scheme are adopted to alleviate the effect of the class‐imbalance. To deal with the task imbalance problem, a new dynamic weighting scheme is proposed to update the weight of each attribute classification task in the training process. We apply four evaluation metrics to evaluate the proposed method. Experimental results show all the proposed strategies are effective, and our approach outperforms state‐of‐the‐art FAC methods on two challenging datasets CelebA and LFWA.
Fan Zhang 0062, Liang Chang 0001, Fuqing Duan
IET Comput. Vis.3
2023 Recurrent 3D Hand Pose Estimation Using Cascaded Pose-Guided 3D Alignments
abstract
3D hand pose estimation is a challenging problem in computer vision due to the high degrees-of-freedom of hand articulated motion space and large viewpoint variation. As a consequence, similar poses observed from multiple views can be dramatically different. In order to deal with this issue, view-independent features are required to achieve state-of-the-art performance. In this paper, we investigate the impact of view-independent features on 3D hand pose estimation from a single depth image, and propose a novel recurrent neural network for 3D hand pose estimation, in which a cascaded 3D pose-guided alignment strategy is designed for view-independent feature extraction and a recurrent hand pose module is designed for modeling the dependencies among sequential aligned features for 3D hand pose estimation. In particular, our cascaded pose-guided 3D alignments are performed in 3D space in a coarse-to-fine fashion. First, hand joints are predicted and globally transformed into a canonical reference frame; Second, the palm of the hand is detected and aligned; Third, local transformations are applied to the fingers to refine the final predictions. The proposed recurrent hand pose module for aligned 3D representation can extract recurrent pose-aware features and iteratively refines the estimated hand pose. Our recurrent module could be utilized for both single-view estimation and sequence-based estimation with 3D hand pose tracking. Experiments show that our method improves the state-of-the-art by a large margin on popular benchmarks with the simple yet efficient alignment and network architectures.
Xiaoming Deng 0001, Dexin Zuo, Yinda Zhang 0001, Zhaopeng Cui, Jian Cheng 0006, Ping Tan 0002, Liang Chang 0001, Marc Pollefeys, Sean Ryan Fanello, Hongan Wang
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 LAGAN: Landmark Aided Text to Face Sketch Generation
Wentao Chao, Liang Chang 0001, Fangfang Xi, Fuqing Duan
PRCV (4)2
2022 Human-object interaction detection via interactive visual-semantic graph learning
Tongtong Wu, Fuqing Duan, Liang Chang 0001, Ke Lu 0002
Sci. China Inf. Sci.3
2021 Multi-branch Graph Network for Learning Human-Object Interaction
Tongtong Wu, Fuqing Duan, Liang Chang 0001
PRCV (4)4
2021 Hand Pose Understanding With Large-Scale Photo-Realistic Rendering Dataset
abstract
Hand pose understanding is essential to applications such as human computer interaction and augmented reality. Recently, deep learning based methods achieve great progress in this problem. However, the lack of high-quality and large-scale dataset prevents the further improvement of hand pose related tasks such as 2D/3D hand pose from color and depth from color. In this paper, we develop a large-scale and high-quality synthetic dataset, PBRHand. The dataset contains millions of photo-realistic rendered hand images and various ground truths including pose, semantic segmentation, and depth. Based on the dataset, we firstly investigate the effect of rendering methods and used databases on the performance of three hand pose related tasks: 2D/3D hand pose from color, depth from color and 3D hand pose from depth. This study provides insights that photo-realistic rendering dataset is worthy of synthesizing and shows that our new dataset can improve the performance of the state-of-the-art on these tasks. This synthetic data also enables us to explore multi-task learning, while it is expensive to have all the ground truth available on real data. Evaluations show that our approach can achieve state-of-the-art or competitive performance on several public datasets.
Xiaoming Deng 0001, Yinda Zhang 0001, Yuying Zhu 0002, Dachuan Cheng, Dexin Zuo, Zhaopeng Cui, Ping Tan 0002, Liang Chang 0001, Hongan Wang
IEEE Trans. Image Process.9
2020 Face-sketch learning with human sketch-drawing order enforcement
Liang Chang 0001, Lihua Jin, Lifen Weng, Wentao Chao, Xiaoming Deng 0001, Qiulei Dong
Sci. China Inf. Sci.1
2020 Leveraging 3D blendshape for facial expression recognition using CNN
Sa Wang, Zhengxin Cheng, Xiaoming Deng 0001, Liang Chang 0001, Fuqing Duan, Ke Lu 0002
Sci. China Inf. Sci.4
2020 Edge-guided single facial depth map super-resolution using CNN
abstract
In recent years, consumer depth cameras have been widely used in digital entertainment and human‐machine interaction due to the advantages of real‐time performance and low cost. Facial depth maps have shown great potential in 3D‐face‐related studies. However, disadvantages of low resolution and precision limit its further applications. In this work, the authors propose an edge‐guided convolutional neural network for single facial depth map super‐resolution. It consists of two parts: an edge prediction sub‐network and a depth reconstruction sub‐network. The edge prediction sub‐network generates an edge guidance map to guide the depth reconstruction sub‐network to recover sharp edges and fine structures. Effective data augmentation methods are proposed as well. The network is patch‐based and able to cope with any size of the input depth maps. In addition, it is insensitive to the face pose since the synthetic training dataset they generated covers a wide range of face poses. The proposed method is validated with three datasets including a synthetic facial depth data set, a real Kinect V2 facial depth data set and Middlebury Stereo Data set. Experimental results show that it outperforms the state‐of‐the‐art methods on all the three data sets.
Fan Zhang 0062, Liang Chang 0001, Fuqing Duan, Xiaoming Deng 0001
IET Image Process.3
2019 Cascaded Point Network for 3D Hand Pose Estimation
abstract
Recent PointNet-family hand pose methods have the advantages of high pose estimation performance and small model size, and it is a key problem to get effective sample points for PointNet-family methods. In this paper, we propose a two-stage coarse to fine hand pose estimation method, which belongs to PointNet-family methods and explores a new sample point strategy. In the first stage, we use 3D coordinate and surface normal of normalized point cloud as input to regress coarse hand joints. In the second stage, we use the hand joints in the first stage as the initial sample points to refine the hand joints. Experiments on widely used datasets demonstrate that using joints as sample points is more effective and our method achieves top-rank performance.
Yikun Dou, Yuying Zhu 0002, Xiaoming Deng 0001, CuiXia Ma, Liang Chang 0001, Hongan Wang
ICASSP6
2019 PGR-Net: A Parallel Network Based on Group and Regression for Age Estimation
abstract
Age is an important biometric feature of human face. Estimating the specific age of facial images is challenging, because of face aging's highly nonlinearity and randomness. Commonly age predictors are based on classification or regression method, which may be affected greatly by the category number or data distribution of the labelled samples. In this paper, we design a parallel deep neural network, called PGR-Net. It is a unified learning model which combines the merits of traditional classification methods and regression methods. The model consists of a classification network and several age regressors. The classification network is designed to divide facial images into several age groups, and a regressor is trained for each group separately. We train the classification network and the re-gressors in parallel, and perform age estimation with the regressor of the group predicted by the classification network. Experiments show that the proposed approach is fairly competitive compared with the state-of-the-art methods on two public datasets.
Liang Chang 0001, Fuqing Duan
ICASSP2
2019 High-Fidelity Face Sketch-To-Photo Synthesis Using Generative Adversarial Network
abstract
Face sketch-photo synthesis has important usage in law enforcement and human authentication. Due to the sparse information (no color or texture), the abstraction level, the diversity of sketches, and the domain gap between sketch and photo, it is challenging to synthesize a photo-realistic photo from an input sketch. Moreover, the deficiency of data also restricts the synthesis performance. In this paper, we present a high-fidelity face sketch-photo synthesis method using Generation Adversarial Network (GAN). Our network adopts a deep residual U-Net as generator and a Patch-GAN with residual blocks as discriminator. We design effective loss functions by enforcing pixels, edges and high-level features of the produced face photos. Moreover, we augment the CUHK sketch dataset using an effective sampling method. With the improved GAN and augmented dataset, we achieve high-fidelity face photos. Qualitative and quantitative experiments demonstrate the approach outperforms other method. Further experiments with a sketch-based photo editing application also validate the performance of our method.
Wentao Chao, Liang Chang 0001, Jian Cheng 0006, Xiaoming Deng 0001, Fuqing Duan
ICIP2
2018 Text2Sketch: Learning Face Sketch from Facial Attribute Text
abstract
Face sketch is the main approach to find suspect in law enforcement, especially in many cases when facial attribute descriptions of suspects by witnesses are available. Face sketch synthesized from facial attribute text can also be used in sketch based face recognition. While most previous work focus on face photo to sketch synthesis, the problem of sketch synthesis with facial attribute text has not been explored yet. The problem is challenging due to two facts: firstly, no database of face attribute text to sketch is available; secondly, it is hard to synthesize high-quality face sketches due to the ambiguity and complexity of text description. In this paper, we propose a face sketch synthesis approach with text using Stagewise-GAN. Our contributions lie in two aspects: 1) we construct the first text to face sketch database. The database, namely Text2Sketch dataset, is annotated with CUFSF dataset of 1194 sketches. For each sketch, an attribute description is labelled; 2) we synthesize vivid face sketches using Stagewise-GAN. We use user study, face retrieval performance with synthesized sketch, and quantitative results for evaluation. Experimental results show the effectiveness of our approach.
Liang Chang 0001, Lihua Jin, Zhengxin Cheng, Xiaoming Deng 0001, Fuqing Duan
ICIP2
2018 User-Invariant Facial Animation with Convolutional Neural Network
Shuiquan Wang, Zhengxin Cheng, Liang Chang 0001, Xuejun Qiao, Fuqing Duan
ICONIP (1)3
2018 Joint Hand Detection and Rotation Estimation Using CNN
abstract
Hand detection is essential for many hand related tasks, e.g., recovering hand pose and understanding gesture. However, hand detection in uncontrolled environments is challenging due to the flexibility of wrist joint and cluttered background. We propose a convolutional neural network (CNN), which formulates in-plane rotation explicitly to solve hand detection and rotation estimation jointly. Our network architecture adopts the backbone of faster R-CNN to generate rectangular region proposals and extract local features. The rotation network takes the feature as input and estimates an in-plane rotation which manages to align the hand, if any in the proposal, to the upward direction. A derotation layer is then designed to explicitly rotate the local spatial feature map according to the rotation network and feed aligned feature map for detection. Experiments show that our method outperforms the state-of-the-art detection models on widely-used benchmarks, such as Oxford and Egohands database. Further analysis show that rotation estimation and classification can mutually benefit each other.
Xiaoming Deng 0001, Yinda Zhang 0001, Shuo Yang 0002, Ping Tan 0002, Liang Chang 0001, Hongan Wang
IEEE Trans. Image Process.5
2015 Face sketch synthesis using non-local means and patch-based seaming
abstract
This paper proposed a face sketch synthesis method by using non-local means (NL-Means), which takes the advantage of the non-local self-similarity of face photo and sketch patches. With a learning database of individuals described by one face photo and one face sketch, we assume that, for a given individual, the NL-Means coefficient of a given face photo patch is the same as its corresponding sketch patch. In order to handle the visible seam due to intensity difference of neighbor overlapping patches, we use patch based optimal seam to enforce the consistency of synthesized overlapping sketch patches. Experimental results on CUHK Face Sketch Database illustrate that our method has the advantage of easy implementation and much less required training samples, meanwhile our method can achieve fairly competitive synthesis results.
Liang Chang 0001, Yves Rozenholc, Xiaoming Deng 0001, Fuqing Duan
ICIP1
2014 Motion estimation of multiple depth cameras using spheres
abstract
Automatic motion estimation of multiple depth cameras has remained a challenging topic in computer vision due to its reliance on the image correspondence problem. In this paper, spherical objects are employed to estimate motion parameters between multiple depth cameras. We move a sphere several times in the common view of depth cameras. We fit the spherical point clouds to get the sphere centers in each depth camera system, and then introduce a factorization based approach to estimate motions between the depth cameras. Both simulated and real experiments show the robustness and effectiveness of our method.
Xiaoming Deng 0001, Jie Liu 0029, Feng Tian 0001, Liang Chang 0001, Hongan Wang
ICIP4
2014 Automatic Gait Motion Capture with Missing-Marker Fillings
abstract
Although marker-based optical motion capture has been a useful method for computer animation during the past decades, automatic and robust motion tracking from multiple video sequences is still very challenging. Several critical issues in practical implementations are not adequately addressed. For example, how to track and identify the reconstructed 3D points after image matching process? How to handle the heavy occlusion problem? This paper gives a careful investigation of the above issues. In particular, we propose a novel way to track and identify proper markers, and a new method of filling missing markers by taking account of the human model constraints. Experiments are presented to show its accuracy and robustness.
Xiaoming Deng 0001, Shihong Xia, Wenzhong Wang, Liang Chang 0001, Hongan Wang
ICPR5
2014 Craniofacial reconstruction based on least square support vector regression
abstract
Craniofacial reconstruction is to get a visual outlook of an individual from its skull. It is an important technology in both forensic medicine and archeology. This paper proposes a novel craniofacial reconstruction method based on least square support vector regression (LSSVR), which has the flexibility for uncovering nonlinear relationships between variables and is easy to solve. We firstly build statistical shape models for skulls and face skins respectively, and then train the LSSVR model in the shape parameter spaces. Given an unknown skull, we project it to the skull shape parameter space, and use the LSSVR model to reconstruct the corresponding face skin. Cross validation is used for parameter selection in LSSVR. Experiments are given on a data set including 150 training pairs of skull and skin samples and 58 testing ones. Comparisons with ridge regression and partial least square regression show that our method can reconstruct the craniofacial effectively and accurately.
Yan Li 0121, Liang Chang 0001, Xuejun Qiao, Fuqing Duan
SMC2
2012 Smoothness-constrained face photo-sketch synthesis using sparse representation
Liang Chang 0001, Xiaoming Deng 0001, Fuqing Duan, Zhongke Wu
ICPR1
2012 Self-calibration of hybrid central catadioptric and perspective cameras
Xiaoming Deng 0001, Fuchao Wu, Yihong Wu 0002, Fuqing Duan, Liang Chang 0001, Hongan Wang
Comput. Vis. Image Underst.5
2011 Calibration of central catadioptric camera with one-dimensional object undertaking general motions
abstract
AID object is a segment with several known-distance markers, and calibration methods with ID objects are more flexible than those with 2D/3D objects. Under the pinhole camera model, it is proved that the calibration with free-moving ID objects is not possible. For a central catadioptric camera setup, can the camera be calibrated by a ID object under general motions? In this paper, we prove that a central catadioptric camera can indeed be calibrated, and propose a catadioptric camera calibration method using ID objects undertaking general motions. The proposed method consists of two steps. Firstly, the principal point is calculated with geometric invariants under catadioptric camera model; Secondly, we use images of ID object to calibrate the focal lengths, skew factor and mirror parameter. The method needs neither prior knowledge of catadioptric parameters nor conic fitting, and it is linear, which makes it easy to implement. Experiments demonstrate its usefulness and stability.
Xiaoming Deng 0001, Fuchao Wu, Yihong Wu 0002, Liang Chang 0001, Wei Liu 0023, Hongan Wang
ICIP4
2010 Face Sketch Synthesis via Sparse Representation
abstract
Face sketch synthesis with a photo is challenging due to that the psychological mechanism of sketch generation is difficult to be expressed precisely by rules. Current learning-based sketch synthesis methods concentrate on learning the rules by optimizing cost functions with low-level image features. In this paper, a new face sketch synthesis method is presented, which is inspired by recent advances in sparse signal representation and neuroscience that human brain probably perceives images using high-level features which are sparse. Sparse representations are desired in sketch synthesis due to that sparseness can adaptively selects the most relevant samples which give best representations of the input photo. We assume that the face photo patch and its corresponding sketch patch follow the same sparse representation. In the feature extraction, we select succinct high-level features by using the sparse coding technique, and in the sketch synthesis process each sketch patch is synthesized with respect to high-level features by solving an l1-norm optimization. Experiments have been given on CUHK database to show that our method can resemble the true sketch fairly well.
Liang Chang 0001, Yanjun Han, Xiaoming Deng 0001
ICPR1
2006 An Improved Gilbert Algorithm with Rapid Convergence
abstract
Gilbert algorithm is a very popular algorithm in collision detection in robotics and also in classification in pattern recognition. However, the major drawback of Gilbert algorithm is that in many cases it becomes very slow as it approaches the final solution and the vertices selection vibrates in these cases. In this paper: a) It is proven theoretically that when the selection of vertices vibrates among several points, the algorithm will converge to the hyperplane determined by these points. b) Based on the above results, an improved Gilbert algorithm for computing the distance between two convex polytopes is presented. The algorithm can avoid the slow convergence of the original one. Numerical simulation results demonstrate the effectiveness and advantage of the improved algorithm
Liang Chang 0001, Hong Qiao, Anhua Wan, John A. Keane
IROS1