EDBT 2026 Demo / reviewers in the wild / expert
Yang Yang 0066
dblp:48/450-66
· DBLP profile ↗
47ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0001-8687-4427ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Diffusion for Sparse-to-Full Human Motion ReconstructionabstractHuman motion generation from sparse observations is an ill-posed problem in AR/VR, where head-mounted devices often capture only head and wrist trajectories. Prior methods usually reconstruct full-body motion in a single stage, forcing inference over a vast solution space and producing inaccurate lower-body motion, weak temporal coherence, and implausible sequences that degrade avatar embodiment. We presentMAGE, aMulti-stageAvatarGEnerator based on hierarchical diffusion. Instead of predicting 22-joint motion at once, MAGE progressively refines motion from a coarse 6-part representation to full joints. Each stage injects stage-specific motion priors and uses intermediate predictions to constrain subsequent refinement, reducing ambiguity and stabilizing dynamics. Experiments on large-scale motion datasets show that MAGE improves reconstruction accuracy, temporal smoothness, and perceptual realism over state-of-the-art baselines, enabling more reliable full-body animation from minimal AR/VR sensing while preserving real-time interaction. Fangyu Du, Yang Yang 0066, Xuehao Gao, Hongye Hou |
IEEE Signal Process. Lett. | 2 |
| 2025 | Jointly Understand Your Command and Intention: Reciprocal Co-Evolution Between Scene-Aware 3D Human Motion Synthesis and AnalysisabstractAs two intimate reciprocal tasks, scene-aware human motion synthesis and analysis require a joint understanding between multiple modalities, including 3D body motions, 3D scenes, and textual descriptions. In this paper, we integrate these two paired processes into a Co-Evolving Synthesis-Analysis (CESA) pipeline and mutually benefit their learning. Specifically, scene aware text-to-human synthesis generates diverse indoor motion samples from the same textual description to enrich human scene interaction intra-class diversity, thus significantly benefiting training a robust human motion analysis system. Reciprocally, human motion analysis would enforce semantic scrutiny on each synthesized motion sample to ensure its semantic consistency with the given textual description, thus improving realistic motion synthesis. Considering that real-world indoor human motions are goal-oriented and path-guided, we propose a cascaded generation strategy that factorizes text-driven scene-specific human motion generation into three stages: goal inferring, path planning, and pose synthesizing. Coupling CESA with this powerful cascaded motion synthesis model, we jointly improve realistic human motion synthesis and robust human motion analysis in 3D scenes. Xuehao Gao, Yang Yang 0066, Shaoyi Du, Guo-Jun Qi, Junwei Han 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Full-Dimensional Optimizable Network: A Channel, Frame and Joint-Specific Network Modeling for Skeleton-Based Action RecognitionabstractRecent human action recognition systems widely adopt graph convolution networks to extract spatial-temporal movement patterns. In graph convolution layers, inter-joint and inter-frame dependencies dominate spatial and temporal feature aggregation and thus are pivotal to representation learning. To enrich learned motion patterns, a powerful feature extractor should introduce its information propagation flexibility into three dimensions: (1) inferring different inter-joint correlations at different frames; (2) inferring different inter-frame correlations at different joints; (3) inferring different inter-joint and interframe correlations at different channels. In this paper, we take a closer look at effective feature aggregation in a skeleton sequence and propose a novel full-dimensional optimizable network with Channel, Frame and Joint-specific Network (CFJ-s Net) modeling for improving action recognition. By promoting dynamic information flows within different channels, frames, and joints, CFJ-s Net significantly extracts richer body posture features and trajectory features from a skeleton sequence. As verified on three large-scale datasets, NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA, CFJ-s Net achieves substantial improvements over state-of-the-art methods. Yang Yang 0066, Xuehao Gao, Shaoyi Du |
IJCNN | 2 |
| 2024 | Lightweight Graph Convolutional Network For Efficient Skeleton Based Action RecognitionabstractGraph convolutional network (GCN) has been widely used by skeleton based action recognition algorithms and achieves remarkable performance. However, recent GCN based State-Of-The-Art (SOTA) models for skeleton based action recognition tend to become increasingly sophisticated and over-parameterized. The low efficiency in model training and inference poses a challenge for their practical implementation in real-world scenarios. To address this issue, we construct a GCN based lightweight model for skeleton based action recognition, termed LightGCN. In this work we introduce an efficient convolutional neural network (CNN) structure to our temporal convolutional (TC) layer to extract temporal dynamics, effectively reducing model complexity. Furthermore, we propose a novel attention module that first extends the multi-spectral channel attention mechanism to the field of skeleton based action recognition, which preserves not only the lowest frequency information, but also useful information encoded by other frequency components, reducing the information loss during the channel compression. In order to further reduce the model complexity, we design a new compound scaling strategy to expand the model’s width and depth to different extent. This strategy enables the model to achieve an excellent balance between complexity and accuracy. On the two large-scale datasets, i.e., NTU RGB+D 60 and 120, our proposed LightGCN achieves 92.5% accuracy on the cross-subject benchmark of NTU 60 dataset, outperforming previous SOTA lightweight models and most heavyweight models, while needing 24.54% fewer parameters and 24.76% fewer flops than EfficientGCN-B4, which is the SOTA lightweight model. Yang Yang 0066, Xuehao Gao |
IJCNN | 2 |
| 2024 | Using Rotation-Invariant Point and Line Features for Image MatchingabstractIn recent years, convolutional neural networks (CNNs) have outperformed traditional approaches in image matching tasks. However, suffering from poor robustness against object rotations, conventional CNNs tend to extract angle-specific feature representations from given images. To this end, group CNNs improve conventional CNNs with symmetric group theory and thus benefit their rotation equivariance for powerful feature learning. Nonetheless, how to enrich extracted features with better discriminability is an under-explored challenge for group CNNs. In this paper, we propose a powerful rotation-invariant image matching method that combines point and line features to jointly improve rotation equivariance and discriminability. Specifically, we first characterize richer features from images by detecting their keypoints and lines. Then, we employ a group convolutional backbone to extract rotation-invariant descriptors from detected keypoints and lines. Finally, we develop inter-image and intra-image attention strategies to integrate point-level and line-level features from two images, significantly facilitating the two-image matching task. Extensive experiments verify that our method achieves state-of-the-art matching accuracy among existing methods on varying rotation image datasets and also shows competitive results when transferred to real-world image matching. Wenpeng Zheng, Yang Yang 0066, Xuehao Gao |
IJCNN | 2 |
| 2024 | Dig into Detailed Structures: Key Context Encoding and Semantic-based Decoding for Point Cloud Completion
Hongye Hou, Xuehao Gao, Yang Yang 0066 |
ACM Multimedia | 4 |
| 2024 | Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion PredictionabstractInferring 3D human motion is fundamental in many applications, including understanding human activity and analyzing one's intention. While many fruitful efforts have been made to human motion prediction, most approaches focus on pose-driven prediction and inferring human motion in isolation from the contextual environment, thus leaving the body location movement in the scene behind. However, real-world human movements are goal-directed and highly influenced by the spatial layout of their surrounding scenes. In this paper, instead of planning future human motion in a "dark" room, we propose a Multi-Condition Latent Diffusion network (MCLD) that reformulates the human motion prediction task as a multi-condition joint inference problem based on the given historical 3D body motion and the current 3D scene contexts. Specifically, instead of directly modeling joint distribution over the raw motion sequences, MCLD performs a conditional diffusion process within the latent embedding space, characterizing the cross-modal mapping from the past body movement and current scene context condition embeddings to the future human motion embedding. Extensive experiments on large-scale human motion prediction datasets demonstrate that our MCLD achieves significant improvements over the state-of-the-art methods on both realistic and diverse predictions. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du, Guo-Jun Qi |
IEEE Trans. Image Process. | 2 |
| 2024 | Learning Heterogeneous Spatial-Temporal Context for Skeleton-Based Action RecognitionabstractGraph convolution networks (GCNs) have been widely used and achieved fruitful progress in the skeleton-based action recognition task. In GCNs, node interaction modeling dominates the context aggregation and, therefore, is crucial for a graph-based convolution kernel to extract representative features. In this article, we introduce a closer look at a powerful graph convolution formulation to capture rich movement patterns from these skeleton-based graphs. Specifically, we propose a novel heterogeneous graph convolution (HetGCN) that can be considered as the middle ground between the extremes of (2 + 1)-D and 3-D graph convolution. The core observation of HetGCN is that multiple information flows are jointly intertwined in a 3-D convolution kernel, including spatial, temporal, and spatial-temporal cues. Since spatial and temporal information flows characterize different cues for action recognition, HetGCN first dynamically analyzes pairwise interactions between each node and its cross-space-time neighbors and then encourages heterogeneous context aggregation among them. Considering the HetGCN as a generic convolution formulation, we further develop it into two specific instantiations (i.e., intra-scale and inter-scale HetGCN) that significantly facilitate cross-space-time and cross-scale learning on skeleton graphs. By integrating these modules, we propose a strong human action recognition system that outperforms state-of-the-art methods with the accuracy of 93.1% on NTU-60 cross-subject (X-Sub) benchmark, 88.9% on NTU-120 X-Sub benchmark, and 38.4% on kinetics skeleton. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | GUESS: GradUally Enriching SyntheSis for Text-Driven Human Motion GenerationabstractIn this article, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strategy sets up generation objectives by grouping body joints of detailed skeletons in close semantic proximity together and then replacing each of such joint group with a single body-part node. Such an operation recursively abstracts a human pose to coarser and coarser skeletons at multiple granularity levels. With gradually increasing the abstraction level, human motion becomes more and more concise and stable, significantly benefiting the cross-modal motion synthesis task. The whole text-driven human motion synthesis problem is then divided into multiple abstraction levels and solved with a multi-stage generation framework with a cascaded latent diffusion model: an initial generator first generates the coarsest human motion guess from a given text description; then, a series of successive generators gradually enrich the motion details based on the textual description and the previous synthesized results. Notably, we further integrate GUESS with the proposed dynamic multi-condition fusion mechanism to dynamically balance the cooperative effects of the given textual condition and synthesized coarse motion prompt in different generation stages. Extensive experiments on large-scale datasets verify that GUESS outperforms existing state-of-the-art methods by large margins in terms of accuracy, realisticness, and diversity. Xuehao Gao, Yang Yang 0066, Zhenyu Xie, Shaoyi Du, Zhongqian Sun, Yang Wu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Decompose More and Aggregate Better: Two Closer Looks at Frequency Representation Learning for Human Motion PredictionabstractEncouraged by the effectiveness of encoding temporal dynamics within the frequency domain, recent human motion prediction systems prefer to first convert the motion representation from the original pose space into the frequency space. In this paper, we introduce two closer looks at effective frequency representation learning for robust motion prediction and summarize them as: decompose more and aggregate better. Motivated by these two insights, we develop two powerful units that factorize the frequency representation learning task with a novel decomposition-aggregation two-stage strategy: (1) frequency decomposition unit unweaves multi-view frequency representations from an input body motion by embedding its frequency features into multiple spaces; (2) feature aggregation unit deploys a series of intra-space and inter-space feature aggregation layers to collect comprehensive frequency representations from these spaces for robust human motion prediction. As evaluated on large-scale datasets, we develop a strong baseline model for the human motion prediction task that outperforms state-of-the-art methods by large margins: 8%∼12% on Human3.6M, 3%∼7% on CMU MoCap, and 7%∼10% on 3DPW. Xuehao Gao, Shaoyi Du, Yang Wu 0001, Yang Yang 0066 |
CVPR | 4 |
| 2023 | Glimpse and focus: Global and local-scale graph convolution network for skeleton-based action recognition
Xuehao Gao, Shaoyi Du, Yang Yang 0066 |
Neural Networks | 3 |
| 2023 | MTL-FaultNet: Seismic Data Reconstruction Assisted Multitask Deep Learning 3-D Fault InterpretationabstractSeismic fault interpretation is of extraordinary significant for hydrocarbon reservoir characterization and drilling hazard mitigation. In recent years, deep learning-based seismic fault detection methods have been conducted actively. Considering efficiency and fault prediction consistency, the most appealing way is to train a 3D segmentation network using synthetic seismic data with ground truth fault structure. However, the differences in signal-to-noise ratio, resolution, and fault strike distribution between synthetic and real data, can lead to inconsistent and unreliable prediction results. In this paper, we propose a multi-task deep learning-based seismic fault detection method, which takes seismic fault detection as the main task and 3D seismic data reconstruction as the auxiliary task, named MTL-FaultNet. The auxiliary branch can provide suggestive information to the main branch thereby improving its performance. We also designed two levels of multi-scale modules and embedded attention mechanisms in the network, so as to improve the network’s ability to focus on multi-scale fault features and learn stable fault structures. Different weights are assigned to the loss for different tasks, with large and small weights on the main and the auxiliary branch respectively. We apply the proposed method to Netherlands offshore F3 seismic data and a land field seismic data collected from Tarim Basin with mainly strike-slip faults, and Poseidon 3D seismic data. The proposed fault detection method is experimentally demonstrated on the improved network generalization and achieves reliable fault interpretation on field seismic data. Weihua Wu, Yang Yang 0066, Bangyu Wu, Debo Ma, Zhanxin Tang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Learning Relative Feature Displacement for Few-Shot Open-Set RecognitionabstractFew-shot learning (FSL) usually assumes that the query is drawn from the same label space as the support set, while queries from unknown classes may emerge unexpectedly in many open-world application scenarios. Such an open-set issue will limit the practical deployment of FSL systems, which remains largely unexplored. In this paper, we investigate the problem of few-shot open-set recognition (FSOR) and propose a novel solution, called Relative Feature Displacement Network (RFDNet), which empowers FSL systems to reject queries from unknown classes while accurately classifying those from known classes. First, we suggest a different relative feature displacement learning (RFDL) paradigm for FSOR, i.e., meta-learning a feature displacement relative to a pretrained reference feature embedding, based on our insightful observations on the randomness drift issue of previous meta-learning based for FSOR methods, as well as the generalization ability of the feature embedding pretrained for general classification. Second, we design the RFDNet framework to implement the RFDL paradigm, which is mainly featured by a task-aware RFD generator and a marginal open-set loss. Comprehensive experiments on three public datasets, i.e., miniImageNet, CIFAR-FS and tieredImageNet, demonstrate that RFDNet can consistently outperform the state-of-the-art methods, achieving improvement of 5.2%, 2.0% and 1.7% respectively, in terms of AUROC for unknown-class rejection under the 5-way 5-shot setting. Shule Deng, Jin-Gang Yu, Zihao Wu 0004, Hongxia Gao, Yansheng Li 0001, Yang Yang 0066 |
IEEE Trans. Multim. | 6 |
| 2023 | Efficient Spatio-Temporal Contrastive Learning for Skeleton-Based 3-D Action RecognitionabstractIn this paper, we propose a simple yet effective self-supervised method called spatio-temporal contrastive learning (ST-CL) for 3D skeleton-based action recognition. ST-CL acquires action-specific features by regarding the spatio-temporal continuity of motion tendency as the supervisory signal. To yield effective representations, ST-CL first designs some novel contrastive proxy tasks by providing different spatio-temporal observation scenes for the same 3D action and pulling them together in the embedding space. Second, three key components are devised in the action encoding to efficiently extract representations in contrastive tasks: (1) Information Representation introduces the awareness of joint type when analyzing motion dynamics. (2) Non-local GCN learns a data-driven graph topology structure and promotes a spatial message passing among long-range joints in each frame. (3) Multi-Scale TCN makes larger receptive fields for capturing richer longe-range temporal dynamics amomg adjacent frames. In ST-CL, these effective proxy tasks yield useful representations and efficient action encoding further enhances the representation capacity. As validated on four large-scale datasets, ST-CL is a strong baseline with high performance and efficiency for the contrastive learning study of the skeleton data. Compared to previous self-supervised methods, the proposed ST-CL achieves significant improvement consistently with a smaller model size and better training efficiency. Xuehao Gao, Yang Yang 0066, Maosen Li, Jin-Gang Yu, Shaoyi Du |
IEEE Trans. Multim. | 2 |
| 2022 | Motion Guided Attention Learning for Self-Supervised 3D Human Action Recognitionabstract3D human action recognition has received increasing attention due to its potential application in video surveillance equipment. To guarantee satisfactory performance, previous studies are mainly based on supervised methods, which have to add a large amount of manual annotation costs. In addition, general deep networks for video sequences suffer from heavy computational costs, thus cannot satisfy the basic requirement of embedded systems. In this paper, a novel Motion Guided Attention Learning (MG-AL) framework is proposed, which formulates the action representation learning as a self-supervised motion attention prediction problem. Specifically, MG-AL is a lightweight network. A set of simple motion priors (e.g., intra-joint variance, inter-frame deviation, intra-joint variance, and cross-joint covariance), which minimizes additional parameters and computational overhead, is regarded as a supervisory signal to guide the attention generation. The encoder is trained via predicting multiple self-attention tasks to capture action-specific feature representations. Extensive evaluations are performed on three challenging benchmark datasets (NTU-RGB+D 60, NTU-RGB+D 120 and NW-UCLA). The proposed method achieves superior performance compared to state-of-the-art methods, while having a very low computational cost. Yang Yang 0066, Xuehao Gao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Edge-Aware Superpixel Segmentation with Unsupervised Convolutional Neural NetworksabstractSuperpixels provide an efficient representation of images, and are applicable for subsequent vision tasks. In this paper, we propose an edge-aware superpixel algorithm based on an unsupervised convolutional neural network (CNN). Noticing that to adhere the boundaries of objects is one of the most essential characteristics of superpixels, we propose an entropy-based edge-aware term, which helps fit the differential model of the pixel-superpixel soft-assignment matrix predicted from CNN to image gradients, i.e. generate boundary-aligning superpixels. The proposed algorithm yields more boundary-adhering superpixels, and experimental results on BSDS500 show the effectiveness of the proposed edge-aware term. Yang Yang 0066, Kezhao Liu |
ICIP | 2 |
| 2021 | DWG-Reg: Deep Weight Global RegistrationabstractIn this paper, we propose a deep weight global registration (DWG-Reg) algorithm for poor initialization and partially overlapping point clouds registration problem. Our DWG-Reg is based on three modules: a bidirectional nearest search strategy for correspondence, a convolutional network for correspondence confidence prediction which consists of Hybird Distance Generator, optimal annealing Parameter Prediction network and a robust kernel function, a weighted optimizer algorithm for closed-form pose estimation. Experimental results show that our DWG-Reg achieves state-of-the-art performance compared to existing non-deep learning and recent deep learning methods. Our source code will open at https://github.com/BiaoBiaoLi/DWG-Reg. Qixing Xie, Shaoyi Du, Wenting Cui, Runzhao Yao, Yang Yang 0066, Jing Yang 0014, Lin Wang 0026 |
IJCNN | 6 |
| 2021 | Parallel Rotated-Invariant Palm Detection Network with Angle EstimationabstractAs one of the basic mission of computer vision, object detection is taken as the preprocessing strategy by an increasing number of algorithms. In the field of biometric information recognition, palm detection results seriously affect the accuracy of subsequent palm-print recognition. To facilitate practical application under the unconstrained condition, we propose a parallel rotated-invariant palm detection network to locate the position, size and especially the angle of in-plane of the region of interest(ROI). In order to obtain the low time complexity meanwhile maintaining the high accuracy, we choose the anchor-free algorithm CenterNet, which detects the object as a point with size, as our baseline. Otherwise, we design the specific Unit Circle Constraint(UCC) loss function to reduce the interference of boundary angle value to network training. The experimental results on the Rotated Palm Dataset demonstrate the effectiveness of our method. Guobin Zhang, Yang Yang 0066, Xuhui Tu |
SMC | 2 |
| 2021 | Robust registration algorithm based on rational quadratic kernel for point sets with outliers and noise
Runzhao Yao, Shaoyi Du, Teng Wan, Wenting Cui, Yang Yang 0066, Yang Jing, Ce Li 0001 |
Multim. Tools Appl. | 5 |
| 2021 | Point Set Registration With Similarity and Affine Transformations Based on Bidirectional KMPE LossabstractRobust point set registration is a challenging problem, especially in the cases of noise, outliers, and partial overlapping. Previous methods generally formulate their objective functions based on the mean-square error (MSE) loss and, hence, are only able to register point sets under predefined constraints (e.g., with Gaussian noise). This article proposes a novel objective function based on a bidirectional kernel mean p -power error (KMPE) loss, to jointly deal with the above nonideal situations. KMPE is a nonsecond-order similarity measure in kernel space and shows a strong robustness against various noise and outliers. Moreover, a bidirectional measure is applied to judge the registration, which can avoid the ill-posed problem when a lot of points converges to the same point. In particular, we develop two effective optimization methods to deal with the point set registrations with the similarity and the affine transformations, respectively. The experimental results demonstrate the effectiveness of our methods. Yang Yang 0066, Shaoyi Du, Muyi Wang, Badong Chen, Yue Gao 0002 |
IEEE Trans. Cybern. | 1 |
| 2020 | 3-D Oral Shape Retrieval Using Registration Algorithm
Wenting Cui, Shaoyi Du, Teng Wan, Yuying Liu 0007, Yang Yang 0066, Qingnan Mou, Mengqi Han, Yu-Cheng Guo |
MMM (2) | 6 |
| 2020 | Pamls Alignment Based On Two-Stage Convolutional Network with a Large in-Plane RotationabstractPalms alignment is an important work for palmprint recognition in uncontrolled environment. Many methods have made progress to achieve alignment. But most of them ignore the palm's angles, which could not satisfy the alignment initialization when the hand has a large in-plane rotation. In this paper, we propose a palms alignment with affine transformation method based on a two-stage convolutional neural network (CNN). The basic idea is to rotate the target palm into the same angle category to avoid the following affine registration has a big matching error at the beginning. At the stage I, the given target palm is classified into two angle categories. At the stage II the upside down palm is firstly rotated 180 degrees, and then inputted into the subsequent feature extraction network, feature matching layer and regression network to achieve the affine alignment. Experimental results have proved the effectiveness of our method. Yang Yang 0066, Guobin Zhang, Wenting Cui, Shaoyi Du |
SMC | 2 |
| 2020 | Multiple Facial Expressions Synthesis Driven by Editable Line MapsabstractFacial expression is an important facial semantics on visual aspect. The facial expressions synthesis has a wide range of applications in human-computer interaction and virtual reality. In recent years, image synthesis base on generative adversarial networks(GANs) is developing rapidly. In the image-to-image translation work, we propose a new facial expression generation method base on the idea of conditional GANs and realize the optimization of the generated results. The main work of this paper includes: Editable facial lines map is utilized as a constraint, combining with neutral face images as inputs of generator, so that a variety of facial expression images can be generated by editing the constraints. Correntropy loss of feature matching is added, which is used to measure the intermediate representation between the real images and the generated images by improving the adversarial loss. Consequently, the generated facial expressions can be more realistic. Base on the ideas above, the proposed method needs only one generator to generate different realistic facial images with various expressions. Dingdong Liu, Yang Yang 0066, Xiangyi Jing |
SMC | 2 |
| 2020 | Robust Point Set Registration Based on Semantic InformationabstractPoint cloud registration a challenging task in situations with poor initial value and scenarios with limited geometric structure. In these cases, the correct correspondence between two point clouds is unknown and difficult to establish. To cope with this problem, the semantic of partial points is introduced in this paper. Firstly, the semantic information is used to find more reasonable correspondence, i.e. semantic point pairs. Secondly, we formulate a novel objective function to integrate the matching error of semantic point pairs as guidance of registration. Thirdly, a hyperparameter is applied to balance the confidence of semantic point pairs. At last, a novel algorithm under the ICP framework is presented to optimize the rigid transformation iteratively. The evaluation of KITTI data set reveals the robustness and accuracy of our method in the complex scenes mentioned above. Qinlong Wang, Yang Yang 0066, Teng Wan, Shaoyi Du |
SMC | 2 |
| 2020 | Robust template matching with large angle localization
Yang Yang 0066, Weili Guan, Dexing Zhong, Meifeng Xu |
Neurocomputing | 1 |
| 2019 | Structured Down-Sampling and Registration Method for 3D Point Cloud of Indoor SceneabstractIn this paper, noted by the regular geometric structure of indoor scene, a biased down-sampling scheme is designed to automatically adjust the local sampling rate according to the local density and distribution. Our down-sampling results can effectively remain the main structure while greatly reduce the data number for the following work. Moreover, an improved Iterative Closest Point (ICP) algorithm for point clouds registration is proposed with the prior of structure information. Sampled structured data is weighted to give their contributions for registration. This leads the parameter estimation to naturally focus on aligning the structures of indoor scenes. The experimental results demonstrate the effectiveness of the proposed method on improving the registration accuracy with the same level of down-sampling data. Yang Yang 0066, Jing Yang 0014, Dexing Zhong |
SMC | 1 |
| 2019 | A Mahalanobis Distance-Based Fitness Approximation Method for Estimation of Distribution Algorithms in Solving Expensive Optimization ProblemsabstractFitness approximation methods have been widely employed in evolutionary algorithms to reduce the number of fitness evaluations in solving expensive optimization problems. As a simple and efficient approximation approach, k-nearest neighbors (kNN) estimates the fitness value of an unknown solution by combining the fitness values of its nearest neighbors according to a similarity measure. kNN generally adopts the Euclidean distance as the similarity measure, which may limit its performance as the solution distribution information is underutilized in the approximation process. Aiming at this issue, this study proposes a Mahalanobis distance-based k-nearest neighbors (MkNN) to improve the approximation accuracy by utilizing the distribution information. Compared to the Euclidean distance-based kNN (EkNN), MkNN adopts the Mahalanobis distance to measure the similarity between solutions, which is capable of capturing the distribution information of solutions and thus can improve the approximation efficiency. Furthermore, considering that the main idea of estimation of distribution algorithms (EDAs) is also to learn the distribution information of solutions, the proposed MkNN as well as EkNN are combined with an EDA and two new algorithms named EDA-MkNN and EDA-EkNN, respectively, are developed for expensive optimization. The performances of EDA-MkNN and EDA-EkNN were comprehensively tested on a set of 28 benchmark functions and compared with that of a typical EDA. Experimental results demonstrate that MkNN and EkNN could effectively improve the performance of EDA in solving different kinds of expensive optimization problems and MkNN can have an edge over EkNN on condition that the distribution information is well captured. Yongsheng Liang 0002, Yang Yang 0066, An Chen 0001, Daofu Guo, Bei Pang |
SMC | 3 |
| 2019 | Boosting Cooperative Coevolution for Large Scale Optimization With a Fine-Grained Computation Resource Allocation StrategyabstractCooperative coevolution (CC) has shown great potential for solving large-scale optimization problems (LSOPs). However, traditional CC algorithms often waste part of the computation resource (CR) as they equally allocate CR among all subproblems. The recently developed contribution-based CC algorithms improve the traditional ones to a certain extent by adaptively allocating CR according to some heuristic rules. Different from existing works, this paper explicitly constructs a mathematical model for the CR allocation (CRA) problem in CC and proposes a novel fine-grained CRA (FCRA) strategy by fully considering both the theoretically optimal solution of the CRA model and the evolution characteristics of CC. FCRA takes a single iteration as a basic CRA unit and always selects the subproblem which is most likely to make the largest contribution to the total fitness improvement to undergo a new iteration, where the contribution of a subproblem at a new iteration is estimated according to its current contribution, current evolution status, as well as the estimation for its current contribution. We verified the efficiency of FCRA by combining it with the success-history-based adaptive differential evolution which is an excellent DE variant but has never been employed in the CC framework. Experimental results on two benchmark suites for LSOPs demonstrate that FCRA significantly outperforms existing CRA strategies and the resulting CC algorithm is highly competitive in solving LSOPs. Yongsheng Liang 0002, Yang Yang 0066, Lin Wang 0026 |
IEEE Trans. Cybern. | 4 |
| 2018 | A global information based adaptive threshold for grouping large scale optimization problemsabstractBy taking the idea of divide-and-conquer, cooperative coevolution (CC) provides a powerful architecture for large scale global optimization (LSGO) problems, but its efficiency highly relies on the decomposition strategy. It has been shown that differential grouping (DG) performs well on decomposing LSGO problems by effectively detecting the interaction among decision variables. However, its decomposition accuracy highly depends on the threshold. To improve the decomposition accuracy of DG, a global information based adaptive threshold setting algorithm (GIAT) is proposed in this paper. On the one hand, by reducing the sensitivities of the indicator in DG to the roundoff error and the magnitude of contribution weight of subcomponent, we proposed a new indicator for two variables which is much more sensitive to their interaction. On the other hand, instead of setting the threshold only based on one pair of variables, the threshold is generated from the interaction information for all pair of variables. By conducting the experiments on two sets of LSGO benchmark functions, the correctness and robustness of this new indicator and GIAT were verified. An Chen 0001, Yang Yang 0066, Yongsheng Liang 0002, Bei Pang |
GECCO | 4 |
| 2018 | Dynamic Facial Expression Synthesis Driven by Deformable Semantic PartsabstractDynamic facial expression synthesis has some wild applications in human-computer interaction and virtual reality. The popular data-driven synthesis method like generative adversarial network (GAN) has made a great progress in generating a single face image, but has not well performed for expression sequences. To solve this problem, we design a series of deformable semantic parts to represent facial geometrical movement. And we synthesize the facial appearance by the geometrical driven under the-state-of-art pix2pixHD framework. In order to maintain the person identity among image sequence, we utilize an encoder to constrain the attributes of target face. With the above efforts, our method is capable to synthesize satisfied dynamic facial expression sequences. Nanxue Gong, Yang Yang 0066, Yuehu Liu, Dingdong Liu |
ICPR | 2 |
| 2018 | Registration of Color Point Cloud by Combining with Color Moments InformationabstractWith the development of RGB-D sensors, the highquality color point cloud can be obtained conveniently. Besides the geometrical information of point cloud, the color has great potential to assist the point cloud registration. In this paper, we propose a registration method by adaptively combining with color moment information to improve the registration accuracy. Firstly, three kinds of central moments are used to characterize the color distribution of two point clouds. And then, we build the correspondence between point clouds by dynamically combining the geometric feature and color moment feature of each point, hence the corresponding points satisfy both geometric similarity and color similarity. Finally, for the partial registration problem in practice, we apply the trimmed ICP algorithm framework to calculate the rigid transformation. Experimental results demonstrate that our algorithm is more robust and accurate in dealing with point cloud geometry defects, missing data and poor initial position. Weile Chen, Yang Yang 0066, Qian Kou |
SMC | 2 |
| 2018 | Precise Point Set Registration with Color Assisted and Correntropy for 3D ReconstructionabstractIterative closest point (ICP) algorithm, as its accuracy and efficiency, is widely used in rigid registration. However, ICP algorithm is easily failed when point sets lack of structure variety, such as semicircles. To solve this problem, a precise point set registration method for RGB-D data is proposed. Firstly, the color information provides a new information for registration, and the correntropy is introduced to deal with the noises and outliers. With color assisted and correntropy, a more robust objective function is built. Secondly, a variant ICP algorithm is used to deal with optimization problem via multiple iterations. Finally, as shown in the experimental results and scene reconstruction, our method obtains more precise results than other ICP algorithms. Teng Wan, Shaoyi Du, Yiting Xu, Guanglin Xu, Yang Yang 0066, Yue Gao 0002, Badong Chen |
SMC | 5 |
| 2018 | Building Correspondence Based on Matching Triangles for Partial RegistrationabstractAs an important problem in point set registration, partial registration has been solved by some variants of Iterative Closest Point (ICP) algorithm under good initial values. However, the initial parameters remained to be solved for partial registration. This paper presents a parameter initialization algorithm based on matching triangles for partial registration. Experimental results demonstrate that the proposed initialization method can find an appropriate initial transformation for next accurate registration, even the initial rotation angle between two sets is large. Based on the initialization of two point sets, the partial registration can be accomplished by auto trimmed ICP (ATICP) algorithm. Yiting Xu, Shaoyi Du, Teng Wan, Yang Yang 0066, Badong Chen, Yue Gao 0002 |
SMC | 4 |
| 2017 | An Iterative Feature-Pair Updating Framework for Rigid Template Matching with OutliersabstractTo deal with the rigid template matching problem in real-world scenarios, we propose a novel iterative feature-pair updating framework which is also robust to high levels of outliers, such as background changing, complex nonrigid deformation and partial occlusion. Given a pair of template image and target image, we first extract a set of corresponding feature-pairs as candidates. Then, we propose a robust objective function under the iterative framework for discriminatively updating these candidates, where the space distance, appearance distance, and the overlapping percentage of feature pairs are integrated simultaneously. Finally, a hierarchical matching strategy is provided with the parameter discussion. Experimental results compared with the-state-of-art methods on public data sets demonstrate the effectiveness of the proposed method. Yang Yang 0066, Qian Kou, Shaoyi Du, Yuehu Liu, Bangyu Wu |
ISM | 1 |
| 2017 | Robust 2D point set matching with Kernel mean P-power error lossabstractIn this paper, we propose a novel point set matching algorithm to improve the matching precision in the presence of non-Gaussian noises and outliers. In our method, a non-second order similarity measure known as Kernel Mean p-Power Error (KMPE) loss is employed as the matching cost function. We introduce a local optimal solution for computing the rigid transform by repeating the correspondence estimation and parameter updating processes. This new algorithm assigns a non-linear distance evaluation in kernel space according to the current estimation of the correspondence to yield a more accurate matching result between two point sets in practice. Experimental results demonstrate that our algorithm is more robust and accurate than the traditional ICP and the state-of-the-art algorithms. Yang Yang 0066, Weile Chen, Badong Chen, Shaoyi Du |
SMC | 1 |
| 2016 | A Modified Non-rigid ICP Algorithm for Registration of Chromosome Images
Qian Kou, Yang Yang 0066, Shaoyi Du, Dongge Cai |
ICIC (2) | 2 |
| 2016 | Robust image registration with rotation, scale and translation using Best-Buddies PairsabstractSince the complex outliers caused by the background and non-rigid deformation, image registration remains a challenging task. This paper proposes a new method for image registration, which is robust to the transformations with rotation, scale and translation. Our method is under the general Iterative Closet Point (ICP) framework which contains two main parts: finding the corresponding points and updating transformation parameters. In our method, to be robust to outliers, Best-Buddies Pairs (BBPs) are used as similarity measure between two images to obtain the corresponding points. While, in order to be invariant to scale transformation, Scaling ICP is employed to update transformation parameters. Besides the image registration, we also apply the method for object tracking in video sequences. Experimental results demonstrate the effectiveness and efficiency of the proposed method. Yang Yang 0066, Shaoyi Du, Qian Kou |
IJCNN | 2 |
| 2016 | A novel image enhancement method using fuzzy Sure entropy
Ce Li 0001, Yang Yang 0066, Limei Xiao, Yannan Zhou, Jizhong Zhao |
Neurocomputing | 2 |
| 2016 | Nonlinear deformation learning for face alignment across expression and pose
Yang Yang 0066, Yuanqi Su, Dongge Cai, Meifeng Xu |
Neurocomputing | 1 |
| 2015 | Pose Estimation for Vehicles Based on Binocular Stereo Vision in Urban Traffic
Fei Wang 0008, Yicong He, Hang Dong 0001, Haiwei Yang, Yang Yang 0066 |
ICIC (1) | 6 |
| 2013 | Automatic face image annotation based on a single template with constrained warping deformationabstractIn this study, an automatic face image annotation method is proposed by aligning faces with different expressions to an annotated neutral face. This work is useful in reducing tedious manual work for labelling image data in large databases. However, it is challenging because of the appearance variations caused by non‐rigid face deformations under various expressions. Unlike some conventional approaches acquiring sufficient image templates to model the query appearance, only a single given template is necessary for the proposed method. The authors address the problem through dense image alignment. Specifically, image warping in the alignment process is constrained by prior knowledge about facial shape deformation. The proposed method is independent of the appearance model, and is available for unseen faces. In addition, to initialise warping parameters, the authors present a robust patch‐based estimation method. Context information for feature points is carefully modelled to propagate the searching path for local patch matching. The face annotation experiments are performed on some large expressions, with noisy image qualities and in low image resolutions. Comparison results with conventional methods demonstrate the proposed method's superiority on both accuracy and robustness. Yang Yang 0066, Yuehu Liu |
IET Comput. Vis. | 1 |
| 2011 | 3D facial mesh detection using geometric saliency of surfaceabstractThis paper proposes a 3D facial mesh detection algorithm based on the geometric saliency of surface. Specifically, the geometric saliency of each vertex on 3D triangle mesh is measured by the combination of Gaussian-weighted curvature and spin-image correlation. Salient vertices with similar properties are clustered into regions on the saliency map, and represented as nodes by the graph model. To detect a 3D facial mesh, initialization and registration steps are applied to match each triangle in the graph model with a reference graph, corresponding to a 3D reference facial mesh. Furthermore, the match error between the graph model of the testing 3D mesh and the reference facial mesh is computed to classify face and non-face meshes. Experimental results demonstrate that the proposed algorithm is effective to detect 3D facial meshes and robust to facial expressions and geometric noises. Yaochen Li, Yuehu Liu, Yuanchun Wang 0003, Zhengwang Wu, Yang Yang 0066 |
ICME | 5 |
| 2011 | Expression transfer for facial sketch animation
Yang Yang 0066, Nanning Zheng 0001, Yuehu Liu, Shaoyi Du, Yuanqi Su, Yoshifumi Nishio |
Signal Process. | 1 |
| 2010 | Optimal Trajectory Space Finding for Nonrigid Structure from Motion
Yuanqi Su, Yuehu Liu, Yang Yang 0066 |
ACIVS (1) | 3 |
| 2009 | A Statistical-Structural Constraint Model for Cartoon Face Wrinkle Representation and Generation
Ping Wei 0001, Yuehu Liu, Nanning Zheng 0001, Yang Yang 0066 |
ACCV (3) | 4 |
| 2009 | Categorization of Multiple Objects in a Scene without Semantic Segmentation
Lei Yang 0063, Nanning Zheng 0001, Yang Yang 0066, Jie Yang 0001 |
ACCV (1) | 4 |
| 2009 | Example-based performance driven facial shape animationabstractA novel performance driven facial shape animation method is presented for mapping the expressions from the source face to the target face automatically. Unlike the prior expression cloning approaches, the proposed method aims to animate a new target face with the help of real facial expression samples. The basic idea is to learn the shape deformation from samples for target face to generate corresponding expressions. The process consists of two main stages. First of all, source motion vectors are transferred by statistic face model to generate a reasonable expression on the target face. And then, local deformation constraints are proposed to refine the animation results. In the second part, the local deformation characters for each target organ are learned from the samples, which preserve the personality as well as the expression styles. Experimental results on different facial animation demonstrate the feasibility and effectiveness of the proposed method. Yang Yang 0066, Nanning Zheng 0001, Yuehu Liu, Shaoyi Du, Yoshifumi Nishio |
ICME | 1 |