Jun Cheng 0002

dblp:78/5816-2 · DBLP profile ↗
← Back
173ranked-venue papers
15as first author
91since 2021 · last 2026
0000-0002-3131-3275ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 4 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 9 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 6 since 2021Systems, architecture and hardware · 11 · 1 first-author · 8 since 2021Computer networks · 10 · 2 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Enhancing robustness of neural ODEs via attractor dynamics
Qieshi Zhang, Jun Cheng 0002
Neurocomputing3
2026 Diffusion-Based Text-Guided Image Generation With Fine-Grained Spatial Object-Attribute Relationships
abstract
Expressing and controlling fine-grained spatial attributes of objects in large-scale models presents significant challenges, as these spatial attributes are often difficult to describe textually and exhaustive enumeration is impractical. This hinders effective alignment with user preferences regarding spatial attribute-object relationships in fine-grained synthesis tasks. To tackle this problem, we propose AttrObjDiff, a novel framework built on the pre-trained Stable Diffusion model to integrate spatial attribute maps. Firstly, AttrObjDiff constrains the denoising step using trainable cross-attention fusion modules, attribute-enhancing cross-attention and LoRAs. The fusion modules take layout features extracted by a frozen ControlNet and corresponding fine-grained attribute maps as inputs to generate joint constraint features of spatial attribute-object relationships. We leverage attribute-enhancing cross-attention within the U-Net to further refine these spatial attributes. Finally, LoRAs are employed to align with these joint constraint features of finegrained relationships. Secondly, AttrObjDiff enhances the reverse process with lightweight noise reranking models to improve spatial object-attribute alignment. The reranking models select semantic noises related to fine-grained relationships, improving synthesis quality without significantly increasing computational costs. Experimental results demonstrate that our method can generate high-quality images guided by fine-grained spatial object-attribute relationships, improving synthesis controllability and semantic consistency.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Ziliang Ren, Dacheng Tao, Xinyu Wu 0001, Jun Cheng 0002
IEEE Trans. Circuits Syst. Video Technol.7
2026 Robust Nonfragile Consensus Control of MASs With Controller Gain Perturbations and Switching Directed Networks
abstract
This article investigates the robust nonfragile leaderless consensus control issues of nonlinear multiagent systems (MASs) in the presence of controller gain perturbations, external interferences, and switching directed networks. A novel distributed nonfragile consensus controller is first devised. Subsequently, on the basis of the property that an MAS directed network's Laplacian matrix can be broken down into the product of two particular matrices, the conversion from the consensus control issue to the asymptotic stability control issue is achieved via two variable substitutions related to the above property. Additionally, a sufficient condition, which can guarantee the MASs' asymptotic stability, is proposed and proved by Lyapunov stability theory and algebraic graph theory. Finally, the validity of the devised method is demonstrated by a simulation example.
Jian Liao 0006, Bin Xin 0002, Qing Wang 0010, Jun Cheng 0002
IEEE Trans. Cybern.5
2026 Phy-HHMR: Physics-Aware Holistic Human Mesh Reconstruction
Binhong Ye, Lei Wang 0018, Baoyu Liu, Xiaoliang Ma 0001, Jun Cheng 0002
IEEE Trans. Hum. Mach. Syst.6
2026 Speech2Blend: A Hybrid Network for Speech-Driven 3-D Facial Animation by Learning Blendshape
abstract
Recent advances in speech-driven facial animation have attracted significant interest across computer graphics, human–computer interaction systems, and immersive virtual reality applications. However, existing methods remain constrained by dependencies on specific reference videos or proprietary face mesh structures, limiting their applicability across diverse production pipelines and reducing compatibility with industry-standard animation workflows. To overcome these fundamental limitations in generalization and deployment flexibility, we propose Speech2Blend—an end-to-end hybrid convolutional-recurrent network that directly learns nonlinear speech-to-blendshape parameter mappings. This novel approach enables markerless speech-driven facial animation generation without restrictive inputs like video references or specialized facial rigs. Trained on the largest available digital human dataset (BEAT) and rigorously evaluated using three benchmark datasets with photorealistic visualization tools, Speech2Blend achieves state-of-the-art performance. It delivers superior audio-visual synchronization through learned temporal dynamics and reduces lip vertex error by 30% compared to existing baseline methods. These advances significantly lower production costs for virtual human speech animation while enabling cross-platform compatibility with common game engines and animation software.
Lei Wang 0018, Gongbin Chen, Feng Liu 0013, Jiaji Wu, Jun Cheng 0002
IEEE Trans. Hum. Mach. Syst.5
2026 NoisePO: Efficient Semantic Noise Generation and Ranking for Diffusion-Based Text-to-Image Synthesis
abstract
Diffusion-based methods have achieved remarkable success in photorealistic image generation, leveraging iterative denoising steps to improve image quality. However, multi-step denoising often suffers from error accumulation-similar to exposure bias in autoregressive models-due to suboptimal noise estimation, which can lead to degraded semantic alignment and image fidelity. To tackle the challenge of suboptimal inner latent representations in generation and improve the inner latent, this paper introduces a novel method NoisePO, an efficient semantic noise preference optimization framework. NoisePO employs a semantic noise preference optimization generative adversarial network (NPO-GAN) and noise ranking methods to search for semantically relevant noises based on textual conditions, thus eliminating undesired semantic features while emphasizing the necessary semantic ones. Specifically, NoisePO utilizes a light NPO-GAN to generate semantic noises that encourage the latent at the previous step to incorporate more semantic information from the caption. Then, light ranking models are employed to filter out low-quality noises and select the best noise. Experimental results demonstrate that NoisePO consistently outperforms the baselines across widely used frameworks, achieving notable improvements in image quality, semantic consistency, and user-specific alignment as measured by IS, FID, CLIP, and other metrics. These results indicate that NoisePO effectively enhances synthesis quality and strengthens text-image alignment.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Chengqun Song, Dacheng Tao, Jun Cheng 0002
IEEE Trans. Image Process.6
2026 Language-Guided Multimodal Spiking Neural Networks for Event-Based Action Recognition
abstract
Event-based human action recognition has gained increasing attention due to its efficiency in dynamic scenarios. Contemporary methodologies for event-based action recognition predominantly treat the problem as a one-hot classification task, which limits their ability to leverage the semantic relationships among various actions. To address this limitation, we propose a Spiking Event-Text Feature Fusion (SETFF) framework, which enhances recognition performance by integrating event and text modalities through a dual-stream architecture. SETFF leverages generative large language models to produce action descriptions, serving as semantic prompts that guide event feature learning. Specifically, a contrastive loss function is employed to align the features of both modalities, enriching the model's capacity to distinguish intricate and subtle actions. Extensive experiments on neuromorphic datasets, including PAF, DailyAction-DVS, DVS128 Gesture, Bullying10K, and UCF101-DVS, demonstrate that SETFF achieves state-of-the-art accuracy, with top-1 accuracy rates of up to 99.65% on the DailyAction-DVS dataset and 98.39% on the PAF dataset. Experimental results underscore the effectiveness of multimodal fusion in SNNs, advancing event-based action recognition while preserving the energy efficiency characteristic of SNNs.
Ziliang Ren, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Trans. Multim.5
2025 Task-Aware Clustering for Prompting Vision-Language Models
abstract
Prompt learning has attracted widespread attention in adapting vision-language models to downstream tasks. Existing methods largely rely on optimization strategies to ensure the task-awareness of learnable prompts. Due to the scarcity of task-specific data, overfitting is prone to occur. The resulting prompts often do not generalize well or exhibit limited task-awareness. To address this issue, we propose a novel Task-Aware Clustering (TAC) framework for prompting vision-language models, which increases the task-awareness of learnable prompts by introducing task-aware pre-context. The key ingredients are as follows: (a) generating task-aware pre-context based on task-aware clustering that can preserve the backbone structure of a downstream task with only a few clustering centers, (b) enhancing the task-awareness of learnable prompts by enabling them to interact with task-aware pre-context via the well-pretrained encoders, and (c) preventing the visual task-aware pre-context from interfering the interaction between patch embeddings by masked attention mechanism. Extensive experiments are conducted on benchmark datasets, covering the base-to-novel, domain generalization, and cross-dataset transfer settings. Ablation studies validate the effectiveness of key ingredients. Comparative results show the superiority of our TAC over competitive counterparts. The code is available at https://github.com/FushengHao/TAC.
Fusheng Hao, Fengxiang He, Fuxiang Wu, Tichao Wang, Chengqun Song, Jun Cheng 0002
CVPR6
2025 YOLO-KED: A Novel Framework for Rotated Object Detection in Complex Environments
abstract
Rotated object detection aims to locate and classify objects with arbitrary orientations. In complex backgrounds, small rotated objects with limited salient features present challenges for standard backbones to extract high-quality, discriminative features. Additionally, traditional single-stage detection heads suffer from spatial prediction biases due to misalignment between classification and localization tasks. To tackle these challenges, We propose a novel network architecture, YOLO-KED, which incorporates the KAGFusion Module (KAGFM), EdgeFusion Module (EFM), and Dynamic Alignment and Rotated Detection Head (DARH) to enhance the feature extraction precision of the backbone and dynamically align the classification and localization tasks. We have also released a new dataset, termed as ICDM, specifically for rotated object detection, containing 3,214 images across 11 common categories of industrial components. Experimental results demonstrate that YOLOKED outperforms existing state-of-the-art methods across multiple detection metrics, especially in complex scenarios and small object detection tasks. ICDM is available at https://github.com/twodian/ICDM.
Zhaoyu Zhuang, Penglei Liu, Dejia Xu, Jun Cheng 0002
ICASSP4
2025 Improved YOLOv11 for Low-illumination Object Detection in Autonomous Driving Scenarios
abstract
With the rapid advancement of autonomous driving technology, environmental perception, as its core component, faces critical challenges such as under low-illumination conditions. In this paper, a low-illumination environmental perception model based on YOLOv11 (You Only Look Once version 11) is proposed to tackle the performance degradation of visual sensors under low-illumination conditions. First, a progressive illumination enhancement module (PIE) is designed and incorporated into the backbone network of YOLOv11 model to capture the incremental local illumination. Then a histogram attention mechanism (HSSA) is introduced to address the illumination consistency issues, by capturing the illumination variations at different granularities. Finally, GSCONV is integrated into the model to implement lightweight improvements. The improved model can achieve a 7.1% increase in detection accuracy over the baseline model on the public dataset.
Kaihong Zhang, Yimin Zhou 0001, Jun Cheng 0002
IECON3
2025 Human-Imperceptible, Machine-Recognizable Images
abstract
Massive human-related data is collected to train neural networks for computer vision tasks. A major conflict is exposed relating to software engineers between better developing AI systems and distancing from the sensitive training data. To reconcile this conflict, the paper proposes an efficient privacy-preserving learning paradigm, where images are encrypted to become ``human-imperceptible, machine-recognizable'' via one of the two encryption strategies: (1) random shuffling equally-sized patches and (2) mixing-up sub-patches. Then, minimal adaptations are made to vision transformer to enable it to learn on the encrypted images for vision tasks, including image classification and object detection. Extensive experiments on ImageNet and COCO show that the proposed paradigm achieves comparable accuracy with the competitive methods. Decrypting the encrypted images requires solving an NP-hard jigsaw puzzle or ill-posed inverse problem, which is empirically shown intractable to be recovered by various attackers, including the powerful vision transformer-based attacker. We thus show that the proposed paradigm can ensure the encrypted images have become human-imperceptible while preserving machine-recognizable information.
Fusheng Hao, Fengxiang He, Yikai Wang 0001, Fuxiang Wu, Jing Zhang 0037, Dacheng Tao, Jun Cheng 0002
IJCAI7
2025 The Sampling-Gaussian for Stereo Matching
abstract
The soft-argmax operation is widely adopted in neural network-based stereo matching methods to enable differentiable regression of disparity. However, networks trained with soft-argmax tend to predict multimodal probability distributions due to the absence of explicit constraints on the shape of the distribution. Previous methods leveraged Laplacian distributions and cross-entropy for training but failed to effectively improve accuracy and even increased the network’s processing time. In this paper, we propose a novel method called Sampling-Gaussian as a substitute for soft-argmax. It improves accuracy without increasing inference time. We innovatively interpret the training process as minimizing the distance in vector space and propose a combined loss of L1 loss and cosine similarity loss. We leveraged the normalized discrete Gaussian distribution for supervision. Moreover, we identified two issues in previous methods and proposed extending the disparity range and employing bilinear interpolation as solutions. We have conducted comprehensive experiments to demonstrate the superior performance of our Sampling-Gaussian method. The experimental results prove that we have achieved better accuracy on five baseline methods across four datasets. Moreover, we have achieved significant improvements on small datasets and models with weaker generalization capabilities. Our method is easy to implement, and the code is available online.
Baiyu Pan, Jichao Jiao, Jianxin Pang, Jun Cheng 0002
IROS5
2025 An Efficient Hand Grasping Method Based on CVAE for Target Pose Estimation
abstract
With the advancement of humanoid robot industrialization, dexterous grasping has become a critical research area. Traditional two-finger grippers, while effective for regular geometries, struggle with complex shapes. Multi-finger dexterous hands offer significant advantages for adapting to diverse objects. This study proposes a grasp pose estimation method for multi-finger dexterous hands utilizing a Conditional Variational Autoencoder (CVAE) framework. Point cloud data and grasp poses from industrial components are used as inputs to the CVAE, with K-Nearest Neighbors (KNN) employed to enhance local feature extraction. Experimental results show that the proposed method achieves robust generalization and stability across various grasping scenarios.
Pengpeng Xu, Huaxi Zhang 0004, Wenlong Qin, Jianxin Pang, Jun Cheng 0002
IWCMC6
2025 Bidirectional Mixed Augmentation Sample Generation under Dual Perturbations in Semi-supervised Medical Image Segmentation
Yibo Feng, Feng Liu 0013, Zhiyi Shan, Lei Wang 0018, Jun Cheng 0002
PRCV (14)5
2025 Two-Stage Modal Feature Enhancement for Multispectral Object Detection
Tichao Wang, Ziliang Ren, Qieshi Zhang, Yimin Zhou 0001, Jun Cheng 0002
PRCV (5)5
2025 A survey of graph neural networks and their industrial applications
Lei Wang 0018, Xiaoliang Ma 0001, Jun Cheng 0002, MengChu Zhou
Neurocomputing4
2025 Imitation learning and interactive game-based prediction-decision planning
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Weiyu Yu, Huiquan Zhang
Neurocomputing5
2025 Neural adaptive delay differential equations
Qieshi Zhang, Jun Cheng 0002
Neurocomputing3
2025 Progressive background-foreground difference enhancement for few-shot 3D point cloud semantic segmentation
Tichao Wang, Fusheng Hao, Qieshi Zhang, Jun Cheng 0002
Image Vis. Comput.4
2025 Textual Embeddings are Good Class-Aware Visual Prompts for Adapting Vision-Language Models
abstract
Due to the parallel nature of the textual and visual encoders, very little attention has been paid to developing prompt learning by using well-pretrained encoders in a serial manner, in which the low-biased high-level semantic information accessible to each other for these encoders is ignored. In this letter, we find that textual embeddings are good class-aware visual prompts for adapting vision-language models, which leads to a new framework called TVPrompt (Textual embeddings as class-aware Visual Prompts). To eliminate the modal gap between text and vision, we design a bridging module, which integrates textual embeddings and class token to produce class-aware visual prompts. To ensure that such prompts could effectively collect class-relevant information, we further propose using masked attention to block the unnecessary interactions. Experimental evidence on benchmark datasets demonstrates that our TVPrompt achieves competitive efficiency and performance.
Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Signal Process. Lett.5
2025 HCNV: Hand-Eye Calibration Based on Surface Normal Optimization and View Selection
abstract
Accurate hand-eye calibration is crucial for robotics and computer vision tasks, such as robotic manipulation and 3D object recognition, as it directly affects operational precision and visual perception. However, existing methods rely solely on point clouds and neglect geometric features like plane normals, leading to suboptimal spatial alignment. Additionally, the lack of an intelligent view selection strategy worsens initial alignment errors, causing cumulative inaccuracies and degraded performance. We propose HCNV, a novel framework that combines optimized point cloud alignment with intelligent view selection to overcome these limitations. Our method uses plane normals and a hybrid genetic algorithm to refine the point cloud transformation matrix, significantly improving spatial alignment accuracy. Furthermore, the intelligent view selection strategy enhances point cloud matching and optimizes view coverage, increasing robustness across various conditions. Experimental results show that HCNV improves calibration accuracy by 20% and reduces computational time by 30% compared to state-of-the-art methods, demonstrating its effectiveness and practicality in real-world applications.
Yu Liu 0046, Hui Ma 0016, Penglei Liu, Jun Cheng 0002
IEEE Trans Autom. Sci. Eng.4
2025 Class-Irrelevant Feature Removal for Few-Shot Image Classification
abstract
Most existing few-shot image classification methods employ global pooling to aggregate class-relevant local features in a data-drive manner. Due to the difficulty and inaccuracy in locating class-relevant regions in complex scenarios, as well as the large semantic diversity of local features, the class-irrelevant information could reduce the robustness of the representations obtained by performing global pooling. Meanwhile, the scarcity of labeled images exacerbates the difficulties of data-hungry deep models in identifying class-relevant regions. These issues severely limit deep models' few-shot learning ability. In this work, we propose to remove the class-irrelevant information by making local features class relevant, thus bypassing the big challenge of identifying which local features are class irrelevant. The resulting class-irrelevant feature removal (CIFR) method consists of three phases. First, we employ the masked image modeling strategy to build an understanding of images' internal structures that generalizes well. Second, we design a semantic-complementary feature propagation module to make local features class relevant. Third, we introduce a weighted dense-connected similarity measure, based on which a loss function is raised to fine-tune the entire pipeline, with the aim of further enhancing the semantic consistency of the class-relevant local features. Visualization results show that CIFR achieves the removal of class-irrelevant information by making local features related to classes. Comparison results on four benchmark datasets indicate that CIFR yields very promising performance.
Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Trans. Neural Networks Learn. Syst.5
2025 Real-Time Semantic Segmentation via Spatial-Detail Guided Context Propagation
abstract
Nowadays, vision-based computing tasks play an important role in various real-world applications. However, many vision computing tasks, e.g., semantic segmentation, are usually computationally expensive, posing a challenge to the computing systems that are resource-constrained but require fast response speed. Therefore, it is valuable to develop accurate and real-time vision processing models that only require limited computational resources. To this end, we propose the spatial-detail guided context propagation network (SGCPNet) for achieving real-time semantic segmentation. In SGCPNet, we propose the strategy of spatial-detail guided context propagation. It uses the spatial details of shallow layers to guide the propagation of the low-resolution global contexts, in which the lost spatial information can be effectively reconstructed. In this way, the need for maintaining high-resolution features along the network is freed, therefore largely improving the model efficiency. On the other hand, due to the effective reconstruction of spatial details, the segmentation accuracy can be still preserved. In the experiments, we validate the effectiveness and efficiency of the proposed SGCPNet model. On the Cityscapes dataset, for example, our SGCPNet achieves 69.5% mIoU segmentation accuracy, while its speed reaches 178.5 FPS on 768 1536 images on a GeForce GTX 1080 Ti GPU card. In addition, SGCPNet is very lightweight and only contains 0.61 M parameters. The code will be released at https://github.com/zhouyuan888888/SGCPNet.
Shijie Hao, Yuan Zhou 0016, Yanrong Guo, Richang Hong, Jun Cheng 0002, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 AAGF: An Efficient Transformer With Mix-Features For Visual Place Recognition
abstract
Visual Place Recognition (VPR) is a task predicting the current location solely based on the visual features of images. It is susceptible to changes in perspective, lighting, and environmental conditions. Now the performance of the VPR method still relies on re-ranking, and the effectiveness of pure global retrieval is not ideal. To address this, we introduce a novel feature aggregation model based on the Transformer architecture, Agent-Attention with Gating Forward, which can aggregate the global relationships from feature maps obtained by a pre-trained backbone into a new global feature. Besides, a valid training strategy, Mix-Features Data Augment, is proposed to enhance the diversity of features and make the model more robust. Through experiments on multiple benchmarks, we demonstrate that our approach outperforms many existing techniques in terms of lightweight pre-trained backbone network aggregation.
Kuan Zhou, Zhenyu Xu 0014, Qieshi Zhang, Jun Cheng 0002, Ziliang Ren, Xiangyang Gao
ICIP4
2024 Dilated Pyramid Attention in Hierarchical Vision Transformer for Texture Recognition
Kangyu Tang, Penglei Liu, Jun Cheng 0002
ICONIP (7)3
2024 Distill-then-prune: An Efficient Compression Framework for Real-time Stereo Matching Network on Edge Devices
abstract
In recent years, numerous real-time stereo matching methods have been introduced, but they often lack accuracy. These methods attempt to improve accuracy by introducing new modules or integrating traditional methods. However, the improvements are only modest. In this paper, we propose a novel strategy by incorporating knowledge distillation and model pruning to overcome the inherent trade-off between speed and accuracy. As a result, we obtained a model that maintains real-time performance while delivering high accuracy on edge devices. Our proposed method involves three key steps. Firstly, we review state-of-the-art methods and design our lightweight model by removing redundant modules from those efficient models through a comparison of their contributions. Next, we leverage the efficient model as the teacher to distill knowledge into the lightweight model. Finally, we systematically prune the lightweight model to obtain the final model. Through extensive experiments conducted on two widely-used benchmarks, Sceneflow and KITTI, we perform ablation studies to analyze the effectiveness of each module and present our state-of-the-art results.
Baiyu Pan, Jichao Jiao, Jianxing Pang, Jun Cheng 0002
ICRA4
2024 Semantic-focused Patch Tokenizer with Multi-branch Mixer for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is critical for navigation and loop closure in autonomous driving tasks, mitigating the impact of shift errors caused by dynamic changes in the environment. Due to the limited ability of backbone networks and extreme environmental changes, current methods fail to capture foundational semantic details that include the distinctive attributes for unique place identification. To address this problem, we propose a new visual token-guided VPR framework that contains a semantic-focused patch tokenizer and a multi-branch Mixer. To mitigate the inference from place-unrelated objects, the semantic-focused patch tokenizer exploits attention-based channel selection and spatial partition, which efficiently captures important semantic information within the channels and preserve spatial relationships among the backbone features. To extract abstract features with spatial structure information, the multi-branch Mixer utilizes a multi-branch structure to aggregate local and global position information, improving the robustness of global representations to environmental changes. Experimental results demonstrate that our method outperforms state-of-the-art methods, achieving 85.3% Recall@1 on the MSLS val dataset and 59.1% Recall@1 on the Nordland dataset when using ResNet18 as the backbone.
Zhenyu Xu 0014, Ziliang Ren, Qieshi Zhang, Jie Lou, Dacheng Tao, Jun Cheng 0002
ICRA6
2024 Realistic and Visually-Pleasing 3D Generation of Indoor Scenes from a Single Image
Lei Wang 0018, Gongbin Chen, Yuhao Qiu, Jiaji Wu, Jun Cheng 0002
PRCV (6)7
2024 A Dense-Sparse Complementary Network for Human Action Recognition based on RGB and Skeleton Modalities
Qin Cheng, Jun Cheng 0002, Zhen Liu 0049, Ziliang Ren
Expert Syst. Appl.2
2024 Two-stage feature distribution rectification for few-shot point cloud semantic segmentation
Tichao Wang, Fusheng Hao, Guosheng Cui, Fuxiang Wu, Mengjie Yang, Qieshi Zhang, Jun Cheng 0002
Pattern Recognit. Lett.7
2024 BDR6D: Bidirectional Deep Residual Fusion Network for 6D Pose Estimation
abstract
Six-dimensional (6D) pose estimation is an important branch in the field of robotics focused on enhancing the ability of robots to manipulate and grasp objects. The latest research trend in 6D pose estimation is to directly predict the positions of two-dimensional (2D) keypoints from a single red, green, and blue (RGB) image through convolutional neural networks (CNNs) and establish a corresponding relationship with the three-dimensional (3D) keypoints of the model. Then, the perspective-n-point (PnP) algorithm is used to recover the 6D pose parameters. Currently, two challenges are encountered in pose estimation based on an RGB image. On the one hand, an RGB image lacks depth information, and it is thus difficult to directly obtain the corresponding geometric object information. On the other hand, when depth information is available, it is difficult to efficiently fuse the features of the RGB image with the features of the corresponding depth image. In this paper, we propose a bidirectional depth residual fusion network with a depth prediction (DP) network to estimate the 6D poses of objects (BDR6D). The BDR6D network predicts the depth information of objects using an RGB image, converts the depth information into point cloud information, and performs feature extraction and representation together with the RGB information during the feature extraction and representation stages. Specifically, the RGB image is fed into the BDR6D network, the DP network predicts the depth information of the objects in the image, and the depth map and RGB image are input into a point cloud network (PCN) and CNN, respectively, for feature extraction and representation. We build the bidirectional depth residual (BDR) structure so that the CNN and PCN can share information during feature extraction and representation. This approach allows the two networks to use each other’s local and global information to improve feature extraction and representation. For the keypoint selection stage, we propose an effective 2D keypoint selection method that considers the appearance and geometric information of the object of interest. We evaluate the proposed method with three benchmark datasets and compare it with other 6D pose estimation algorithms. The experimental results show that our method outperforms the state-of-the-art approach. Finally, we deploy our proposed method in conjunction with the Universal Robots 5 manipulator (UR5) robot to grasp and manipulate objects.Note to Practitioners—The purpose of this paper is to solve the problem of 6D pose estimation for robot grasping. The existing RGB image-based pose estimation approach faces two challenges. On the one hand, a single RGB image lacks depth information, so that it is difficult to directly obtain the corresponding geometric object information. On the other hand, when depth information is available, it is difficult to efficiently fuse the features of the RGB image with the features of the corresponding depth image. To solve the above problems, a novel network that can predict the depth information of objects from an RGB image and fuse the depth information with the RGB information to estimate the 6D pose of objects is proposed. Furthermore, we propose an effective 2D keypoint selection method that considers the appearance and geometric information of objects of interest. We evaluate the proposed approach based on three benchmark datasets and the UR5 robot platform and verify that our method is effective.
Penglei Liu, Qieshi Zhang, Jun Cheng 0002
IEEE Trans Autom. Sci. Eng.3
2024 Simultaneous Scheduling of Processing Machines and Automated Guided Vehicles via a Multi-View Modeling-Based Hybrid Algorithm
abstract
The flexible job-shop co-scheduling problem (FJCSP) for processing machines and automated guided vehicles (AGVs) in a flexible manufacturing system (FMS) has attracted more attention with the aim of improving production efficiency. In FMS, AGVs in charge of transporting jobs realize the flexible linkage of operations between different processing machines. The added interdependence between transporting and processing tasks brings more difficulties than the traditional flexible job-shop scheduling problem (FJSP). In this paper, the mathematical model of FJCSP is formulated to minimize the makespan. Considering the feature similarity of FJCSP with FJSP and AGV-routing problem in different cases, a multi-view modeling-based hybrid algorithm consisting of an estimation of distribution algorithm (EDA) and an ant colony optimization (ACO) is proposed. In EDA, a probability model abstracts the information in superior solutions about the operation sequencing and the rule selection for scheduling machines and AGVs. In ACO, a job-path pheromone model and an AGV-path pheromone model are designed to jointly select the job-machine-AGV combination with shorter processing time and transportation time. In the proposed hybrid algorithm, EDA and ACO generate solutions independently and achieve cooperation by sharing elites. An adaptive parameter is designed to regulate the use of the two methods to adapt to the varying demands of multi-view modeling in different cases and search stages. Furthermore, a local search with a three-layer operator based on the critical path method is proposed to balance exploration and exploitation in solution space. Finally, computational experiments involving a case study verified the advantage of the multi-view modeling-based hybrid algorithm in comparison with the state-of-the-art approaches.Note to Practitioners—This paper was motivated by the optimization problem of scheduling machines and automated guided vehicles (AGVs) in flexible manufacturing system (FMS). In FMS with AGVs, the transportation stages for jobs by AGVs significantly impact the overall production efficiency of the FMS and cannot be overlooked. This paper suggested a hybrid evolutionary algorithm using an estimation of distribution algorithm (EDA), an ant colony optimization (ACO) and a local search algorithm based on the critical path method. In the proposed hybrid algorithm, an adaptive parameter is introduced to regulate the utilization of EDA and ACO in generating a new population. This paper presents a mathematical characterization of the scheduling problem and subsequently outlines the step-by-step design of the hybrid algorithm. Computational experiments, including a case study, demonstrate that the hybrid algorithm exhibits adaptability to various instances and outperforms state-of-the-art approaches.
Bin Xin 0002, Sai Lu, Qing Wang 0010, Fang Deng, Jun Cheng 0002, Yuhang Kang
IEEE Trans Autom. Sci. Eng.6
2024 HQDec: Self-Supervised Monocular Depth Estimation Based on a High-Quality Decoder
abstract
Decoders play significant roles in recovering scene depths. However, the decoders used in previous works ignore the propagation of multilevel lossless fine-grained information, cannot adaptively capture local and global information in parallel, and cannot perform sufficient global statistical analyses on the final output disparities. In addition, the process of mapping from a low-resolution (LR) feature space to a high-resolution (HR) feature space is a one-to-many problem that may have multiple solutions. Therefore, the quality of the recovered depth map is low. To this end, we propose a high-quality decoder (HQDec), with which multilevel near-lossless fine-grained information, obtained by the proposed adaptive axial-normalized position-embedded channel attention sampling module (AdaAxialNPCAS), can be adaptively incorporated into a LR feature map with high-level semantics utilizing the proposed adaptive information exchange scheme. In the HQDec, we leverage the proposed adaptive refinement module (AdaRM) to model the local and global dependencies between pixels in parallel and utilize the proposed disparity attention module to model the distribution characteristics of disparity values from a global perspective. To recover fine-grained HR features with maximal accuracy, we adaptively fuse the high-frequency information obtained by constraining the upsampled solution space utilizing the local and global dependencies between pixels into the HR feature map generated from the nonlearning method. Extensive experiments demonstrate that each proposed component improves the quality of the depth estimation results over the baseline results, and the developed approach achieves state-of-the-art results on the KITTI and DDAD datasets. The code and models will be publicly available at HQDec.
Fei Wang 0066, Jun Cheng 0002
IEEE Trans. Circuits Syst. Video Technol.2
2023 Reject Decoding via Language-Vision Models for Text-to-Image Synthesis
abstract
Transformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, the common practice is drawing multi-paths from the transformer-based models and re-ranking the multi-images decoded from multi-paths to find the best one and filter out others. Therefore, the computing procedure of excluding images may be inefficient. To improve the effectiveness and efficiency of decoding, we exploit a reject decoding algorithm with tiny multi-modal models to enlarge the searching space and exclude the useless paths as early as possible. Specifically, we build tiny multi-modal models to evaluate the similarities between the partial paths and the caption at multi scales. Then, we propose a reject decoding algorithm to exclude some lowest quality partial paths at the inner steps. Thus, under the same computing load as the original decoding, we could search across more multi-paths to improve the decoding efficiency and synthesizing quality. The experiments conducted on the MS-COCO dataset and large-scale datasets show that the proposed reject decoding algorithm can exclude the useless paths and enlarge the searching paths to improve the synthesizing quality by consuming less time.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Lei Wang 0018, Jun Cheng 0002
AAAI6
2023 Efficiently Fusing Sparse Lidar for Enhanced Self-Supervised Monocular Depth Estimation
abstract
Monocular self-supervised depth estimation with a low-cost sensor is the mainstream solution to gathering dense depth maps for robots and autonomous driving. In this paper, based on the philosophy "less is more" (i.e., focusing only on valid pixels in sparse LiDAR), we propose a novel framework, Efficient Sparse Depth (EffisDepth), for predicting dense depth. The Sparse Feature Extractor (SFE) embedded in the proposed framework effectively handles sparse LiDAR by forming sparse tensors. The Slender Group Block (SGB) is the main building block in SFE, which extracts features from sparse tensors via a structure of two branches. Extensive experiments show that our method achieves state-of-the-art performance on the KITTI benchmark, demonstrating the effectiveness of each proposed component and the self-supervised learning framework.
Mingrong Gong, Qieshi Zhang, Jun Cheng 0002
ICASSP5
2023 Class-Aware Patch Embedding Adaptation for Few-Shot Image Classification
abstract
"A picture is worth a thousand words", significantly beyond mere a categorization. Accompanied by that, many patches of the image could have completely irrelevant meanings with the categorization if they were independently observed. This could significantly reduce the efficiency of a large family of few-shot learning algorithms, which have limited data and highly rely on the comparison of image patches. To address this issue, we propose a Class-aware Patch Embedding Adaptation (CPEA) method to learn "class-aware embeddings" of the image patches. The key idea of CPEA is to integrate patch embeddings with class-aware embeddings to make them class-relevant. Furthermore, we define a dense score matrix between class-relevant patch embeddings across images, based on which the degree of similarity between paired images is quantified. Visualization results show that CPEA concentrates patch embeddings by class, thus making them class-relevant. Extensive experiments on four benchmark datasets, miniImageNet, tieredImageNet, CIFAR-FS, and FC-100, indicate that our CPEA significantly outperforms the existing state-of-the-art methods. The source code is available at https://github.com/FushengHao/CPEA.
Fusheng Hao, Fengxiang He, Liu Liu 0014, Fuxiang Wu, Dacheng Tao, Jun Cheng 0002
ICCV6
2023 GSNet: Model Reconstruction Network for Category-level 6D Object Pose and Size Estimation
abstract
Category-level 6D pose and size estimation is to estimate the rotation, translation and size of the observed instance objects from an arbitrary angle in a cluttered scene. Compared with instance-level 6D pose estimation, there are two main challenges for category-level 6D pose estimation. One is that the algorithm needs to estimate the 6D pose and size of unseen objects, and no 3D models are available. Another is that different instance objects of the same class of objects differ greatly in shape. This paper propose a novel method to estimate the 6D pose and size of unseen objects from an RGB-D image. To handle intra-class shape variation, we propose an autoencoder-decoder that is trained on a set of object models to learn structural feature-invariant and shape-variant features of intra-class objects, and constructs a category-level priori model containing the structure feature and shape feature. To solve the problem of 3D model, this paper proposes a model reconstruction network including 3D graph convolution and spherical convolution (GSNet), which can reconstruct the 3D model of the observed instance object from the input RGB-D image and the priori model, and establish a dense correspon-dence between the 3D model and the observed instance object. Finally, random sample consensus (RANSAC) algorithm and Umeyama algorithm are used to estimate the 6D pose and size of the object. Extensive experiments on benchmark datasets show that the proposed method achieves state-of-the-art performance in category-level 6D object pose estimation. In order to prove that our method can be applied to the grasping and operation tasks of robots in industry and life, we deploy our method to a physical UR5 robot to perform grasping tasks on unseen but category known instances, and the results validate the efficacy of our proposed method.
Penglei Liu, Qieshi Zhang, Jun Cheng 0002
ICRA3
2023 Trajectory Tracking Control for Unmanned Aerial Manipulator with Unknown Object Grasping
abstract
In the process of grasping objects with an unmanned aerial manipulator (UAM), the weight of the payload is not always known in advance. Using model-based high-performance controllers to capture unknown objects introduces new interference to the system, which has a negative impact on the closed-loop control performance. This paper proposes a trajectory tracking controller for UAM during the object grasping with unknown mass. An adaptive module using back-stepping is developed for online estimation of the object mass, and an improved super-twisting extended state observer (ISTESO) is designed to cope with external disturbances and counteract uncertainties. Lya-punov stability theory is used for stability analysis to ensure that the trajectory tracking converges with relative estimation error approaching zero. Simulation experiments have been performed to verify the effectiveness of the proposed controller.
Huaxing Lin, Jun Cheng 0002, Yimin Zhou 0001
IECON2
2023 Blendshape-Based Migratable Speech-Driven 3D Facial Animation with Overlapping Chunking-Transformer
Jixi Chen, Xiaoliang Ma 0001, Lei Wang 0018, Jun Cheng 0002
PRCV (2)4
2023 MixPose: 3D Human Pose Estimation with Mixed Encoder
Jisheng Cheng, Qin Cheng, Mengjie Yang, Zhen Liu 0049, Qieshi Zhang, Jun Cheng 0002
PRCV (8)6
2023 Autoencoder and Masked Image Encoding-Based Attentional Pose Network
Long-Hua Hu, Xiaoliang Ma 0001, Lei Wang 0018, Jun Cheng 0002
PRCV (2)5
2023 UAV swarm formation reconfiguration control based on variable-stepsize MPC-APCMPIO algorithm
Jian Liao 0006, Jun Cheng 0002, Bin Xin 0002, Lihui Zheng, Yuhang Kang, Shaolei Zhou
Sci. China Inf. Sci.2
2023 Prototype expansion and feature calibration for few-shot point cloud semantic segmentation
Qieshi Zhang, Tichao Wang, Fusheng Hao, Fuxiang Wu, Jun Cheng 0002
Neurocomputing5
2023 Distilled representation using patch-based local-to-global similarity strategy for visual place recognition
Qieshi Zhang, Zhenyu Xu 0014, Yuhang Kang, Fusheng Hao, Ziliang Ren, Jun Cheng 0002
Knowl. Based Syst.6
2023 Semantic-Aware Feature Aggregation for Few-Shot Image Classification
Fusheng Hao, Fuxiang Wu, Fengxiang He, Qieshi Zhang, Chengqun Song, Jun Cheng 0002
Neural Process. Lett.6
2023 Multi-scale spatial-temporal convolutional neural network for skeleton-based action recognition
Qin Cheng, Jun Cheng 0002, Ziliang Ren, Qieshi Zhang
Pattern Anal. Appl.2
2023 Cross-scale cascade transformer for multimodal human action recognition
Zhen Liu 0049, Qin Cheng, Chengqun Song, Jun Cheng 0002
Pattern Recognit. Lett.4
2023 A Progressive Quadric Graph Convolutional Network for 3D Human Mesh Recovery
abstract
Human mesh recovery from one single image has achieved rapid progress recently, but many methods suffer from the image appearance overfitting since the training data are collected along with accurate 3D annotations in controlled settings of monotonous backgrounds or simple clothes. Some methods regress human mesh vertices from poses to tackle the above problem. However the mesh topologies have not been well exploited, and artifacts are often generated. In this paper, we aim to find an efficient low-cost solution to human mesh reconstruction. To this end, we propose a Progressive Quadric Graph Convolutional Network (PQ-GCN), and design a simple and fast method for 3D human mesh recovery from a single image in the wild. Specifically, we apply quadric-based surface simplification to human meshes and design a progressive graph convolution network, accompanied by mesh feature up-sampling, to deal with the mesh topologies. We carry out a series of studies to validate our method. The results prove that our method achieves superior performance on a challenging in-the-wild dataset, while using 66% fewer parameters than the existing method, Pose2Mesh. Artifacts have also been eliminated and better visual quality has been obtained without any further post-processing and model fitting. Besides, the recovery can be stopped at an earlier stage by adding a decoder head. Consequently, the computational complexity can be reduced greatly.
Lei Wang 0018, Xun-Yu Liu, Xiaoliang Ma 0001, Jiaji Wu, Jun Cheng 0002, MengChu Zhou
IEEE Trans. Circuits Syst. Video Technol.5
2023 Robust Leaderless Time-Varying Formation Control for Nonlinear Unmanned Aerial Vehicle Swarm System With Communication Delays
abstract
This article investigates the tracking-oriented robust leaderless time-varying formation (TVF) control problem for unmanned aerial vehicle swarm systems (UAVSSs) with Lipschitz nonlinear dynamics under directed topology, where external disturbances are random and bounded, and communication delays (CDs) are bounded. In this article, a state-feedback control approach is adopted to make sure that a UAVSS forms a desired TVF and follows a specified trajectory when CDs and external disturbances occur. First, a novel PD-like formation control protocol with several unknown parameters and CDs is designed. The protocol contains the information of the local neighborhood status and its differential quantities. Second, the tracking-oriented robust leaderless TVF control problem with Lipschitz dynamics, external disturbances, and CDs is transformed into a problem about asymptotic stability of a lower dimensional closed-loop control system through a special matrix decomposition. Third, a theorem is proposed to determine the unknown parameters of the control protocol and the upper bound of CDs. In the theorem, sufficient conditions for a UAVSS to attain the anticipated TVF and trajectory tracking are obtained. A Lyapunov-Krasovskii (LK) functional is constructed to verify that the error among the practical flight state of UAVs, the anticipant TVF configuration, and tracking trajectory can asymptotically converge to 0. Finally, with the presentation of a simulation case, the effectiveness of the theoretical results is illustrated.
Yuhang Kang, Bin Xin 0002, Jun Cheng 0002, Tangwen Yang, Shaolei Zhou
IEEE Trans. Cybern.4
2023 CbwLoss: Constrained Bidirectional Weighted Loss for Self-Supervised Learning of Depth and Pose
abstract
Photometric differences are widely used as supervision signals to train neural networks for estimating depth and camera pose from unlabeled monocular videos. However, this approach is detrimental for model optimization because occlusions and moving objects in a scene violate the underlying static scenario assumption. In addition, pixels in textureless regions or less discriminative pixels hinder model training. To address these problems, in this paper, we deal with moving objects and occlusions by utilizing the differences between the flow fields, and the differences between the depth structure generated by affine transformation and view synthesis, respectively. Secondly, we mitigate the effect of textureless regions on model optimization by measuring the differences between features with more semantic and contextual information without requiring additional networks. In addition, although the bidirectionality component is used in each sub-objective function, a pair of images is reasoned about only once, which helps reduce overhead. Extensive experiments and visual analysis demonstrate the effectiveness of the proposed method, which outperforms existing state-of-the-art self-supervised methods under the same conditions and without introducing additional auxiliary information.
Fei Wang 0066, Jun Cheng 0002, Penglei Liu
IEEE Trans. Intell. Transp. Syst.2
2023 Mixer-Based Semantic Spread for Few-Shot Learning
abstract
Key semantics can come from everywhere on an image. Semantic alignment is a key part of few-shot learning but still remains challenging. In this paper, we design a Mixer-Based Semantic Spread (MBSS) algorithm that employs amixermodule to spread the key semantic on the whole image, so that one can directly compare the processed image pairs. We first adopt a convolutional neural network to extract features from both support and query images and separate each of them into multiple Local Descriptor-based Representations (LDRs). The LDRs are then fed into themixerfor semantic spread, where every LDR attracts complementary information from its peers. In this way, the objective semantic is made spread on the whole image in a data-driven manner. The overall pipeline is supervised by a voting-based loss, guaranteeing a goodmixer. Visualization results validate the feasibility of ourmixer. Comprehensive experiments on three benchmark datasets, miniImageNet, tieredImageNet, and CUB, show that our algorithm achieves the state-of-the-art performance in both 5-way 1-shot and 5-way 5-shot settings.
Jun Cheng 0002, Fusheng Hao, Fengxiang He, Liu Liu 0014, Qieshi Zhang
IEEE Trans. Multim.1
2023 InDecGAN: Learning to Generate Complex Images From Captions via Independent Object-Level Decomposition and Enhancement
abstract
Text-to-image synthesis is a challenging problem, in which a complex scene contains diverse objects of various sizes and sub-images of objects belonging to the same class have diverse forms from different perspectives. Thus, synthesis models have difficulty in capturing varied objects in the complex scene. To alleviate these problems, we devise an independent object-level decomposing and enhancing generative adversarial networks, denoted as InDecGAN, to synthesize complex images and capture varied objects in a complex scene. Specifically, InDecGAN fully utilizes the independent object-level information, bounding boxes and high-resolution images of objects in training, by employing independent object-level pathways to synthesize varied objects. The independent object-level pathway integrates an independent object-level adversarial loss and the bounding box information to learn the visual features of objects independently, then, the main pathway exploits the features provided by the object-level pathway to compose the full scene and synthesize images. In addition, we analyze the generalization properties of the proposed InDecGAN and demonstrate the improvement from the perspective of the model architecture. Moreover, extensive experiments conducted on a widely used dataset are presented to demonstrate that the proposed model with an independent object-level pathway produces synthesized images of significantly improved quality.
Jun Cheng 0002, Fuxiang Wu, Liu Liu 0014, Qieshi Zhang, Leszek Rutkowski, Dacheng Tao
IEEE Trans. Multim.1
2023 Contextual Attention Network for Emotional Video Captioning
abstract
This paper investigates an emerging and challenging task—emotional video captioning. Formally, given a video, the task aims to not only describe the factual content of the video, but also discover the emotional clues in the video. We propose a novel Contextual Attention Network (CANet), which recognizes and describes the fact and emotion in the video by semantic-rich context learning. To be specific, at each time step, we first extract visual and textual features from both input video and previously generated words. Then, we apply the attention mechanism to these features to capture informative contexts for captioning. We train the CANet model with the joint optimization of cross-entropy loss$\mathcal {L}_{CE}$and contrastive loss$\mathcal {L}_{CL}$, where$\mathcal {L}_{CE}$constrains the semantics of the generated sentence to be close to human annotation and$\mathcal {L}_{CL}$encourages discriminative representation learning from positive and negative pairs of video and caption. Experiments on two emotional video captioning datasets (i.e., EmVidCap and EmVidCap-S) demonstrate the superiority of CANet compared to the state-of-the-art approaches.
Peipei Song, Dan Guo 0001, Jun Cheng 0002, Meng Wang 0001
IEEE Trans. Multim.3
2023 Language-Based Image Manipulation Built on Language-Guided Ranking
abstract
Text-based image manipulation is a popular subject and has many applications. However, it is a challenging task because there is no ground-truth edited dataset and textual descriptions have abstractive and ambiguous properties. To alleviate the difficult issues, we propose a manipulation framework consisting of the proposal attentional GANs, language-related semantic mask, and language-guided ranker. Specially, we construct an editing proposal generator to generate the suitable edited proposals with and without semantic conditions, which supports the reorganization of sub-generators to output proposals in various aspects as many as possible. To distinguish the text-relevant and the text-irrelevant regions, we introduce a language-related semantic mask based on the source image and target caption. Then, we exploit a language-guided ranker to retrieve the best edited result from the edited proposals through using the multi-modal similarity and the language-related semantic mask. Extensive experiments on widely-used datasets demonstrate that our model could manipulate images interactively and improve the editing quality effectively.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Jun Cheng 0002
IEEE Trans. Multim.5
2023 Prototypical Contrast and Reverse Prediction: Unsupervised Skeleton Based Action Recognition
abstract
We focus on unsupervised representation learning for skeleton based action recognition. Existing unsupervised approaches usually learn action representations by motion prediction but they lack the ability to fully learn inherent semantic similarity. In this paper, we propose a novel framework named Prototypical Contrast and Reverse Prediction (PCRP) to address this challenge. Different from plain motion prediction, PCRP performs reverse motion prediction based on encoder-decoder structure to extract more discriminative temporal pattern, and derives action prototypes by clustering to explore the inherent action similarity within the action encoding. Specifically, we regard action prototypes as latent variables and formulate PCRP as an expectation-maximization (EM) task. PCRP iteratively runs (1) E-step as to determine the distribution of action prototypes by clustering action encoding from the encoder while estimating concentration around prototypes, and (2) M-step as optimizing the model by minimizing the proposed ProtoMAE loss, which helps simultaneously pull the action encoding closer to its assigned prototype by contrastive learning and perform reverse motion prediction task. Besides, the sorting can also serve as a temporal task similar as reverse prediction in the proposed framework. Extensive experiments on N-UCLA, NTU 60, and NTU 120 dataset present that PCRP outperforms main stream unsupervised methods and even achieves superior performance over many supervised methods. The codes are available at:https://github.com/LZUSIAT/PCRP.
Haocong Rao, Xiping Hu, Jun Cheng 0002, Bin Hu 0001
IEEE Trans. Multim.4
2022 RA Loss: Relation-Aware Loss for Robust Person Re-identification
Kan Wang 0004, Shuping Hu, Jun Cheng 0002, Jianxin Pang, Huan Tan
ACCV (2)3
2022 Text-to-Image Synthesis based on Object-Guided Joint-Decoding Transformer
abstract
Object-guided text-to-image synthesis aims to generate images from natural language descriptions built by two-step frameworks, i.e., the model generates the layout and then synthesizes images from the layout and captions. However, such frameworks have two issues: 1) complex structure, since generating language-related layout is not a trivial task; 2) error propagation, because the inappropriate layout will mislead the image synthesis and is hard to be revised. In this paper, we propose an object-guided joint-decoding module to simultaneously generate the image and the corresponding layout. Specially, we present the joint-decoding transformer to model the joint probability on images tokens and the corresponding layouts tokens, where layout tokens provide additional observed data to model the complex scene better. Then, we describe a novel Layout-Vqgan for layout encoding and decoding to provide more information about the complex scene. After that, we present the detail-enhanced module to enrich the language-related details based on two facts: 1) visual details could be omitted in the compression of VQGANs; 2) the joint-decoding transformer would not have sufficient generating capacity. The experiments show that our approach is competitive with previous object-centered models and can generate diverse and high-quality objects under the given layouts.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Jun Cheng 0002
CVPR5
2022 Mask-Vit: an Object Mask Embedding in Vision Transformer for Fine-Grained Visual Classification
abstract
Fine-grained visual classification (FGVC) targets to accurately identify the subordinate categories from a target class. Convolutional neural network (CNN) based methods prove that the attention mechanism can enhance the representation of local regions and improve the recognition accuracy. Recently, vision transformer (ViT) has shown great application potential in image classification tasks by taking advantage of its inherent self-attention mechanism and early global information acquisition capability. However, this global information acquisition approach involves an irrelevant environment in the interaction process, which makes it difficult for fine-grained tasks that rely on local differences to quickly learn discriminant features. To this end, we propose a hybrid network termed Mask-ViT, which can effectively avoid environmental interference and express more robust features by focusing on the instance itself. Specifically, Contour Knowledge Embedding (CKE) is employed to transferred prior location information to ViT and guided the subsequent recognition. The experiments on three benchmarks demonstrate the effectiveness of the proposed method.
Shuo Ye, Chengqun Song, Jun Cheng 0002
ICIP4
2022 Indoor Target-Driven Visual Navigation based on Spatial Semantic Information
abstract
Target-driven visual navigation is a widely focused learning-based approach in the field of computer vision. However, it faces two major challenges: poor generalization ability to unknown scenes, and poor navigation performance for increased number of scenes. In this paper, an end-to-end target-driven visual navigation method, which uses Spatial Semantic Information (SSI) to navigate the agent to the target, is presented. To fully integrate the spatial and semantic information in the scene, visual information is encoded into an 8-D spatial context vector. In addition, the size of detected bounding box is used to improve the reward function in end-to-end learning to solve the problem of sparse rewards. Experiments in interactive environment dataset AI2-THOR show that compared with state-of-the-art approaches, our approach has a higher success rate and a better route to target.
Jiaojie Yan, Qieshi Zhang, Jun Cheng 0002, Ziliang Ren
ICIP3
2022 Structure-Preserving View-Invariant Skeleton Representation for Action Detection
abstract
Skeleton-based action detection has attracted increasing attention in recent years due to its action-focusing and compactness. To enable the usability of convolutional neural networks, many methods convert a skeleton sequence to a pseudo image by stacking the skeleton joints based on a predefined order. However, this practice ignores the skeletons structure and the influence of viewpoints, thus limiting the performance of learned models. In this paper, we propose a novel representation, which preserves the structure information while being view-invariant. To achieve this, we first generate a structure-preserving chain order by leveraging the depth-first traversal algorithm. Then, to eliminate the influence of viewpoints, we propose a reference joint-based encoding approach. Finally, we improve YOLOv5 by introducing an attention feature learning network, which enables YOLOv5 to automatically select the most informative joints. Comprehensive experiments on the PKU-MMD dataset demonstrate our method achieves state-of-the-art performance while maintaining high efficiency.
Hushan Qin, Jun Cheng 0002, Chengqun Song, Fusheng Hao, Qin Cheng
ICPR2
2022 Triplet Ratio Loss for Robust Person Re-identification
Shuping Hu, Kan Wang 0004, Jun Cheng 0002, Huan Tan, Jianxin Pang
PRCV (1)3
2022 EEP-Net: Enhancing Local Neighborhood Features and Efficient Semantic Segmentation of Scale Point Clouds
Fuxiang Wu, Qieshi Zhang, Ziliang Ren, Jun Cheng 0002
PRCV (3)5
2022 Character animation and retargeting from video streams
abstract
Virtual character animation is widely used in 3D games and virtual reality. Traditional character animation can be achieved through key-frame animation or motion capture technology. These methods have limited applications due to expensive equipments or sophisticated operations. Aiming at a lower-cost solution for this issue, in this paper we propose a method of virtual character animations and retargeting from RGB video streams based on human pose reconstruction. We conduct extensive experiments with different videos and virtual characters, and the resulting character animation is well represented in the virtual scene. The proposed method greatly has reduced the production cost of character animation, which has potential applications in virtual reality.
Gongbin Chen, Lei Wang 0018, Xun-Yu Liu, Long-Hua Hu, Jun Cheng 0002
SMC5
2022 Data association and loop closure in semantic dynamic SLAM using the table retrieval method
Chengqun Song, Jun Cheng 0002
Appl. Intell.5
2022 Spatial-temporal 3D dependency matching with self-supervised deep learning for monocular visual sensing
Chengqun Song, Maolong Niu, Zhaopeng Liu, Jun Cheng 0002, Luoying Hao
Neurocomputing4
2022 Dual-stream cross-modality fusion transformer for RGB-D action recognition
Zhen Liu 0049, Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Chengqun Song
Knowl. Based Syst.2
2022 A Self-Supervised Gait Encoding Approach With Locality-Awareness for 3D Skeleton Based Person Re-Identification
abstract
Person re-identification (Re-ID) via gait features within 3D skeleton sequences is a newly-emerging topic with several advantages. Existing solutions either rely on hand-crafted descriptors or supervised gait representation learning. This paper proposes a self-supervised gait encoding approach that can leverage unlabeled skeleton data to learn gait representations for person Re-ID. Specifically, we first create self-supervision by learning to reconstruct unlabeled skeleton sequences reversely, which involves richer high-level semantics to obtain better gait representations. Other pretext tasks are also explored to further improve self-supervised learning. Second, inspired by the fact that motion's continuity endows adjacent skeletons in one skeleton sequence and temporally consecutive skeleton sequences with higher correlations (referred as locality in 3D skeleton data), we propose a locality-aware attention mechanism and a locality-aware contrastive learning scheme, which aim to preserve locality-awareness on intra-sequence level and inter-sequence level respectively during self-supervised learning. Last, with context vectors learned by our locality-aware attention mechanism and contrastive learning scheme, a novel feature named Constrastive Attention-based Gait Encodings (CAGEs) is designed to represent gait effectively. Empirical evaluations show that our approach significantly outperforms skeleton-based counterparts by 15-40 percent Rank-1 accuracy, and it even achieves superior performance to numerous multi-modal methods with extra RGB or depth information. Our codes are available at https://github.com/Kali-Hac/Locality-Awareness-SGE.
Haocong Rao, Siqi Wang 0001, Xiping Hu, Mingkui Tan, Yi Guo 0007, Jun Cheng 0002, Xinwang Liu 0002, Bin Hu 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Covered Style Mining via Generative Adversarial Networks for Face Anti-spoofing
Yiqiang Wu, Dapeng Tao, Yong Luo 0002, Jun Cheng 0002, Xuelong Li 0001
Pattern Recognit.4
2022 Cross-Modality Compensation Convolutional Neural Networks for RGB-D Action Recognition
abstract
RGB-D-based human action recognition has attracted much attention recently because it can provide more complementary information than a single modality. However, it is difficult for two modalities to effectively learn spatial-temporal information from each other. To facilitate information interaction between different modalities, a cross-modality compensation convolutional neural network (ConvNet) is proposed for human action recognition, which enhances the discriminative ability by jointly learning compensation features from the RGB and depth modalities. Moreover, we design a cross-modality compensation block (CMCB) to extract compensation features from the RGB and depth modalities. Specifically, CMCB is incorporated into two typical network architectures, ResNet and VGG, to verify the ability to improve the performance of our model. The proposed architecture has been evaluated on three challenging datasets: NTU RGB+D 120, THU-READ and PKU-MMD. We experimentally verify that our proposed model with CMCB is effective for different input types, such as pairs of raw images and dynamic images constructed from the entire RGB-D sequence, and the experimental results show that the proposed framework achieves state-of-the-art performance on all three datasets.
Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Fusheng Hao
IEEE Trans. Circuits Syst. Video Technol.1
2022 RiFeGAN2: Rich Feature Generation for Text-to-Image Synthesis From Constrained Prior Knowledge
abstract
Text-to-image synthesis is a challenging task that generates realistic images from a textual description. The description contains limited information compared with the corresponding image and is ambiguous and abstract, which will complicate the generation and lead to low-quality images. To address this problem, we propose a novel generation text-to-image synthesis method, called RiFeGAN2, to enrich the given description. To improve the enrichment quality while accelerating the enrichment process, RiFeGAN2 exploits a domain-specific constrained model to limit the search scope and then uses an attention-based caption matching model to refine the compatible candidate captions based on constrained prior knowledge. To improve the semantic consistency between the given description and the synthesized results, RiFeGAN2 employs improved SAEMs, SAEM2s, to compact better features of the retrieved captions and effectively emphasize the descriptions via incorporating centre-attention layers. Finally, multi-caption attentional GANs are exploited to synthesize images from those features. Experiments performed on widely-used datasets show that the models can generate vivid images from enriched captions and effectually improve the semantic consistency.
Jun Cheng 0002, Fuxiang Wu, Yanling Tian, Lei Wang 0018, Dapeng Tao
IEEE Trans. Circuits Syst. Video Technol.1
2022 Global-Local Interplay in Semantic Alignment for Few-Shot Learning
abstract
Few-shot learning aims to recognize novel classes from only a few labeled training examples. Aligning semantically relevant local regions has shown promise in effectively comparing a query image with support images. However, global information is usually overlooked in the existing approaches, resulting in a higher possibility of learning semantics unrelated to the global information. To address this issue, we propose a Global-Local Interplay Metric Learning (GLIML) framework to employ the interplay between global features and local features to guide semantic alignment. We first design a Global-Local Information Concurrent Learning (GLICL) module to extract both global features and local features and perform global-local interplay. We then design a Global-Local Information Cross-Covariance Estimator (GLICCE) to learn the similarity on the global-local interplay, in contrast to the current practice where only local features are considered. Visualizations show that the global-local interplay decreases (1) the weights placed on the semantics that are irrelevant to the global information and (2) the variability of the learned features within every class in the feature space. Quantitative experiments on three benchmark datasets demonstrate that GLIML achieves state-of-the-art performance while maintaining high efficiency.
Fusheng Hao, Fengxiang He, Jun Cheng 0002, Dacheng Tao
IEEE Trans. Circuits Syst. Video Technol.3
2022 Adversarial UV-Transformation Texture Estimation for 3D Face Aging
abstract
Face aging aims to estimate aged facial textures given a certain face image. A number of 2D face-aging methods have been developed, but there have been few studies on 3D face aging, which would be valuable in several real-world applications. The lack of 3D face-aging data has had a significant impact on the development of 3D face aging, but we hypothesized that the large amounts of 2D face-aging data on the internet could be leveraged for 3D aged facial textures. In this paper, we propose a novel 3D aging framework, which we call UV-transformation texture estimation based on generative adversarial networks (UVTE-GAN), to achieve 3D face aging. Specifically, the proposed framework has three parts: 1) a 3D vertex and texture estimator, which accurately estimates the face’s spatial vertices and textures; 2) a texture-aging GAN, which is responsible for aging the estimated texture map via adversarial learning; and 3) a 2D & 3D rendering rebuilder, which recovers 2D & 3D faces using the estimated facial vertex map and aged facial texture map. In addition, we also design a plugin layer that allows us to train the whole model in an end-to-end manner. Experimental results demonstrate the effectiveness of the proposed method in synthesizing visually pleasing 3D aged face pictures, and state-of-the-art performance is achieved on several public datasets.
Yiqiang Wu, Ruxin Wang 0002, Mingming Gong, Jun Cheng 0002, Zhengtao Yu 0001, Dapeng Tao
IEEE Trans. Circuits Syst. Video Technol.4
2022 Visual Relationship Detection: A Survey
abstract
Visual relationship detection (VRD) is one newly developed computer vision task, aiming to recognize relations or interactions between objects in an image. It is a further learning task after object recognition, and is important for fully understanding images even the visual world. It has numerous applications, such as image retrieval, machine vision in robotics, visual question answer (VQA), and visual reasoning. However, this problem is difficult since relationships are not definite, and the number of possible relations is much larger than objects. So the complete annotation for visual relationships is much more difficult, making this task hard to learn. Many approaches have been proposed to tackle this problem especially with the development of deep neural networks in recent years. In this survey, we first introduce the background of visual relations. Then, we present categorization and frameworks of deep learning models for visual relationship detection. The high-level applications, benchmark datasets, as well as empirical analysis are also introduced for comprehensive understanding of this task.
Jun Cheng 0002, Lei Wang 0018, Jiaji Wu, Xiping Hu, Gwanggil Jeon, Dacheng Tao, MengChu Zhou
IEEE Trans. Cybern.1
2022 Image Hallucination From Attribute Pairs
abstract
Recent image-generation methods have demonstrated that realistic images can be produced from captions. Despite the promising results achieved, existing caption-based generation methods confront a dilemma. On the one hand, the image generator should be provided with sufficient details for realistic hallucination, meaning that longer sentences with rich content are preferred, but on the other hand, the generator is meanwhile fragile to long sentences due to their complex semantics and syntax like long-range dependencies and the combinatorial explosion of object visual features. Toward alleviating this dilemma, a novel approach is proposed in this article to hallucinate images from attribute pairs, which can be extracted from natural language processing (NLP) toolsets in the presence of complex semantics and syntax. Attribute pairs, therefore, enable our image generator to tackle long sentences handily and alleviate the combinatorial explosion, and at the same time, allow us to enlarge the training dataset and to produce hallucinations from randomly combined attribute pairs at ease. Experiments on widely used datasets demonstrate that the proposed approach yields results superior to the state of the art.
Fuxiang Wu, Jun Cheng 0002, Xinchao Wang, Lei Wang 0018, Dapeng Tao
IEEE Trans. Cybern.2
2022 Imposing Semantic Consistency of Local Descriptors for Few-Shot Learning
abstract
Few-shot learning suffers from the scarcity of labeled training data. Regarding local descriptors of an image as representations for the image could greatly augment existing labeled training data. Existing local descriptor based few-shot learning methods have taken advantage of this fact but ignore that the semantics exhibited by local descriptors may not be relevant to the image semantic. In this paper, we deal with this issue from a new perspective of imposing semantic consistency of local descriptors of an image. Our proposed method consists of three modules. The first one is a local descriptor extractor module, which can extract a large number of local descriptors in a single forward pass. The second one is a local descriptor compensator module, which compensates the local descriptors with the image-level representation, in order to align the semantics between local descriptors and the image semantic. The third one is a local descriptor based contrastive loss function, which supervises the learning of the whole pipeline, with the aim of making the semantics carried by the local descriptors of an image relevant and consistent with the image semantic. Theoretical analysis demonstrates the generalization ability of our proposed method. Comprehensive experiments conducted on benchmark datasets indicate that our proposed method achieves the semantic consistency of local descriptors and the state-of-the-art performance.
Jun Cheng 0002, Fusheng Hao, Liu Liu 0014, Dacheng Tao
IEEE Trans. Image Process.1
2022 Time-Varying Trajectory Tracking Formation H∞ Control for Multiagent Systems With Communication Delays and External Disturbances
abstract
Time-varying formation (TVF) and trajectory tracking$H_{\infty }$control problem of multiagent systems (MASs) subject to communication delays and external disturbances under the directed communication topology is studied. This article’s objective is for all agents to attain the desired TVF and track the pregiven formation center trajectory simultaneously. First, a distributed TVF and trajectory tracking control protocol employing neighborhood interaction information is developed in the presence of communication delays. Second, since the Laplacian matrix of a graph can be decomposed into the product of two specific matrices, the TVF and trajectory tracking$H_{\infty }$control problem is converted into the lower dimension asymptotic stability problem of a closed-loop system by applying an appropriate variable conversion. Third, a Lyapunov–Krasovskii functional is constructed to analyze the stability of MASs. Sufficient conditions are obtained in the form of linear matrix inequalities (LMIs) to ensure the completion of the TVF and formation center trajectory tracking of MASs. In the meantime, the maximum allowable communication delay can be calculated by the LMIs. Finally, the results of numerical simulations are presented to verify the validity of the approach this article proposes.
Jun Cheng 0002, Yuhang Kang, Bin Xin 0002, Qieshi Zhang, Shaolei Zhou
IEEE Trans. Syst. Man Cybern. Syst.1
2021 MFPN-6D : Real-time One-stage Pose Estimation of Objects on RGB Images
abstract
6D pose estimation of objects is an important part of robot grasping. The latest research trend on 6D pose estimation is to train a deep neural network to directly predict the 2D projection position of the 3D key points from the image, establish the corresponding relationship, and finally use Pespective-n-Point (PnP) algorithm performs pose estimation. The current challenge of pose estimation is that when the object texture-less, occluded and scene clutter, the detection accuracy will be reduced, and most of the existing algorithm models are large and cannot take the real-time requirements. In this paper, we introduce a Multi-directional Feature Pyramid Network, MFPN, which can efficiently integrate and utilize features. We combined the Cross Stage Partial Network (CSPNet) with MFPN to design a new network for 6D pose estimation, MFPN-6D. At the same time, we propose a new confidence calculation method for object pose estimation, which can fully consider spatial information and plane information. At last, we tested our method on the LINEMOD and Occluded-LINEMOD datasets. The experimental results demonstrate that our algorithm is robust to textureless materials and occlusion, while running more efficiently compared to other methods.
Penglei Liu, Qieshi Zhang, Jin Zhang 0013, Fei Wang 0066, Jun Cheng 0002
ICRA5
2021 Multi-Level Graph Encoding with Structural-Collaborative Relation Learning for Skeleton-Based Person Re-Identification
abstract
Skeleton-based person re-identification (Re-ID) is an emerging open topic providing great value for safety-critical applications. Existing methods typically extract hand-crafted features or model skeleton dynamics from the trajectory of body joints, while they rarely explore valuable relation information contained in body structure or motion. To fully explore body relations, we construct graphs to model human skeletons from different levels, and for the first time propose a Multi-level Graph encoding approach with Structural-Collaborative Relation learning (MG-SCR) to encode discriminative graph features for person Re-ID. Specifically, considering that structurally-connected body components are highly correlated in a skeleton, we first propose a multi-head structural relation layer to learn different relations of neighbor body-component nodes in graphs, which helps aggregate key correlative features for effective node representations. Second, inspired by the fact that body-component collaboration in walking usually carries recognizable patterns, we propose a cross-level collaborative relation layer to infer collaboration between different level components, so as to capture more discriminative skeleton graph features. Finally, to enhance graph dynamics encoding, we propose a novel self-supervised sparse sequential prediction task for model pre-training, which facilitates encoding high-level graph semantics for person Re-ID. MG-SCR outperforms state-of-the-art skeleton-based methods, and it achieves superior performance to many multi-modal methods that utilize extra RGB or depth features. Our codes are available at https://github.com/Kali-Hac/MG-SCR.
Haocong Rao, Xiping Hu, Jun Cheng 0002, Bin Hu 0001
IJCAI4
2021 SM-SGE: A Self-Supervised Multi-Scale Skeleton Graph Encoding Framework for Person Re-Identification
abstract
Person re-identification via 3D skeletons is an emerging topic with great potential in security-critical applications. Existing methods typically learn body and motion features from the body-joint trajectory, whereas they lack a systematic way to model body structure and underlying relations of body components beyond the scale of body joints. In this paper, we for the first time propose a Self-supervised Multi-scale Skeleton Graph Encoding (SM-SGE) framework that comprehensively models human body, component relations, and skeleton dynamics from unlabeled skeleton graphs of various scales to learn an effective skeleton representation for person Re-ID. Specifically, we first devise multi-scale skeleton graphs with coarse-to-fine human body partitions, which enables us to model body structure and skeleton dynamics at multiple levels. Second, to mine inherent correlations between body components in skeletal motion, we propose a multi-scale graph relation network to learn structural relations between adjacent body-component nodes and collaborative relations among nodes of different scales, so as to capture more discriminative skeleton graph features. Last, we propose a novel multi-scale skeleton reconstruction mechanism to enable our framework to encode skeleton dynamics and high-level semantics from unlabeled skeleton graphs, which encourages learning a discriminative skeleton representation for person Re-ID. Extensive experiments show that SM-SGE outperforms most state-of-the-art skeleton-based methods. We further demonstrate its effectiveness on 3D skeleton data estimated from large-scale RGB videos. Our codes are open at https://github.com/Kali-Hac/SM-SGE.
Haocong Rao, Xiping Hu, Jun Cheng 0002, Bin Hu 0001
ACM Multimedia3
2021 VGG-CAE: Unsupervised Visual Place Recognition Using VGG16-Based Convolutional Autoencoder
Zhenyu Xu 0014, Qieshi Zhang, Fusheng Hao, Ziliang Ren, Yuhang Kang, Jun Cheng 0002
PRCV (2)6
2021 Weighted ensemble networks for multiview based tiny object quality assessment
abstract
Summary As demand for intelligent manufacturing continues to grow, tiny object quality assessment (TOQA) is becoming increasingly importance in industrial automation. Recently, visual‐based TOQA has attracted an increasing attention, since the physical appearance is the foremost assessment index for evaluating the tiny object quality. It is exhausted and challenging to determine the quality of tiny object by manual visual inspection, and thus some machine vision systems are developed for automatic TOQA. Existing systems often use a limited number of cameras to capture the image of fallen tiny object, and thus may be not reliable since the tiny object may be unsound (such as cracked or damaged) in an invisible side. In this article, we develop a novel system for automatic TOQA that captures images of tiny object from multiple (more than two) view points, and propose a novel method termed weighted ensemble network (WENet) to effectively integrate the information of different views. In particular, convolutional neural networks (CNNs) are adopted to extract features from the images of different views. Then the multiview features are weighted combined for tiny object quality prediction. Traditional ensemble approaches usually directly applying average or voting to the prediction results of different views, or learn fixed weights to combine the results. Different from these approaches, the weights are adaptively determined in our method according to the quality of the captured image, since the features extracted from a low‐quality (e.g., blurred) image should contribute less to the final prediction. Handcrafted features and deep features are integrated in a sophisticated way in our method, and we empirically demonstrate the effectiveness of our method on grain quality assessment by investigating different CNN architectures for feature extraction and comparing with the conventional ensemble approaches.
Wanyin Wu, Jianwang Qiao, Jun Cheng 0002
Concurr. Comput. Pract. Exp.5
2021 Segment spatial-temporal representation and cooperative learning of convolution neural networks for multimodal-based action recognition
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Fusheng Hao, Xiangyang Gao
Neurocomputing3
2021 Visual relationship detection with recurrent attention and negative sampling
Lei Wang 0018, Peizhen Lin, Jun Cheng 0002, Feng Liu 0013, Xiaoliang Ma 0001, Jian Yin 0004
Neurocomputing3
2021 Gate-ID: WiFi-Based Human Identification Irrespective of Walking Directions in Smart Home
abstract
Research has shown the potential of device-free WiFi sensing for human identification. Each and every human has a unique gait and prior works suggest WiFi devices are able to capture the unique signature of a person's gait. In this article, we show for the first time that the monitored gait could be inconsistent and have mirror-like perturbations when individuals walk through WiFi devices in different directions, provided that the WiFi antenna array is horizontal to the walking path. Such inconsistent mirrored patterns are to negatively affect the uniqueness of gait and accuracy of human identification. Therefore, we propose a system called Gate-ID for accurately identifying individuals' identities irrespective of different walking directions. Gate-ID employs theoretical communication model and real measurements to demonstrate that antenna array orientations and walking directions contribute to the mirror-like patterns in WiFi signals. A novel heuristic algorithm is proposed to infer individual's walking directions. A set of methods are employed to extract and augment the representative spatial-temporal features of gait and enable the system performing irrespective of walking directions. We further propose a novel attention-based deep learning model that fuses various weighted features and ignores ineffective noises to uniquely identify individuals. We implement Gate-ID on commercial off-the-shelf devices. Extensive experiments demonstrate that our system can uniquely identify people with average accuracy of 90.7%-75.7% from a group of 6-20 people, respectively, and improve the accuracy by 12.5%-43.5% compared with baselines.
Jin Zhang 0013, Bo Wei 0003, Fuxiang Wu, Limeng Dong, Wen Hu 0001, Salil S. Kanhere, Chengwen Luo 0001, Shui Yu 0001, Jun Cheng 0002
IEEE Internet Things J.9
2021 Data Augmentation and Dense-LSTM for Human Activity Recognition Using WiFi Signal
abstract
Recent research has devoted significant efforts on the utilization of WiFi signals to recognize various human activities. An individual's limb motions in the WiFi coverage area could interfere with wireless signal propagation, that manifested as unique patterns for activity recognition. Existing approaches though yielding reasonable performance in certain cases, are ignorant of two major challenges. The performed activities of the individual normally have inconsistent speed in different situations and time. Besides that the wireless signal reflected by human bodies normally carries substantial information that is specific to that subject. The activity recognition model trained on a certain individual may not work well when being applied to predict another individual's activities. Since only recording activities of limited subjects in a certain speed and scale, recent works commonly have a moderate amount of activity data for training the recognition model. The small-size data could often incur the overfitting issue that negative affect the traditional classification model. To address these challenges, we propose a WiFi-based human activity recognition system that synthesizes variant activities data through eight channel state information (CSI) transformation methods to mitigate the impact of activity inconsistency and subject-specific issues, and also design a novel deep-learning model that caters to the small-size WiFi activity data. We conduct extensive experiments and show synthetic data improve performance by up to 34.6% and our system achieves around 90% of accuracy with well robustness in adapting to small-size CSI data.
Jin Zhang 0013, Fuxiang Wu, Bo Wei 0003, Qieshi Zhang, Hui Huang 0014, Syed Wajid Ali Shah, Jun Cheng 0002
IEEE Internet Things J.7
2021 Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
Haocong Rao, Xiping Hu, Jun Cheng 0002, Bin Hu 0001
Inf. Sci.4
2021 HAR-sEMG: A Dataset for Human Activity Recognition on Lower-Limb sEMG
Yu Luan, Yuhang Shi, Wanyin Wu, Zhiyao Liu, Hai Chang, Jun Cheng 0002
Knowl. Inf. Syst.6
2021 Exploiting spatio-temporal representation for 3D human action recognition from depth map sequences
Xiaopeng Ji, Jun Cheng 0002, Chenfei Ma
Knowl. Based Syst.3
2021 Multi-modality learning for human action recognition
Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Pengyi Hao, Jun Cheng 0002
Multim. Tools Appl.5
2021 Meta-learning based relation and representation learning networks for single-image deraining
Xinjian Gao, Yang Wang 0023, Jun Cheng 0002, Mingliang Xu 0001, Meng Wang 0001
Pattern Recognit.3
2021 Deep features for person re-identification on metric learning
Wanyin Wu, Dapeng Tao, Zhao Yang 0001, Jun Cheng 0002
Pattern Recognit.5
2020 RiFeGAN: Rich Feature Generation for Text-to-Image Synthesis From Prior Knowledge
abstract
Text-to-image synthesis is a challenging task that generates realistic images from a textual sequence, which usually contains limited information compared with the corresponding image and so is ambiguous and abstractive. The limited textual information only describes a scene partly, which will complicate the generation with complementing the other details implicitly and lead to low-quality images. To address this problem, we propose a novel rich feature generating text-to-image synthesis, called RiFeGAN, to enrich the given description. In order to provide additional visual details and avoid conflicting, RiFeGAN exploits an attention-based caption matching model to select and refine the compatible candidate captions from prior knowledge. Given enriched captions, RiFeGAN uses self-attentional embedding mixtures to extract features across them effectually and handle the diverging features further. Then it exploits multi-captions attentional generative adversarial networks to synthesize images from those features. The experiments conducted on widely-used datasets show that the models can generate images from enriched captions effectually and improve the results significantly.
Jun Cheng 0002, Fuxiang Wu, Yanling Tian, Lei Wang 0018, Dapeng Tao
CVPR1
2020 Multiple Time Scale Motion Images for Action Recognition
abstract
This paper proposes a simple and effective approach for RGB-based action recognition using multiple streams Convolutional Neural Networks (ConvNets). We utilize a new method to represent temporal structure of RGB videos, named Motion Image (MI), which is constructed from the difference between frames with a certain time scale. Considering the different duration of different actions, multiple time scale sampling MIs can obtain more temporal information. Furthermore, we adopt multiple streams ConvNets, including MIs and RGB streams, to learn spatial-temporal features for action recognition. Our approach has been evaluated on UCF-101 and HMDB-51, and the experimental results demonstrate the effectiveness and significantly improve action recognition rate at a small computational cost.
Qin Cheng, Ziliang Ren, Jun Cheng 0002
HealthCom4
2020 Phase-Sensitive Model for Temporal Action Proposal Generation
abstract
Temporal action proposal generation is an important and challenging task, aiming to localize the position where an action or event may occur in an untrimmed video. In this paper, we propose an efficient and end-to-end framework to generate temporal action proposals, named Phase-Sensitive Model (PSM), which fully understands all phases of temporal information. In particular, the PSM consists two modules: Boundary Phase Classification (BPC) and Action Phase Classification (APC). The BPC aims to provide two temporal boundary phase confidence maps by rich local information, while the APC is designed to generate an action phase confidence map by global features. Moreover, we introduce a new method boundary probability calculation to get the final score. Our experiments on ActivityNet-1.3 show a significant improvement with remarkable efficiency and generalizability.
Ziliang Ren, Lei Wang 0018, Jun Cheng 0002
HealthCom5
2020 ST-LSTM: Spatio-Temporal Graph Based Long Short-Term Memory Network For Vehicle Trajectory Prediction
abstract
Autonomous vehicles need the ability to predict the trajectory of surrounding vehicles, so as to make a rational decision planning, improve driving safety and ride comfort. In this paper, a new hierarchical Long Short-Term Memory (LSTM) based on Spatio-Temporal (ST) graph is proposed for vehicle trajectory prediction. Our ST-LSTM uses three layers of different LSTMs to capture the information of spatial, temporal and trajectory data, and LSTM-based encoder-decoder model as a whole, which is capable of accurately predicting future trajectories for vehicles on the highway. Our model trained and validated on the publicly available NGSIM US-101 and I-80 datasets. In comparison to state-of-art methods, our method could achieve a more accurate prediction trajectory over 5s time horizon.
Guangxi Chen, Qieshi Zhang, Ziliang Ren, Xiangyang Gao, Jun Cheng 0002
ICIP6
2020 Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification
abstract
Gait-based person re-identification (Re-ID) is valuable for safety-critical applications, and using only 3D skeleton data to extract discriminative gait features for person Re-ID is an emerging open topic. Existing methods either adopt hand-crafted features or learn gait features by traditional supervised learning paradigms. Unlike previous methods, we for the first time propose a generic gait encoding approach that can utilize unlabeled skeleton data to learn gait representations in a self-supervised manner. Specifically, we first propose to introduce self-supervision by learning to reconstruct input skeleton sequences in reverse order, which facilitates learning richer high-level semantics and better gait representations. Second, inspired by the fact that motion's continuity endows temporally adjacent skeletons with higher correlations (“locality”), we propose a locality-aware attention mechanism that encourages learning larger attention weights for temporally adjacent skeletons when reconstructing current skeleton, so as to learn locality when encoding gait. Finally, we propose Attention-based Gait Encodings (AGEs), which are built using context vectors learned by locality-aware attention, as final gait representations. AGEs are directly utilized to realize effective person Re-ID. Our approach typically improves existing skeleton-based methods by 10-20% Rank-1 accuracy, and it achieves comparable or even superior performance to multi-modal methods with extra RGB or depth information.
Haocong Rao, Siqi Wang 0001, Xiping Hu, Mingkui Tan, Huang Da, Jun Cheng 0002, Bin Hu 0001
IJCAI6
2020 Corrections to "A Cooperative Quality-Aware Service Access System for Social Internet of Vehicles"
Zhaolong Ning, Xiping Hu, Zhikui Chen, MengChu Zhou, Bin Hu 0001, Jun Cheng 0002, Mohammad S. Obaidat
IEEE Internet Things J.6
2020 Embedded adaptive cross-modulation neural network for few-shot learning
Jun Cheng 0002, Fusheng Hao, Wei Feng 0009
Neural Comput. Appl.2
2020 Online learning using projections onto shrinkage closed balls for adaptive brain-computer interface
Jun Cheng 0002, Dapeng Tao
Pattern Recognit.2
2020 A Cuboid CNN Model With an Attention Mechanism for Skeleton-Based Action Recognition
abstract
The introduction of depth sensors such as Microsoft Kinect have driven research in human action recognition. Human skeletal data collected from depth sensors convey a significant amount of information for action recognition. While there has been considerable progress in action recognition, most existing skeleton-based approaches neglect the fact that not all human body parts move during many actions, and they fail to consider the ordinal positions of body joints. Here, and motivated by the fact that an action's category is determined by local joint movements, we propose a cuboid model for skeleton-based action recognition. Specifically, a cuboid arranging strategy is developed to organize the pairwise displacements between all body joints to obtain a cuboid action representation. Such a representation is well structured and allows deep CNN models to focus analyses on actions. Moreover, an attention mechanism is exploited in the deep model, such that the most relevant features are extracted. Extensive experiments on our new Yunnan University-Chinese Academy of Sciences-Multimodal Human Action Dataset (CAS-YNU MHAD), the NTU RGB+D dataset, the UTD-MHAD dataset, and the UTKinect-Action3D dataset demonstrate the effectiveness of our method compared to the current state-of-the-art.
Kaijun Zhu, Ruxin Wang 0002, Jun Cheng 0002, Dapeng Tao
IEEE Trans. Multim.4
2019 Collect and Select: Semantic Alignment Metric Learning for Few-Shot Learning
abstract
Few-shot learning aims to learn latent patterns from few training examples and has shown promises in practice. However, directly calculating the distances between the query image and support image in existing methods may cause ambiguity because dominant objects can locate anywhere on images. To address this issue, this paper proposes a Semantic Alignment Metric Learning (SAML) method for few-shot learning that aligns the semantically relevant dominant objects through a "collect-and-select'' strategy. Specifically, we first calculate a relation matrix (RM) to "`collect" the distances of each local region pairs of the 3D tensor extracted from a query image and the mean tensor of the support images. Then, the attention technique is adapted to "select" the semantically relevant pairs and put more weights on them. Afterwards, a multi-layer perceptron (MLP) is utilized to map the reweighted RMs to their corresponding similarity scores. Theoretical analysis demonstrates the generalization ability of SAML and gives a theoretical guarantee. Empirical results demonstrate that semantic alignment is achieved. Extensive experiments on benchmark datasets validate the strengths of the proposed approach and demonstrate that SAML significantly outperforms the current state-of-the-art methods. The source code is available at https://github.com/haofusheng/SAML.
Fusheng Hao, Fengxiang He, Jun Cheng 0002, Lei Wang 0018, Jianzhong Cao, Dacheng Tao
ICCV3
2019 Embedded Block Residual Network: A Recursive Restoration Model for Single-Image Super-Resolution
abstract
Single-image super-resolution restores the lost structures and textures from low-resolved images, which has achieved extensive attention from the research community. The top performers in this field include deep or wide convolutional neural networks, or recurrent neural networks. However, the methods enforce a single model to process all kinds of textures and structures. A typical operation is that a certain layer restores the textures based on the ones recovered by the preceding layers, ignoring the characteristics of image textures. In this paper, we believe that the lower-frequency and higher-frequency information in images have different levels of complexity and should be restored by models of different representational capacity. Inspired by this, we propose a novel embedded block residual network (EBRN) which is an incremental recovering progress for texture super-resolution. Specifically, different modules in the model restores information of different frequencies. For lower-frequency information, we use shallower modules of the network to recover; for higher-frequency information, we use deeper modules to restore. Extensive experiments indicate that the proposed EBRN model achieves superior performance and visual improvements against the state-of-the-arts.
Yajun Qiu, Ruxin Wang 0002, Dapeng Tao, Jun Cheng 0002
ICCV4
2019 Emotion Recognition Based on Multi-View Body Gestures
abstract
Body gesture, a crucial component of "body language", remains less explored to recognize emotion while face expression-based and speech-based approaches are widely investigated. In this paper, we introduce an exploratory experiment to recognize emotion using deep learning only from body gestures. 43,200 multi-view RGB videos of simplified body gestures and their neutral control groups are captured from 80 humans using Hikvision network cameras to support the experiment. A novel approach is proposed to use deep neural network fuse skeleton and RGB features only using single-modality RGB video data. Experimental results show our approach achieves substantial improvements both in individual categories and overall and is provided with stronger generalization capability as well.
Zhijuan Shen, Jun Cheng 0002, Xiping Hu
ICIP2
2019 End-to-End Panoptic Segmentation with Pixel-Level Non-Overlapping Embedding
abstract
Recent panoptic segmentation even instance segmentation methods usually rely on the region-based method or highly-specialized combination with heuristics module, followed by post-processing techniques. While most of the recent methods neglect low-fill rate linear objects and cannot recognize pixels located in bounding box margins. We propose a branched, end-to-end trainable multi-task architecture focusing on pixel-level grouping problems for panoptic segmentation. The embedding branch regress pixels into an embedding space, so that pixels from the same group are at close range while those from different groups have a specified margin. Every pixel can be considered in an image without overlapping. And semantic branch produces best seed scores with labels as clustering center. The further-embedding branch disentangles each pixel in pixel embedding space. Thus, we are able to segment both thing and stuff classes, and explain all the pixels in the image. We obtain state-of-the-art results on Pascal VOC2012 and Cityscapes.
Qieshi Zhang, Jun Cheng 0002, Cong Bai, Pengyi Hao
ICME3
2019 WiEnhance: Towards Data Augmentation in Human Activity Recognition Using WiFi Signal
abstract
Recent research have devoted significant efforts on the utilization of WiFi signals to recognize various human activities. An individual's limb motions in the WiFi spectrum could interfere wireless signal propagation which manifested as unique patterns for activities recognition. Existing approaches though yielding reasonable performance in certain cases, are ignorant of a major challenge. The performed activities of the individual normally have inconsistent speed in different situations and time. Besides that the wireless signal reflected by human bodies normally carry substantial information that is specific to that subject. The activity recognition model trained on a certain individual may not work well when being applied to predict another individual's activities. To address this challenge, we propose WiEnhance, a WiFi based activity recognition system that synthesize variant activities data and mitigate the impact of activity inconsistency and subject-specific issues. We conduct extensive experiments and show an average 15.6% performance improvement on activity recognition.
Jin Zhang 0013, Fuxiang Wu, Wen Hu 0001, Qieshi Zhang, Weitao Xu, Jun Cheng 0002
MSN6
2019 Joint Resource Allocation for Latency-Sensitive Services Over Mobile Edge Computing Networks With Caching
abstract
Mobile edge computing (MEC) has risen as a promising paradigm to provide high quality of experience via relocating the cloud server in close proximity to smart mobile devices (SMDs). In MEC networks, the MEC server with computation capability and storage resource can jointly execute the latency-sensitive offloading tasks and cache the contents requested by SMDs. In order to minimize the total latency consumption of the computation tasks, we jointly consider computation offloading, content caching, and resource allocation as an integrated model, which is formulated as a mixed integer nonlinear programming (MINLP) problem. We design an asymmetric search tree and improve the branch and bound method to obtain a set of accurate decisions and resource allocation strategies. Furthermore, we introduce the auxiliary variables to reformulate the proposed model and apply the modified generalized benders decomposition method to solve the MINLP problem in polynomial computation complexity time. Simulation results demonstrate the superiority of the proposed schemes.
Jiao Zhang 0001, Xiping Hu, Zhaolong Ning, Edith C. H. Ngai, Li Zhou 0002, Jibo Wei, Jun Cheng 0002, Bin Hu 0001, Victor C. M. Leung
IEEE Internet Things J.7
2019 Mobile crowdsourcing based context-aware smart alarm sound for smart living
Yanxiang Guo, Wenhan Han, Jianbo Zheng, Hong Peng 0003, Xiping Hu, Jun Cheng 0002
Pervasive Mob. Comput.7
2019 A tensor framework for geosensor data forecasting of significant societal events
Lihua Zhou, Guowang Du, Ruxin Wang 0002, Dapeng Tao, Lizhen Wang 0001, Jun Cheng 0002
Pattern Recognit.6
2019 Single-Image De-Raining With Feature-Supervised Generative Adversarial Network
abstract
De-raining, which aims at rain-steak removal from images, is a practical task in computer vision. However, it is difficult due to its ill-posed nature. In this letter, we propose a deep neural network architecture, feature-supervised generative adversarial network (FS-GAN) for single-image rain removal. Its main idea is to train a generative adversarial network (GAN) for which the supervision from ground truth is imposed on different layers of the generator network. We design a feature-supervised generator, a discriminator, an optimization target, as well as the detailed structure of FS-GAN. Experiments show that the proposed FS-GAN achieves better performance than state-of-the-art de-raining methods on both synthetic and real-world images in terms of quantitative and visual quality.
Lei Wang 0018, Fuxiang Wu, Jun Cheng 0002, MengChu Zhou
IEEE Signal Process. Lett.4
2019 $p$ -Laplacian Regularization for Scene Recognition
abstract
The explosive growth of multimedia data on the Internet makes it essential to develop innovative machine learning algorithms for practical applications especially where only a small number of labeled samples are available. Manifold regularized semi-supervised learning (MRSSL) thus received intensive attention recently because it successfully exploits the local structure of data distribution including both labeled and unlabeled samples to leverage the generalization ability of a learning model. Although there are many representative works in MRSSL, including Laplacian regularization (LapR) and Hessian regularization, how to explore and exploit the local geometry of data manifold is still a challenging problem. In this paper, we introduce a fully efficient approximation algorithm of graph p -Laplacian, which significantly saving the computing cost. And then we propose p -LapR (pLapR) to preserve the local geometry. Specifically, p -Laplacian is a natural generalization of the standard graph Laplacian and provides convincing theoretical evidence to better preserve the local structure. We apply pLapR to support vector machines and kernel least squares and conduct the implementations for scene recognition. Extensive experiments on the Scene 67 dataset, Scene 15 dataset, and UC-Merced dataset validate the effectiveness of pLapR in comparison to the conventional manifold regularization methods.
Weifeng Liu 0001, Xueqi Ma, Yicong Zhou, Dapeng Tao, Jun Cheng 0002
IEEE Trans. Cybern.5
2019 Enhancing the Robustness of Neural Collaborative Filtering Systems Under Malicious Attacks
abstract
Recommendation systems have become ubiquitous in online shopping in recent decades due to their power in reducing excessive choices of customers and industries. Recent collaborative filtering methods based on the deep neural network are studied and introduce promising results due to their power in learning hidden representations for users and items. However, it has revealed its vulnerabilities under malicious user attacks. With the knowledge of a collaborative filtering algorithm and its parameters, the performance of this recommendation system can be easily downgraded. Unfortunately, this problem is not addressed well, and the study on defending recommendation systems is insufficient. In this paper, we aim to improve the robustness of recommendation systems based on two concepts - stage-wise hints training and randomness. To protect a target model, we introduce noise layers in the training of a target model to increase its resistance to adversarial perturbations. To reduce the noise layers' influence on model performance, we introduce intermediate layer outputs as hints from a teacher model to regularize the intermediate layers of a student target model. We consider white box attacks under which attackers have the knowledge of the target model. The generalizability and robustness properties of our method have been analytically inspected in experiments and discussions, and the computational cost is comparable to training a standard neural network-based collaborative filtering model. Through our investigation, the proposed defensive method can reduce the success rate of malicious user attacks and keep the prediction accuracy comparable to standard neural recommendation systems.
Yali Du 0001, Jinfeng Yi, Chang Xu 0002, Jun Cheng 0002, Dacheng Tao
IEEE Trans. Multim.5
2019 Domain-Weighted Majority Voting for Crowdsourcing
abstract
Crowdsourcing labeling systems provide an efficient way to generate multiple inaccurate labels for given observations. If the competence level or the "reputation," which can be explained as the probabilities of annotating the right label, for each crowdsourcing annotators is equal and biased to annotate the right label, majority voting (MV) is the optimal decision rule for merging the multiple labels into a single reliable one. However, in practice, the competence levels of annotators employed by the crowdsourcing labeling systems are often diverse very much. In these cases, weighted MV is more preferred. The weights should be determined by the competence levels. However, since the annotators are anonymous and the ground-truth labels are usually unknown, it is hard to compute the competence levels of the annotators directly. In this paper, we propose to learn the weights for weighted MV by exploiting the expertise of annotators. Specifically, we model the domain knowledge of different annotators with different distributions and treat the crowdsourcing problem as a domain adaptation problem. The annotators provide labels to the source domains and the target domain is assumed to be associated with the ground-truth labels. The weights are obtained by matching the source domains with the target domain. Although the target-domain labels are unknown, we prove that they could be estimated under mild conditions. Both theoretical and empirical analyses verify the effectiveness of the proposed method. Large performance gains are shown for specific data sets.
Dapeng Tao, Jun Cheng 0002, Zhengtao Yu 0001, Kun Yue, Lizhen Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 UFSM VO: Stereo Odometry Based on Uniformly Feature Selection and Strictly Correspondence Matching
abstract
Robust visual feature plays a critical role in improving camera localization performance. However, it will cost much computation time for feature extracting and matching, such as SIFT or SURF. In this paper, we present a novel visual odometry (VO) algorithm based on stereo image sequences by performing uniformly feature selection and strict correspondence matching. Firstly, the stable and uniform feature selection is performed by setting adaptive feature thresholds and selecting limited number of features in each local region. Secondly, the precise correspondence matching is achieved by double verification based on the motion model. Finally, the translation vector and rotation matrix of camera are computed, the five-point method is combined with RANSAC-based outlier rejection scheme for initial rotation estimation. And then all inliers are used for minimizing reprojection error to get final camera pose. The experimental results show that the proposed method can achieve the average translational error lower than 1.16% with 12Hz on the public KITTI dataset [1].
Liangliang Pan, Jun Cheng 0002, Qieshi Zhang
ICIP2
2018 Co-consistent Regularization with Discriminative Feature for Zero-Shot Learning
Yanling Tian, Qieshi Zhang, Jun Cheng 0002, Pengyi Hao
ICONIP (1)4
2018 The Impact of Digital Alarm Sound to Human Emotions: A Case Study
abstract
In many people's daily life, alarm sounds play an important role, which reflects the fast paced life in the modern society. On most occasions, people uses alarm sounds to wake them up in the morning. Improper alarm sounds could make people feel terrible. In this paper, we mainly propose a smart alarm sound recommendation system and construct an application to study how alarm sounds can impact human emotions. The recommendation system is deployed on the cloud, working with smartphones to deliver smart alarm sounds by considering not only sleep patterns, but also context information such as weather. The designed system can recommend smart alarm sounds to users, orchestrate sensing data collected by multiple sensors on smartphones, and collaborate with cloud computing to recommend preferable alarm sounds. An application is developed to demonstrate system effectiveness, which consists of the fore-end on Android OS and the back-end on the cloud. Experiments demonstrate that our system can recommend smart alarm sounds to wake participants up in the morning and the participants give feedback about their emotional states. The results show the system can improve people's emotion states by about 14.57%, compared to traditional alarm sound delivery.
Wenhan Han, Xiping Hu, Hanshu Cai, Jun Cheng 0002, Zhaolong Ning
SMC5
2018 A Cooperative Quality-Aware Service Access System for Social Internet of Vehicles
abstract
Because of the enormous potential to guarantee road safety and improve driving experience, social Internet of Vehicle (SIoV) is becoming a hot research topic in both academic and industrial circles. As the ever-increasing variety, quantity, and intelligence of on-board equipment, along with the evergrowing demand for service quality of automobiles, the way to provide users with a range of security-related and user-oriented vehicular applications has become significant. This paper concentrates on the design of a service access system in SIoVs, which focuses on a reliability assurance strategy and quality optimization method. First, in lieu of the instability of vehicular devices, a dynamic access service evaluation scheme is investigated, which explores the potential relevance of vehicles by constructing their social relationships. Next, this work studies a trajectory-based interaction time prediction algorithm to cope with an unstable network topology and high rate of disconnection in SIoVs. At last, a cooperative quality-aware system model is proposed for service access in SIoVs. Simulation results demonstrate the effectiveness of the proposed scheme.
Zhaolong Ning, Xiping Hu, Zhikui Chen, MengChu Zhou, Bin Hu 0001, Jun Cheng 0002, Mohammad S. Obaidat
IEEE Internet Things J.6
2018 Energy-Latency Tradeoff for Energy-Aware Offloading in Mobile Edge Computing Networks
abstract
Mobile edge computing (MEC) brings computation capacity to the edge of mobile networks in close proximity to smart mobile devices (SMDs) and contributes to energy saving compared with local computing, but resulting in increased network load and transmission latency. To investigate the tradeoff between energy consumption and latency, we present an energy-aware offloading scheme, which jointly optimizes communication and computation resource allocation under the limited energy and sensitive latency. In this paper, single and multicell MEC network scenarios are considered at the same time. The residual energy of smart devices' battery is introduced into the definition of the weighting factor of energy consumption and latency. In terms of the mixed integer nonlinear problem for computation offloading and resource allocation, we propose an iterative search algorithm combining interior penalty function with D.C. (the difference of two convex functions/sets) programming to find the optimal solution. Numerical results show that the proposed algorithm can obtain lower total cost (i.e., the weighted sum of energy consumption and execution latency) comparing with the baseline algorithms, and the energy-aware weighting factor is of great significance to maintain the lifetime of SMDs.
Jiao Zhang 0001, Xiping Hu, Zhaolong Ning, Edith C. H. Ngai, Li Zhou 0002, Jibo Wei, Jun Cheng 0002, Bin Hu 0001
IEEE Internet Things J.7
2018 Reinforcement online learning for emotion prediction by using physiological signals
Weifeng Liu 0001, Lianbo Zhang, Dapeng Tao, Jun Cheng 0002
Pattern Recognit. Lett.4
2018 Skeleton embedded motion body partition for human action recognition using depth sequences
Xiaopeng Ji, Jun Cheng 0002, Wei Feng 0009, Dapeng Tao
Signal Process.2
2018 Ensemble One-Dimensional Convolution Neural Networks for Skeleton-Based Action Recognition
abstract
This letter proposes an ensemble neural network (Ensem-NN) for skeleton-based action recognition. The Ensem-NN is introduced based on the idea of ensemble learning, “two heads are better than one.” According to the property of skeleton sequences, we design one-dimensional convolution neural network with residual structure asBase-Net. From entirety to local, from focus to motion, we designed four different subnets based on theBase-Netto extract diverse features. The first subnet is aTwo-stream Entirety Net, which performs on the entirety skeleton and explores both temporal and spatial features. The second is aBody-part Net, which can extract fine-grained spatial and temporal features. The third is anAttention Net, in which a channel-wised attention mechanism can learn important frames and feature channels.Frame-difference Net, as the fourth subnet, aims at exploring motion features. Finally, the four subnets are fused as one ensemble network. Experimental results show that the proposed Ensem-NN performs better than state-of-the-art methods on three widely used datasets.
Yangyang Xu 0004, Jun Cheng 0002, Lei Wang 0018, Haiying Xia, Feng Liu 0013, Dapeng Tao
IEEE Signal Process. Lett.2
2018 Guest Editorial Special Issue on Advancing Intelligent Automation in Sharing Economy
abstract
Sharing economy refers to peer-based activities of obtaining, giving, or sharing the access to goods and services, coordinated through community-based online services. It is known as collaborative consumption that people share the services rather than having individual ownership. By leveraging idle resources to produce more goods and services, sharing economy significantly drives green consumption and sustainable development in our human society. Using information technology to provide individuals with information enables the optimization of resources through the mutualization of excess capacity in goods and services. A common premise is that when information is shared, the value of the goods may increase for businesses, for individuals, for communities, and for the whole society in general. Currently, sharing economy has potentially resulted in a great impact on citizens’ everyday life and generated huge economic benefits, e.g., Airbnb, Uber, and Amazon Mechanical Turk. A host of enabling technologies has reached the mainstream for the rise of sharing economy, including open data, the ubiquity of low-cost mobile phones, and social media. These technologies dramatically reduce the friction of share-based business and organizational models.
Xiping Hu, Xitong Li, Wei Tan 0001, Jun Cheng 0002, MengChu Zhou, Yu-Kwong Kwok
IEEE Trans Autom. Sci. Eng.4
2018 Intelligent Smoke Alarm System with Wireless Sensor Network Using ZigBee
abstract
The conflagration of fire is still a serious problem caused by humans, and houses are at a high risk of fire. Recently, people have used smoke alarms which only have one sensor to detect fire. Smoke is emitted in several forms in daily life. A single sensor is not a reliable way to detect fire. With the rapid advancement in Internet technology, people can monitor their houses remotely to determine the current condition of the house. This paper introduces an intelligent smoke alarm system that uses ZigBee transmission technology to build a wireless network, uses random forest to identify smoke, and uses E‐charts for data visualization. By combining the real‐time dynamic changes of various environmental factors, compared to the traditional smoke alarm, the accuracy and controllability of the fire warning are increased, and the visualization of the data enables users to monitor the room environment more intuitively. The proposed system consists of a smoke detection module, a wireless communication module, and intelligent identification and data visualization module. At present, the collected environmental data can be classified into four statuses, that is, normal air, water mist, kitchen cooking, and fire smoke. Reducing the frequency of miscalculations also means improving the safety of the person and property of the user.
Jiashuo Cao, Shin-Ming Cheng, Jun Cheng 0002, Guanghui Pan
Wirel. Commun. Mob. Comput.7
2017 A Robust RGB-D Image-Based SLAM System
Liangliang Pan, Jun Cheng 0002, Wei Feng 0009, Xiaopeng Ji
ICVS2
2017 Poster: Emotion-Aware Smart Tips for Healthy and Happy Sleep
abstract
People spend up to one-third of lives asleep, and healthy sleep habits can make a big difference in their quality of life. But in modern society, many people have unhealthy sleep diaries and suffer from various sleep disorders, which may result in irregular mood fluctuations or even mental health problems such as anxiety and depression. We propose the Emotion-Aware Smart Tips (EAST), a novel approach that could help to inform users about their irregular emotional states with smart tips to improve their sleep qualities. EAST aims at helping users keep healthy sleep schedules and emotional states by providing smart tips through a novel model that combines multivariate regression, random forest, and neural network to quantify the relations between sleep patterns and emotional states. Prototype implementation and initial experiments of EAST in mobile phones have demonstrated its desired functionality and practicality for real-world deployment.
Yanxiang Guo, Jiao Zhang 0001, Chunbin Zhong, Xiping Hu, Bin Hu 0001, Jun Cheng 0002, Zhaolong Ning
MobiCom8
2017 Canonical correlation analysis networks for two-view image recognition
Xinghao Yang, Weifeng Liu 0001, Dapeng Tao, Jun Cheng 0002
Inf. Sci.4
2017 Support vector machine active learning by Hessian regularization
Weifeng Liu 0001, Lianbo Zhang, Dapeng Tao, Jun Cheng 0002
J. Vis. Commun. Image Represent.4
2017 The spatial Laplacian and temporal energy pyramid representation for human action recognition using depth sequences
Xiaopeng Ji, Jun Cheng 0002, Dapeng Tao, Xinyu Wu 0001, Wei Feng 0009
Knowl. Based Syst.2
2017 Multiview Canonical Correlation Analysis Networks for Remote Sensing Image Recognition
abstract
In the past decade, deep learning (DL) algorithms have been widely used for remote sensing (RS) image recognition tasks. As the most typical DL model, convolutional neural networks (CNNs) achieves outstand performance for big RS data classification. Recently, a variant of CNN, dubbed canonical correlation analysis network (CCANet), was proposed to abstract the two-view image features. Extensive experiments conducted on several benchmark databases validate the effectiveness of CCANet. However, the CCANet structure is powerless when the observations arrive from more than two sources. To serve the multiview purpose, in this letter, we propose multiview CCANets (MCCANets). Particularly, the MCCANet model learns the stacked multiperspective filter banks by the MCCA method and builds a deep convolutional structure. In the output stage, the binarization and the blockwise histogram are employed as nonlinear processing and feature pooling, respectively. To access the effectiveness of the MCCANet, we conduct a host of experiments on the RSSCN7 RS database. Extensive experimental results demonstrate that the MCCANet outperforms the two-view CCANet.
Xinghao Yang, Weifeng Liu 0001, Dapeng Tao, Jun Cheng 0002
IEEE Geosci. Remote. Sens. Lett.4
2017 LMAE: A large margin Auto-Encoders for classification
Weifeng Liu 0001, Tengzhou Ma, Qiangsheng Xie, Dapeng Tao, Jun Cheng 0002
Signal Process.5
2017 Robust Sparse Coding for Mobile Image Labeling on the Cloud
abstract
With the rapid development of the mobile service and online social networking service, a large number of mobile images are generated and shared on the social networks every day. The visual content of these images contains rich knowledge for many uses, such as social categorization and recommendation. Mobile image labeling has, therefore, been proposed to understand the visual content and received intensive attention in recent years. In this paper, we present a novel mobile image labeling scheme on the cloud, in which mobile images are first and efficiently transmitted to the cloud by Hamming compressed sensing, such that the heavy computation for image understanding is transferred to the cloud for quick response to the queries of the users. On the cloud, we design a sparse correntropy framework for robustly learning the semantic content of mobile images, based on which the relevant tags are assigned to the query images. The proposed framework (called maximum correntropy-based mobile image labeling) is very insensitive to the noise and the outliers, and is optimized by a half-quadratic optimization technique. We theoretically show that our image labeling approach is more robust than the squared loss, absolute loss, Cauchy loss, and many other robust loss function-based sparse coding methods. To further understand the proposed algorithm, we also derive its robustness and generalization error bounds. Finally, we conduct experiments on the PASCAL VOC’07 data set and empirically demonstrate the effectiveness of the proposed robust sparse coding method for mobile image labeling.
Dapeng Tao, Jun Cheng 0002, Xinbo Gao 0001, Xuelong Li 0001, Cheng Deng 0002
IEEE Trans. Circuits Syst. Video Technol.2
2017 Multiview Cauchy Estimator Feature Embedding for Depth and Inertial Sensor-Based Human Action Recognition
abstract
The ever-growing popularity of Kinect and inertial sensors has prompted intensive research efforts on human action recognition. Since human actions were extracted from Kinect and inertial sensors, they can be characterized by multiple feature representations. By encoding the multiview features into a unified space, it could be optimal for human action recognition. In this paper, we propose a new unsupervised feature fusion method termed multiview Cauchy estimator feature embedding (MCEFE) for human action recognition. By minimizing empirical risk, MCEFE integrates the encoded complementary information in multiple views to find the unified data representation and the projection matrices. To enhance robustness to outliers, the Cauchy estimator is imposed on the reconstruction error. Furthermore, ensemble manifold regularization is enforced on the projection matrices to encode the correlations between different views and avoid overfitting. Experiments are conducted on the new Chinese Academy of Sciences—Yunnan University—multimodal human action database to demonstrate the effectiveness and robustness of MCEFE for human action recognition.
Yanan Guo 0003, Dapeng Tao, Weifeng Liu 0001, Jun Cheng 0002
IEEE Trans. Syst. Man Cybern. Syst.4
2017 A Crowdsensing-Based Real-Time System for Finger Interactions in Intelligent Transport System
abstract
Crowdsensing leverages human intelligence/experience from the general public and social interactions to create participatory sensor networks, where context-aware and semantically complex information is gathered, processed, and shared to collaboratively solve specific problems. This paper proposes a real-time projector-camera finger system based on the crowdsensing, in which user can interact with a computer by bare hand touching on arbitrary surfaces. The interaction process of the system can be completely carried out automatically, and it can be used as an intelligent device in intelligent transport system where the driver can watch and interact with the display information while driving, without causing visual distractions. A single camera is used in the system to recover 3D information of fingertip for hand touch detection. A linear-scanning method is used in the system to determine the touch for increasing the users’ collaboration and operationality. Experiments are performed to show the feasibility of the proposed system. The system is robust to different lighting conditions. The average percentage of correct hand touch detection of the system is 92.0% and the average time of processing one video frame is 30 milliseconds.
Chengqun Song, Jun Cheng 0002, Wei Feng 0009
Wirel. Commun. Mob. Comput.2
2016 Dimensionality reduction of data sequences for human activity recognition
Yen-Lun Chen, Xinyu Wu 0001, Teng Li 0001, Jun Cheng 0002, Yongsheng Ou, Mingliang Xu 0001
Neurocomputing4
2016 Quaternion discrete cosine transformation signature analysis in crowd scenes for abnormal event detection
Huiwen Guo, Xinyu Wu 0001, Shibo Cai, Nannan Li 0001, Jun Cheng 0002, Yen-Lun Chen
Neurocomputing5
2016 Cauchy estimator discriminant analysis for face recognition
Xipeng Yang, Jun Cheng 0002, Wei Feng 0009, Zhengyao Bai, Dapeng Tao
Neurocomputing2
2016 Online tracking based on efficient transductive learning with sample matching costs
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Dapeng Tao, Jun Cheng 0002
Neurocomputing5
2016 Realtime and robust object matching with a large number of templates
Jianke Zhu, Jun Yu 0002, Jun Cheng 0002
Multim. Tools Appl.4
2016 Data-driven facial animation via semi-supervised local patch alignment
Jian Zhang 0026, Jun Yu 0002, Jane You, Dapeng Tao, Jun Cheng 0002
Pattern Recognit.6
2016 Superpixel-guided nonlocal means for image denoising and super-resolution
Ruxin Wang 0002, Jun Cheng 0002
Signal Process.4
2016 Tensor Manifold Discriminant Projections for Acceleration-Based Human Activity Recognition
abstract
With the rapid development of wearable sensors and pervasive computing technologies including smartphones, acceleration-based human activity recognition is receiving increased attention for medical research applications. Motivated by the “weightlessness” feature, here we apply a bidirectional feature during the feature extraction phase of activity recognition; however, since the bidirectional feature has two components, they cannot simply be concatenated into a long vector, but can be naturally treated as a second-order tensor. Therefore, we propose a new tensor-based feature selection method termed tensor manifold discriminant projections (TMDP). TMDP simultaneously considers: 1) applying an optimization criterion that can directly process the tensor spectral analysis problem, thereby decreasing the computational cost compared to traditional tensor-based feature selection methods; 2) extracting local rank information by finding a tensor subspace that preserves the rank order information of the within-class input samples; and 3) extracting discriminant information by maximizing the sum of distances between every sample and their interclass sample mean. Experiments on the naturalistic mobile devices-based human activity 2.0 dataset are performed to demonstrate the effectiveness and robustness of TMDP.
Yanan Guo 0003, Dapeng Tao, Jun Cheng 0002, Alan William Dougherty, Yaotang Li, Kun Yue, Bob Zhang 0001
IEEE Trans. Multim.3
2016 Manifold Ranking-Based Matrix Factorization for Saliency Detection
abstract
Saliency detection is used to identify the most important and informative area in a scene, and it is widely used in various vision tasks, including image quality assessment, image matching, and object recognition. Manifold ranking (MR) has been used to great effect for the saliency detection, since it not only incorporates the local spatial information but also utilizes the labeling information from background queries. However, MR completely ignores the feature information extracted from each superpixel. In this paper, we propose an MR-based matrix factorization (MRMF) method to overcome this limitation. MRMF models the ranking problem in the matrix factorization framework and embeds query sample labels in the coefficients. By incorporating spatial information and embedding labels, MRMF enforces similar saliency values on neighboring superpixels and ranks superpixels according to the learned coefficients. We prove that the MRMF has good generalizability, and develops an efficient optimization algorithm based on the Nesterov method. Experiments using popular benchmark data sets illustrate the promise of MRMF compared with the other state-of-the-art saliency detection methods.
Dapeng Tao, Jun Cheng 0002, Mingli Song
IEEE Trans. Neural Networks Learn. Syst.2
2015 Local mean spatio-temporal feature for depth image-based speed-up action recognition
abstract
With the promptly growing population of the low-cost Microsoft Kinect sensor, action recognition, which is a hard yet important problem in computer vision, has been received substantial attention. However, most existing approaches in action recognition spend much time on feature detection even though these methods can achieve high recognition rates. In this paper, we propose a local mean spatio-temporal feature (LMSF) to speed up depth image based action recognition. In particular, we solve the problem from three aspects: (1) associate the 4D normals by a local mean spatio-temporal neighborhood; (2) extract motion frames by detecting the differences between consecutive frames; (3) reduce redundant normals extracted from depth cloud points by sparse coding. The proposed approach is tested on two public benchmark datasets, i.e., MSRAction3D and MSRGesture3D. Experimental results demonstrate the advantages of our improvement method and the state-of-the-art performance on processing speed.
Xiaopeng Ji, Jun Cheng 0002, Dapeng Tao
ICIP2
2015 Local structure preserving discriminative projections for RGB-D sensor-based scene classification
Dapeng Tao, Jun Cheng 0002
Inf. Sci.2
2015 Road Boundaries Detection Based on Local Normal Saliency From Mobile Laser Scanning Data
abstract
The accurate extraction of roads is a prerequisite for the automatic extraction of other road features. This letter describes a method for detecting road boundaries from mobile laser scanning (MLS) point clouds in an urban environment. The key idea of our method is directly constructing a saliency map on 3-D unorganized point clouds to extract road boundaries. The method consists of four major steps, i.e., road partition with the assistance of the vehicle trajectory, salient map construction and salient points extraction, curb detection and curb lowest points extraction, and road boundaries fitting. The performance of the proposed method is evaluated on the point clouds of an urban scene collected by a RIEGL VMX-450 MLS system. The completeness, correctness, and quality of the extracted road boundaries are 95.41%, 99.35%, and 94.81%, respectively. Experimental results demonstrate that our method is feasible for detecting road boundaries in MLS point clouds.
Hanyun Wang, Huan Luo 0001, Chenglu Wen, Jun Cheng 0002, Peng Li 0064, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2015 Fingertip-based interactive projector-camera system
Jun Cheng 0002, Rui Song 0002, Xinyu Wu 0001
Signal Process.1
2014 Position-Based Action Recognition Using High Dimension Index Tree
abstract
Most current approaches in action recognition face difficulties that cannot handle recognition of multiple actions, fusion of multiple features, and recognition of action in frame by frame model, incremental learning of new action samples and application of position information of space-time interest points to improve performance simultaneously. In this paper, we propose a novel approach based on Position-Tree that takes advantage of the relationship of the position of joints and interest points. The normalized position of interest points indicates where the movement of body part has occurred. The extraction of local feature encodes the shape of the body part when performing action, justifying body movements. Additionally, we propose a new local descriptor calculating the local energy map from spatial-temporal cuboids around interest point. In our method, there are three steps to recognize an action: (1) extract the skeleton point and space-time interest point, calculating the normalized position according to their relationships with joint position, (2) extract the LEM (Local Energy Map) descriptor around interest point, (3) recognize these local features through non-parametric nearest neighbor and label an action by voting those local features. The proposed approach is tested on publicly available MSRAction3D dataset, demonstrating the advantages and the state-of-art performance of the proposed method.
Jun Cheng 0002, Wei Feng 0009
ICPR2
2014 Multiview Hessian discriminative sparse coding for image annotation
Weifeng Liu 0001, Dacheng Tao, Jun Cheng 0002, Yuan Yan Tang
Comput. Vis. Image Underst.3
2014 Disparity prediction between adjacent frames for dynamic scenes
Jun Cheng 0002, Baowen Chen, Xinyu Wu 0001
Neurocomputing2
2014 Conditional simultaneous localization and mapping: A robust visual SLAM system
Jigang Liu, Dongquan Liu, Jun Cheng 0002, Yuan Yan Tang
Neurocomputing3
2014 Semantic preserving distance metric learning and applications
Jun Yu 0002, Dapeng Tao, Jonathan Li 0001, Jun Cheng 0002
Inf. Sci.4
2013 Real-time hand detection based on multi-stage HOG-SVM classifier
abstract
In this paper, we propose a real-time hand detection method with multi-stage HOG-SVM classifier. Unlike traditional methods based on learning which make decomposition of feature vector or combination of different types of features or classifiers, upon the division of background into several categories, we propose a multi-stage classifier which combines several SVM classifies each of which is trained to distinguish corresponding divisions of background and target. Furthermore, in order to improve speed performance, skin color information and integral histogram are also applied. Experiment results demonstrate that the proposed algorithm works well under multiple challenging backgrounds in real-time speed (16 frames per second).
Jun Cheng 0002, Jianxin Pang
ICIP2
2013 Multiple instance learning via distance metric optimization
abstract
Multiple Instance Learning (MIL) has been widely applied in practice, such as drug activity prediction, content-based image retrieval. In MIL, a sample, comprised of a set of instances, is called a bag. Labels are assigned to bags instead of instances. The uncertainty of labels on instances makes MIL different from conventional supervised single instance learning (SIL) tasks. Therefore, it is critical to learn an effective mapping to convert an MIL task to an SIL task. In this paper, we present OptMILES by learning the optimal transformation on the bag-to-instance similarity measure, exploring the optimal distance metric between instances, by an alternating minimization training procedure. We thoroughly evaluate the proposed method on both a synthetic dataset and real world datasets by comparing with representative MIL algorithms. The experimental results suggest the effectiveness of OptMILES.
Haifeng Zhao 0002, Jun Cheng 0002, Dacheng Tao
ICIP2
2013 A defects detection system for the surfaces of stampings
abstract
Detecting defects on the surfaces of stampings plays a critical role in the manufacturing process. Many methods have been proposed to detect and identify simple defects on stampings. However, these methods suffer from large system size, high cost, and low speed for inspection. This paper proposes a new visual system for detecting defects on the surfaces of stampings. A set of LED bar lights are used to illuminate the stamping surface from the four sides. This can ensure that the irradiation directions are parallel to the surface. Thus, it can enhance the imaging of the defects and punching edges in the vertical orientation of the surface, which facilitates the location of the defects such as scratch and pitting and the measurement of the punching sizes. Thereby, the defects can be classified using simple shape and dimension analysis. The proposed system is a part of the automated sorting system. Practical operations verify the effectiveness of the proposed system.
Baowen Chen, Jun Cheng 0002, Sanming Shen
ICMV3
2013 Classification-based learning by particle swarm optimization for wall-following robot navigation
Yen-Lun Chen, Jun Cheng 0002, Xinyu Wu 0001, Yongsheng Ou, Yangsheng Xu
Neurocomputing2
2013 LF-EME: Local features with elastic manifold embedding for human action recognition
Xiaoyu Deng 0002, Xiao Liu 0012, Mingli Song, Jun Cheng 0002, Jiajun Bu, Chun Chen 0001
Neurocomputing4
2013 Locally regularized sliced inverse regression based 3D hand gesture recognition on a dance robot
Jun Cheng 0002, Wei Bian 0003, Dacheng Tao
Inf. Sci.1
2013 Pairwise constraints based multiview features fusion for scene classification
Jun Yu 0002, Dacheng Tao, Yong Rui, Jun Cheng 0002
Pattern Recognit.4
2013 Structured light-based shape measurement system
Jun Cheng 0002, Shiguang Zheng, Xinyu Wu 0001
Signal Process.1
2012 Stereo Matching Based on Random Speckle Projection for Dynamic 3D Sensing
abstract
Real-time 3D sensing has many important applications in areas such as robotic navigation, virtual reality and human-computer interaction. A variety of techniques have been developed for the determination of 3D geometry information such as binocular vision, structured light and their combination. However, existing non-contact optical 3D sensing approaches have their own limitations in the process of 3D information calculation. The reliability of binocular vision is limited to textures of the object surface, and the measuring accuracy of the method based on structured-light projection is limited to the stability of the light generator and the number of projected images. In this paper, we will combine random speckle projection with stereo matching algorithms to study the related problems on 3D measurement in the scene. An unique random speckle pattern is projected to encode the object surface to reduce the influence of projector flickering, and improve the accuracy of stereo matching. Besides, we take advantage of temporal consistency to reduce the range of disparity updating to improve the speed of 3D sensing further. The tests performed on real captured images confirm the validity of our approach.
Jun Cheng 0002, Haifeng Zhao 0002
ICMLA (1)2
2012 Transductive Cartoon Retrieval by Multiple Hypergraph Learning
Jun Yu 0002, Jun Cheng 0002, Jianmin Wang 0009, Dacheng Tao
ICONIP (3)2
2012 Clustering-based discriminative locality alignment for face gender recognition
abstract
To facilitate human-robot interactions, human gender information is very important. Motivated by the success of manifold learning for visual recognition, we present a novel clustering-based discriminative locality alignment (CDLA) algorithm to discover the low-dimensional intrinsic submanifold from the embedding high-dimensional ambient space for improving the face gender recognition performance. In particular, CDLA exploits the global geometry through k-means clustering, extracts the discriminative information through margin maximization and explores the local geometry through intra cluster sample concentration. These three properties uniquely characterize CDLA for face gender recognition. The experimental results obtained from the FERET data sets suggest the superiority of the proposed method in terms of recognition speed and accuracy by comparing with several representative methods.
Duo Chen 0001, Jun Cheng 0002, Dacheng Tao
IROS2
2012 Segment-Based Features for Time Series Classification
abstract
In this paper, we propose an approach termed segment-based features (SBFs) to classify time series. The approach is inspired by the success of the component- or part-based methods of object recognition in computer vision, in which a visual object is described as a number of characteristic parts and the relations among the parts. Utilizing this idea in the problem of time series classification, a time series is represented as a set of segments and the corresponding temporal relations. First, a number of interest segments are extracted by interest point detection with automatic scale selection. Then, a number of feature prototypes are collected by random sampling from the segment set, where each feature prototype may include single segment or multiple ordered segments. Subsequently, each time series is transformed to a standard feature vector, i.e. SBF, where each entry in the SBF is calculated as the maximum response (maximum similarity) of the corresponding feature prototype to the segment set of the time series. Based on the original SBF, an incremental feature selection algorithm is conducted to form a compact and discriminative feature representation. Finally, a multi-class support vector machine is trained to classify the test time series. Extensive experiments on different time series datasets, including one synthetic control dataset, two sign language datasets and one gait dynamics dataset, have been performed to evaluate the proposed SBF method. Compared with other state-of-the-art methods, our approach achieves superior classification performance, which clearly validates the advantages of the proposed method.
Zhang Zhang 0001, Jun Cheng 0002, Jun Li 0010, Wei Bian 0003, Dacheng Tao
Comput. J.2
2012 An energy model approach to people counting for abnormal crowd behavior detection
Guogang Xiong, Jun Cheng 0002, Xinyu Wu 0001, Yen-Lun Chen, Yongsheng Ou, Yangsheng Xu
Neurocomputing2
2012 Graph based transductive learning for cartoon correspondence construction
Jun Yu 0002, Wei Bian 0003, Mingli Song, Jun Cheng 0002, Dacheng Tao
Neurocomputing4
2012 Feature fusion for 3D hand gesture recognition by learning a shared hidden space
Jun Cheng 0002, Can Xie, Wei Bian 0003, Dacheng Tao
Pattern Recognit. Lett.1
2012 Interactive cartoon reusing by transfer learning
Jun Yu 0002, Jun Cheng 0002, Dacheng Tao
Signal Process.2
2011 Semi-automatic cartoon generation by motion planning
Jun Yu 0002, Dacheng Tao, Meng Wang 0001, Jun Cheng 0002
Multim. Syst.4
2011 3D human posture segmentation by spectral clustering with surface normal constraint
Jun Cheng 0002, Maoying Qiao, Wei Bian 0003, Dacheng Tao
Signal Process.1
2009 Biased isomap projections for interactive reranking
abstract
Image search has recently gained more and more attention for various applications. To capture users' intensions and to bridge the gap between the low level visual features and the high level semantics, a dozen of interactive reranking (IR) or relevance feedback (RF) algorithms have been developed and achieved significant performance improvements. In this paper, we develop a novel subspace learning based IR algorithm by using the patch alignment framework, termed the biased ISOMap projections or BIP for short. BIP models both the intraclass local geometry for query relevant images and the interclass discrimination between query relevant images and irrelevant images. In addition, BIP never meets the small samples size problem. We present experimental evidence suggesting that BIP is effective for targeting the intensions of users and reducing the semantic gaps for image search.
Wei Bian 0003, Jun Cheng 0002, Dacheng Tao
ICME2
2009 Registration for 3-D point cloud using angular-invariant feature
Jun Cheng 0002, Xinglin Chen
Neurocomputing2
2008 Handling of multi-reflections in wafer bump 3D reconstruction
abstract
In advanced electronic manufacturing that involves say die-to-die bonding, microscopic surfaces like solder bumps on wafers have to be inspected in 3D. However, because the bumps are of hemispherical shape, light projected onto the bumps could be reflected and illuminate other regions. Such multi-reflections could greatly disturb the intensity distribution in the image data and limit the use of gray level intensities for accurate 3D reconstruction of the bumps. In a previous work, we described a new solution mechanism that was based upon the concept of binary pattern projection, but unlike the traditional mechanisms which use an array of light sources it uses only a single light source. The light source in combination with a binary fringe grating could induce binary pattern on the target surface to be imaged, and the displacements of the binary fringe grating could allow the binary pattern to be varied. In this work, we describe under that solution framework how multi-reflections could be detected and the correct binary signals could be restored. Experimental results on solder bumps validate the feasibility of the proposed approach.
Jun Cheng 0002, Ronald Chung, Edmund Y. Lam, Kenneth S. M. Fung
SMC1
2007 A New Solder Paste Inspection Device: Design and Algorithm
abstract
In this paper, we present an innovative design of a solder paste inspection device which can be practically integrated into existing solder paste printing machines. Since solder paste inspection systems usually occupy a large space in vertical direction, we designed a mirror box that can re-direct the transmission of fringe pattern. In this way, a new parallel solder paste inspection device with a significant reduction in the vertical constraint is developed. We also developed a hybrid weighting algorithm that applied the distance and fringe contrast to acquire the height of solder pastes. Furthermore, we developed an algorithm that generates the 2-D image from the fringe pattern images during the 4-steps algorithm. It gives benefit (time for solder paste inspection) to traditional approach that uses some special lighting systems to create the 2-D image. Experimental results show our device can inspect the 20mm times 20mm PCB area within 2 seconds and the maximum standard deviation for the average height is 3 mum.
Xinyu Wu 0001, Wingkwong Chung, Hang Tong, Jun Cheng 0002, Yangsheng Xu
ICRA4
2004 Learning Human Tracking and Intercepting Skill
abstract
Robot tracking and intercepting fast-maneuvering object is a classical and important issue. Many research results were published in recent years. Most of them employed model-based methods which require robot's model in advance. However, it is difficult and time-consuming to obtain robot's mathematical model. In this paper, we present a novel approach which needs no mathematical model. The proposed approach is based on learning tracking strategy from human beings. With human's demonstrations, the robot can learn and abstract human tracking and intercepting skill using cascade neural network. Preliminarily simulation results attest the feasibility of this novel approach. Furthermore, experiment is done on a real-time human face tracking system and the results verify the validity and efficiency of the approach.
Jun Cheng 0002, Yangsheng Xu, Ronald Chung
ICRA1