Jian Cheng 0003

dblp:14/6145-3 · DBLP profile ↗
← Back
55ranked-venue papers
2as first author
36since 2021 · last 2026
0000-0001-6966-0531ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 12 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-feature collaboration with spatial-frequency learning guided by vision foundation model for remote sensing image captioning
Jian Cheng 0003, Ziying Xia, Siyu Liu 0003, Changjian Deng, Zunni Zhu, Zichong Chen
Neurocomputing2
2026 WideTopo: Improving foresight neural network pruning through training dynamics preservation and wide topologies exploration
Changjian Deng, Jian Cheng 0003, Yanzhou Su, Zeyu An, Ziying Xia, Shiguang Wang
Neural Networks2
2026 Robust multi-camera tracking in terminal environments: a spatio-temporal-appearance fusion approach
Wanli Dang, Jian Cheng 0003, Huaiyu Zheng
Vis. Comput.2
2025 Region Confidence Refinement with Progressive Semantic Mining for Source-Free Domain Adaptive Object Detection
abstract
Source-Free Domain Adaptive (SFDA) object detection addresses the challenges of detection in scenarios where both source domain data and target domain labels are unavailable. Due to the lack of data supervision, pseudo label learning has become the key to SFDA object detection. However, prevailing SFDA methods primarily concentrate on pseudo labels exhibiting exceptionally high or low confidence, without simultaneously considering false negative samples in high confidence and false positive samples in low confidence. We summarize this issue as the double-sided problem of pseudo labels. To address this issue, we propose the Region Confidence Refinement (RCR) aimed at refining the quality of pseudo labels via progressive semantic mining. Specifically, we bolster the semantic representation capacity of the detector across both pixel and image levels. Firstly, we design the Multi-Channel Style Filter (MSF) module to enrich pixel-level semantic representation by eliminating background-induced noise. Secondly, we design the Cross-Modal Semantic Enhancement (CSE) to enhance the classification efficacy of the detector amidst supervised information scarcity by aligning textual and image features, thereby amplifying image-level semantic representation. Finally, we design a Semantic Aggregation Strategy (SAS) for reconstructing region-level confidence. Extensive experiments demonstrate our proposed RCR achieves the state-of-the-art (SOTA) performance.
Zichong Chen, Zeyu An, Jian Cheng 0003
ICME3
2025 Adaptive Uncertainty Masking Network for Action Anticipation
Chenyue Jiang, Ziying Xia, Rinchen Dongrub, Gadeng Luosang, Jian Cheng 0003, Nyima Tashi
PRCV (11)5
2025 Cross-Modal Semantic Alignment via Concept Enrichment for Temporal Action Detection
Siyu Liu 0003, Rinchen Dongrub, Ziying Xia, Gadeng Luosang, Jian Cheng 0003, Nyima Tashi
PRCV (11)5
2025 Focusing on feature-level domain alignment with text semantic for weakly-supervised domain adaptive object detection
Zichong Chen, Jian Cheng 0003, Ziying Xia, Yongxiang Hu 0004, Zhicheng Dong 0003, Nyima Tashi
Neurocomputing2
2025 Trajectory tracking of QUAV based on cascade DRL with feedforward control
Shuliang He, Haoran Han, Jian Cheng 0003
Neurocomputing3
2025 CITAL: Counterfactual intervention for temporal action localization with point-level annotation
Yongxiang Hu 0004, Ziying Xia, Zichong Chen, Thupten Tsering, Jian Cheng 0003, Nyima Tashi
Neurocomputing5
2025 QUAV flight control based on axially symmetric DRL
Haoran Han, Jian Cheng 0003
Neurocomputing3
2025 Enhancing Collision-Free Formation Control in Multiagent Systems: An Approach Based on Time-Derivative of Artificial Potential Functions
abstract
The artificial potential function (APF) is a widely applied algorithm in collision-free formation control in multiagent systems (MASs). However, it suffers from oscillations and acceleration surges, particularly when the current formation and the desired one conflict. To address this problem and enhance collision-free formation control in MAS, this article introduces the time-derivative of APFs. This approach unifies attractive and repulsive APFs. The gradients of the APFs transform potential and kinetic energy, and the time-derivative of the APF gradients serve as damping terms to dissipate energy. This article discusses the general properties of APFs and introduces a time-variant formation tracking scheme that encompasses existing algorithms as specific instances. Then, a collision-free formation control algorithm is presented. This article gives proof of its Lyapunov stability and collision avoidance ability, followed by a maneuverability analysis from the geometry perspective. By incorporating the time-derivatives of repulsive APF gradients as damping terms, the proposed method mitigates oscillations and acceleration surges caused by conflicting attractive and repulsive effects.
Haoran Han, Jian Cheng 0003, Maolong Lv, Choon Ki Ahn
IEEE Trans. Cybern.2
2024 Analyzing Generalization in Policy Networks: A Case Study with the Double-Integrator System
abstract
Extensive utilization of deep reinforcement learning (DRL) policy networks in diverse continuous control tasks has raised questions regarding performance degradation in expansive state spaces where the input state norm is larger than that in the training environment. This paper aims to uncover the underlying factors contributing to such performance deterioration when dealing with expanded state spaces, using a novel analysis technique known as state division. In contrast to prior approaches that employ state division merely as a post-hoc explanatory tool, our methodology delves into the intrinsic characteristics of DRL policy networks. Specifically, we demonstrate that the expansion of state space induces the activation function $\tanh$ to exhibit saturability, resulting in the transformation of the state division boundary from nonlinear to linear. Our analysis centers on the paradigm of the double-integrator system, revealing that this gradual shift towards linearity imparts a control behavior reminiscent of bang-bang control. However, the inherent linearity of the division boundary prevents the attainment of an ideal bang-bang control, thereby introducing unavoidable overshooting. Our experimental investigations, employing diverse RL algorithms, establish that this performance phenomenon stems from inherent attributes of the DRL policy network, remaining consistent across various optimization algorithms.
Ruining Zhang, Haoran Han, Maolong Lv, Qisong Yang, Jian Cheng 0003
AAAI5
2024 Realigning Confidence with Temporal Saliency Information for Point-Level Weakly-Supervised Temporal Action Localization
abstract
Point-level weakly-supervised temporal action localization (P- TAL) aims to localize action instances in untrimmed videos through the use of single-point annotations in each instance. Existing methods predict the class activation se-quences without any boundary information, and the unreli-able sequences result in a significant misalignment between the quality of proposals and their corresponding confidence. In this paper, we surprisingly observe the most salientframe tend to appear in the central region of the each instance and is easily annotated by humans. Guided by the temporal saliency information, we present a novel proposal-level plug-in framework to relearn the aligned confidence of proposals generated by the base locators. The proposed approach consists of Center Score Learning (CSL) and Alignment-based Boundary Adaptation (ABA). In CSL, we design a novel center label generated by the point annotations for predicting aligned center scores. During inference, we first fuse the center scores with the predicted action probabilities to obtain the aligned confidence. ABA utilizes the both aligned confidence and IoU information to enhance localization completeness. Extensive experiments demon-strate the generalization and effectiveness of the proposed framework, showcasing state-of-the-art or competitive per-formances across three benchmarks. Our code is available at https://github.com/zyxial0091CVPR2024-TspNet.
Ziying Xia, Jian Cheng 0003, Siyu Liu 0003, Yongxiang Hu 0004, Shiguang Wang, Liwan Dang
CVPR2
2024 Tsdlinknet: D-Linknet with Transformer and scSE Module for High Resolution Remote Sensing Images Road Extraction
abstract
Road extraction for the high-resolution remote sensing images has important applications in urban planning, traffic management, geographic information systems(GIS) and other fields. However, The impact of the environment such as pedestrians and farms causes discontinuity and incompleteness in extraction results. In this paper, We present D-Linknet with transformer and Concurrent Spatial and Channel Squeeze and Channel Excitation(scSE) module(TSDlinknet) which is built with D-Linknet architecture. Firstly, we use transformer encoders in the central part instead of D-Block to handle long-range dependencies efficiently. Secondly, scSE module is added after each convolutional layer in the decoding stage to refine spatial information and channel information of the feature map. Experimental results on CHN6-CUG road dataset demonstrate that our method outperforms D-Linknet101 by 5.6% IOU higher, in addition, our extraction results are more continuous and complete.
Wenqi Yin, Ruiyu Zhang, Nyima Tashi, Jian Cheng 0003
IGARSS6
2024 High-Order Transformer Semantic Segmentation Network for High-Resolution Remote Sensing Images
abstract
Semantic segmentation of high-resolution remote sensing (HRRS) images is an important task in the field of remote sensing image analysis. However, the presence of a large number of complex ground objects in HRRS images poses challenges for its semantic segmentation. In this paper, we propose a high-order transformer semantic segmentation network (HOT-Net) for HRRS images. The network uses ResNet-50 as backbone to capture local features at the encoder stage. In order to expand the receptive field of the network and enhance its ability to perceive global contextual information, several proposed high-order transformer blocks are uesd to establish global dependencies in local features. Global enhancement attention modules (GEAM) are used to enhance the representation of global features during the decoder stage. We conducte ablation and comparative experiments on the LoveDA dataset, and the experimental results show that our method has excellent performance compared to other popular methods.
Zunni Zhu, Ziying Xia, Changjian Deng, Nyima Tashi, Jian Cheng 0003
IGARSS6
2024 TS-ILM: Class Incremental Learning for Online Action Detection
abstract
Online action detection aims to identify ongoing actions within untrimmed video streams, with extensive applications in real-life scenarios. However, in practical applications, video frames are received sequentially over time and new action categories continually emerge, giving rise to the challenge of catastrophic forgetting - a problem that remains inadequately explored. Generally, in the field of video understanding, researchers address catastrophic forgetting through class-incremental learning. Nevertheless, online action detection is based solely on historical observations, thus demanding higher temporal modeling capabilities for class-incremental learning methods. In this paper, we conceptualize this task as Class-Incremental Online Action Detection (CIOAD) and propose a novel framework, TS-ILM, to address it. Specifically, TS-ILM consists of two components: task-level temporal pattern extractor and temporal-sensitive exemplar selector. The former extracts the temporal patterns of actions in different tasks and saves them, allowing the data to be comprehensively observed on a temporal level before it is input into the backbone. The latter selects a set of frames with the highest causal relevance and minimum information redundancy for subsequent replay, enabling the model to learn the temporal information of previous tasks more effectively. We benchmark our approach against SoTA class-incremental learning methods applied in the image and video domains on THUMOS'14 and TVSeries datasets. Our method outperforms the previous approaches.
Jian Cheng 0003, Ziying Xia, Zichong Chen, Junhao Shi, Zhicheng Dong 0003, Nyima Tashi
ACM Multimedia2
2024 Interpretable DRL-Based Maneuver Decision of UCAV Dogfight
abstract
This paper proposes a three-layer unmanned combat aerial vehicle (UCAV) dogfight frame where Deep reinforcement learning (DRL) is responsible for high-level maneuver decision. A four-channel low-level control law is firstly constructed, followed by a library containing eight basic flight maneuvers (BFMs). Double deep Q network (DDQN) is applied for BFM selection in UCAV dogfight, where the opponent strategy during the training process is constructed with DT. Our simulation result shows that, the agent can achieve a win rate of 85.75% against the DT strategy, and positive results when facing various unseen opponents. Based on the proposed frame, interpretability of the DRL-based dogfight is significantly improved. The agent performs yo-yo to adjust its turn rate and gain higher maneuverability. Emergence of “Dive and Chase” behavior also indicates the agent can generate a novel tactic that utilizes the drawback of its opponent.
Haoran Han, Jian Cheng 0003, Maolong Lv
SMC2
2024 Semantic consistency knowledge transfer for unsupervised cross domain object detection
Zichong Chen, Ziying Xia, Junhao Shi, Nyima Tashi, Jian Cheng 0003
Appl. Intell.6
2024 PSE-Net: Channel pruning for Convolutional Neural Networks with parallel-subnets estimator
Shiguang Wang, Tao Xie 0010, Haijun Liu 0001, Xingcheng Zhang, Jian Cheng 0003
Neural Networks5
2024 Global Adaptive Second-Order Transformer for Remote Sensing Image Semantic Segmentation
abstract
In the domain of remote sensing (RS) image analysis, capturing global context is the key for precise semantic segmentation. Current vision transformer (ViT) advance this field by addressing convolutional neural network’s (CNN) local receptive field limitations. However, ViT predominantly rely on the first-order information in image to establish global relationships, often overlooking the potential of second-order information, which is crucial for enhancing the discrimination of ground objects that exhibit high similarity and constant changes. To address this issue, we propose a global adaptive second-order transformer network (GASOT-Net). Specifically, the proposed global adaptive second-order transformer (GASOT) enhances the existing ViT structure by mining second-order information and adaptively fusing it with the first-order information during the process of establishing global dependency relationships. This approach enables the extraction of more discriminative features, thereby enriching the representation of global features. In addition, the local feature aggregation module (LFAM) is proposed to effectively aggregate features from different stages of CNN as input to the GASOT blocks. Moreover, to refine boundaries of complex ground objects, the global feature enhancement module (GFEM) is used in the decoder stage. In particular, GFEM includes two sub modules—feature shift module (FSM) and hierarchical feature fusion module (HFFM). FSM is used to enhance the local feature representation at first, and then, HFFM hierarchically aggregates local and global features from different stages. We conduct extensive experiments on four benchmark RS datasets, and the results show that our GASOT-Net outperforms other state-of-the-art methods. The code will be available at:https://github.com/j136812832/GASOT-Net.
Jian Cheng 0003, Yanzhou Su, Changjian Deng, Ziying Xia, Nyima Tashi
IEEE Trans. Geosci. Remote. Sens.2
2024 HCM: Online Action Detection With Hard Video Clip Mining
abstract
Online action detection plays a vital role in video action understanding and can be widely used in various video analysis applications. This task aims to detect actions at the current moment within long untrimmed video streams. However, accurately identifying action-background transitions that are ambiguous in terms of time during detection can be challenging due to the similarity between the action and background clips, adding to the difficulty in finding a suitable division between them. To address this issue, we propose a hard video clip mining method based on deep metric learning for online action detection named HCM. The HCM method first selects video clips that are hard to distinguish to determine the optimization objects. Then, a hard clip mining loss is adopted to push the features toward the centers of the categories to which they belong and away from others. Furthermore, we introduce an intra-class feature compaction loss to constrain the divergence of action features, ensuring the stability of their distribution. We evaluated the proposed method on two challenging online action detection datasets, THUMOS14 and TVSeries. The results show that HCM is effective and efficient in online action detection and action anticipation tasks.
Siyu Liu 0003, Jian Cheng 0003, Ziying Xia, Zhilong Xi, Qin Hou, Zhicheng Dong 0003
IEEE Trans. Multim.2
2023 FLRKD: Relational Knowledge Distillation Based on Channel-wise Feature Quality Assessment
Zeyu An, Changjian Deng, Wanli Dang, Zhicheng Dong 0003, Jian Cheng 0003
BMVC6
2023 MDL-NAS: A Joint Multi-domain Learning Framework for Vision Transformer
abstract
In this work, we introduce MDL-NAS, a unified frame-work that integrates multiple vision tasks into a manageable supernet and optimizes these tasks collectively under diverse dataset domains. MDL-NAS is storage-efficient since multiple models with a majority of shared parameters can be deposited into a single one. Technically, MDL-NAS constructs a coarse-to-fine search space, where the coarse search space offers various optimal architectures for different tasks while the fine search space provides fine-grained parameter sharing to tackle the inherent obstacles of multi-domain learning. In the fine search space, we suggest two parameter sharing policies, i.e., sequential sharing policy and mask sharing policy. Compared with previous works, such two sharing policies allow for the partial sharing and non-sharing of parameters at each layer of the network, hence attaining real fine-grained parameter sharing. Finally, we present a joint-subnet search algorithm that finds the optimal architecture and sharing parameters for each task within total resource constraints, challenging the traditional practice that downstream vision tasks are typically equipped with backbone networks designed for image classification. Experimentally, we demonstrate that MDL-NAS families fitted with non-hierarchical or hierarchical transformers deliver competitive performance for all tasks compared with state-of-the-art methods while maintaining efficient storage deployment and computation. We also demonstrate that MDL-NAS allows incremental learning and evades catastrophic forgetting when generalizing to a new task.
Shiguang Wang, Tao Xie 0010, Jian Cheng 0003, Xingcheng Zhang, Haijun Liu 0001
CVPR3
2023 Poly-PC: A Polyhedral Network for Multiple Point Cloud Tasks at Once
abstract
In this work, we show that it is feasible to perform multiple tasks concurrently on point cloud with a straightforward yet effective multi-task network. Our framework, Poly-PC, tackles the inherent obstacles (e.g., different model architectures caused by task bias and conflicting gradients caused by multiple dataset domains, etc.) of multi-task learning on point cloud. Specifically, we propose a residual set abstraction (Res-SA) layer for efficient and effective scaling in both width and depth of the network, hence accommodating the needs of various tasks. We develop a weight-entanglement- based one-shot NAS technique to find optimal architectures for all tasks. Moreover, such technique entangles the weights of multiple tasks in each layer to offer task-shared parameters for efficient storage deployment while providing ancillary task-specific parameters for learning task-related features. Finally, to facilitate the training of Poly-PC, we introduce a task-prioritization-based gradient balance algorithm that leverages task prioritization to reconcile conflicting gradients, ensuring high performance for all tasks. Benefiting from the suggested techniques, models optimized by Poly-PC collectively for all tasks keep fewer total FLOPs and parameters and outperform previous methods. We also demonstrate that Poly-PC allows incremental learning and evades catastrophic forgetting when tuned to a new task.
Tao Xie 0010, Shiguang Wang, Ke Wang 0028, Linqi Yang, Xingcheng Zhang, Ruifeng Li 0001, Jian Cheng 0003
CVPR9
2023 Revisiting Feature Propagation and Aggregation in Polyp Segmentation
Yanzhou Su, Yiqing Shen 0003, Jin Ye 0002, Junjun He, Jian Cheng 0003
MICCAI (5)5
2023 Symmetric actor-critic deep reinforcement learning for cascade quadrotor flight control
Haoran Han, Jian Cheng 0003, Zhilong Xi, Maolong Lv
Neurocomputing2
2023 Accurate polyp segmentation through enhancing feature fusion and boosting boundary performance
Yanzhou Su, Jian Cheng 0003, Chuqiao Zhong, Chengzhi Jiang, Jin Ye 0002, Junjun He
Neurocomputing2
2022 Fe-LinkNet: Enhanced D-LinkNet with Attention and Dense Connection for Road Extraction in High-Resolution Remote Sensing Images
abstract
Extracting roads from high-resolution remote sensing images automatically is more efficient than field acquisition and manual annotation by experts. However, most road extraction methods based on deep-learning have problems of poor connectivity, due to occlusion of buildings and trees or confusion of backgrounds that have similar texture. In this paper, we proposed a novel feature-enhanced D-LinkNet (FE-LinkNet) to deal with the problem that road information is vulnerable to loss. Firstly, we introduced the idea of dense connection in the down-sampling stage to provide enhanced information for subsequent modules. Then we redesigned the D-Block to DP-Block in another cascade way to extract densely multi-scale contexts for road extraction. Finally, we adopted the self-attention mechanism in the up-sampling stage to learn long-distance pixel dependence to improve the connectivity of roads. Experimental results on CHN6-CUG Road Dataset prove that our FE-LinkNet performs better in accuracy and connectivity than D-LinkNet.
Qi Wang 0048, Haiwei Bai, Changtao He, Jian Cheng 0003
IGARSS4
2022 Global and Multi-Scale Feature Learning for Remote Sensing Scene Classification
abstract
Although the convolutional neural network (CNN)-based and vision transformer (ViT)-based methods have achieved effective remote sensing scene classification results in the past few years, the CNN's inductive bias and the ViT's single spatial scale limit the further improvement of accuracy. To address these problems, in this paper, we proposed a novel multi-scale vision transformer (MS-ViT) for remote sensing scene image classification, consists of two different spatial scale information streams that extract spatial multi-scale information, a hybrid attention module that fuses the information and a classifier to predict the scene category. We evaluated the effectiveness of our proposed method on two different remote sensing datasets, namely NWPU-RESISC45 and AID. The experimental results also show that our method outperforms CNN-based methods and the original ViT-based method in performance.
Ziying Xia, Guolong Gan, Siyu Liu 0003, Jian Cheng 0003
IGARSS5
2022 HCANet: A Hierarchical Context Aggregation Network for Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
Many practical applications of high-resolution remote sensing images (HRRSIs) are based on semantic segmentation. However, due to the complex ground object information contained in remote sensing images, it is difficult to make precise semantic segmentation of HRRSIs. In this letter, we proposed a hierarchical context aggregation network (HCANet) for the semantic segmentation of HRRSIs. The HCANet has an encoder-decoder structure which is similar to UNet. In the HCANet, we designed two Compact Atrous Spatial Pyramid Pooling (CASPP and CASPP+) modules. The CASPP modules replace the copy and crop operation in UNet to extract the multiscale context information of the multisemantic features of ResNet. The CASPP+ module is embedded in the middle layer of HCANet’s decoder to provide a strong aggregation path of contextual information. In the decoder of HCANet, the multiscale context information obtained by CASPP modules is hierarchically merged layer by layer for the semantic segmentation of HRRSIs. We compared our method with several of the most advanced methods on the ISPRS Vaihingen and Potsdam data sets. The final results demonstrate that our method can achieve outstanding performance.
Haiwei Bai, Jian Cheng 0003, Xia Huang 0005, Siyu Liu 0003, Changjian Deng
IEEE Geosci. Remote. Sens. Lett.2
2022 Semantic Segmentation for High-Resolution Remote-Sensing Images via Dynamic Graph Context Reasoning
abstract
Semantic segmentation for high-resolution remote-sensing (HRRS) images is one of the most challenging tasks in remote-sensing images understanding. Capturing long-range dependencies in feature representations is crucial for semantic segmentation. Recent graph-based global reasoning networks (GloRe) focus on modeling the global contextual relationship between latent nodes based on fully connected graph in interaction space. However, such a dense operation is susceptible to redundant features. Most importantly, it treats each node equally, ignoring the contextual relationship between nodes in graphs. In this work, we propose to explore more effective contextual representations in semantic segmentation by introducing dynamic graph contextual reasoning module overGloRe, dubbed DGCR. It incorporates local semantic information that represents the relationships between nodes to perform long-range contextual reasoning. More specifically, to provide effectively and flexible reasoning in graph-based reasoning approaches, we construct$k$-nearest neighbor (KNN) graphs rather than fully connected graphs using only the$k$closest nodes depends on pairwise semantic distance. Extensive experiments on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam datasets demonstrate the effectiveness and superiority of our proposed DGCR module over other state-of-the-art methods.
Yanzhou Su, Jian Cheng 0003, Wen Wang 0012, Haiwei Bai, Haijun Liu 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Multilevel Feature Fusion and Attention Network for High-Resolution Remote Sensing Image Semantic Labeling
abstract
Semantic labeling of high-resolution remote sensing images(HRRSIs) has always been an important research field in remote sensing images analysis. However, remote sensing images contain substantial low-level features and high-level features, which makes them quite difficult to be recognized. In this letter, we proposed a multi-level feature fusion and attention network(MFANet) to adaptively capture and fuse mutil-level features in a more effective and efficient manner. Specifically, the backbone of our network is divided into two branches – the detail branch and the semantic branch, where the detail branch extracts low-level features and the semantic branch extracts high-level features. The Deep Atrous Spatial Pyramid(DASPP) module is embedded in the end of the semantic branch to capture multiscale features as a supplement to high-level features. It is worth noting that the feature alignment and fusion (FAF) module is used to align and fuse features from different stages to enhance feature representation. Furthermore, the context attention (CA) module is employed to process feature map from the two branches to establish contextual dependencies in the spatial dimension and channel dimension, which can help network focus on more meaningful features. The experiments are carried out on the ISPRS Vaihingen and Potsdam datasets, and the results show that our proposed method has achieved better performance than other state-of-art methods.
Jian Cheng 0003, Haiwei Bai, Qi Wang 0048, Xingyu Liang
IEEE Geosci. Remote. Sens. Lett.2
2021 Semantic Segmentation for High-Resolution Remote Sensing Images by Light-Weight Network
abstract
Accurate segmentation of high-resolution remote sensing images is increasingly demanded, yet poses significant challenges for algorithm efficiency. Most current approaches pursue accuracy by employing global context information to enhance the overall consistency or utilize multi-scale features or attention mechanisms to optimize object details, without considering the network complexity uniformly. In this paper, we propose a light-weight semantic segmentation network for HRRS images by way of explicitly supervising the objects' body and edge features to optimize the overall consistency and object details of semantic segmentation at the same time. Furthermore, we introduce a score-based feature fusion module to establish the long-range dependency between pixels in the final stage of feature fusion (to combine the body and edge features) effectively. Experiments on ISPRS Vaihingen dataset show an obvious advantage of the proposed approach compared with the existing approaches. Specifically, it achieves 89.60% overall accuracy with only 2.83M parameters and 2.37GFLOPs computation costs.
Changjian Deng, Leikun Liang, Yanzhou Su, Changtao He, Jian Cheng 0003
IGARSS5
2021 Dual Lightweight Network with Attention and Feature Fusion for Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
Semantic segmentation of high-resolution remote sensing (HRRS) images has been a long-term research topic in the field of remote sensing. Nowdays, many excellent networks based on deep learning have been applied in various remote sensing fields. However, these networks always have a large number of network parameters and rely on extensive computing resources. To solve the above problems, we propose a lightweight dual branches network with the attention modules and the feature fusion module. The backbone networks of dual branches, which have fewer parameters, are used to obtain the detail information and the context information respectively. The attention modules are used to establish full-image dependencies over the local feature representations. The feature fusion module is used to fuse the low-leve features and the high-leve features effectively. Compared to other popular networks, our network has better results evaluated on ISPRS Vaihingen Dataset while with fewer parameters(8M).
Yulan Chen, Qijun Ma, Changtao He, Jian Cheng 0003
IGARSS5
2021 TTPP: Temporal Transformer with Progressive Prediction for efficient action anticipation
Wen Wang 0012, Xiaojiang Peng, Yanzhou Su, Yu Qiao 0001, Jian Cheng 0003
Neurocomputing5
2021 DCT-net: A deep co-interactive transformer network for video temporal grounding
Wen Wang 0012, Jian Cheng 0003, Siyu Liu 0003
Image Vis. Comput.2
2020 Light-Weight Attention Semantic Segmentation Network for High-Resolution Remote Sensing Images
abstract
Semantic segmentation of high-resolution remote sensing (HRRS) images becomes more and more important at present. Popular approaches use deep learning to solve this task, which depends on a large amount of labeled data and powerful computing resources. When computing resources or the labeled data are insufficient, their performance will be severely degraded. To deal with this problem, we proposed a light-weight network with attention modules for semantic segmentation of HRRS images. The depth and width of the network are designed, which has a small number of parameters to ensure the efficiency of training. The network adopts an encoder-decoder architecture. The feature maps of different scales from the encoder are concatenated together after resizing to carry out multi-scale feature fusion. To capture the global semantic information from the context, the attention mechanism is employed in the decoder. With one GTX2080Ti GPU and only 15 MB parameters the model owns, our light-weight network has quality results evaluated on ISPRS Vaihingen Dataset with fewer parameters compared to other popular approaches.
Siyu Liu 0003, Changtao He, Haiwei Bai, Jian Cheng 0003
IGARSS5
2020 Enhancing the discriminative feature learning for visible-thermal cross-modality person re-identification
Haijun Liu 0001, Jian Cheng 0003, Wen Wang 0012, Yanzhou Su, Haiwei Bai
Neurocomputing2
2020 Multi-Scale Based Context-Aware Net for Action Detection
abstract
We address the problem of action detection in continuous untrimmed video streams, based on the two-stage framework: one stage for action proposals generation and the other for proposals classification and refinement. The context features inside and outside a candidate region (proposal) are critical for classification in action detection. Therefore, effective integration of these features with different scales has become a fundamental problem. We contend that different action instances and candidate proposals may need different context features. To address this issue, we present a novel multiple scales based context-aware net (MSCA-Net) to effectively classify the action proposals for action detection in this paper. For each candidate action proposal, MSCA-Net takes its multiple regions with different temporal scales as input and then generates suitable context features. Based on the “candidate-control” mechanism of LSTM, the proposed MSCA-Net specially adopts the two-branch structure: Branch1 generates multi-scale context features for each candidate proposal, whereas Branch2 utilizes the context-aware gate function to control the message passing. Extensive experiments on THUMOS’14, Charades daily and ActivityNet action detection datasets, demonstrate the effectiveness of the designed structure and show how these context features influence the detection results.
Haijun Liu 0001, Shiguang Wang, Wen Wang 0012, Jian Cheng 0003
IEEE Trans. Multim.4
2019 Semantic Segmentation of High Resolution Remote Sensing Image Based on Batch-Attention Mechanism
abstract
Deep convolution neural network has been widely used in recent works for semantic segmentation of High Resolution Remote Sensing(HRRS) images. Because of the limitation of GPU memory, HRRS images are usually split into several sub-images for training convolutional neural networks. For each sub-image, the segmentation model may not have enough information to predict the segmentation map very well. In order to alleviate this problem, we propose to apply a batch-attention module to capture the discriminative information from similar objects, which come from other sub-images in a mini-batch. We also utilize global attention upsample module as the decoder to provide global context and fuse high and low level information better. We evaluate our model on the Potsdam dataset and achieve 88.30% pixAcc and 73.78% mIoU.
Yanzhou Su, Feng Wang 0015, Jian Cheng 0003
IGARSS5
2019 A Discriminatively Learned CNN Embedding For Remote Sensing Image Scene Classification
abstract
In this work, a discriminatively learned CNN embedding is proposed for remote sensing image scene classification. Our proposed siamese network simultaneously computes the classification loss function and the metric learning loss function of the two input images. Specifically, for the classification loss, we use the standard cross-entropy loss function to predict the classes of the images. For the metric learning loss, our siamese network learns to map the intra-class and inter-class input pairs to a feature space where intra-class inputs are close and inter-class inputs are separated by a margin. Concretely, for remote sensing image scene classification, we would like to map images from the same scene to feature vectors that are close, and map images from different scenes to feature vectors that are widely separated. Experiments are conducted on three different remote sensing image datasets to evaluate the effectiveness of our proposed approach. The results demonstrate that the proposed method achieves an excellent classification performance.
Wen Wang 0012, Lijun Du, Yinxing Gao, Yanzhou Su, Feng Wang 0015, Jian Cheng 0003
IGARSS6
2019 Gallery based k-reciprocal-like re-ranking for heavy cross-camera discrepancy in person re-identification
Haijun Liu 0001, Jian Cheng 0003
Neurocomputing2
2019 Deep Continuous Conditional Random Fields With Asymmetric Inter-Object Constraints for Online Multi-Object Tracking
abstract
Online multi-object tracking (MOT) is a challenging problem and has many important applications including intelligence surveillance, robot navigation, and autonomous driving. In existing MOT methods, individual object's movements and inter-object relations are mostly modeled separately and relations between them are still manually tuned. In addition, inter-object relations are mostly modeled in a symmetric way, which we argue is not an optimal setting. To tackle those difficulties, in this paper, we propose a deep continuous conditional random field (DCCRF) for solving the online MOT problem in a track-by-detection framework. The DCCRF consists of unary and pairwise terms. The unary terms estimate tracked objects' displacements across time based on visual appearance information. They are modeled as deep convolution neural networks, which are able to learn discriminative visual features for tracklet association. The asymmetric pairwise terms model inter-object relations in an asymmetric way, which encourages high-confidence tracklets to help correct errors of low-confidence tracklets and not to be affected by low-confidence ones much. The DCCRF is trained in an end-to-end manner for better adapting the influences of visual information as well as inter-object relations. Extensive experimental comparisons with state-of-the-arts as well as detailed component analysis of our proposed DCCRF on two public benchmarks demonstrate the effectiveness of our proposed MOT framework.
Hui Zhou 0005, Wanli Ouyang, Jian Cheng 0003, Xiaogang Wang 0001, Hongsheng Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 Temporal Action Detection by Joint Identification-Verification
abstract
Temporal action detection aims at not only recognizing action category but also detecting start time and end time for each action instance in an untrimmed video. The key challenge of this task is to accurately classify the actions and determine the temporal boundaries of each action instance. In temporal action detection benchmark: THUMOS 2014, large variations exist in the same action category while many similarities exist in different action categories, which always limit the performance of temporal action detection. To address this problem, we propose to use joint Identification-Verification network to reduce the intra-action variations and enlarge inter-action differences. The joint Identification-Verification network is a siamese network based on 3D ConvNets, which can simultaneously predict the action categories and the similarity scores for the input pairs of video proposal segments. Extensive experimental results on the challenging THUMOS 2014 dataset demonstrate the effectiveness of our proposed method compared to the existing state-of-art methods for temporal action detection in untrimmed videos. We further demonstrate that our model is a general framework by evaluating our approach on Charades dataset.
Wen Wang 0012, Haijun Liu 0001, Shiguang Wang, Jian Cheng 0003
ICPR5
2018 Visualizing deep neural network by alternately image blurring and deblurring
Feng Wang 0015, Haijun Liu 0001, Jian Cheng 0003
Neural Networks3
2018 Additive Margin Softmax for Face Verification
abstract
In this letter, we propose a conceptually simple and intuitive learning objective function, i.e., additive margin softmax, for face verification. In general, face verification tasks can be viewed as metric learning problems, even though lots of face verification models are trained in classification schemes. It is possible when a large-margin strategy is introduced into the classification model to encourage intraclass variance minimization. As one alternative, angular softmax has been proposed to incorporate the margin. In this letter, we introduce another kind of margin to the softmax loss function, which is more intuitive and interpretable. Experiments on LFW and MegaFace show that our algorithm performs better when the evaluation criteria are designed for very low false alarm rate.
Feng Wang 0015, Jian Cheng 0003, Weiyang Liu, Haijun Liu 0001
IEEE Signal Process. Lett.2
2018 Sequential Subspace Clustering via Temporal Smoothness for Sequential Data Segmentation
abstract
This paper develops a novel sequential subspace clustering method for sequential data. Inspired by the state-of-the-art methods, ordered subspace clustering, and temporal subspace clustering, we design a novel local temporal regularization term based on the concept of temporal predictability. Through minimizing the short-term variance on historical data, it can recover the temporal smoothness relationships in sequential data. Moreover, we claim that the local temporal regularization is more important than the global structural regularization for a specific task, such as sequential subspace clustering, which leads to a concise minimization objective function. To solve the bi-convex objective function, a simple and efficient optimization algorithm based on the alternate convex search method is devised to jointly learn the coding matrix and the dictionary. Furthermore, five baseline methods are also devised for comparison with our proposed method from different aspects. Extensive experimental results and comparisons with the state-of-the-art methods on three data sets demonstrate the effectiveness of the proposed temporal smoothness sequential subspace clustering method for sequential data.
Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015
IEEE Trans. Image Process.2
2018 Pedestrian Detection via Body Part Semantic and Contextual Information With DNN
abstract
Pedestrian detection has achieved great improve-ments in recent years, while complex occlusion handling and high-accurate localization are still the most important problems. To take advantage of the body part semantic information and the contextual information for pedestrian detection, we propose the part and context network (PCN) in this paper. A PCN is composed of three branches: the basic branch; the part branch; and the context branch. It specially utilizes two branches to detect the pedestrians through the body part semantic information and the contextual information, respectively. In the part branch, the semantic information of body parts can communicate with each other via long short-term memory (LSTM). In the context branch, we adopt a local competition mechanism (maxout) for adaptive context scale selection. By combining the outputs of all branches, we develop a strong complementary pedestrian detector with a lower miss rate and higher localization accuracy, especially for the occlusion pedestrian. The combination of the body part semantic information and the contextual information in pedestrian detection is fully explored in this paper. Comprehensive evaluations on three challenging pedestrian detection datasets (i.e., Caltech, INRIA and KITTI) well demonstrate the effectiveness of our proposed PCN. Code for PCN is publicly available on GitHub https://github.com/sunnyxiaohu/pcn_pedestrian.
Shiguang Wang, Jian Cheng 0003, Haijun Liu 0001, Feng Wang 0015, Hui Zhou 0005
IEEE Trans. Multim.2
2017 Sequential Subspace Clustering via Temporal Smoothness
abstract
This paper develops a novel sequential subspace clustering method for sequential data. Inspired by state-of-the-art methods ordered subspace clustering (OSC) and temporal subspace clustering (TSC), we design a novel local temporal regularization term based on the concept of temporal predictability, which is measured by short-term variance against long-term variance, to recover the temporal smoothness relationships in sequential data. To solve the bi-convex objective function, a simple and efficient optimization algorithm based on the alternate convex search (ACS) method is devised to jointly learn the codings matrix and dictionary. Extensive experimental results and comparisons with state-of-the-art methods on gesture and face datasets demonstrate the effectiveness of the proposed temporal smoothness sequential subspace clustering method for sequential data.
Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015
FG2
2017 Kinship verification based on status-aware projection learning
abstract
Kinship verification for parent-child is considered to be an asymmetric metric process, in which parents and children are associated with different status where the parents are priorly known to be significantly older than the children. To address the asymmetric metric learning, a status-aware projection learning (SaPL) method is proposed for facial image-based kinship verification, especially for the parent-child kinship. SaPL learns two status-specific projections to capture the significant appearance commonality between parents and children, respectively. Each status-specific projection consists of two components: a common component shared by the two status projections and a status-specific component. SaPL generally outperforms the one Mahalanobis distance metric. Extensive experimental results and comparisons with state-of-the-art approaches demonstrate the effectiveness of the proposed SaPL for kinship verification.
Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015
ICIP2
2017 Regularizing face verification nets for pain intensity regression
abstract
Limited labeled data are available for the research of estimating facial expression intensities. For instance, the ability to train deep networks for automated pain assessment is limited by small datasets with labels of patient-reported pain intensities. Fortunately, fine-tuning from a data-extensive pre-trained domain, such as face verification, can alleviate this problem. In this paper, we propose a network that fine-tunes a state-of-the-art face verification network using a regularized regression loss and additional data with expression labels. In this way, the expression intensity regression task can benefit from the rich feature representations trained on a huge amount of data for face verification. The proposed regularized deep regressor is applied to estimate the pain expression intensity and verified on the widely-used UNBC-McMaster Shoulder-Pain dataset, achieving the state-of-the-art performance. A weighted evaluation metric is also proposed to address the imbalance issue of different pain intensities.
Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran, Austin Reiter, Gregory D. Hager, Harry Quon, Jian Cheng 0003, Alan L. Yuille
ICIP8
2017 NormFace: L2 Hypersphere Embedding for Face Verification
abstract
Thanks to the recent developments of Convolutional Neural Networks, the performance of face verification methods has increased rapidly. In a typical face verification method, feature normalization is a critical step for boosting performance. This motivates us to introduce and study the effect of normalization during training. But we find this is non-trivial, despite normalization being differentiable. We identify and study four issues related to normalization through mathematical analysis, which yields understanding and helps with parameter settings. Based on this analysis we propose two strategies for training using normalized features. The first is a modification of softmax loss, which optimizes cosine similarity instead of inner-product. The second is a reformulation of metric learning by introducing an agent vector for each class. We show that both strategies, and small variants, consistently improve performance by between 0.2% to 0.4% on the LFW dataset based on two models. This is significant because the performance of the two models on LFW dataset is close to saturation at over 98%.
Feng Wang 0015, Xiang Xiang 0001, Jian Cheng 0003, Alan L. Yuille
ACM Multimedia3
2015 Silhouette Analysis for Human Action Recognition Based on Supervised Temporal t-SNE and Incremental Learning
abstract
This paper develops a human action recognition method for human silhouette sequences based on supervised temporal t-stochastic neighbor embedding (ST-tSNE) and incremental learning. Inspired by the SNE and its variants, ST-tSNE is proposed to learn the underlying relationship between action frames in a manifold, where the class label information and temporal information are introduced to well represent those frames from the same action class. As to the incremental learning, an important step for action recognition, we introduce three methods to perform the low-dimensional embedding of new data. Two of them are motivated by local methods, locally linear embedding and locality preserving projection. Those two techniques are proposed to learn explicit linear representations following the local neighbor relationship, and their effectiveness is investigated for preserving the intrinsic action structure. The rest one is based on manifold-oriented stochastic neighbor projection to find a linear projection from high-dimensional to low-dimensional space capturing the underlying pattern manifold. Extensive experimental results and comparisons with the state-of-the-art methods demonstrate the effectiveness and robustness of the proposed ST-tSNE and incremental learning methods in the human action silhouette analysis.
Jian Cheng 0003, Haijun Liu 0001, Feng Wang 0015, Hongsheng Li 0001, Ce Zhu
IEEE Trans. Image Process.1
2014 Silhouette analysis for human action recognition based on maximum spatio-temporal dissimilarity embedding
Jian Cheng 0003, Haijun Liu 0001, Hongsheng Li 0001
Mach. Vis. Appl.1
2014 Solving a Special Type of Jigsaw Puzzles: Banknote Reconstruction From a Large Number of Fragments
abstract
In this paper, we propose a method to solve a special type of jigsaw puzzles, reconstructing banknotes from a large number of fragments based on fragments' images. Existing jigsaw puzzle assembly algorithms have difficulty solving this problem effectively. A main limitation of these methods is that they do not leverage the following important observations: 1) an intact banknote's image is known and thus can be used as prior information; 2) if two aligned fragments overlap each other, they must not be from a same banknote. Based on these two important observations, a three-step method is proposed to reconstruct banknotes from their fragments. Each fragment is first aligned to its original position on the banknote by a RANSAC method. After evaluating every two aligned fragments' relationships, all fragments are embedded into a lower dimensional space and then clustered into small groups using a modified agglomerative clustering method. Fragments in a same cluster are likely to be from a same banknote. Experiments on both synthetic and real data demonstrate the effectiveness of our proposed method.
Hongsheng Li 0001, Yuanjie Zheng, Shaoting Zhang 0001, Jian Cheng 0003
IEEE Trans. Multim.4