Chenggang Yan 0001

dblp:146/1605 · also Chenggang Clarence Yan · DBLP profile ↗
← Back
252ranked-venue papers
23as first author
175since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 161 · 18 first-author · 108 since 2021Artificial intelligence and machine learning · 96 · 1 first-author · 73 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 4 first-author · 12 since 2021Computer networks · 11 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Forgetting Knowledge Localization and Isolation for Continual Forgetting of Pre-trained Vision Models
abstract
Continual forgetting task aims to continuously remove multiple target knowledge subsets from pre-trained models while maintaining the integrity of remaining knowledge. Existing methods suffer from both incomplete forgetting of target knowledge and unintended forgetting of indistinguishable remaining knowledge. To address these challenges, we propose the forgetting knowledge localization and isolation for continual forgetting in pre-trained vision models which precisely forgets target knowledge while reducing over-forgetting of remaining knowledge. To achieve precise forgetting, we first propose the forgetting knowledge layer localization to explore layers in the model which are more related to forgetting knowledge. Then, we design the forgetting knowledge parameter isolation to isolate the parameters sensitive to forgetting knowledge in these selected layers, mitigating over-forgetting of remaining knowledge. Finally, we fine-tune these isolated parameters and freeze the remaining parameters to achieve efficient forgetting while maintaining high performance on retained datasets. Extensive experimental results demonstrate that our method achieves superior performance over state-of-the-art methods across multiple continual forgetting tasks.
Zhiwen Yang 0003, Chenggang Yan 0001, Zongpeng Li, Xichun Sheng, Liang Li 0003
AAAI3
2026 2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information Propagation
abstract
Despite recent progress in adapting State Space Models such as Mamba to vision tasks, their intrinsic 1D scanning mechanism imposes limitations when applied to inherently 2D-structured data like images. Existing adaptations, including VMamba and 2DMamba, either suffer from inconsistency between scanning order and spatial locality or restrict inter-patch communication to singular paths, hindering effective information propagation. In this paper, we propose 2D-CrossScan, a novel 2D-compatible scan framework that enables spatially consistent, multi-path hidden state propagation by integrating modified state equations over two-dimensional neighborhoods. Furthermore, we mitigate redundant information accumulation due to overlapping paths via cross-directional subtraction. To fully align with the 2D spatial structure, we introduce a multi-directional scanning strategy that starts simultaneously from all four corners of the image, enabling diverse propagation paths and better feature integration. Our approach maintains efficiency, requiring only minimal architectural changes to existing Mamba variants. Experimental results demonstrate substantial improvements in multiple visual tasks, including object detection and semantic segmentation on PANDA and COCO datasets. Compared to baseline SSM-based methods, 2D-CrossScan consistently yields better spatial representations, as confirmed by extensive effective receptive field visualizations and attention analyses. These results highlight the importance of geometry-aware state propagation and validate 2D-CrossScan as a simple yet powerful extension to SSMs for vision.
Longlong Yu 0001, Wenxi Li, Yaoqi Sun, Chenggang Yan 0001
AAAI5
2026 Temporal Calibrating and Distilling for Scene-Text Aware Text-Video Retrieval
abstract
Existing text-video retrieval methods mainly focus on singlemodal video content (i.e., visual entities), often overlooking heterogeneous scene text that is ubiquitous in human environments. Although scene text in videos provides finegrained semantics for cross-modal retrieval, effectively utilizing it presents two key challenges: (1) Temporally dense scene text disrupts sync with sparse video frames, obstructing video understanding;(2) Redundant scene text and irrelevant video frames hinder the learning of discriminative temporal clues for retrieval. To address them, we propose a temporal scene-text calibrating and distilling (TCD) network for textvideo retrieval. Specifically, we first design a window-OCR captioner that aggregates dense scene text into OCR captions to facilitate feature interaction. Next, we devise a heterogeneous semantics calibration module that leverages scene text as a self-supervised signal to temporally align window-level OCR captions and frame-level video features. Further, we introduce a context-guided temporal clue distillation module to learn the complementary and relevant details between scene text and video modalities, thereby obtaining discriminative temporal clues for retrieval. Extensive experiments show that our TCD achieves state-of-the-art performance on three scene-text related benchmarks.
Zhiqian Zhao, Liang Li 0003, Xichun Sheng, Yaoqi Sun, Fang Kang, Chenggang Yan 0001
AAAI7
2026 A Comprehensive Survey on the Research and Development of RGB-T Salient Object Detection
abstract
Salient object detection (SOD) aims to mimic human visual perception by identifying the most eye-catching objects within a scene, and has remained a popular research topic for many years. The introduction of thermal (T) images offers additional information for challenging scenarios such as those with low light and complex backgrounds, and thus enhance performance when combined with RGB images. In this paper, we have, to the best of our ability, conducted the first comprehensive survey of dual-modality RGB-T SOD. We summarize and categorize published RGB-T SOD models, emphasizing their characteristics and features. Important components of these models are classified and elaborated, such as feature extraction, modality fusion, and loss function design. Following this, we analyze existing RGB-T SOD datasets and evaluation metrics. We evaluate a selection of representative SOD models using unified protocols and statistical analysis. We also present comparative experiments on the impact of the choice of loss function and the usage of datasets on performance. Finally, we consider several key issues and potential solutions in RGB-T SOD research, revealing promising directions for future efforts. We hope this survey will offer an effective way to understand the current state of the technology and, more importantly, stimulate discussion within the community.
Hongfa Wen, Qiang Zhao 0005, Junbo Ma, Zongpeng Li, Shuai Wang 0003, Chenggang Yan 0001
Comput. Vis. Media7
2026 Adaptive eco-cooperative adaptive cruise control for heterogeneous Vehicle platoons using online identification-informed deep reinforcement learning
Chenhao Xiong, Chunjie Zhai, Yuyuan Li 0001, Xiongding Liu, Chuqiao Chen, Chenggang Yan 0001, Yahong Chen
Eng. Appl. Artif. Intell.7
2026 Feature boosting and scale-aware network with multi-modal information for underwater salient object detection
Tingyu Wang 0002, Junzhe Lu 0002, Bin Wan, Rongfeng Lu, Yaoqi Sun, Duanpo Wu, Chenggang Yan 0001
Eng. Appl. Artif. Intell.8
2026 Feature distribution learning based on variance transfer and center shift for long-tailed classification
Chenxi Hong, Qiang Zhao 0005, Tao Tan 0002, Chenggang Yan 0001
Neurocomputing5
2026 AND-GS: Adaptive supervision of normal and depth in Gaussian splatting for accurate and efficient surface reconstruction
Xiang Le, Qiang Zhao 0005, Haofan Ren, Zhongtian Zheng, Tingyu Wang 0002, Jiyong Zhang 0001, Chenggang Yan 0001
Neurocomputing7
2026 Adaptive Resilient Control for Vehicle Platoons Under Hybrid Cyberattacks, Actuator Saturation, and Parameter Uncertainties
abstract
This paper presents a unified adaptive resilient control framework for heterogeneous connected vehicle platoons subjected to hybrid cyber-attacks, actuator saturation, parameter uncertainties, and communication delays. Unlike existing methods that often rely on simplified attack models or predefined parameters, the proposed strategy introduces a two-tiered architecture for comprehensive security and robustness. First, a hybrid attack detection mechanism leveraging a Random Forest classifier is developed, which operates under a novel DoS-first, FDI-second sequential logic to accurately identify both Denial-of-Service (DoS) and False Data Injection (FDI) attacks in real-time. Subsequently, an adaptive distributed control protocol, integrated with fuzzy logic systems and self-tuning laws, is designed to ensure platoon stability without prior knowledge of attack patterns or system parameters. The framework is rigorously underpinned by theoretical analysis, providing guarantees for both exponential stability and string stability. Extensive simulations demonstrate that the proposed approach not only maintains internal and string stability under hybrid attacks and system constraints but also exhibits superior transient performance, including faster recovery and enhanced robustness, compared to conventional baseline methods.
Xiyan Chen, Chunjie Zhai, Chuqiao Chen, Yuliang Ma 0002, Bo Wang 0031, Chenggang Yan 0001, Ya-Hong Chen
IEEE Internet Things J.6
2026 Multidimensional Hypergraph Fusion Network: A Novel Approach for Pediatric Seizure Detection Using Multimodal Physiological Signals
abstract
Accurate seizure detection from multimodal physiological signals is a key clinical imperative for improving patient care and diagnostic outcomes. However, conventional methods often fail to adequately model the complex, higher-order relationships within and across signal modalities, limiting performance. To address this limitation, we propose the Multi-dimensional Hypergraph Fusion Network (MHFN), a hypergraph-based multimodal fusion framework that explicitly models these intricate correlations. MHFN constructs three complementary hypergraphs: (1) an intra-modal hypergraph based on cosine similarity to capture fine-grained feature dependencies; (2) an inter-modal hypergraph that represents synergistic interactions across modalities; and (3) a temporal hypergraph incorporating dynamic time warping (DTW) to model evolutionary signal dynamics. By applying hypergraph convolutional networks (HGCN) to these structures, MHFN learns discriminative feature embeddings that are fused to generate an initial prediction. The clinical utility of this output is enhanced by a temporal correction pipeline, which smooths predictions and suppresses short-duration artifacts to produce coherent event-level classification. Comprehensive evaluations on clinical datasets demonstrate state-of-the-art performance, achieving an accuracy of 96.87%, precision of 98.06%, sensitivity of 96.17%, and F1-Score of 97.19%. Furthermore, we validate deployment feasibility on a Raspberry Pi 4B edge node, demonstrating low inference latency of 0.49 s and efficient power consumption (2.82 W). These results confirm that MHFN offers a robust, high-performance, and deployable paradigm for ambulatory clinical monitoring.
Qi Weng, Duanpo Wu, Tiejia Jiang, Yixuan Yuan, Xiaolong Ye, Chenggang Yan 0001
IEEE Internet Things J.8
2026 A Seizure Warning System Based on Multidimensional Attention Entropy and Improved Binary Mantis Search Algorithm
abstract
Seizure warning system (SWS) based on electroencephalography (EEG) is a prominent research focus in the internet of medical things (IoMT). An effective SWS enables epilepsy patients to proactively implement interventions before seizures. Effective feature representation of EEG reduces communication load from IoMT processing devices to cloud servers and enhances subsequent machine/deep learning classification performance. This paper proposes a seizure prediction system based on multidimensional attention entropy (MAE) and an improved binary mantis search algorithm (IBMSA). First, the MAE is calculated at the edge computing gateway by extracting attention entropy (AE) from each subband of single channel (AE-ESSC). Based on this, the algorithm further computes the AE of single channel with multiple subbands (AE-SCMS), the AE of single subband with multiple channels (AE-SSMC) and the AE of the original signal in single channel (AE-OSSC). Second, the IBMSA which is deployed on cloud servers uses a composite S-shaped and V-shaped function (CSVF) and dynamic stage selection strategy (DSSS) to find the optimal feature subset. The proposed algorithm is evaluated on data from the CHB-MIT dataset via leave-one-out cross-validation, with experimental results demonstrating an average sensitivity of 94.74% and a false prediction rate of 0.045 per hour. Additionally, edge deployment testing on Raspberry Pi 4 verifies the lightweight feature of the proposed system.
Duanpo Wu, Shuchang Zhang, Tiejia Jiang, Yixuan Yuan, Xiaolong Ye, Chenggang Yan 0001
IEEE Internet Things J.8
2026 Brain network construction and analysis for epilepsy: A methodology review
Yuge Yang, Duanpo Wu, Tiejia Jiang, Chenggang Yan 0001, Yixuan Yuan, Samaneh Kashi, Peiwu Qin
Neural Networks5
2026 ThermalGaussian++: Improving Alignment and Resolution for ThermalGaussian
abstract
Thermography is especially valuable for the military and other users of surveillance cameras. Some recent methods based on Neural Radiance Fields (NeRF) have been proposed to reconstruct thermal scenes in 3D from a set of thermal and RGB images. However, unlike NeRF, 3D Gaussian splatting (3DGS) prevails due to its rapid training and real-time rendering. In this work, we propose ThermalGaussian, the first thermal 3DGS approach capable of rendering high-quality images in RGB and thermal modalities. We first calibrate the RGB camera and the thermal camera to ensure that both modalities are accurately aligned. Subsequently, we use the registered images to learn the multimodal 3D Gaussians. To prevent the overfitting of any single modality, we introduce several multimodal regularization constraints. We also develop smoothing constraints tailored to the physical characteristics of the thermal modality. Besides, we contribute a real-world dataset named RGBT-Scenes, captured by a handheld thermal-infrared camera, facilitating future research on thermal scene reconstruction. Based on ThermalGaussian, we further introduce ThermalGaussian++ to improve the alignment and resolution of ThermalGaussian. To improve multimodal alignment, we design a multimodal pose optimization module. This module enables direct processing of non-aligned multimodal image pairs, reducing the need for professional calibration before each use. To improve thermal resolution, we also propose a multimodal joint super-resolution reconstruction module, which enhances the quality of low-resolution thermal fields. Additionally, we contribute a new dataset: RGBT-Scenes++, which offers higher-resolution thermal images. We conduct comprehensive experiments demonstrating that ThermalGaussian++ achieves photorealistic thermal rendering and improves RGB rendering quality. It significantly enhances both alignment and resolution, enabling better practical deployment. In addition, our multimodal regularization constraints reduce the model's storage requirements. The code and datasets will be released.
Rongfeng Lu, Ming Lu 0002, Tingyu Wang 0002, Haofan Ren, Yitian Xue, Chenggang Yan 0001
IEEE Trans. Pattern Anal. Mach. Intell.10
2026 Semantic-decoupled spatial partition guided point-supervised oriented object detection
Xinyuan Liu 0003, Yike Ma, Chenggang Yan 0001
Pattern Recognit.5
2026 Event-aware temporal modeling and semantic alignment for long-form video question answering
Xichun Sheng, Haibo Gong, Liang Li 0003, Chenggang Yan 0001, Tao Tan 0002
Pattern Recognit.6
2026 Understanding gait recognition through silhouette sequence disentanglement and fine-grained visualization
Shaoxiong Zhang 0001, Yixiu Liu, Jinkai Zheng, Liangqiong Qu, Ming Li 0073, Chenggang Yan 0001
Pattern Recognit.6
2026 SGM-Net: 3D Point Cloud Class-Incremental Segmentation via Semantic-Aware Global Modeling
Jinshuo Liu, Bingtao Ma, Zhidong Zhao, Chenggang Yan 0001, Shuai Wang 0003
IEEE Signal Process. Lett.4
2026 Spatial-Temporal Clue Reasoning Chain for Long Video Question Answering
Haibo Gong, Chenggang Yan 0001, Yaoqi Sun, Liang Li 0003
IEEE Trans. Circuits Syst. Video Technol.2
2026 Empirical Study on Fusion Strategy in RGB-T Salient Object Detection
abstract
In the research field of RGB-Thermal saliency object detection (RGB-T SOD), the effective exploitation of the complementary characteristics of the two modalities represents a major challenge for enhancing detection performance. Current fusion methodologies can be roughly classified into early fusion and middle fusion strategies, with prevalent techniques primarily encompassing concatenation, summation, and multiplication of the two modalities. To in depth assess the efficacy of these fusion strategies, we took an empirical investigation on them. Our findings demonstrate that the concatenation of middle features constitutes a more advantageous fusion strategy, yielding superior performance and demonstrating enhanced stability. Furthermore, observing the unique properties of thermal (T) images, we introduced gamma correction as a novel data augmentation methodology to RGB-T SOD. We subsequently evaluated the responses across varying correction parameter ranges, revealing that while the response to this data augmentation technique differs across various models, data augmentation is found to be effective in general. Building upon these findings, we proposed the Gamma Correction Network (GaCNet). Specifically, we also integrated image pyramid mechanism in a lightweight manner, which facilitates a more effective recovery of fine-grained image details. Significant improvement was achieved on commonly used RGB-T testing datasets, especially in VT821 dataset, manifesting the effectiveness of our method.
Shuai Wang 0003, Qiang Zhao 0005, Junbo Ma, Xichun Sheng, Yaoqi Sun, Hongfa Wen, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.8
2026 LLFeat: Noise-Aware Feature Matching Under Various Low-Light Conditions
Longjian Zeng, Zunjie Zhu, Ming Lu 0002, Bolun Zheng, Rongfeng Lu, Tingyu Wang 0002, Zhongtian Zheng, Yaoqi Sun, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.9
2026 Constituency-Tree-Induced Vision-Language Alignment for Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) integrate sophisticated large vision models (LVMs) to empower large language models (LLMs) with vision ability to perceive, reason, and interact in vision-language (V-L) tasks, while the modality bridge between two specialists becomes the bottleneck that translates visual signals into linguistic representations. However, most of the existing methods train the modality bridge with coarse-grained image-text pairs, neglecting the structural mapping between V-L semantics that facilitates modality translation from LVMs to LLMs. To mitigate this, we propose a Constituency-Tree-Induced Multimodal Bridging mechanism (CTIMB) that learns the fine-grained connection from LVMs to LLMs by the structural guidance from multi-modal constituency tree. Our approach consists of: 1) the multi-modal constituency-tree parser that jointly exploits the semantic structure of vision and language; 2) the lightweight connector that translates visual signals into linguistic representation and re-arranges them according to the constituency-tree structure; 3) the dynamic construction loss that aids in aligning the semantic structures derived from the tree parser and the connector. The CTIMB can learn the fine-grained mapping between visual and linguistic semantics, seamlessly bridge the LVMs and LLMs to enhance V-L tasks, and is more cost-efficient compared with current methods. Extensive experiments have demonstrated that our method more accurately interprets the visual features, enabling LLMs to conduct downstream tasks more effectively, and achieve superior performance with less training cost.
Yingchen Zhai, Ning Xu 0003, Hongshuo Tian, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.5
2026 Anatomy-Aware MR-Imaging-Only Radiotherapy
abstract
The synthesis of computed tomography images can supplement electron density information and eliminate MR-CT image registration errors. Consequently, an increasing number of MR-to-CT image translation approaches are being proposed for MR-only radiotherapy planning. However, due to substantial anatomical differences between various regions, traditional approaches often require each model to undergo independent development and use. In this paper, we propose a unified model driven by prompts that dynamically adapt to the different anatomical regions and generates CT images with high structural consistency. Specifically, it utilizes a region-specific attention mechanism, including a region-aware vector and a dynamic gating factor, to achieve MRI-to-CT image translation for multiple anatomical regions. Qualitative and quantitative results on three datasets of anatomical parts demonstrate that our models generate clearer and more anatomically detailed CT images than other state-of-the-art translation models. The results of the dosimetric analysis also indicate that our proposed model generates images with dose distributions more closely aligned to those of the real CT images. Thus, the proposed model demonstrates promising potential for enabling MR-only radiotherapy across multiple anatomical regions. we have released the source code for our RSAM model. The repository is accessible to the public at: https://github.com/yhyumi123/RSAM.
Hao Yang 0026, Yue Sun 0001, Chi Kin Lam, Qiang Zhao 0005, Xiangyu Xiong, Kunyan Cai, Behdad Dashtbozorg, Chenggang Yan 0001, Tao Tan 0002
IEEE Trans. Image Process.10
2026 Prompt Learning With Knowledge Regularization for Pre-Trained Vision-Language Models
abstract
Prompt learning is an effective way to adapt pre-trained models to downstream tasks by training a small number of additional learnable prompts. Recent studies address several early challenges by combining generalized knowledge from frozen pre-trained VL models with task-specific knowledge from training data as guidance for prompt learning. However, existing methods still struggle with the generalization-adaptation (GA) trade-off dilemma: excessive reliance on generalized knowledge hinders adaptation to downstream tasks, while overemphasis on task-specific knowledge undermines the inherent generalization capabilities of pre-trained models. To address this issue, we propose a novel prompt learning method called Prompt Learning with Knowledge Regularization (PLKR). PLKR effectively mitigates the GA trade-off dilemma by offering greater flexibility in adapting to task-specific knowledge while minimizing the disruption of pre-trained knowledge. Specifically, we propose category-invariant and topology-invariant knowledge regularization to preserve generalized knowledge: the former enhances category-level discriminative capabilities while allowing flexible task-specific learning, and the latter maintains global topological stability during adaptation to new tasks. Through the proposed regularization, PLKR improves the performance on both base and new tasks. We evaluate the effectiveness of our approach on four representative tasks over 11 datasets. Experimental results show our method outperforms existing SOTA methods by a large margin.
Boyang Guo, Liang Li 0003, Yaoqi Sun, Chenggang Yan 0001, Xichun Sheng
IEEE Trans. Multim.5
2026 Hybrid Debiasing Transformer With Adaptive Regularization for Video Moment Localization
abstract
Video Moment Localization (VML) is a task that seeks to pinpoint the most pertinent segment within an untrimmed video using a linguistic query. Previous works expose the severe data bias issues in VML and note that models avoid understanding visual-textual content by adapting the timestamp distribution. The work investigates data biases from both intrinsic and extrinsic perspectives: The former arises primarily from moment boundary ambiguity and the inputoutput information imbalance. The latter is attributed to the longtail distribution and the semantic bias resulting from the limited tail samples. To reduce the issues, we develop a hybrid multimodal debiasing network with a temporal consistency constraint for VML. Firstly, we propose a multi-temporal Transformer to alleviate boundary ambiguity by merging frame-wise features into segment-wise representations and dynamically aligning with moment boundaries. Subsequently, we implement a temporal consistency constraint to accentuate action information from complex moment context and mitigate the intrinsic bias caused by information imbalance. Moreover, we develop a hybrid linguistic activation module to mitigate the long-tail bias, which offers prior guidance to emphasize distinguishing clues from tail samples. Additionally, we introduce the prior-guided Transformer to alleviate the semantic bias by learning the global semantics of sentences, thereby circumventing the tail-sample overfitting issue. Comprehensive experiments demonstrate the efficacy of our proposed method across three datasets.
Jiong Yin, Liang Li 0003, Chenggang Yan 0001, Hongkui Wang, Yaoqi Sun, Zunjie Zhu
IEEE Trans. Multim.4
2026 Enhance Panoramic Object Detection Using Planar Image Datasets
abstract
Panoramic images have been used in various applications because of their ability to provide comprehensive spatial information. However, the high cost of obtaining panoramic images and the complexity of annotation pose serious obstacles to enhancing the performance of panoramic tasks by restricting the size and quality of datasets. The use of annotated planar images to synthesize panoramic images proves to be an effective approach to narrow this gap. Prior synthesis methods introduced new distortions that lead to inconsistent object shapes before and after synthesis. Concurrently, since the angle of view of the planar image is much smaller than that of the panoramic image, there is an issue of missing spatial information in the synthesized panoramic image. To address these challenges, we introduce a novel approach for converting planar images into panoramic ones with reduced deformation. For annotating targets in synthetic images, we develop a new algorithm based on determining the minimum spherical area to calculate spherical bounding boxes that closely adhere to object boundaries, rather than relying on an estimated target center point which results in inevitable calculation errors as in previous studies. Subsequently, we propose a new method for effectively filling any blank areas in the synthetic panoramic images to compensate for the loss of precision caused by absence of spatial information. Additionally, with a focus on the distortion characteristics of panoramic images, an innovative data augmentation strategy is devised to further enhance the model's ability to recognize objects in different positions. These methods are conducive to a more effective utilization of the rich planar image datasets for panoramic object detection tasks. In the experiments, we generated two synthetic panoramic datasets based on COCO for training models. Experimental results demonstrate that training with these synthetic datasets significantly improves prediction accuracy, far surpasses the state-of-the-art methods, and helps to unleash the full potential of panoramic object detection models. Source code is available athttps://github.com/longlong-yu/official-panorama-coco.git.
Longlong Yu 0001, Qiang Zhao 0005, Xinyuan Liu 0003, Tingyu Wang 0002, Chenggang Yan 0001
IEEE Trans. Multim.9
2026 Semantic Distribution and Authenticity Discrepancy Alignment for AI-Generated Image Detection
abstract
Generative models have achieved remarkable success in producing vivid images. Compared with real images, generated ones still show different semantic structures that features with different semantic classes collapse as a single cluster. Pioneer works leverage the discrepancy of semantic structure in fixed high-level semantic feature space to identify forgery images. Nevertheless, such frozen pre-trained representation models are insensitive to subtle forgery traces. Meanwhile, vanilla fine-tuning methods can distort the pre-trained semantic knowledge and collapse to the real-fake binary distribution, losing generalization capability in newly emerged generative models. In this paper, we propose thesemantic distribution and authenticity discrepancy alignment algorithm (STERM), which learns high-level semantic structures of real-world categories and low-level forgery traces for detecting AI-generated images from unseen generative models and frameworks. Specifically, we first capture semantic features of images by the frozen CLIP and further extract forgery features by a forgery encoder. Then, we propose semantic distribution alignment (SDA) to align the semantic structure of real-world categories by enforcing forgery feature distribution shifting towards the semantic feature space. Next, we introduce authenticity discrepancy alignment (ADA) to minimize the authenticity discrepancy between forgery and semantic features, constraining forgery features from collapsing into the source domain-biased distribution and learning the semantic structure of real-world categories. Extensive experiments on GAN-based and diffusion model-based datasets demonstrate the generalization capability of the proposed method. The source code is publicly available athttps://github.com/freshjh/STERM.
Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong
IEEE Trans. Multim.3
2026 QoS-Oriented Task Offloading in NOMA-Based Multi-UAV Cooperative MEC Systems
abstract
As resource-intensive and latency-sensitive applications continue to expand, the integration of unmanned aerial vehicles (UAVs) with mobile edge computing (MEC) has emerged as a viable solution, offering flexible, on-demand services for mobile users (MUs) without reliance on terrestrial infrastructure. The adoption of non-orthogonal multiple access (NOMA) further reduces latency by allowing MUs to offload tasks simultaneously over a single subchannel. However, many existing offloading methods do not explicitly incorporate a priority-based task scheduling mechanism and instead optimize task execution based on system constraints such as latency or energy consumption. To bridge this gap, we propose a QoS-oriented task offloading scheme that systematically optimizes task scheduling. We formulate an average system utility maximization problem that jointly optimizes UAVs’ 3D trajectories, MU association, task offloading ratios, and resource allocation. The optimization problem is inherently complex due to its non-convex nature and multiple constraints. To address this, we first employ Lagrange duality to decouple constraints, reducing computational complexity. Subsequently, we propose a novel improved soft actor-critic (ISAC) algorithm, which incorporates a perturbation term into the loss function to guide the training process away from local minima and toward globally optimal solutions. Through extensive simulation, we demonstrate that the ISAC algorithm guarantees convergence and significantly outperforms benchmark methods on offloading transmission rates, task completion rates, and overall system utility.
Lailong Luo, Deke Guo, Jiaju Wu 0004, Kaikai Chi, Chenggang Yan 0001, Xu-dong Dong 0001
IEEE Trans. Wirel. Commun.6
2026 Knowledge and multi-detail enhanced GAN for human-driven text-to-image synthesis
abstract
Human-driven text-to-image synthesis aims to create controllable images, which not only adhere to the semantic of given text but also incorporate the visual characteristics of given human. For example, given “a man on the beach” (text) along with a photo of human, the model aims to generate an image depicting the human on the beach. Although current diffusion-based methods have shown promise in this task, they face two major limitations: (1) The generated images appear to be a bit stiff and unnatural, almost like collages of human and backgrounds; (2) The details of human in the generated image are inconsistent with those in the input, losing the original identity. To address these issues, we present the Knowledge and Multi-Detail Enhanced GAN for the task of human-driven text-to-image synthesis. It employs external knowledge as references to improve the harmony between human and backgrounds, and uses CLIP’s multi-layer features to intensify human details. First, we search the database to retrieve external images that are similar to the given text, serving as our knowledge. Second, to preserve the human details, we present the Multi-Detail Enhancer, which uses the image encoder of CLIP to extract human representation at multiple levels. Third, to enhance the human-background naturalness, we present the Knowledge Attention Enhancer, which can seamlessly blend human, text, and knowledge by attentively retain useful information and filter out noise from knowledge. Finally, we introduce the dual discriminators to guide the entire network, which can facilitate the accurate capture of human details and generation of images. Extensive experiments demonstrate the superiority of our method with its efficiency and lower computational demands. It is about 300 times faster than diffusion-based models, uses only 5% of the parameters, and completes training in just two days on three V100 GPUs.
Ning Xu 0003, Zhewen Shen, Hongshuo Tian, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu
Vis. Informatics5
2025 Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change Captioning
abstract
Change captioning aims to describe the differences between two similar images using natural language, significantly aiding in understanding and monitoring changes. This challenging task requires a fine-grained understanding of subtle changes while resisting disturbances like viewpoint shifts and illumination variations. Existing methods often rely solely on global difference features and lack comprehensive alignment of linguistic and visual information, leading to overlooking fine-grained details and generating semantic hallucinated sentences. To address these limitations, we propose the region-aware difference distilling (RDD) network with attribute-guided contrastive regularization (ACR). The RDD uses global difference features to progressively distill regional difference features using learnable vectors, allowing for more precise identification of changed regions. The ACR enhances comprehensive alignment between linguistic and visual information by formulating Nouns-to-Objects (N2O) and Verbs-to-Actions (V2A) alignment losses to regularize the regional difference features. Promising results on three datasets demonstrate that our method outperforms the state-of-the-art change captioning methods.
Liang Li 0003, Qiang Zhao 0005, Hongkui Wang, Chenggang Yan 0001
AAAI6
2025 SdalsNet: Self-Distilled Attention Localization and Shift Network for Unsupervised Camouflaged Object Detection
abstract
Unsupervised camouflaged object detection (UCOD) poses significant challenges, primarily attributed to the absence of human labels. Existing UCOD methodologies, leveraging attention mechanisms, often struggle to achieve precise localization of camouflaged objects. To overcome this limitation, we introduce a groundbreaking fully unsupervised algorithm for attention-guided camouflaged object localization, shift, and inference, termed the self-distilled attention localization and shift network (SdalsNet). In this study, we formulate an attention localization methodology aimed at accurately identifying the central coordinate of the camouflaged object. Furthermore, we propose four distinct loss functions tailored to refine the precision of attentional positioning. These loss functions effectively constrain the distances between three types of class tokens, facilitating seamless attentional shifting across the input sample. Additionally, we design a sophisticated prediction inference technique to reconstruct the binary output of an attention map, thereby providing a comprehensive understanding of the detected camouflaged objects. Experimental results on four challenging COD benchmark datasets corroborate the effectiveness of our proposed approach, demonstrating notable superiority over state-of-the-art methods.
Peiyao Shou, Yixiu Liu, Wei Wang 0335, Yaoqi Sun, Zhigao Zheng 0001, Shangdong Zhu, Chenggang Yan 0001
AAAI7
2025 Frequency Dynamic Convolution for Dense Image Prediction
abstract
While Dynamic Convolution (DY-Conv) has shown promising performance by enabling adaptive weight selection through multiple parallel weights combined with an attention mechanism, the frequency response of these weights tends to exhibit high similarity, resulting in high parameter costs but limited adaptability. In this work, we introduce Frequency Dynamic Convolution (FDConv), a novel approach that mitigates these limitations by learning a fixed parameter budget in the Fourier domain. FDConv divides this budget into frequency-based groups with disjoint Fourier indices, enabling the construction of frequency-diverse weights without increasing the parameter cost. To further enhance adaptability, we propose Kernel Spatial Modulation (KSM) and Frequency Band Modulation (FBM). KSM dynamically adjusts the frequency response of each filter at the spatial level, while FBM decomposes weights into distinct frequency bands in the frequency domain and modulates them dynamically based on local content. Extensive experiments on object detection, segmentation, and classification validate the effectiveness of FD-Conv. We demonstrate that when applied to ResNet-50, FDConv achieves superior performance with a modest increase of +3.6M parameters, outperforming previous methods that require substantial increases in parameter budgets (e.g., CondConv +90M, KW +76.5M). Moreover, FD-Conv seamlessly integrates into a variety of architectures, including ConvNeXt, Swin-Transformer, offering a flexible and efficient solution for modern vision tasks. The code is made publicly available at https://github.com/Linwei-Chen/FDConv.
Lin Gu 0003, Liang Li 0003, Chenggang Yan 0001, Ying Fu 0001
CVPR4
2025 Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain Adaptation
abstract
This paper explores the Class-Incremental Source-Free Unsupervised Domain Adaptation (CI-SFUDA) problem, where the unlabeled target data come incrementally without access to labeled source instances. This problem poses two challenges, the interference of similar source-class knowledge in target-class representation learning and the shocks of new target knowledge to old ones. To address them, we propose the Multi-Granularity Class Prototype Topology Distillation (GROTO) algorithm, which effectively transfers the source knowledge to the class-incremental target domain. Concretely, we design the multi-granularity class prototype self-organization module and the prototype topology distillation module. First, we mine the positive classes by modeling accumulation distributions. Next, we introduce multi-granularity class prototypes to generate reliable pseudo-labels, and exploit them to promote the positive-class target feature self-organization. Second, the positive-class prototypes are leveraged to construct the topological structures of source and target feature spaces. Then, we perform the topology distillation to continually mitigate the shocks of new target knowledge to old ones. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on three public datasets.
Peihua Deng, Xichun Sheng, Chenggang Yan 0001, Yaoqi Sun, Ying Fu 0001, Liang Li 0003
CVPR4
2025 SMTPD: A New Benchmark for Temporal Prediction of Social Media Popularity
abstract
Social media popularity prediction task aims to predict the popularity of posts on social media platforms, which has a positive driving effect on application scenarios such as content optimization, digital marketing and online advertising. Though many studies have made significant progress, few of them pay much attention to the integration between popularity prediction with temporal alignment. In this paper, with exploring YouTube’s multilingual and multi-modal content, we construct a new social media temporal popularity prediction benchmark, namely SMTPD, and suggest a baseline framework for temporal popularity prediction. Through data analysis and experiments, we verify that temporal alignment and early popularity play crucial roles in social media popularity prediction for not only deepening the understanding of temporal dynamics of popularity in social media but also offering a suggestion about developing more effective prediction models in this field. Code is available at https://github.com/zhuwei321/SMTPD
Yijie Xu, Bolun Zheng, Hangjia Pan, Yuchen Yao, Ning Xu 0003, Anan Liu, Chenggang Yan 0001
CVPR9
2025 Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
abstract
Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker’s voice demonstrated in a short reference audio clip. This task demands the model bridge character performances and complicated prosody structures to build a high-quality video-synchronized dubbing track. The limited scale of movie dubbing datasets, along with the background noise inherent in audio data, hinder the acoustic modeling performance of trained models. To address these issues, we propose an acoustic-prosody disentangled two-stage method to achieve high-quality dubbing generation with precise prosody alignment. First, we propose a prosody-enhanced acoustic pre-training to develop robust acoustic modeling capabilities. Then, we freeze the pre-trained acoustic system and design an acoustic-disentangled framework to model prosodic text features and dubbing style while maintaining acoustic quality. Additionally, we incorporate an in-domain emotion analysis module to reduce the impact of visual domain shifts across different movies, thereby enhancing emotion-prosody alignment. Extensive experiments show that our method performs favorably against the state-of-the-art models on two primary benchmarks. The project is available at https://zzdoog.github.io/ProDubber/.
Zhedong Zhang, Liang Li 0003, Chenggang Yan 0001, Chunshan Liu, Anton van den Hengel, Yuankai Qi
CVPR3
2025 Debiased Teacher for Day-to-Night Domain Adaptive Object Detection
Liang Li 0003, Haibing Yin, Yaoqi Sun, Chenggang Yan 0001
ICCV6
2025 Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning
Jiong Yin, Liang Li 0003, Chenggang Yan 0001, Xichun Sheng
ICCV5
2025 ThermalGaussian: Thermal 3D Gaussian Splatting
abstract
Thermography is especially valuable for the military and other users of surveillance cameras. Some recent methods based on Neural Radiance Fields (NeRF) are proposed to reconstruct the thermal scenes in 3D from a set of thermal and RGB images. However, unlike NeRF, 3D Gaussian splatting (3DGS) prevails due to its rapid training and real-time rendering. In this work, we propose ThermalGaussian, the first thermal 3DGS approach capable of rendering high-quality images in RGB and thermal modalities. We first calibrate the RGB camera and the thermal camera to ensure that both modalities are accurately aligned. Subsequently, we use the registered images to learn the multimodal 3D Gaussians. To prevent the overfitting of any single modality, we introduce several multimodal regularization constraints. We also develop smoothing constraints tailored to the physical characteristics of the thermal modality. Besides, we contribute a real-world dataset named RGBT-Scenes, captured by a hand-hold thermal-infrared camera, facilitating future research on thermal scene reconstruction. We conduct comprehensive experiments to show that ThermalGaussian achieves photorealistic rendering of thermal images and improves the rendering quality of RGB images. With the proposed multimodal regularization constraints, we also reduced the model's storage cost by 90\%. Our project page is at https://thermalgaussian.github.io/.
Rongfeng Lu, Zunjie Zhu, Yuhang Qin, Ming Lu 0002, Chenggang Yan 0001, Anke Xue
ICLR7
2025 K-Buffers: A Plug-in Method for Enhancing Neural Fields with Multiple Buffers
abstract
Neural fields are now the central focus of research in 3D vision and computer graphics. Existing methods mainly focus on various scene representations, such as neural points and 3D Gaussians. However, few works have studied the rendering process to enhance the neural fields. In this work, we propose a plug-in method named K-Buffers that leverages multiple buffers to improve the rendering performance. Our method first renders K buffers from scene representations and constructs K pixel-wise feature maps. Then, We introduce a K-Feature Fusion Network (KFN) to merge the K pixel-wise feature maps. Finally, we adopt a feature decoder to generate the rendering image. We also introduce an acceleration strategy to improve rendering speed and quality. We apply our method to well-known radiance field baselines, including neural point fields and 3D Gaussian Splatting (3DGS). Extensive experiments demonstrate that our method effectively enhances the rendering performance of neural point fields and 3DGS.
Haofan Ren, Zunjie Zhu, Xiang Chen 0015, Ming Lu 0002, Rongfeng Lu, Chenggang Yan 0001
IJCAI6
2025 Region-Based Text-Consistent Augmentation for Multimodal Medical Segmentation
Kunyan Cai, Chenggang Yan 0001, Liangqiong Qu, Shuai Wang 0003, Tao Tan 0002
MICCAI (3)2
2025 Hypergraph-Guided Federated Distillation Learning for Efficient and Robust Multi-center fMRI Data Analysis
Yidan Xu, Xichun Sheng, Chenggang Yan 0001, Yaoqi Sun, Xiangmin Han, Yue Gao 0002
MICCAI (11)5
2025 VGNC: Reducing the Overfitting of Sparse-view 3DGS via Validation-guided Gaussian Number Control
abstract
Sparse-view 3D reconstruction is a fundamental yet challenging task in practical 3D reconstruction applications. Recently, many methods based on 3D Gaussian Splatting (3DGS) have been proposed to address sparse-view 3D reconstruction. Although these methods have made considerable advancements, they still show significant issues with overfitting. To reduce the overfitting, we introduce VGNC, a novel Validation-guided Gaussian Number Control approach based on generative novel view synthesis (NVS) models. To the best of our knowledge, this is the first attempt to alleviate the overfitting issue of sparse-view 3DGS with generative validation images. Specifically, we first introduce a validation image generation method based on a generative NVS model. We then propose a Gaussian number control strategy that utilizes generated validation images to determine optimal Gaussian numbers, thereby reducing the issue of overfitting. We conducted detailed experiments on various sparse-view 3DGS baselines and datasets to evaluate the effectiveness of VGNC. Extensive experiments show that our approach not only reduces overfitting but also improves rendering quality on the test set while decreasing the number of Gaussians. This reduction lowers storage demands and accelerates both training and rendering. Our code is available at: https://github.com/LinLif1869/VGNC.
Rongfeng Lu, Haofan Ren, Ming Lu 0002, Yaoqi Sun, Chenggang Yan 0001, Anke Xue
ACM Multimedia7
2025 DepthDark: Robust Monocular Depth Estimation for Low-Light Environments
abstract
In recent years, foundation models for monocular depth estimation have received increasing attention. Current methods mainly address typical daylight conditions, but their effectiveness notably decreases in low-light environments. There is a lack of robust foundational models for monocular depth estimation specifically designed for low-light scenarios. This largely stems from the absence of large-scale, high-quality paired depth datasets for low-light conditions and the effective parameter-efficient fine-tuning (PEFT) strategy. To address these challenges, we propose DepthDark, a robust foundation model for low-light monocular depth estimation. We first introduce a flare-simulation module and a noise-simulation module to accurately simulate the imaging process under nighttime conditions, producing high-quality paired depth datasets for low-light conditions. Additionally, we present an effective low-light PEFT strategy that utilizes illumination guidance and multiscale feature fusion to enhance the model's capability in low-light environments. Our method achieves state-of-the-art depth estimation performance on the challenging nuScenes-Night and RobotCar-Night datasets, validating its effectiveness using limited training data and computing resources.
Longjian Zeng, Zunjie Zhu, Rongfeng Lu, Ming Lu 0002, Bolun Zheng, Chenggang Yan 0001, Anke Xue
ACM Multimedia6
2025 Frequency-aware Correlation Discovering and Spatial Forgery Clue Distilling for Synthetic Image Detection
abstract
Recent text-to-image generative models facilitate creating vivid images with arbitrary contents that are indistinguishable from authentic ones by naked eyes. Despite progress in synthetic image detection, detecting the image from new generators remains challenging. Because advanced generators leave fewer visible forgery traces, while different generative frameworks produce varied forgery patterns. We notice that generative models consistently struggle with fine-detailed content generation, creating abnormal spatial dependencies among neighboring pixels in complex texture regions. In this paper, we propose a methodology of gazing local detail of forgery (GLDF) for generator agnostic synthetic image detection, which identifies prominent spatial dependencies to capture subtle forgery. Concretely, we design frequency-aware correlation discovering (FACD) module to learn dynamic filters by instance-adaptive frequency masking block for identifying prominent spatial deficiencies, which distributed in different spatial positions with various patterns. Furthermore, we introduce the spatial forgery clue distilling module (SFCD) to iteratively aggregate and refine spatial dependencies from different positions by spatial aggregating and prototype global interacting blocks. Extensive experiments demonstrate that GLDF outperforms state-of-the-art methods on detecting synthetic images from different generators.
Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong
ACM Multimedia3
2025 DASC-SPT: Towards Self-Supervised Panoramic Semantic Segmentation
abstract
Self-Supervised Semantic Segmentation, aiming to leverage masses of unlabeled data for boosting semantic segmentation, has been rapidly emerging as an active task in recent years. However, existing self-supervised semantic segmentation approaches mainly focus on planar images, leaving multiple distorted objects encountered in panoramic images unexplored due to the formidable challenge of handling heterogeneous degrees of distortions across different locations. In this paper, we propose a novel Self-Supervised Panoramic Semantic Segmentation model, termed DASC-SPT, built upon the mainstream contrastive learning framework. Towards distortions in panoramic images, we present two structures to better learn from distorted features by applying planar images. For the input images of self-supervision, we design a Spherical Projection Transformation (SPT) strategy that involves randomly projecting planar images onto various locations of the sphere to introduce the distortions. For pixel-wise distorted features, we construct a Deformation-aware Sampling Consistency (DASC) framework to further utilize the shared content and discrepancies caused by different distortions of paired views, where the deformation-aware consistency can be quantified on pixel-wise features. Both of the two components facilitate the model to adapt to distortions and boost panoramic semantic segmentation. Extensive comprehensive experiments on three panoramic datasets demonstrate the effectiveness and superiority of DASC-SPT approach.
Tianlong Tan, Bin Chen 0021, Hongliang Cao, Chenggang Yan 0001, Yike Ma
WACV4
2025 Channel pruning on frequency response
Lin Bie, Chenggang Yan 0001, Xibin Zhao, Yue Gao 0002
Sci. China Inf. Sci.4
2025 Consistency perception network for 360° omnidirectional salient object detection
Hongfa Wen, Zunjie Zhu, Xiaofei Zhou 0003, Jiyong Zhang 0001, Chenggang Yan 0001
Neurocomputing5
2025 Lightweight three-stream encoder-decoder network for multi-modal salient object detection
Junzhe Lu 0002, Tingyu Wang 0002, Bin Wan, Qiang Zhao 0005, Shuai Wang 0003, Yaoqi Sun, Yang Zhou 0052, Chenggang Yan 0001
J. Vis. Commun. Image Represent.8
2025 Counterfactual GAN for debiased text-to-image synthesis
Xianghua Kong, Ning Xu 0003, Zefang Sun, Zhewen Shen, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu
Multim. Syst.6
2025 Dynamic domain generalization for medical image segmentation
Zhiming Cheng, Mingxia Liu 0001, Chenggang Yan 0001, Shuai Wang 0003
Neural Networks3
2025 Latent Diffusion Enhanced Rectangle Transformer for Hyperspectral Image Restoration
abstract
The restoration of hyperspectral image (HSI) plays a pivotal role in subsequent hyperspectral image applications. Despite the remarkable capabilities of deep learning, current HSI restoration methods face challenges in effectively exploring the spatial non-local self-similarity and spectral low-rank property inherently embedded with HSIs. This paper addresses these challenges by introducing a latent diffusion enhanced rectangle Transformer for HSI restoration, tackling the non-local spatial similarity and HSI-specific latent diffusion low-rank property. In order to effectively capture non-local spatial similarity, we propose the multi-shape spatial rectangle self-attention module in both horizontal and vertical directions, enabling the model to utilize informative spatial regions for HSI restoration. Meanwhile, we propose a spectral latent diffusion enhancement module that generates the image-specific latent dictionary based on the content of HSI for low-rank vector extraction and representation. This module utilizes a diffusion model to generatively obtain representations of global low-rank vectors, thereby aligning more closely with the desired HSI. A series of comprehensive experiments were carried out on four common hyperspectral image restoration tasks, including HSI denoising, HSI super-resolution, HSI reconstruction, and HSI inpainting. The results of these experiments highlight the effectiveness of our proposed method, as demonstrated by improvements in both objective metrics and subjective visual quality.
Miaoyu Li, Ying Fu 0001, Tao Zhang 0042, Ji Liu 0003, Dejing Dou, Chenggang Yan 0001, Yulun Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Calibration-Free Raw Image Denoising via Fine-Grained Noise Estimation
abstract
Image denoising has progressed significantly due to the development of effective deep denoisers. To improve the performance in real-world scenarios, recent trends prefer to formulate superior noise models to generate realistic training data, or estimate noise levels to steer non-blind denoisers. In this paper, we bridge both strategies by presenting an innovative noise estimation and realistic noise synthesis pipeline. Specifically, we integrates a fine-grained statistical noise model and contrastive learning strategy, with a unique data augmentation to enhance learning ability. Then, we use this model to estimate noise parameters on evaluation dataset, which are subsequently used to craft camera-specific noise distribution and synthesize realistic noise. One distinguishing feature of our methodology is its adaptability: our pre-trained model can directly estimate unknown cameras, making it possible to unfamiliar sensor noise modeling using only testing images, without calibration frames or paired training data. Another highlight is our attempt in estimating parameters for fine-grained noise models, which extends the applicability to even more challenging low-light conditions. Through empirical testing, our calibration-free pipeline demonstrates effectiveness in both normal and low-light scenarios, further solidifying its utility in real-world noise synthesis and denoising tasks.
Yunhao Zou, Ying Fu 0001, Yulun Zhang 0001, Tao Zhang 0042, Chenggang Yan 0001, Radu Timofte
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Domain generalization for image classification with dynamic decision boundary
Zhiming Cheng, Mingxia Liu 0001, Defu Yang, Zhidong Zhao, Chenggang Yan 0001, Shuai Wang 0003
Pattern Recognit.5
2025 Content adaptive JND profile by leveraging HVS inspired channel modeling and perception oriented energy allocation optimization
Haibing Yin, Xia Wang 0006, Guangtao Zhai, Xiaofei Zhou 0003, Chenggang Yan 0001
Signal Process.5
2025 GLA: Global-Local Awareness for 3D Point Cloud Class-Incremental Semantic Segmentation
abstract
Semantic segmentation of 3D point clouds has garnered considerable attention in academic research and industrial applications. However, existing methods typically assume fixed semantic classes, which is unrealistic for practical scenarios where new classes emerge incrementally. This leads to catastrophic forgetting of previous knowledge and semantic shifts where annotations treat previous classes as background in new tasks. Moreover, the unordered and unstructured properties of point clouds further exacerbate catastrophic forgetting. To address these, we propose the Global-Local Awareness for 3D point cloud class-incremental semantic segmentation (i.e., GLA). Specifically, we introduce a global-aware modeling network (GAMN) based on the state space model, providing robust feature representations. Furthermore, to alleviate catastrophic forgetting, we propose a Graph Attention Knowledge Distillation (GAKD) module. GAKD adaptively integrates local geometric features through an attention mechanism, emphasizing the distillation of geometric structural knowledge. Notably, our approach integrates global context-awareness and local attention through the state space model and GAKD module, respectively. In addition, we introduce pseudo-labeling to alleviate semantic shift. Experiments on the S3DIS dataset validate the superiority of our approach.
Bingtao Ma, Shitong Zhang, Chenggang Yan 0001, Shuai Wang 0003
IEEE Signal Process. Lett.4
2025 SDRS: Sentiment-Aware Disentangled Representation Shifting for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) aims to leverage the complementary information from multiple modalities for affective understanding of user-generated videos. Existing methods mainly focused on designing sophisticated feature fusion strategies to integrate the separately extracted multimodal representations, ignoring the interference of the information irrelevant to sentiment. In this paper, we propose to disentangle the unimodal representations into sentiment-specific and sentiment-independent features, the former of which are fused for the MSA task. Specifically, we design a novel Sentiment-aware Disentangled Representation Shifting framework, termed SDRS, with two components.Interactive sentiment-aware representation disentanglementaims to extract sentiment-specific feature representations for each nonverbal modality by considering the contextual influence of other modalities with the newly developed cross-attention autoencoder.Attentive cross-modal representation shiftingtries to shift the textual representation in a latent token space using the nonverbal sentiment-specific representations after projection. The shifted representation is finally employed to fine-tune a pre-trained language model for multimodal sentiment analysis. Extensive experiments are conducted on three public benchmark datasets, i.e., CMU-MOSI, CMU-MOSEI, and CH-SIMS. The results demonstrate that the proposed SDRS framework not only obtains state-of-the-art results based solely on multimodal labels but also outperforms the methods that additionally require the labels of each modality.
Sicheng Zhao, Zhenhua Yang, Henglin Shi, Lingpengkun Meng, Bing Qin 0001, Chenggang Yan 0001, Jianhua Tao 0001, Guiguang Ding
IEEE Trans. Affect. Comput.7
2025 Pyramid Learnable Bandpass Filters for Ultra-High-Definition Image Demoiréing
abstract
Moiré patterns usually depend on the style of display grids and the position of shooting camera, appearing in the form of stripes, meshes or ripples, with various and irregular colors. Compared with low-resolution moiré images, high-definition (HD) and ultra-high-definition (UHD) moiré images exhibit more complex moiré patterns, e.g., wider distribution of moiré frequencies and higher coupling degree of moirés of different scales, which poses a greater challenge to the modeling capabilities of the model. To address these challenges, we propose a novel Pyramid Learnable Bandpass Filtering Network (PBNet) for demoiréing UHD images. Specifically, we propose a pyramid learnable bandpass filter (P-LBF) to perform multi-scale filtering in the same semantic context to obtain richer frequency domain information. The P-LBF contains three stages: aligning, filtering and fusing. First, we introduce a pyramid alignment (DA) to align neighbor pixels for eliminating the deviations raised by different styles of display grids and relative position of the shooting camera. Then, a pyramid filtering (PF) is conducted to model the complex and variable moiré patterns with aligned neighbor pixels. Finally, the frequency domain responses of these different scales are fused with a multi-dimensional feature fusion (MFF). The PBNet is constructed based on the P-LBF, incorporating a cross-layer feature fusion (CLF) module to facilitate more effective information interaction between features at different depths. Extensive experiments on four public datasets show that our model achieves state-of-the-art performance for both high- and low-resolution moiré images. The code is publicly available at:https://github.com/liuzhongqi1/PBNet.
Zhongqi Liu, Bolun Zheng, Qianyu Zhang 0002, Xu Jia 0012, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 Monocular Depth Estimation on Adverse Weathers With Curriculum Domain Distribution Alignment
abstract
Despite the remarkable success of monocular depth estimation, most works focus on ideal experiment conditions, such as favorable weather, where there is few environmental factors impacting the depth estimation system. In practical, when suffering from adverse weather conditions, such as fog and rain, the model trained on favorable weather degrades sharply as the domain shift, caused by the decreasing of visibility. To solve this problem, in this paper, we propose a Curriculum Domain Distribution Alignment (CDA) algorithm to learn the domain-invariant representation, progressively aligning data distributions across favorable weather and adverse weather in the feature space. Concretely, to construct a domain adaptation curriculum, we first separate the target domain into several subsets with increased domain discrepancy based on an optical model. Then, we bridge the distribution discrepancy between domains from easier to harder data by matching the source and target representation subspace. Furthermore, to control the distribution aligning pace, we introduce self-paced learning to learn a dynamic domain adaptation weight, promoting the generalization ability of monocular depth estimation networks against environmental factors. We conduct experiments with six monocular depth estimation frameworks on FoggyCityScapes, RainCityScapes, SnowCityscapes, and All-day Cityscapes, improving RMSE with 8.5 %, 30.5 %, 30.9 %, 20.9 %. The extraordinary performance demonstrates the effectiveness and generalizability of our method under adverse weather conditions.
Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong
IEEE Trans. Circuits Syst. Video Technol.3
2025 P2FCN: Environment-Independent UAV-View Geo-Localization via Pixel-to-Feature Co-Enhancement
abstract
This paper investigates the challenges of UAV-view geo-localization under extreme environmental changes, where significant cross-domain style differences can lead to degradation of model performance. Existing methods primarily focus on mitigating domain shift issues caused by environmental factors but generally overlook the direct interference of environmental noise. We argue that mitigating environmental noise is equally critical for extracting discriminative cross-view features and introduce a pixel-to-feature co-enhancement network (P2FCN). P2FCN comprises a style-noise dual suppression module (SNDS) and a part-based multi-dimensional feature learning strategy (PMDFL). Specifically, the SNDS module mitigates the stylistic discrepancies between cross-environment images through pixel-level dynamic adjustment, while reducing the introduction of environmental noise to a certain degree. PMDFL improves the representation generalization by purifying discriminative information within partitioned channel-and spatial-wise feature subspaces. Extensive experiments on two widely used benchmarks,i.e., University-1652 and SUES-200, demonstrate that the proposed method achieves state-of-the-art performance in multiple environments compared to existing methods.
Qiang Zhao 0005, Tingyu Wang 0002, Rongfeng Lu, Chenggang Yan 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 IFENet: Interaction, Fusion, and Enhancement Network for V-D-T Salient Object Detection
abstract
Visible-depth-thermal (VDT) salient object detection (SOD) aims to highlight the most visually attractive object by utilizing the triple-modal cues. However, existing models don't give sufficient exploration of the multi-modal correlations and differentiation, which leads to unsatisfactory detection performance. In this paper, we propose an interaction, fusion, and enhancement network (IFENet) to conduct the VDT SOD task, which contains three key steps including the multi-modal interaction, the multi-modal fusion, and the spatial enhancement. Specifically, embarking on the Transformer backbone, our IFENet can acquire multi-scale multi-modal features. Firstly, the inter-modal and intra-modal graph-based interaction (IIGI) module is deployed to explore inter-modal channel correlation and intra-modal long-term spatial dependency. Secondly, the gated attention-based fusion (GAF) module is employed to purify and aggregate the triple-modal features, where multi-modal features are filtered along spatial, channel, and modality dimensions, respectively. Lastly, the frequency split-based enhancement (FSE) module separates the fused feature into high-frequency and low-frequency components to enhance spatial information (i.e., boundary details and object location) of the salient object. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art models. Our code and results are available at https://github.com/Lx-Bao/IFENet.
Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Runmin Cong, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Image Process.7
2025 Source-Free Object Detection With Detection Transformer
abstract
Source-Free Object Detection (SFOD) enables knowledge transfer from a source domain to an unsupervised target domain for object detection without access to source data. Most existing SFOD approaches are either confined to conventional object detection (OD) models like Faster R-CNN or designed as general solutions without tailored adaptations for novel OD architectures, especially Detection Transformer (DETR). In this paper, we introduce Feature Reweighting ANd Contrastive Learning NetworK (FRANCK), a novel SFOD framework specifically designed to perform query-centric feature enhancement for DETRs. FRANCK comprises four key components: 1) an Objectness Score-based Sample Reweighting (OSSR) module that computes attention-based objectness scores on multi-scale encoder feature maps, reweighting the detection loss to emphasize less-recognized regions; 2) a Contrastive Learning with Matching-based Memory Bank (CMMB) module that integrates multi-level features into memory banks, enhancing class-wise contrastive learning; 3) an Uncertainty-weighted Query-fused Feature Distillation (UQFD) module that improves feature distillation through prediction quality reweighting and query feature fusion; and 4) an improved self-training pipeline with a Dynamic Teacher Updating Interval (DTUI) that optimizes pseudo-label quality. By leveraging these components, FRANCK effectively adapts a source-pre-trained DETR model to a target domain with enhanced robustness and generalization. Extensive experiments on several widely used benchmarks demonstrate that our method achieves state-of-the-art performance, highlighting its effectiveness and compatibility with DETR-based SFOD models.
Huizai Yao, Sicheng Zhao, Shuo Lu, Hui Chen 0013, Tengfei Xing, Chenggang Yan 0001, Jianhua Tao 0001, Guiguang Ding
IEEE Trans. Image Process.8
2025 TrackletGait: A Robust Framework for Gait Recognition in the Wild
abstract
Gait recognition aims to identify individuals based on their body shape and walking patterns. Though much progress has been achieved driven by deep learning, gait recognition in real-world surveillance scenarios remains quite challenging to current methods. Conventional approaches, which rely on periodic gait cycles and controlled environments, struggle with the non-periodic and occluded silhouette sequences encountered in the wild. In this paper, we propose a novel framework,TrackletGait, designed to address these challenges in the wild. We propose Random Tracklet Sampling, a generalization of existing sampling methods, which strikes a balance between robustness and representation in capturing diverse walking patterns. Next, we introduce Haar Wavelet-based Downsampling to preserve information during spatial downsampling. Finally, we present a Hardness Exclusion Triplet Loss, designed to exclude low-quality silhouettes by discarding hard triplet samples. TrackletGait achieves state-of-the-art results, with 77.8% and 80.4% rank-1 accuracy on the Gait3D and GREW datasets, respectively, while using only 10.3M backbone parameters. Extensive experiments are also conducted to further investigate the factors affecting gait recognition in the wild.
Shaoxiong Zhang 0001, Jinkai Zheng, Shangdong Zhu, Chenggang Yan 0001
IEEE Trans. Multim.4
2025 Progressive Decision Boundary Shifting for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) is attracting more attention from researchers for boosting the task-specific generalization on target domain. It focuses on addressing the domain shift between the labeled source domain and the unlabeled target domain. Recent biclassifier-based UDA models perform category-level alignment to reduce domain shift, and meanwhile, self-training is used for improving the discriminability of target instances. However, the error accumulation problem of instances with high semantic uncertainty may cause discriminability degradation and category-level misalignment. To solve this issue, we design the progressive decision boundary shifting algorithm, where stable category information of target instances is explored for learning a discriminability structure on target domain. Specifically, we first model the semantic uncertainty of instances by progressively shifting decision boundaries of category. Then, we introduce the uncertainty decoupling in a contrastive manner, where the discriminative information is learned from the source domain for instance with low semantic uncertainty. Furthermore, we minimize the predictive entropy of instances with high semantic uncertainty to reduce their prediction confidence. Extensive experiments on three popular datasets show that our model outperforms the current state-of-the-art (SOTA) UDA methods.
Liang Li 0003, Tongyu Lu, Yaoqi Sun, Chenggang Yan 0001, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.5
2025 Multi-Objective Unlearning in Recommender Systems via Preference Guided Pareto Exploration
abstract
Recommender systems typically collect and analyze user data, which raises the risk of privacy invasion. User-sensitive information can be leaked from the user portrait, e.g., user embedding, within recommender models. Therefore, the task of recommendation unlearning has been widely studied, aiming to eliminate the influence of target data on recommender models. This paper explores the extended concept of unlearning, which seeks to remove sensitive user information while retaining the essential information for recommendation purposes. Previous studies have primarily focused on extended unlearning in isolation, e.g., attribute unlearning. However, users often need to fulfill multiple unlearning objectives simultaneously. Therefore, we bridge this gap by introducing post-training multi-objective unlearning, which allows the concurrent fulfillment of multiple unlearning objectives while preserving recommendation performance. Note that the objectives may conflict with each other, leading to the compromise of one objective when minimizing the overall objective value. To address this challenge, we introduce a Pareto exploration approach that incorporates the recommendation performance as optimization guidance, allowing us to obtain the Pareto optimal solution through the trade-off between conflicting objectives. To adapt to practical scenarios where data is not accessible post-training, we utilize a data-free regularization to guide recommendation performance. We conducted extensive experiments on three real-world datasets, which demonstrate the effectiveness of our proposed method.
Yuyuan Li 0001, Yizhao Zhang, Weiming Liu 0005, Xiaohua Feng 0002, Zhongxuan Han, Chaochao Chen 0001, Chenggang Yan 0001
IEEE Trans. Serv. Comput.7
2025 Unpaired semantic neural person image synthesis
Yixiu Liu, Pengju Si, Shangdong Zhu, Chenggang Yan 0001, Shuai Wang 0003, Haibing Yin
Vis. Comput.5
2025 Loose-tight cluster regularization for unsupervised person re-identification
Yixiu Liu, Long Zhan, Pengju Si, Shaowei Jiang, Qiang Zhao 0005, Chenggang Yan 0001
Vis. Comput.7
2024 Coupled Confusion Correction: Learning from Crowds with Sparse Annotations
abstract
As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sourcing has been widely adopted to alleviate the cost of collecting labels, which also inevitably introduces label noise and eventually degrades the performance of the model. To learn from crowd-sourcing annotations, modeling the expertise of each annotator is a common but challenging paradigm, because the annotations collected by crowd-sourcing are usually highly-sparse. To alleviate this problem, we propose Coupled Confusion Correction (CCC), where two models are simultaneously trained to correct the confusion matrices learned by each other. Via bi-level optimization, the confusion matrices learned by one model can be corrected by the distilled data from the other. Moreover, we cluster the ``annotator groups'' who share similar expertise so that their confusion matrices could be corrected together. In this way, the expertise of the annotators, especially of those who provide seldom labels, could be better captured. Remarkably, we point out that the annotation sparsity not only means the average number of labels is low, but also there are always some annotators who provide very few labels, which is neglected by previous works when constructing synthetic crowd-sourcing annotations. Based on that, we propose to use Beta distribution to control the generation of the crowd-sourcing labels so that the synthetic annotations could be more consistent with the real-world ones. Extensive experiments are conducted on two types of synthetic datasets and three real-world datasets, the results of which demonstrate that CCC significantly outperforms state-of-the-art approaches. Source codes are available at: https://github.com/Hansong-Zhang/CCC.
Hansong Zhang 0003, Shikun Li, Dan Zeng 0001, Chenggang Yan 0001, Shiming Ge
AAAI4
2024 Quad Bayer Joint Demosaicing and Denoising Based on Dual Encoder Network with Joint Residual Learning
abstract
The recent imaging technology Quad Bayer CFA brings better imaging PSNR and higher visual quality compared to traditional Bayer CFA, but also serious challenges for demosaicing and denoising during the ISP pipeline. In this paper, we propose a novel dual encoder network, namely DRNet, to achieve joint demosaicing and denoising for Quad Bayer CFA. The dual encoders are carefully designed in that one is mainly constructed by a joint residual block to jointly estimate the residuals for demosaicing and denoising separately. In contrast, the other one is started with a pixel modulation block which is specially designed to match the characteristics of Quad Bayer pattern for better feature extraction. We demonstrate the effectiveness of each proposed component through detailed ablation investigations. The comparison results on public benchmarks illustrate that our DRNet achieves an apparent performance gain~(0.38dB to the 2nd best) from the state-of-the-art method and balances performance and efficiency well. The experiments on real-world images show that the proposed method could enhance the reconstruction quality from the native ISP algorithm.
Bolun Zheng, Haoran Li 0025, Tingyu Wang 0002, Xiaofei Zhou 0003, Chenggang Yan 0001
AAAI7
2024 Context-aware Difference Distilling for Multi-change Captioning
abstract
Multi-change captioning aims to describe complex and coupled changes within an image pair in natural language.Compared with singlechange captioning, this task requires the model to have higher-level cognition ability to reason an arbitrary number of changes.In this paper, we propose a novel context-aware difference distilling (CARD) network to capture all genuine changes for yielding sentences.Given an image pair, CARD first decouples context features that aggregate all similar/dissimilar semantics, termed common/difference context features.Then, the consistency and independence constraints are designed to guarantee the alignment/discrepancy of common/difference context features.Further, the common context features guide the model to mine locally unchanged features, which are subtracted from the pair to distill locally difference features.Next, the difference context features augment the locally difference features to ensure that all changes are distilled.In this way, we obtain an omni-representation of all changes, which is translated into linguistic sentences by a transformer decoder.Extensive experiments on three public datasets show CARD performs favourably against state-of-the-art methods.
Yunbin Tu, Liang Li 0003, Li Su 0003, Zhengjun Zha, Chenggang Yan 0001, Qingming Huang
ACL (1)5
2024 Rethinking Boundary Discontinuity Problem for Oriented Object Detection
abstract
Oriented object detection has been developed rapidly in the past few years, where rotation equivariance is crucial for detectors to predict rotated boxes. It is expected that the prediction can maintain the corresponding rotation when objects rotate, but severe mutation in angular prediction is sometimes observed when objects rotate near the boundary angle, which is well-known boundary discontinuity problem. The problem has been long believed to be caused by the sharp loss increase at the angular boundary, and widely used joint-optim IoU-like methods deal with this problem by loss-smoothing. However, we experimentally find that even state-of-the-art IoU-like methods actually fail to solve the problem. On further analysis, we find that the key to solution lies in encoding mode of the smoothing function rather than in joint or independent optimization. In existing IoU-like methods, the model essentially attempts to fit the angular relationship between box and object, where the break point at angular boundary makes the predictions highly unstable. To deal with this issue, we propose a dual-optimization paradigm for angles. We decouple reversibility and joint-optim from single smoothing function into two distinct entities, which for the first time achieves the objectives of both correcting angular boundary and blending angle with other parameters. Extensive experiments on multiple datasets show that boundary discontinuity problem is well-addressed. More-over, typical IoU-like methods are improved to the same level without obvious performance gap. The code is available at https://github.com/hangxu-cv/cvpr24acm.
Xinyuan Liu 0003, Yike Ma, Zunjie Zhu, Chenggang Yan 0001
CVPR6
2024 Distractors-Immune Representation Learning with Cross-Modal Contrastive Regularization for Change Captioning
Yunbin Tu, Liang Li 0003, Li Su 0003, Chenggang Yan 0001, Qingming Huang
ECCV (43)4
2024 A Cross-modal Fusion Method for Multispectral Small Ship Detection
abstract
The fusion module of RGB and infrared (IR) remote sensing images is the key of multispectral ship detection. Existing works have shown that the cross-attention-based feature fusion can achieve good performance by extracting the complementary information of RGB and IR modalities. However, the existing commonly used cross-attention mechanisms introduce lots of redundancy parameters and mainly focus on global feature interaction of multispectral images, ignoring local detail information that is also important for small ship detection. In this paper, we propose a novel multispectral ship detection approach named LoGFusion. In LoGFusion, we design the cross stage partial module with partial convolution (CSPMPC) to reduce feature redundancy and utilize the local cross-modal fusion module (LoCFM) and global cross-modal fusion module (GCFM) to capture both local and global cross-modal features. Furthermore, we introduce a Multispectral Small Ship Dataset (MSSD) containing over 5k ship targets for small target detection. Experiments on MSSD validate the effectiveness of our method in terms of small ship detection in multispectral images.
Yang Liu 0119, Yu Liu 0005, Xueqian Wang 0002, Linping Zhang, Zhizhuo Jiang, Yaowen Li, Chenggang Yan 0001, Ying Fu 0001, Tao Zhang 0042
FUSION7
2024 Improving Radiology Report Generation with D2-Net: When Diffusion Meets Discriminator
abstract
Radiology report generation (RRG) aims to automatically provide observations and insight into a patient’s condition based on radiology images, which is able to greatly reduce the workload of physicians on the premise of ensuring the quality of medical treatment. Existing works leverage the Transformer decoder to generate reports word-by-wordly. However, unlike image captioning, radiology reports are long text containing many semantic words. The autoregressive method, such as the Transformer-base method, will accumulate errors in the generation process and generate unsatisfied reports. Benefiting from the recent success of Diffusion, we propose a novel Diffusion-based paradigm for RRG, which leverages visual information as a condition, making the generation process focus on pathological features within the radiology image. Meanwhile, we integrate a discriminator into each layer of the Diffusion to actively judge whether the generated words are meaningful, which, on the one hand, controls the length of predicted reports and, on the other hand, calibrates confidence scores and token generation results, improving the quality of the generated reports. Extensive experiment results demonstrate the superiority of our proposed method. Source code is available at: https://github.com/Yuda-Jin/D-2-Net.
Yuda Jin, Weidong Chen 0013, Yuanhe Tian, Yan Song 0004, Chenggang Yan 0001, Zhendong Mao 0001
ICASSP5
2024 Generating High-Quality Symbolic Music Using Fine-Grained Discriminators
Zhedong Zhang, Liang Li 0003, Hongkui Wang, Chenggang Yan 0001, Jian Yang 0001, Yuankai Qi
ICPR (20)6
2024 Language-Assisted Siamese Contrastive Framework for Fine-Grained Remote Sensing Ship Image Retrieval
abstract
As the number of remote sensing (RS) images increases, it is crucial to retrieval ship targets according to specific demands. The existing ship image retrieval methods only extract features from the image modality, which may not fully utilize the rich text information available and ignore the high-level hierarchical relations between ship classes. In this paper, we propose a language-assisted siamese contrastive framework, namely LASCF, for fine-grained ship retrieval in RS images. In the new LASCF, the siamese vision models are employed to measure the similarity between images. Moreover, a label text encoder with a pretrained language model is designed to extract the high-level semantic information from labels, and thus the information of the hierarchical relations between ship classes are fused in LASCF. Finally, the multimodal similarity measurement module based on contrastive learning is proposed to optimize the siamese vision models. The experimental results show that the proposed LASCF outperforms several existing state-of-the-art methods.
Zhizhuo Jiang, Yu Liu 0005, Yaowen Li, Xueqian Wang 0002, Chenggang Yan 0001
IGARSS7
2024 Stochastic Context Consistency Reasoning for Domain Adaptive Object Detection
Liang Li 0003, Chenggang Yan 0001, Hongkui Wang, Shuai Wang 0003, Heng Jin
ACM Multimedia4
2024 Domain Shared and Specific Prompt Learning for Incremental Monocular Depth Estimation
Zhiwen Yang 0003, Liang Li 0003, Tingyu Wang 0002, Yaoqi Sun, Chenggang Yan 0001
ACM Multimedia6
2024 From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency Learning
Zhedong Zhang, Liang Li 0003, Gaoxiang Cong 0001, Haibing Yin, Chenggang Yan 0001, Anton van den Hengel, Yuankai Qi
ACM Multimedia6
2024 It Takes Two: Accurate Gait Recognition in the Wild via Cross-granularity Alignment
abstract
Existing studies for gait recognition primarily utilized sequences of either binary silhouette or human parsing to encode the shapes and dynamics of persons during walking. Silhouettes exhibit accurate segmentation quality and robustness to environmental variations, but their low information entropy may result in sub-optimal performance. In contrast, human parsing provides fine-grained part segmentation with higher information entropy, but the segmentation quality may deteriorate due to the complex environments. To discover the advantages of silhouette and parsing and overcome their limitations, this paper proposes a novel cross-granularity alignment gait recognition method, named XGait, to unleash the power of gait representations of different granularity. To achieve this goal, the XGait first contains two branches of backbone encoders to map the silhouette sequences and the parsing sequences into two latent spaces, respectively. Moreover, to explore the complementary knowledge across the features of two representations, we design the Global Cross-granularity Module (GCM) and the Part Cross-granularity Module (PCM) after the two encoders. In particular, the GCM aims to enhance the quality of parsing features by leveraging global features from silhouettes, while the PCM aligns the dynamics of human parts between silhouette and parsing features using the high information entropy in parsing sequences. In addition, to effectively guide the alignment of two representations with different granularity at the part level, an elaborate-designed learnable division mechanism is proposed for the parsing features. Finally, comprehensive experiments on two large-scale gait datasets not only show the superior performance of XGait with the Rank-1 accuracy of 80.5% on Gait3D and 88.3% CCPG but also reflect the robustness of the learned features even under challenging conditions like occlusions and cloth changes
Jinkai Zheng, Xinchen Liu, Boyue Zhang 0004, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Yongdong Zhang 0001
ACM Multimedia4
2024 CURE4Rec: A Benchmark for Recommendation Unlearning with Deeper Influence
abstract
With increasing privacy concerns in artificial intelligence, regulations have mandated the right to be forgotten, granting individuals the right to withdraw their data from models. Machine unlearning has emerged as a potential solution to enable selective forgetting in models, particularly in recommender systems where historical data contains sensitive user information. Despite recent advances in recommendation unlearning, evaluating unlearning methods comprehensively remains challenging due to the absence of a unified evaluation framework and overlooked aspects of deeper influence, e.g., fairness. To address these gaps, we propose CURE4Rec, the first comprehensive benchmark for recommendation unlearning evaluation. CURE4Rec covers four aspects, i.e., unlearning Completeness, recommendation Utility, unleaRning efficiency, and recommendation fairnEss, under three data selection strategies, i.e., core data, edge data, and random data. Specifically, we consider the deeper influence of unlearning on recommendation fairness and robustness towards data with varying impact levels. We construct multiple datasets with CURE4Rec evaluation and conduct extensive experiments on existing recommendation unlearning methods. Our code is released at https://github.com/xiye7lai/CURE4Rec.
Chaochao Chen 0001, Jiaming Zhang 0009, Yizhao Zhang, Lingjuan Lyu, Yuyuan Li 0001, Biao Gong, Chenggang Yan 0001
NeurIPS8
2024 Aerial-view geo-localization based on multi-layer local pattern cross-attention network
Haoran Li 0025, Tingyu Wang 0002, Qiang Zhao 0005, Shaowei Jiang, Chenggang Yan 0001, Bolun Zheng
Appl. Intell.6
2024 Enhanced local distribution learning for real image super-resolution
Yaoqi Sun, Aiai Huang, Chenggang Yan 0001, Bolun Zheng
Comput. Vis. Image Underst.5
2024 ADNet: Anti-noise dual-branch network for road defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001
Eng. Appl. Artif. Intell.8
2024 A streamlined framework for BEV-based 3D object detection with prior masking
Qinglin Tong, Junjie Zhang 0002, Chenggang Yan 0001, Dan Zeng 0001
Image Vis. Comput.3
2024 Coplane-constrained sparse depth sampling and local depth propagation for depth estimation
Zhiwen Yang 0003, Chuqiao Chen, Hongkui Wang, Tingyu Wang 0002, Chenggang Yan 0001, Yihong Gong
Image Vis. Comput.6
2024 GoLDFormer: A global-local deformable window transformer for efficient image restoration
Bolun Zheng, Chenggang Yan 0001, Zunjie Zhu, Tingyu Wang 0002, Gregory Slabaugh, Shanxin Yuan
J. Vis. Commun. Image Represent.3
2024 Efficient 2D transform hardware architecture for the versatile video coding standard
Qinghua Sheng, Changcai Lai, Xiaofeng Huang, Haibing Yin, Chenggang Yan 0001
J. Vis. Commun. Image Represent.8
2024 Feature rectification and enhancement for no-reference image quality assessment
Daoquan Huang, Zhuonan Shen, Chenggang Yan 0001, Bolun Zheng
J. Vis. Commun. Image Represent.6
2024 Learning degradation priors for reliable no-reference image quality assessment
Zhuonan Shen, Bolun Zheng, Dingguo Yu, Chenggang Yan 0001
J. Vis. Commun. Image Represent.7
2024 GINet:Graph interactive network with semantic-guided spatial refinement for salient object detection in optical remote sensing images
Chenwei Zhu, Xiaofei Zhou 0003, Liuxin Bao, Hongkui Wang, Shuai Wang 0003, Zunjie Zhu, Chenggang Yan 0001, Jiyong Zhang 0001
J. Vis. Commun. Image Represent.7
2024 Body Joint Boundary Prototype Match for Few-Shot Remote Sensing Semantic Segmentation
abstract
Deep networks require a large number of samples for optimization, so few-shot segmentation in remote sensing scenes is still an open problem. However, this challenge is exacerbated by the feature blurring and aliasing of bodies (low frequency) and boundaries (high frequency). The existing methods usually only focus on the body part of the class, that is, the low-frequency part, and ignore the critical role of boundary information, that is, high-frequency details, on feature representation. In this letter, we propose a novel body joint boundary prototype match (B2PM) approach that aims to enable prior learning of low- and high-frequency information by explicitly modeling the body and boundary features of objects. First, body-aware prototype learning (BodyPL) realizes the adaptive modeling of the body part of the object through a precise farthest point sampling (FPS) initialization algorithm and an adaptive part shift (APS) strategy, which alleviates the feature ambiguity of the body. Second, boundary-aware prototype learning (BoundPL) explicitly models boundary prototypes by building a patch division and assignment strategy to alleviate feature aliasing at boundaries. Finally, prototype match performs prior knowledge aggregation by computing the affinity between query features and support prototypes. Extensive experiments on commonly used benchmarks (iSAID and PASCAL VOC) demonstrate that B2PM improves the state of the art by significant margins.
Yongqiang Mao, Zhizhuo Jiang, Yu Liu 0005, Yaowen Li, Chenggang Yan 0001, Bolun Zheng
IEEE Geosci. Remote. Sens. Lett.6
2024 Dynamic interactive refinement network for camouflaged object detection
Yaoqi Sun, Lidong Ma, Peiyao Shou, Hongfa Wen, Yixiu Liu, Chenggang Yan 0001, Haibing Yin
Neural Comput. Appl.7
2024 Frequency-Aware Feature Fusion for Dense Image Prediction
abstract
Dense image prediction tasks demand features with strong category information and precise spatial boundary details at high resolution. To achieve this, modern hierarchical models often utilize feature fusion, directly adding upsampled coarse features from deep layers and high-resolution features from lower levels. In this paper, we observe rapid variations in fused feature values within objects, resulting in intra-category inconsistency due to disturbed high-frequency features. Additionally, blurred boundaries in fused features lack accurate high frequency, leading to boundary displacement. Building upon these observations, we propose Frequency-Aware Feature Fusion (FreqFusion), integrating an Adaptive Low-Pass Filter (ALPF) generator, an offset generator, and an Adaptive High-Pass Filter (AHPF) generator. The ALPF generator predicts spatially-variant low-pass filters to attenuate high-frequency components within objects, reducing intra-class inconsistency during upsampling. The offset generator refines large inconsistent features and thin boundaries by replacing inconsistent features with more consistent ones through resampling, while the AHPF generator enhances high-frequency detailed boundary information lost during downsampling. Comprehensive visualization and quantitative analysis demonstrate that FreqFusion effectively improves feature consistency and sharpens object boundaries. Extensive experiments across various dense prediction tasks confirm its effectiveness.
Ying Fu 0001, Lin Gu 0003, Chenggang Yan 0001, Tatsuya Harada, Gao Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Deep Joint Semantic Adaptation Network for Multi-source Unsupervised Domain Adaptation
Zhiming Cheng, Shuai Wang 0003, Defu Yang, Mang Xiao, Chenggang Yan 0001
Pattern Recognit.6
2024 TMNet: Triple-modal interaction encoder and multi-scale fusion decoder network for V-D-T salient object detection
Bin Wan, Chengtao Lv, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Hongkui Wang, Chenggang Yan 0001
Pattern Recognit.7
2024 Multiple-environment Self-adaptive Network for aerial-view geo-localization
Tingyu Wang 0002, Zhedong Zheng, Yaoqi Sun, Chenggang Yan 0001, Yi Yang 0001, Tat-Seng Chua
Pattern Recognit.4
2024 Rethinking Pooling for Multi-Granularity Features in Aerial-View Geo-Localization
abstract
Vision-based aerial-view geo-localization aims to match drone- and satellite-views of the same geographical location. Several feature partition strategies divide spatial features to mine contextual information. However, the compression from fine-grained features to visual descriptors is ill-considered, that is, classical pooling destroys discriminative features while increasing the sensitivity of networks to contextual information. In order to clarify this, we first review existing pooling layer and analyze their pros and cons when applied in feature compression. Inspired by the appearance of aerial views, we then summarize an ideal feature compression operation, i.e., precisely highlighting the central target while maximizing the use of environmental information in a feature-smoothing manner. To achieve the above process, we propose a distance-dependent parameter initialization strategy and form a novel pooling called$D^{2}$-GeM pooling, which can explicitly guide the network to compress fine-grained features in multiple patterns. Extensive experiments on public benchmark University-1652 substantiate that our strategy attains more appealing results without additional costs.
Tingyu Wang 0002, Yaoqi Sun, Chenggang Yan 0001
IEEE Signal Process. Lett.5
2024 SDPL: Shifting-Dense Partition Learning for UAV-View Geo-Localization
abstract
Cross-view geo-localization aims to match images of the same target from different platforms, e.g., drone and satellite. It is a challenging task due to the changing appearance of targets and environmental content from different views. Most methods focus on obtaining more comprehensive information through feature map segmentation, while inevitably destroying the image structure, and are sensitive to the shifting and scale of the target in the query. To address the above issues, we introduce simple yet effective part-based representation learning, shifting-dense partition learning (SDPL). We propose a dense partition strategy (DPS), dividing the image into multiple parts to explore contextual information while explicitly maintaining the global structure. To handle scenarios with non-centered targets, we further propose the shifting-fusion strategy, which generates multiple sets of parts in parallel based on various segmentation centers, and then adaptively fuses all features to integrate their anti-offset ability. Extensive experiments show that SDPL is robust to position shifting, and performs competitively on two prevailing benchmarks, University-1652 and SUES-200. In addition, SDPL shows satisfactory compatibility with a variety of backbone networks (e.g., ResNet and Swin).https://github.com/C-water/SDPL_release.
Tingyu Wang 0002, Haoran Li 0025, Rongfeng Lu, Yaoqi Sun, Bolun Zheng, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.8
2024 Raw Image Based Over-Exposure Correction Using Channel-Guidance Strategy
abstract
Most existing methods for over-exposure in image correction are developed based on sRGB images, which can result in complex and non-linear degradation due to the image signal processing pipeline. By contrast, data-driven approaches based on RAW image data offer natural advantages for image processing tasks. RAW images, characterized by their near-linear correlation with scene radiance and enriched information content due to higher bit depth, demonstrate superior performance compared to sRGB-based techniques. Further, the spectral sensitivity characteristics intrinsic to digital camera sensors indicate that the blue and red channels in a Bayer pattern RAW image typically encompass more contextual information than the green channels. This property renders them less susceptible to over-exposure, thereby making them more effective for data extraction in high dynamic range scenes. In this paper, we introduce a Channel-Guidance Network (CGNet) that leverages the benefits of RAW images for over-exposure correction. The CGNet estimates the properly-exposed sRGB image directly from the over-exposed RAW image in an end-to-end manner. Specifically, we introduce a RAW-based channel-guidance branch to the U-net-based backbone, which exploits the color channel intensity prior of RAW images to achieve superior over-exposure correction performance. To further facilitate research in over-exposure correction, we present synthetic and real-world over-exposure correction benchmark datasets. These datasets comprise a large set of paired RAW and sRGB images across a variety of scenarios. Experiments on our RAW-sRGB datasets validate the advantages of our RAW-based channel guidance strategy and proposed CGNet over state-of-the-art sRGB-based methods on over-exposure correction. Our code and dataset are publicly available athttps://github.com/whiteknight-WJN/CGNet.
Ying Fu 0001, Yunhao Zou, Qiankun Liu 0001, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Learning Cross-View Geo-Localization Embeddings via Dynamic Weighted Decorrelation Regularization
abstract
In the domain of cross-view geo-localization, the challenge lies in accurately matching images captured from distinct perspectives, such as aerial drone imagery and satellite imagery of the same geographical location. Existing methods predominantly concentrate on minimizing distances between feature embeddings in the representational space, inadvertently overlooking the significance of reducing embedding redundancy. This oversight potentially hampers the extraction of diverse and distinctive visual patterns critical for precise localization. This work argues that minimizing embedding redundancy is a pivotal factor in enhancing a model’s ability to discriminate diverse scene characteristics. To support this claim, we introduce a straightforward yet effective regularization technique, termed dynamic weighted decorrelation regularization (DWDR). DWDR serves to actively promote the learning of orthogonal feature channels within neural networks. By dynamically adjusting weights, DWDR targets the minimization of interchannel correlations, guiding the correlation matrix toward diagonality, indicative of independence among channels. The dynamic weighting mechanism adaptively prioritizes the decorrelation of channels that remain highly correlated throughout training. Additionally, we devise a symmetrical sampling strategy for cross-view scenarios to ensure that the training examples are balanced across different imaging platforms in a batch. Despite its simplicity, the integration of DWDR and the proposed sampling scheme yields remarkable performance across four extensive benchmark datasets: University-1652, CVUSA, CVACT, and VIGOR. Notably, in stringent conditions, such as when constrained to exceedingly compact feature dimensions of 64, our methodology significantly outperforms conventional baselines, thereby affirming its efficacy and robustness under challenging constraints.
Tingyu Wang 0002, Zhedong Zheng, Zunjie Zhu, Yaoqi Sun, Chenggang Yan 0001, Yi Yang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Quality-Aware Selective Fusion Network for V-D-T Salient Object Detection
abstract
Depth images and thermal images contain the spatial geometry information and surface temperature information, which can act as complementary information for the RGB modality. However, the quality of the depth and thermal images is often unreliable in some challenging scenarios, which will result in the performance degradation of the two-modal based salient object detection (SOD). Meanwhile, some researchers pay attention to the triple-modal SOD task, namely the visible-depth-thermal (VDT) SOD, where they attempt to explore the complementarity of the RGB image, the depth image, and the thermal image. However, existing triple-modal SOD methods fail to perceive the quality of depth maps and thermal images, which leads to performance degradation when dealing with scenes with low-quality depth and thermal images. Therefore, in this paper, we propose a quality-aware selective fusion network (QSF-Net) to conduct VDT salient object detection, which contains three subnets including the initial feature extraction subnet, the quality-aware region selection subnet, and the region-guided selective fusion subnet. Firstly, except for extracting features, the initial feature extraction subnet can generate a preliminary prediction map from each modality via a shrinkage pyramid architecture, which is equipped with the multi-scale fusion (MSF) module. Then, we design the weakly-supervised quality-aware region selection subnet to generate the quality-aware maps. Concretely, we first find the high-quality and low-quality regions by using the preliminary predictions, which further constitute the pseudo label that can be used to train this subnet. Finally, the region-guided selective fusion subnet purifies the initial features under the guidance of the quality-aware maps, and then fuses the triple-modal features and refines the edge details of prediction maps through the intra-modality and inter-modality attention (IIA) module and the edge refinement (ER) module, respectively. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art methods with a large margin. Our code and results are available at https://github.com/Lx-Bao/QSFNet.
Liuxin Bao, Xiaofei Zhou 0003, Xiankai Lu, Yaoqi Sun, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Image Process.8
2024 Penalized Flow Hypergraph Local Clustering
abstract
In recent years, hypergraph analysis have attracted increasing attention due to their ability to model complex data correlation, with hypergraph clustering being one of the most important tasks. However, when the scale of hypergraph is large enough, clustering is difficult based on global consistency. Existing flow-based hypergraph local clustering methods have good theoretical cut improvements and runtime guarantees. However, these methods exhibit poor performance when the initial reference node set is small and are prone to causing the output set to shrink into a small subset, resulting in local minima. To address this issue, we propose the Penalized Flow Hypergraph Local Clustering(PFHLC) and provide new conductance guarantees and runtime analyses for our method. First, we use the random walk method to grow the initial seed set, and introduce the random walk information of nodes as penalized flow into the flow-based framework to optimize the output. Second, we propose a generalized objective function containing random walk information, which takes full advantage of the semi-supervised information of the target cluster to protect important nodes. This feature can avoid the local minima of previous flow-based methods. Importantly, our method is strongly-local and can run efficiently on large-scale hypergraphs. We contribute a real-world dataset and the experiments on real-world large-scale datasets show that PFHLC achieves the state-of-the-art significantly.
Yubo Zhang 0006, Chenggang Yan 0001, Zuxing Xuan, Ting Yu 0004, Ji Zhang 0001, Shihui Ying, Yue Gao 0002
IEEE Trans. Knowl. Data Eng.3
2024 Learning Shape-Biased Representations for Infrared Small Target Detection
abstract
Typically, infrared small target detection aims to accurately localize objects from complex backgrounds where the object textures are often dim and the object shapes are varying. A feasible solution is learning discriminative representations with deep convolutional neural networks (CNNs). However, the representations learned by traditional deep CNNs often suffer from low shape bias. In this work, we propose a unified framework to learn shape-biased representations for facilitating infrared small target detection by explicitly incorporating shape information into model learning. The framework cascades a large-kernel encoder and a shape-guided decoder to learn discriminative shape-biased representations in an end-to-end manner. The large-kernel encoder describes infrared images into shape-preserving representations by using a few convolutions whose kernel size is as large as$9\times 9$, in contrast to commonly used$3\times 3$. The shape-guided decoder simultaneously addresses two tasks: decodes the encoder representations via upsampling reconstruction to reconstruct the segmentation, and hierarchically fuses the decoder representations and edge information via cascaded gated ResNet blocks to reconstruct the contour. In this way, the learned shape-biased representations are effective for identifying infrared small targets. Extensive experiments show our approach outperforms 18 state-of-the-arts.
Fanzhao Lin, Shiming Ge, Kexin Bao, Chenggang Yan 0001, Dan Zeng 0001
IEEE Trans. Multim.4
2024 MFFNet: Multi-Modal Feature Fusion Network for V-D-T Salient Object Detection
abstract
This article discusses the limitations of single- and two-modal salient object detection (SOD) methods and the emergence of multi-modal SOD techniques that integrate Visible, Depth, or Thermal information. However, current multi-modal methods often rely on simple fusion techniques such as addition, multiplication and concatenation, to combine the different modalities, which is ineffective for challenging scenes, such as low illumination and background messy. To address this issue, we propose a novel multi-modal feature fusion network (MFFNet) for V-D-T salient object detection, where the two key points are the triple-modal deep fusion encoder and the progressive feature enhancement decoder. The MFFNet's triple-modal deep fusion (TDF) module is designed to integrate the features of the three modalities and explore their complementarity by utilizing mutual optimization during the encoding phase. In addition, the progressive feature enhancement decoder consists of the weighted context-enhanced feature (WCF) module, region optimization (RO) module and boundary perception (BP) module to produce region-aware and contour-aware features. After that, a multi-scale fusion (MF) module is proposed to integrate these features and generate high-quality saliency maps. We conduct extensive experiments on the VDT-2048 dataset, and our results show that the proposed MFFNet outperforms 12 state-of-the-art multi-modal methods.
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001
IEEE Trans. Multim.8
2024 STAT: Multi-Object Tracking Based on Spatio-Temporal Topological Constraints
abstract
The mainstream tracking-by-detection paradigm for multi-object tracking generally conducts detection first, followed by Re-IDentification (Re-ID) and motion estimation. The associations between the predicted boxes and existing tracks are then performed via visual and motion association. However, challenges such as irregular motion patterns, similar appearances, and frequent occlusions often arise, making object tracking a nontrivial task. In this article, we propose a multi-object tracker based on Spatio-TemporAl Topological (STAT) constraints to address the above issues. More specifically, we design the Feature Adaptive Association Module (FAAM) to establish the association between motion and appearance regionally, completing a complementary combination of appearance and motion features. Among these, the Appearance Feature Update Module (AFUM) is proposed to manage the appearance updates of tracked objects by imposing constraints based on the spatial locations and the degree of object occlusion, while temporal consistency is adopted to smooth the appearance states of tracks to mitigate the accumulation of appearance noise. Moreover, the Robust Motion Tracking Module (RMTM) is established to reduce the impact of irregular motions and certain unreliable detection results. The proposed module includes a higher weighted momentum term to accommodate the excessive motion amplitude and considers low-confidence boxes accompanied by the stage-wise association strategy for high-confidence boxes. Extensive experiments on DanceTrack and benchmark MOT datasets verify the effectiveness of our STAT tracker, especially the state-of-the-art results on DanceTrack, which is characterized by irregular motion and indistinguishable appearance attributes.
Junjie Zhang 0002, Xinyu Zhang 0015, Chenggang Yan 0001, Dan Zeng 0001
IEEE Trans. Multim.5
2024 Multi-stage reasoning on introspecting and revising bias for visual question answering
abstract
Visual Question Answering (VQA) is a task that involves predicting an answer to a question depending on the content of an image. However, recent VQA methods have relied more on language priors between the question and answer rather than the image content. To address this issue, many debiasing methods have been proposed to reduce language bias in model reasoning. However, the bias can be divided into two categories: good bias and bad bias. Good bias can benefit to the answer prediction, while the bad bias may associate the models with the unrelated information. Therefore, instead of excluding good and bad bias indiscriminately in existing debiasing methods, we proposed a bias discrimination module to distinguish them. Additionally, bad bias may reduce the model’s reliance on image content during answer reasoning and thus attend little on image features updating. To tackle this, we leverage Markov theory to construct a Markov field with image regions and question words as nodes. This helps with feature updating for both image regions and question words, thereby facilitating more accurate and comprehensive reasoning about both the image content and question. To verify the effectiveness of our network, we evaluate our network on VQA v2 and VQA cp v2 datasets and conduct extensive quantity and quality studies to verify the effectiveness of our proposed network. Experimental resu- lts show that our network achieves significant performance against the previous state-of-the-art methods.
Anan Liu, Zimu Lu, Ning Xu 0003, Min Liu 0008, Chenggang Yan 0001, Bolun Zheng, Yulong Duan, Xuanya Li
ACM Trans. Web5
2024 GLCSA-Net: global-local constraints-based spectral adaptive network for hyperspectral image inpainting
Jia Li 0032, Junjie Zhang 0002, Chenggang Yan 0001, Dan Zeng 0001
Vis. Comput.5
2023 Improving Dynamic HDR Imaging with Fusion Transformer
abstract
Reconstructing a High Dynamic Range (HDR) image from several Low Dynamic Range (LDR) images with different exposures is a challenging task, especially in the presence of camera and object motion. Though existing models using convolutional neural networks (CNNs) have made great progress, challenges still exist, e.g., ghosting artifacts. Transformers, originating from the field of natural language processing, have shown success in computer vision tasks, due to their ability to address a large receptive field even within a single layer. In this paper, we propose a transformer model for HDR imaging. Our pipeline includes three steps: alignment, fusion, and reconstruction. The key component is the HDR transformer module. Through experiments and ablation studies, we demonstrate that our model outperforms the state-of-the-art by large margins on several popular public datasets.
Rufeng Chen, Bolun Zheng, Chenggang Yan 0001, Gregory Slabaugh, Shanxin Yuan
AAAI5
2023 Gaussian Label Distribution Learning for Spherical Image Object Detection
abstract
Spherical image object detection emerges in many applications from virtual reality to robotics and automatic driving, while many existing detectors use$l_{n}$-norms loss for regression of spherical bounding boxes. There are two intrinsic flaws for$l_{n}$-norms loss, i.e., independent optimization of parameters and inconsistency between metric (dominated by IoU) and loss. These problems are common in planar image detection but more significant in spherical image detection. Solution for these problems has been extensively discussed in planar image detection by using IoU loss and related variants. However, these solutions cannot be migrated to spherical image object detection due to the undifferentiable of the Spherical IoU (SphIoU). In this paper, we design a simple but effective regression loss based on Gaussian Label Distribution Learning (GLDL) for spherical image object detection. Besides, we observe that the scale of the object in a spherical image varies greatly. The huge differences among objects from different categories make the sample selection strategy based on SphIoU challenging. Therefore, we propose GLDL-ATSS as a better training sample selection strategy for objects of the spherical image, which can alleviate the drawback of IoU threshold-based strategy of scale-sample imbalance. Extensive results on various two datasets with different baseline detectors show the effectiveness of our approach.
Xinyuan Liu 0003, Qiang Zhao 0005, Yike Ma, Chenggang Yan 0001
CVPR5
2023 Hybrid Spectral Denoising Transformer with Guided Attention
abstract
In this paper, we present a Hybrid Spectral Denoising Transformer (HSDT) for hyperspectral image denoising. Challenges in adapting transformer for HSI arise from the capabilities to tackle existing limitations of CNN-based methods in capturing the global and local spatial-spectral correlations while maintaining efficiency and flexibility. To address these issues, we introduce a hybrid approach that combines the advantages of both models with a Spatial-Spectral Separable Convolution (S3Conv), Guided Spectral Self-Attention (GSSA), and Self-Modulated Feed-Forward Network (SM-FFN). Our S3Conv works as a lightweight alternative to 3D convolution, which extracts more spatial-spectral correlated features while keeping the flexibility to tackle HSIs with an arbitrary number of bands. These features are then adaptively processed by GSSA which performs 3D self-attention across the spectral bands, guided by a set of learnable queries that encode the spectral signatures. This not only enriches our model with powerful capabilities for identifying global spectral correlations but also maintains linear complexity. Moreover, our SM-FFN proposes the self-modulation that intensifies the activations of more informative regions, which further strengthens the aggregated features. Extensive experiments are conducted on various datasets under both simulated and real-world noise, and it shows that our HSDT significantly outperforms the existing state-of-the-art methods while maintaining low computational overhead. Code is at https://github.com/Zeqiang-Lai/HSDT.
Zeqiang Lai, Chenggang Yan 0001, Ying Fu 0001
ICCV2
2023 Self-supervised Cross-view Representation Reconstruction for Change Captioning
abstract
Change captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruction (SCORER) network. Concretely, we first design a multi-head token-wise matching to model relationships between cross-view features from similar/dissimilar images. Then, by maximizing cross-view contrastive alignment of two similar images, SCORER learns two view-invariant image representations in a self-supervised way. Based on these, we reconstruct the representations of unchanged objects by cross-attention, thus learning a stable difference representation for caption generation. Further, we devise a cross-modal backward reasoning to improve the quality of caption. This module reversely models a "hallucination" representation with the caption and "before" representation. By pushing it closer to the "after" representation, we enforce the caption to be informative about the difference in a self-supervised manner. Extensive experiments show our method achieves the state-of-the-art results on four datasets. The code is available at https://github.com/tuyunbin/SCORER.
Yunbin Tu, Liang Li 0003, Li Su 0003, Zhengjun Zha, Chenggang Yan 0001, Qingming Huang
ICCV5
2023 RawHDR: High Dynamic Range Image Reconstruction from a Single Raw Image
abstract
High dynamic range (HDR) images capture much more intensity levels than standard ones. Current methods predominantly generate HDR images from 8-bit low dynamic range (LDR) sRGB images that have been degraded by the camera processing pipeline. However, it becomes a formidable task to retrieve extremely high dynamic range scenes from such limited bit-depth data. Unlike existing methods, the core idea of this work is to incorporate more informative Raw sensor data to generate HDR images, aiming to recover scene information in hard regions (the darkest and brightest areas of an HDR scene). To this end, we propose a model tailor-made for Raw images, harnessing the unique features of Raw data to facilitate the Raw-to-HDR mapping. Specifically, we learn exposure masks to separate the hard and easy regions of a high dynamic scene. Then, we introduce two important guidances, dual intensity guidance, which guides less informative channels with more informative ones, and global spatial guidance, which extrapolates scene specifics over an extended spatial domain. To verify our Raw-to-HDR approach, we collect a large Raw/HDR paired dataset for both training and testing. Our empirical evaluations validate the superiority of the proposed Raw-to-HDR reconstruction model, as well as our newly captured dataset in the experiments.
Yunhao Zou, Chenggang Yan 0001, Ying Fu 0001
ICCV2
2023 Iterative Denoiser and Noise Estimator for Self-Supervised Image Denoising
abstract
With the emergence of powerful deep learning tools, more and more effective deep denoisers have advanced the field of image denoising. However, the huge progress made by these learning-based methods severely relies on large-scale and high-quality noisy/clean training pairs, which limits the practicality in real-world scenarios. To overcome this, researchers have been exploring self-supervised approaches that can denoise without paired data. However, the unavailable noise prior and inefficient feature extraction take these methods away from high practicality and precision. In this paper, we propose a Denoise-Corrupt-Denoise pipeline (DCD-Net) for self-supervised image denoising. Specifically, we design an iterative training strategy, which iteratively optimizes the denoiser and noise estimator, and gradually approaches high denoising performances using only single noisy images without any noise prior. The proposed self-supervised image denoising framework provides very competitive results compared with state-of-the-art methods on widely used synthetic and real-world image denoising benchmarks.
Yunhao Zou, Chenggang Yan 0001, Ying Fu 0001
ICCV2
2023 Towards Confidence-Aware Commonsense Knowledge Integration for Scene Graph Generation
abstract
Commonsense knowledge has been widely explored to improve Scene Graph Generation (SGG). Existing methods simply incorporate the described relations of knowledge bases into each part of the scene for a concrete understanding. However, they ignore the discussion about whether a visual scene needs to associate commonsense knowledge for making inferences. Specifically, the difficulty of relation recognition varies from its type. Some frequent spatial relations (e.g. on) usually produce less perception error even without any prior information, while others involved many rules and patterns (e.g. throwing) possess few samples and require to combine with some commonsense knowledge as supplementary. In this paper, we propose a novel confidence-aware commonsense knowledge integration for SGG. Firstly, we depend on mutual information maximization to design a hybrid-attention module, which decreases the uncertainty in representation learning given external knowledge. Second, we introduce an extra branch for SGG network to perform confidence estimation independent of any ground truth labels, in which the output scalar explicitly reflects the difficulty of visual recognition. This value is equipped with the ability to balance the demand for commonsense knowledge in a given scene. Experiments are conducted with the backbone of MOTIFS on Visual Genome (VG) and our method effectively promotes the metric of mRecall with little performance hit for metric Recall, especially for predicting unseen relations.
Hongshuo Tian, Ning Xu 0003, Yanhui Wang 0001, Chenggang Yan 0001, Bolun Zheng, Xuanya Li, Anan Liu
ICME4
2023 Semantic Embedding Uncertainty Learning for Image and Text Matching
abstract
Image and text matching measures the semantic similarity for cross-modal retrieval. The core of this task is semantic embedding, which mines the intrinsic characteristics of visual and textual for discriminative representation. However, cross-modal ambiguity of image and text (the existence of one-to-many associations) is prone to semantic diversity. The mainstream approaches utilized the fixed point embedding to represent semantics, which ignored the embedding uncertainty caused by semantic diversity leading to incorrect results. To address this issue, we propose a novel Semantic Embedding Uncertainty Learning (SEUL), which represents the embedding uncertainty of image and text as Gaussian distributions and simultaneously learns the salient embedding (mean) and uncertainty (variance) in the common space. We design semantic uncertainty embedding for facilitating the robustness of the representation in the semantic diversity context. A combined objective function is proposed, which optimizes the semantic uncertainty and maintains discriminability to enhance cross-modal associations. Extended experiments are performed on two datasets to demonstrate advanced performance.
Yan Wang 0114, Yuting Su 0001, Wenhui Li 0001, Chenggang Yan 0001, Bolun Zheng, Xuanya Li, Anan Liu
ICME4
2023 Sph2Pob: Boosting Object Detection on Spherical Images with Planar Oriented Boxes Methods
abstract
Object detection on panoramic/spherical images has been developed rapidly in the past few years, where IoU-calculator is a fundamental part of various detector components, i.e. Label Assignment, Loss and NMS. Due to the low efficiency and non-differentiability of spherical Unbiased IoU, spherical approximate IoU methods have been proposed recently. We find that the key of these approximate methods is to map spherical boxes to planar boxes. However, there exists two problems in these methods: (1) they do not eliminate the influence of panoramic image distortion; (2) they break the original pose between bounding boxes. They lead to the low accuracy of these methods. Taking the two problems into account, we propose a new sphere-plane boxes transform, called Sph2Pob. Based on the Sph2Pob, we propose (1) an differentiable IoU, Sph2Pob-IoU, for spherical boxes with low time-cost and high accuracy and (2) an agent Loss, Sph2Pob-Loss, for spherical detection with high flexibility and expansibility. Extensive experiments verify the effectiveness and generality of our approaches, and Sph2Pob-IoU and Sph2Pob-Loss together boost the performance of spherical detectors. The source code is available at https://github.com/AntXinyuan/sph2pob.
Xinyuan Liu 0003, Bin Chen 0021, Qiang Zhao 0005, Yike Ma, Chenggang Yan 0001
IJCAI6
2023 Dynamic Contrastive Learning with Pseudo-samples Intervention for Weakly Supervised Joint Video MR and HD
abstract
Joint video moment retrieval (MR) and highlight detection (HD) aims to find relevant video moments according to the query text. Existing methods are fully supervised based on manual annotation, and their coarse multi-modal information interactions easily lose details about video and text. In addition, some tasks introduce weakly supervised learning with random masks, while the single masking forces the model to focus on masked words and ignore multi-modal contextual information. In view of this, we attempt weakly supervised joint tasks (MR+HD) and propose Dynamic Contrastive Learning with Pseudo-Sample Intervention (CPI) for better multi-modal video comprehension. First, we design pseudo-samples over random masks for a more efficient contrastive learning manner. We introduce a proportional sampling strategy for pseudo-samples to ensure the semantic difference between the pseudo-samples and the query text. This balances the over-reliance from single random mask to global text semantics and makes the model learn multimodal context from each word fairly. Second, we design dynamic intervention contrastive loss to enhance the core feature-matching ability of the model dynamically. We add pseudo-sample intervention when negative proposals are close to positive proposals. This can help the model overcome the vision confusion phenomenon and achieve semantic similarity instead of word similarity. Extensive experiments demonstrate the effectiveness of CPI and the potential of weakly supervised joint tasks.
Shuhan Kong, Liang Li 0003, Beichen Zhang 0006, Bin Jiang 0011, Chenggang Yan 0001, Changhao Xu
ACM Multimedia6
2023 Reducing Intrinsic and Extrinsic Data Biases for Moment Localization with Natural Language
abstract
Moment Localization with Natural Language (MLNL) aims to locate the target moment from an untrimmed video by a linguistic query. Recent works reveal the severe data bias problem in MLNL and point out that the multi-modal content may not be understood by fitting the timestamp distribution. In this paper, we study the data biases on the intrinsic and extrinsic aspects: the former is mainly caused by the ambiguity of the moment boundary and the information imbalance between input and output; The latter results from the long-tail distribution of moments in MLNL datasets. To alleviate this, we propose a hybrid multi-modal debiasing network with temporal consistency constraint for MLNL. Specifically, we first design the multi-temporal Transformer to mitigate the ambiguity of boundary by integrating frame-wise features into segment-wise and dynamically matching with moment boundaries. Then, we introduce the temporal consistency constraint that highlights the action information in complex moment content to overcome the intrinsic bias from information imbalance.Furthermore, we design the hybrid linguistic activating module with external knowledge to relieve the extrinsic bias, which introduces a prior guidance to focus the discriminative information from the tail samples. Extensive experiments on three public datasets demonstrate that our model outperforms the existing methods.
Jiong Yin, Liang Li 0003, Chenggang Yan 0001, Lei Zhang 0119, Zunjie Zhu
ACM Multimedia4
2023 Parsing is All You Need for Accurate Gait Recognition in the Wild
abstract
Binary silhouettes and keypoint-based skeletons have dominated human gait recognition studies for decades since they are easy to extract from video frames. Despite their success in gait recognition for in-the-lab environments, they usually fail in real-world scenarios due to their low information entropy for gait representations. To achieve accurate gait recognition in the wild, this paper presents a novel gait representation, named Gait Parsing Sequence (GPS). GPSs are sequences of fine-grained human segmentation, i.e., human parsing, extracted from video frames, so they have much higher information entropy to encode the shapes and dynamics of fine-grained human parts during walking. Moreover, to effectively explore the capability of the GPS representation, we propose a novel human parsing-based gait recognition framework, named ParsingGait. ParsingGait contains a Convolutional Neural Network (CNN)-based backbone and two light-weighted heads. The first head extracts global semantic features from GPSs, while the other one learns mutual information of part-level features through Graph Convolutional Networks to model the detailed dynamics of human walking. Furthermore, due to the lack of suitable datasets, we build the first parsing-based dataset for gait recognition in the wild, named Gait3D-Parsing, by extending the large-scale and challenging Gait3D dataset. Based on Gait3D-Parsing, we comprehensively evaluate our method and existing gait recognition methods. Specifically, ParsingGait achieves a 17.5% Rank-1 increase compared with the state-of-the-art silhouette-based method. In addition, by replacing silhouettes with GPSs, current gait recognition methods achieve about 12.5% ~ 19.2% improvements in Rank-1 accuracy. The experimental results show a significant improvement in accuracy brought by the GPS representation and the superiority of ParsingGait.
Jinkai Zheng, Xinchen Liu, Shuai Wang 0003, Chenggang Yan 0001, Wu Liu 0005
ACM Multimedia5
2023 Personalized Federated Learning via Backbone Self-Distillation
abstract
In practical scenarios, federated learning frequently necessitates training personalized models for each client using heterogeneous data. This paper proposes a backbone self-distillation approach to facilitate personalized federated learning. In this approach, each client trains its local model and only sends the backbone weights to the server. These weights are then aggregated to create a global backbone, which is returned to each client for updating. However, the client’s local backbone lacks personalization because of the common representation. To solve this problem, each client further performs backbone self-distillation by using the global backbone as a teacher and transferring knowledge to update the local backbone. This process involves learning two components: the shared backbone for common representation and the private head for local personalization, which enables effective global knowledge transfer. Extensive experiments and comparisons with 12 state-of-the-art approaches demonstrate the effectiveness of our approach.
Bochao Liu, Dan Zeng 0001, Chenggang Yan 0001, Shiming Ge
MMAsia4
2023 GFNet: gated fusion network for video saliency prediction
Songhe Wu, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001
Appl. Intell.7
2023 AutoDeconJ: a GPU-accelerated ImageJ plugin for 3D light-field deconvolution with optimal iteration numbers predicting
abstract
MOTIVATION: Light-field microscopy (LFM) is a compact solution to high-speed 3D fluorescence imaging. Usually, we need to do 3D deconvolution to the captured raw data. Although there are deep neural network methods that can accelerate the reconstruction process, the model is not universally applicable for all system parameters. Here, we develop AutoDeconJ, a GPU-accelerated ImageJ plugin for 4.4× faster and more accurate deconvolution of LFM data. We further propose an image quality metric for the deconvolution process, aiding in automatically determining the optimal number of iterations with higher reconstruction accuracy and fewer artifacts. RESULTS: Our proposed method outperforms state-of-the-art light-field deconvolution methods in reconstruction time and optimal iteration numbers prediction capability. It shows better universality of different light-field point spread function (PSF) parameters than the deep learning method. The fast, accurate and general reconstruction performance for different PSF parameters suggests its potential for mass 3D reconstruction of LFM data. AVAILABILITY AND IMPLEMENTATION: The codes, the documentation and example data are available on an open source at: https://github.com/Onetism/AutoDeconJ.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Changqing Su, Yaoqi Sun, Chenggang Yan 0001, Haibing Yin
Bioinform.5
2023 SMINet: Semantics-aware multi-level feature interaction network for surface defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Haibing Yin, Ji Hu 0002, Jiyong Zhang 0001, Chenggang Yan 0001
Eng. Appl. Artif. Intell.8
2023 Aggregating transformers and CNNs for salient object detection in optical remote sensing images
Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Haibing Yin, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001
Neurocomputing7
2023 CRB Weighted Source Localization Method Based on Deep Neural Networks in Multi-UAV Network
abstract
With the advent of the Internet of Things (IoT) era, the multiunmanned aerial vehicle (UAV) networks have attracted great attention in the fields of source detection and localization. However, as the real-time signal processing performance of the UAV is limited by the computing speed and accuracy of the embedded hardware, the effectiveness of source localization is greatly reduced. Aiming at improving the accuracy and computational efficiency of source localization, a Cramer–Rao bound (CRB) weighted multi-UAV network source localization method is proposed based on the deep neural networks (DNNs) and spatial-spectrum fitting (SSF). The proposed source localization system is composed of UAVs equipped with a radar array. The source location can be achieved using the direction of arrival (DOA) of the source signals of UAVs, but the accuracy and real-time performance of the conventional DOA estimation algorithms are not satisfactory, and the data fusion strategy of the conventional cross-location framework needs further improvement. In the proposed method, a DNN-based SSF, denoted as the deep SSF (DeepSSF), is designed to achieve accurate DOA estimation. In the DeepSSF, the DOA estimation performance is guaranteed by the DNN’s strong nonlinear fitting ability and highly parallel structure. In addition, based on the obtained DOA information, the source is located once by every two UAVs. Finally, the source localization is realized based on the weighted CRB according to the principle that the more the DOA distribution deviates from zero, the lower the estimation accuracy. The simulation results verify the efficiency of the proposed method.
Jingyu Cong, Xianpeng Wang 0001, Chenggang Yan 0001, Laurence T. Yang, Mianxiong Dong, Kaoru Ota
IEEE Internet Things J.3
2023 Multi-stage affine motion estimation fast algorithm for versatile video coding using decision tree
Xiaofeng Huang, Fangtao Zhou, Weihong Niu, Yang Zhou 0052, Haibing Yin, Chenggang Yan 0001
J. Vis. Commun. Image Represent.8
2023 SRI-Net: Similarity retrieval-based inference network for light field salient object detection
Chengtao Lv, Xiaofei Zhou 0003, Deyang Liu, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001
J. Vis. Commun. Image Represent.7
2023 CANet: Context-aware Aggregation Network for Salient Object Detection of Surface Defects
Bin Wan, Xiaofei Zhou 0003, Mang Xiao, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001
J. Vis. Commun. Image Represent.8
2023 Fast all zero block detection algorithm for versatile video coding
Weihong Niu, Xiaofeng Huang, Haibing Yin, Yang Zhou 0052, Chenggang Yan 0001
Multim. Tools Appl.6
2023 Depth-guided deep filtering network for efficient single image bokeh rendering
Bolun Zheng, Xiaofei Zhou 0003, Aiai Huang, Yaoqi Sun, Chuqiao Chen, Chenggang Yan 0001, Shanxin Yuan
Neural Comput. Appl.7
2023 ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Spotting
abstract
Scene text spotting is of great importance to the computer vision community due to its wide variety of applications. Recent methods attempt to introduce linguistic knowledge for challenging recognition rather than pure visual classification. However, how to effectively model the linguistic rules in end-to-end deep networks remains a research challenge. In this paper, we argue that the limited capacity of language models comes from 1) implicit language modeling; 2) unidirectional feature representation; and 3) language model with noise input. Correspondingly, we propose an autonomous, bidirectional and iterative ABINet++ for scene text spotting. First, the autonomous suggests enforcing explicitly language modeling by decoupling the recognizer into vision model and language model and blocking gradient flow between both models. Second, a novel bidirectional cloze network (BCN) as the language model is proposed based on bidirectional feature representation. Third, we propose an execution manner of iterative correction for the language model which can effectively alleviate the impact of noise input. Additionally, based on an ensemble of the iterative predictions, a self-training method is developed which can learn from unlabeled images effectively. Finally, to polish ABINet++ in long text recognition, we propose to aggregate horizontal features by embedding Transformer units inside a U-Net, and design a position and content attention module which integrates character order and content to attend to character features precisely. ABINet++ achieves state-of-the-art performance on both scene text recognition and scene text spotting benchmarks, which consistently demonstrates the superiority of our method in various environments especially on low-quality images. Besides, extensive experiments including in English and Chinese also prove that, a text spotter that incorporates our language modeling method can significantly improve its performance both in accuracy and speed compared with commonly used attention-based recognizers. Code is available at https://github.com/FangShancheng/ABINet-PP.
Shancheng Fang, Zhendong Mao 0001, Hongtao Xie 0001, Yuxin Wang 0002, Chenggang Yan 0001, Yongdong Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 STORM: Structure-Based Overlap Matching for Partial Point Cloud Registration
abstract
Partial point cloud registration aims to transform partial scans into a common coordinate system. It is an important preprocessing step to generate complete 3D shapes. Although previous registration methods have made great progress in recent decades, traditional registration methods, such as Iterative Closest Point (ICP) and its variants, all these methods highly depend on the sufficient overlaps between two point clouds, because they cannot distinguish outlier correspondences. Note that the overlap between point clouds could always be small, which limits the application of these methods. To tackle this problem, we present a StrucTure-based OveRlap Matching (STORM) method for partial point cloud registration. In our method, an overlap prediction module with differentiable sampling is designed to detect points in overlap utilizing structure information, and facilitates exact partial correspondence generation, which is based on discriminative pointwise feature similarity. The pointwise features which contain effective structural information are extracted by graph-based methods. Experimental results and comparison with state-of-the-art methods demonstrate that STORM can achieve better performance. Moreover, most registration methods perform worse when the overlap ratio decreases, while STORM can still achieve satisfactory performance when the overlap ratio is small.
Chenggang Yan 0001, Yutong Feng, Shaoyi Du, Qionghai Dai, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Transformer-Based Multi-Scale Feature Integration Network for Video Saliency Prediction
abstract
Most cutting-edge video saliency prediction models rely on spatiotemporal features extracted by 3D convolutions due to its local contextual cues acquirement ability. However, the shortage of 3D convolutions is that it cannot effectively capture long-term spatiotemporal dependencies in videos. To address this limitation, we propose a novel Transformer-based Multi-scale Feature Integration Network (TMFI-Net) for video saliency prediction, where the proposed TMFI-Net consists of a semantic-guided encoder and a hierarchical decoder. Firstly, embarking on the Transformer-based multi-level spatiotemporal features, the semantic-guided encoder enhances the features by inserting the high-level feature into each level feature via a top-down pathway and a longitudinal connection, which endows the multi-level spatiotemporal features with rich contextual information. In this way, the features are steered to give more concerns to saliency regions. Secondly, the hierarchical decoder employs a multi-dimensional attention (MA) module to elevate features along channel, temporal, and spatial dimensions jointly. Successively, the hierarchical decoder deploys a progressive decoding block to conduct an initial saliency prediction, which provides a coarse localization of saliency regions. Lastly, considering the complementarity of different saliency predictions, we integrate all initial saliency prediction results into the final saliency map. Comprehensive experimental results on four video saliency datasets firmly demonstrate that our model achieves superior performance when compared with the state-of-the-art video saliency models. The code is available athttps://github.com/wusonghe/TMFI-Net.
Xiaofei Zhou 0003, Songhe Wu, Bolun Zheng, Shuai Wang 0003, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.8
2023 Edge-Guided Recurrent Positioning Network for Salient Object Detection in Optical Remote Sensing Images
abstract
Optical remote sensing images (RSIs) have been widely used in many applications, and one of the interesting issues about optical RSIs is the salient object detection (SOD). However, due to diverse object types, various object scales, numerous object orientations, and cluttered backgrounds in optical RSIs, the performance of the existing SOD models often degrade largely. Meanwhile, cutting-edge SOD models targeting optical RSIs typically focus on suppressing cluttered backgrounds, while they neglect the importance of edge information which is crucial for obtaining precise saliency maps. To address this dilemma, this article proposes an edge-guided recurrent positioning network (ERPNet) to pop-out salient objects in optical RSIs, where the key point lies in the edge-aware position attention unit (EPAU). First, the encoder is used to give salient objects a good representation, that is, multilevel deep features, which are then delivered into two parallel decoders, including: 1) an edge extraction part and 2) a feature fusion part. The edge extraction module and the encoder form a U-shape architecture, which not only provides accurate salient edge clues but also ensures the integrality of edge information by extra deploying the intraconnection. That is to say, edge features can be generated and reinforced by incorporating object features from the encoder. Meanwhile, each decoding step of the feature fusion module provides the position attention about salient objects, where position cues are sharpened by the effective edge information and are used to recurrently calibrate the misaligned decoding process. After that, we can obtain the final saliency map by fusing all position attention cues. Extensive experiments are conducted on two public optical RSIs datasets, and the results show that the proposed ERPNet can accurately and completely pop-out salient objects, which consistently outperforms the state-of-the-art SOD models.
Xiaofei Zhou 0003, Kunye Shen, Li Weng, Runmin Cong, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Cybern.7
2023 Progressive Recurrent Neural Network for Multispectral Remote Sensing Image Destriping
abstract
An unstable imaging system often introduces additional stripe noise in multispectral remote sensing images during the data acquisition process given a variety of factors. The complicated stripe distributions lead to the residual stripe in the results of existing methods, thus increasing the difficulty of destriping in practice. Mainstream deep learning-based methods show the encouraging destriping performance on multispectral remote sensing images. However, they often require the model to handle the varying degrees of stripe noise in a single shot for each image, which results in the poor destriping performance when facing practical cases with diverse stripe distributions. To address the above issue, we propose a Progressive Recurrent Neural Network (PRNet) to remove the stripe noise for each degraded image in an iterative manner. More specifically, a progressive destriping strategy is designed to gradually restore the clean image, in which the Main Recurrent Module (MRM) is introduced to iteratively process the stripe removal results generated from previous timesteps until the clean image is obtained. Furthermore, since the uniformity of the entire image is supposed to be significantly enhanced after the destriping, it is necessary to take the local spatial correlation into account during the destriping. Therefore, we present the Patch-based Sequence Module (PSM) to leverage the local spatial correlation by splitting the image into multi-scale patch sequences and capturing the relationship among different patches. Extensive experimental results on different datasets demonstrate that the proposed model yields superior destriping performance compared to other methods, especially for removing the stripe noise with complex distributions.
Jia Li 0032, Junjie Zhang 0002, Jungong Han, Chenggang Yan 0001, Dan Zeng 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Ecological Cooperative Adaptive Cruise Control for Heterogenous Vehicle Platoons Subject to Time Delays and Input Saturations
abstract
Public concerns about energy crisis and environmental issues lead to higher fuel economy standards and more stringent limitations on greenhouse gas emissions for ground vehicles, and ecological cooperative adaptive cruise control (eco-CACC) is considered to be effective in reducing fuel consumption and greenhouse gas emissions of vehicle platoons. In this paper, an eco-CACC strategy is presented for heterogeneous platoons with time delays and input saturations to achieve platoon stability, fuel economy, riding comfort and driving efficiency. The proposed eco-CACC strategy includes distributed linear feedback control (DLFC) for following vehicles and model predictive control (MPC) for the leading one. To obtain platoon stability, the DLFC protocol is designed using state errors between an ego vehicle and its preceding one, and the frequency domain method is employed to derive the sufficient conditions of platoon stability for the DLFC protocol. Then, to achieve fuel economy, passenger comfort and driving efficiency without violating the different input saturations of vehicles, the constrained fuel consumption optimization problem of the leading vehicle (CFCO-LV) is formulated based on delay MPC, where the weighted sum of the overall fuel consumption and velocity band-stop functions of vehicle platoon is minimized. To quickly obtain the MPC controller, an improved particle swarm optimization (PSO) algorithm is used to solve the CFCO-LV problem. Simulations are implemented to validate the effectiveness of the proposed eco-CACC strategy, and the simulation results demonstrate that, compared with benchmark, the proposed strategy can save 2.19%~8.24% of fuel for the heterogeneous platoon under different velocity ranges.
Chunjie Zhai, Chuqiao Chen, Xinlei Zheng, Zhimin Han, Chenggang Yan 0001, Fei Luo 0001, Jianmin Xu
IEEE Trans. Intell. Transp. Syst.6
2023 Adaptive Hypergraph Auto-Encoder for Relational Data Clustering
abstract
The embedded representation and clustering tasks both play important roles in relational data analysis and mining. Traditional methods mainly employ graph structure to describe relational data, but intuitive pairwise connections among nodes are insufficient to model high-order data in the real-world, such as the relations between proteins and polypeptide chains. Hypergraphs are a generalization of graphs, and hypergraphs can well model high-order data. When modeling relational data in the real world, hypergraphs are often accompanied by node attributes, i.e. attributed hypergraphs. Besides this, how to integrate the structural information and attribute information appropriately is another important task, while has not been investigated systematically. In this paper, we propose Adaptive Hypergraph Auto-Encoder(AHGAE) to learn node embeddings in low-dimensional space. Our method can utilize the high-order relation to generate embedding for clustering. It is composed of two procedures, i.e. the adaptive hypergraph Laplacian smoothing filter and the relational reconstruction auto-encoder. It has the advantage of integrating more complex data relations compared with graph-based methods, which leads to better modeling and clustering performance. The proposed method has been evaluated on hypergraph datasets and benchmark graph datasets. Experimental results and comparison with the state-of-the-art methods have demonstrated the effectiveness of our proposed method.
Youpeng Hu, Xunkai Li, Chenggang Yan 0001, Jian Yin 0003, Yue Gao 0002
IEEE Trans. Knowl. Data Eng.6
2022 Unbiased IoU for Spherical Image Object Detection
abstract
As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while it is challenging for spherical ones. Some existing methods utilize planar bounding boxes to represent spherical objects. However, they are biased due to the distortions of spherical objects. Others use spherical rectangles as unbiased representations, but they adopt excessive approximate algorithms when computing the IoU. In this paper, we propose an unbiased IoU as a novel evaluation criterion for spherical image object detection, which is based on the unbiased representations and utilize unbiased analytical method for IoU calculation. This is the first time that the absolutely accurate IoU calculation is applied to the evaluation criterion, thus object detection algorithms can be correctly evaluated for spherical images. With the unbiased representation and calculation, we also present Spherical CenterNet, an anchor free object detection algorithm for spherical images. The experiments show that our unbiased IoU gives accurate results and the proposed Spherical CenterNet achieves better performance on one real-world and two synthetic spherical object detection datasets than existing methods.
Bin Chen 0021, Yike Ma, Bailan Feng, Chenggang Yan 0001, Qiang Zhao 0005
AAAI8
2022 Gait Recognition in the Wild with Dense 3D Representations and A Benchmark
abstract
Existing studies for gait recognition are dominated by 2D representations like the silhouette or skeleton of the human body in constrained scenes. However, humans live and walk in the unconstrained 3D space, so projecting the 3D human body onto the 2D plane will discard a lot of crucial information like the viewpoint, shape, and dynamics for gait recognition. Therefore, this paper aims to explore dense 3D representations for gait recognition in the wild, which is a practical yet neglected problem. In particular, we propose a novel framework to explore the 3D Skinned Multi-Person Linear (SMPL) model of the human body for gait recognition, named SMPLGait. Our framework has two elaborately-designed branches of which one extracts appearance features from silhouettes, the other learns knowledge of 3D viewpoints and shapes from the 3D SMPL model. In addition, due to the lack of suitable datasets, we build the first large-scale 3D representation-based gait recognition dataset, named Gait3D. It contains 4,000 subjects and over 25,000 sequences extracted from 39 cameras in an unconstrained indoor scene. More importantly, it provides 3D SMPL models recovered from video frames which can provide dense 3D information of body shape, viewpoint, and dynamics. Based on Gait3D, we comprehensively compare our method with existing gait recognition approaches, which reflects the superior performance of our framework and the potential of 3D representations for gait recognition in the wild. The code and dataset are available at: https://gait3d.github.io.
Jinkai Zheng, Xinchen Liu, Wu Liu 0005, Lingxiao He, Chenggang Yan 0001, Tao Mei 0001
CVPR5
2022 PANDORA: A Panoramic Detection Dataset for Object with Orientation
Qiang Zhao 0005, Yike Ma, Bailan Feng, Chenggang Yan 0001
ECCV (8)7
2022 Multi-task Optimization Based Co-training for Electricity Consumption Prediction
abstract
Real-world electricity consumption prediction may involve different tasks, e.g., prediction for different time steps ahead or different geo-locations. These tasks are often solved independently without utilizing some common problem-solving knowledge that could be extracted and shared among these tasks to augment the performance of solving each task. In this work, we propose a multi-task optimization (MTO) based co-training (MTO-CT) framework, where the models for solving different tasks are co-trained via an MTO paradigm in which solving each task may benefit from the knowledge gained from when solving some other tasks to help its solving process. MTO-CT leverages long short-term memory (LSTM) based model as the predictor where the knowledge is represented via connection weights and biases. In MTO-CT, an inter-task knowledge transfer module is designed to transfer knowledge between different tasks, where the most helpful source tasks are selected by using the probability matching and stochastic universal selection, and evolutionary operations like mutation and crossover are performed for reusing the knowledge from selected source tasks in a target task. We use electricity consumption data from five states in Australia to design two sets of tasks at different scales: a) one-step ahead prediction for each state (five tasks) and b) 6-step, 12-step, 18-step, and 24-step ahead prediction for each state (20 tasks). The performance of MTO-CT is evaluated on solving each of these two sets of tasks in comparison to solving each task in the set independently without knowledge sharing under the same settings, which demonstrates the superiority of MTO-CT in terms of prediction accuracy.
A. K. Qin 0001, Chenggang Yan 0001
IJCNN3
2022 Gait Recognition in the Wild with Multi-hop Temporal Switch
abstract
Existing studies for gait recognition are dominated by in-the-lab scenarios. Since people live in real-world senses, gait recognition in the wild is a more practical problem that has recently attracted the attention of the community of multimedia and computer vision. Current methods that obtain state-of-the-art performance on in-the-lab benchmarks achieve much worse accuracy on the recently proposed in-the-wild datasets because these methods can hardly model the varied temporal dynamics of gait sequences in unconstrained scenes. Therefore, this paper presents a novel multi-hop temporal switch method to achieve effective temporal modeling of gait patterns in real-world scenes. Concretely, we design a novel gait recognition network, named Multi-hop Temporal Switch Network (MTSGait), to learn spatial features and multi-scale temporal features simultaneously. Different from existing methods that use 3D convolutions for temporal modeling, our MTSGait models the temporal dynamics of gait sequences by 2D convolutions. By this means, it achieves high efficiency with fewer model parameters and reduces the difficulty in optimization compared with 3D convolution-based models. Based on the specific design of the 2D convolution kernels, our method can eliminate the misalignment of features among adjacent frames. In addition, a new sampling strategy, i.e., non-cyclic continuous sampling, is proposed to make the model learn more robust temporal features. Finally, the proposed method achieves superior performance on two public gait in-the-wild datasets, i.e., GREW and Gait3D, compared with state-of-the-art methods.
Jinkai Zheng, Xinchen Liu, Xiaoyan Gu 0001, Yaoqi Sun, Chuang Gan 0001, Jiyong Zhang 0001, Wu Liu 0005, Chenggang Yan 0001
ACM Multimedia8
2022 DomainPlus: Cross Transform Domain Learning towards High Dynamic Range Imaging
abstract
High dynamic range (HDR) imaging by combining multiple low dynamic range (LDR) images of different exposures provides a promising way to produce high quality photographs. However, the misalignment between the input images leads to ghosting artifacts in the reconstructed HDR image. In this paper, we propose a cross-transform domain neural network for efficient HDR imaging. Our approach consists of two modules: a merging module and a restoration module. For the merging module, we propose a Multiscale Attention with Fronted Fusion (MAFF) mechanism to achieve coarse-to-fine spatial fusion. For the restoration module, we propose fronted Discrete Wavelet Transform (DWT) and Discrete Cosine Transform (DCT)-based learnable bandpass filters to formulate a cross-transform domain learning block, dubbed DomainPlus Block (DPB) for effective ghosting removal. Our ablation study and comprehensive experiments show that DomainPlus outperforms the existing state-of-the-art on several datasets.
Bolun Zheng, Xiaokai Pan, Xiaofei Zhou 0003, Gregory Slabaugh, Chenggang Yan 0001, Shanxin Yuan
ACM Multimedia6
2022 SHREC'22 track: Open-Set 3D Object Retrieval
Yifan Feng 0001, Yue Gao 0002, Xibin Zhao, Yandong Guo, Nihar Bagewadi, Nhat-Tan Bui, Hieu Dao, Shankar Gangisetty, Ripeng Guan, Xie Han 0001, Cong Hua, Chidambar Hunakunti, Yu Jiang 0006, Shichao Jiao, Yuqi Ke, Liqun Kuang, Anan Liu, Dinh-Huan Nguyen, Hai-Dang Nguyen, Weizhi Nie, Bang-Dang Pham, Karthik Raikar, Qingmei Tang, Minh-Triet Tran, Jialong Wan, Chenggang Yan 0001, Haoxuan You, Difei Zhu
Comput. Graph.26
2022 Bidirectional difference locating and semantic consistency reasoning for change captioning
abstract
Change captioning is an emerging task to describe the changes between a pair of images. The difficulty in this task is to discover the differences between the two images. Recently, some methods have been proposed to address this problem. However, they all employ unidirectional difference localization to identify the changes. This can lead to ambiguity about the nature of the changes. Instead, we propose a framework with bidirectional difference localization and semantic consistency reasoning to describe the image changes. First, we locate the changes in the two images by capturing bidirectional differences. Then we design a decoder with spatial-channel attention to generate the change caption. Finally, we introduce semantic consistency reasoning to constrain our bidirectional difference localization module and spatial-channel attention module. Extensive experiments on three public data sets show that the performance of our proposed model outperforms the state-of-the-art change captioning models by a large margin.
Yaoqi Sun, Liang Li 0003, Tongyv Lu, Bolun Zheng, Chenggang Yan 0001, Yongjun Bao, Guiguang Ding, Gregory Slabaugh
Int. J. Intell. Syst.6
2022 Group-Wise Hub Identification by Learning Common Graph Embeddings on Grassmannian Manifold
abstract
Human brain is a complex yet economically organized system, where a small portion of critical hub regions support the majority of brain functions. The identification of common hub nodes in a population of networks is often simplified as a voting procedure on the set of identified hub nodes across individual brain networks, which ignores the intrinsic data geometry and partially lacks the reproducible findings in neuroscience. Hence, we propose a first-ever group-wise hub identification method to identify hub nodes that are common across a population of individual brain networks. Specifically, the backbone of our method is to learn common graph embedding that can represent the majority of local topological profiles. By requiring orthogonality among the graph embedding vectors, each graph embedding as a data element is residing on the Grassmannian manifold. We present a novel Grassmannian manifold optimization scheme that allows us to find the common graph embeddings, which not only identify the most reliable hub nodes in each network but also yield a population-based common hub node map. Results of the accuracy and replicability on both synthetic and real network data show that the proposed manifold learning approach outperforms all hub identification methods employed in this evaluation.
Defu Yang, Jiazhou Chen 0001, Chenggang Yan 0001, Minjeong Kim 0001, Paul J. Laurienti, Martin Styner, Guorong Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Learning Frequency Domain Priors for Image Demoireing
abstract
Image demoireing is a multi-faceted image restoration task involving both moire pattern removal and color restoration. In this paper, we raise a general degradation model to describe an image contaminated by moire patterns, and propose a novel multi-scale bandpass convolutional neural network (MBCNN) for single image demoireing. For moire pattern removal, we propose a multi-block-size learnable bandpass filters (M-LBFs), based on a block-wise frequency domain transform, to learn the frequency domain priors of moire patterns. We also introduce a new loss function named Dilated Advanced Sobel loss (D-ASL) to better sense the frequency information. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, and then performs local fine tuning of the color per pixel. To determine the most appropriate frequency domain transform, we investigate several transforms including DCT, DFT, DWT, learnable non-linear transform and learnable orthogonal transform. We finally adopt the DCT. Our basic model won the AIM2019 demoireing challenge. Experimental results on three public datasets show that our method outperforms state-of-the-art methods by a large margin.
Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Xiang Tian 0002, Jiyong Zhang 0001, Yaoqi Sun, Lin Liu 0016, Ales Leonardis, Gregory Slabaugh
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 FANet: Feature aggregation network for RGBD saliency detection
Xiaofei Zhou 0003, Hongfa Wen, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001
Signal Process. Image Commun.6
2022 Self-Supervised Synthesis Ranking for Deep Metric Learning
abstract
The core purpose of deep metric learning is to construct an embedding space, where objects belonging to the same class are gathered together and the ones from different classes are pushed apart. Most existing approaches typically insist to inter-class characteristics,e.g., class-level information or instance-level similarity, to obtain semantic relevance of data points and get a large margin between different classes in the embedding space. However, the intra-class characteristics,e.g., local manifold structure or relative relationship within the same class, are usually overlooked in the learning process. Hence the output embeddings have limitation in retrieving a good ranking result if existing multiple positive samples. And the local data structure of embedding space cannot be fully exploited since lack of relative ranking information. As a result, the model is prone to overfitting on a train set and get low generalization on the test set (unseen classes) when losing sight of intra-class variance. This paper presents a novel self-supervised synthesis ranking auxiliary framework, which captures intra-class characteristics as well as inter-class characteristics for better metric learning. Our method designs a synthetic samples generation of polar coordinates to generate measurable intra-class variance with different strength and diversity in the latent space, which can simulate the various local structure change of intra-class in the initial data domain. And then formulates a self-supervised learning procedure to fully exploit this property and preserve it in the embedding space. As a result, the learned embedding space not only keeps inter-class discrimination but also owns subtle intra-class diversity, leading to better global and local embedding structures. Extensive experiments on five benchmarks show that our method significantly improves and outperforms the state-of-the-art methods on the performances of both retrieval and ranking by 2%-4% (personal use of this material is permitted. However, permission to use this material for any other purposes must be obtained from the IEEE by sending an email to [email protected]).
Zheren Fu, Zhendong Mao 0001, Chenggang Yan 0001, Anan Liu, Hongtao Xie 0001, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Joint Local Correlation and Global Contextual Information for Unsupervised 3D Model Retrieval and Classification
abstract
Unsupervised 3D model analysis has attracted tremendous attentions with the increasing growth of 3D model data and the extensive human annotations. Many effective methods have been designed to address the 3D model analysis with labeled information, while rare methods devote to unsupervised deep learning due to the difficulty of mining reliable information. In this paper, we propose a novel unsupervised deep learning method named joint local correlation and global contextual information (LCGC) for 3D model retrieval and classification, which mines the reliable triplet set and uses triplet loss to optimize the deep neural network. Our method proposes two schemes: 1) Local self-correlation information learning, which adopts the intra and inter information to construct the view-level triplet set. 2) Global neighbor contextual information learning, which employs the neighbor contextual information to explore the reliable relations among 3D models and construct the model-level triplet set. The above schemes encourage that the selected triple set can been used to improve the discrimination of learned features. Extensive evaluations on two large-scale datasets, ModelNet40 and ShapeNet55, have demonstrated the effectiveness of our proposed method.
Wenhui Li 0001, Zhenlan Zhao, Anan Liu, Zan Gao 0002, Chenggang Yan 0001, Zhendong Mao 0001, Haipeng Chen 0002, Weizhi Nie
IEEE Trans. Circuits Syst. Video Technol.5
2022 Each Part Matters: Local Patterns Facilitate Cross-View Geo-Localization
abstract
Cross-view geo-localization is to spot images of the same geographic target from different platforms,e.g., drone-view cameras and satellites. It is challenging in the large visual appearance changes caused by extreme viewpoint variations. Existing methods usually concentrate on mining the fine-grained feature of the geographic target in the image center, but underestimate the contextual information in neighbor areas. In this work, we argue that neighbor areas can be leveraged as auxiliary information, enriching discriminative clues for geo-localization. Specifically, we introduce a simple and effective deep neural network, called Local Pattern Network (LPN), to take advantage of contextual information in an end-to-end manner. Without using extra part estimators, LPN adopts a square-ring feature partition strategy, which provides the attention according to the distance to the image center. It eases the part matching and enables the part-wise representation learning. Owing to the square-ring partition design, the proposed LPN has good scalability to rotation variations and achieves competitive results on three prevailing benchmarks,i.e., University-1652, CVUSA and CVACT. Besides, we also show the proposed LPN can be easily embedded into other frameworks to further boost performance.
Tingyu Wang 0002, Zhedong Zheng, Chenggang Yan 0001, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng, Yi Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Task-Adaptive Attention for Image Captioning
abstract
Attention mechanisms are now widely used in image captioning models. However, most attention models only focus on visual features. When generating syntax related words, little visual information is needed. In this case, these attention models could mislead the word generation. In this paper, we propose Task-Adaptive Attention module for image captioning, which can alleviate this misleading problem and learn implicit non-visual clues which can be helpful for the generation of non-visual words. We further introduce a diversity regularization to enhance the expression ability of the Task-Adaptive Attention module. Extensive experiments on the MSCOCO captioning dataset demonstrate that by plugging our Task-Adaptive Attention module into a vanilla Transformer-based image captioning model, performance improvement can be achieved.
Chenggang Yan 0001, Yiming Hao, Liang Li 0003, Jian Yin 0003, Anan Liu, Zhendong Mao 0001, Zhenyu Chen 0003, Xingyu Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 CBREN: Convolutional Neural Networks for Constant Bit Rate Video Quality Enhancement
abstract
Constant bit rate (CBR) videos are widely used in streaming playback applications. However, the image quality of the CBR video is often unstable, especially for scenes with large motion. To this end, we design a new model to represent the distortion of High Efficiency Video Coding (HEVC) constant bit rate video, and propose a neural network for a constant bit rate video quality enhancement (CBREN). We propose a dual-domain restoration module (DRM) to jointly learn the prior knowledge in the pixel domain and the frequency domain. To address the degradation resulting from compression, we propose a two-step quantization degradation estimation strategy. The Inverse DCT (IDCT) Translation Unit (ITU) is used to constrain the quantization table of the constant bit rate video to a suitable range, and the Dynamic Alpha Unit (DAU) is used to fine-tune the quantization table according to the content of each frame. In order to effectively reduce the block distortion of different sizes produced in the compression process, we adopt a multi-scale network. Extensive experiments show that our approach can greatly enhance the quality of CBR compressed video. Moreover, our method can also be applied to constant quantization parameter (CQP) video enhancement tasks, and is certainly superior to existing methods.
Hengrun Zhao, Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Liang Li 0003, Gregory Slabaugh
IEEE Trans. Circuits Syst. Video Technol.5
2022 Domain-Adversarial-Guided Siamese Network for Unsupervised Cross-Domain 3-D Object Retrieval
abstract
Recent advances in 3-D sensors and 3-D modeling have led to the availability of massive amounts of 3-D data. It is too onerous and time consuming to manually label a plentiful of 3-D objects in real applications. In this article, we address this issue by transferring the knowledge from the existing labeled data (e.g., the annotated 2-D images or 3-D objects) to the unlabeled 3-D objects. Specifically, we propose a domain-adversarial guided siamese network (DAGSN) for unsupervised cross-domain 3-D object retrieval (CD3DOR). It is mainly composed of three key modules: 1) siamese network-based visual feature learning; 2) mutual information (MI)-based feature enhancement; and 3) conditional domain classifier-based feature adaptation. First, we design a siamese network to encode both 3-D objects and 2-D images from two domains because of its balanced accuracy and efficiency. Besides, it can guarantee the same transformation applied to both domains, which is crucial for the positive domain shift. The core issue for the retrieval task is to improve the capability of feature abstraction, but the previous CD3DOR approaches merely focus on how to eliminate the domain shift. We solve this problem by maximizing the MI between the input 3-D object or 2-D image data and the high-level feature in the second module. To eliminate the domain shift, we design a conditional domain classifier, which can exploit multiplicative interactions between the features and predictive labels, to enforce the joint alignment in both feature level and category level. Consequently, the network can generate domain-invariant yet discriminative features for both domains, which is essential for CD3DOR. Extensive experiments on two protocols, including the cross-dataset 3-D object retrieval protocol (3-D to 3-D) on PSB/NTU, and the cross-modal 3-D object retrieval protocol (2-D to 3-D) on MI3DOR-2, demonstrate that the proposed DAGSN can significantly outperform state-of-the-art CD3DOR methods.
Anan Liu, Fu-Bin Guo, Heyu Zhou, Chenggang Yan 0001, Zan Gao 0002, Xuanya Li, Wenhui Li 0001
IEEE Trans. Cybern.4
2022 Edge-Aware Multiscale Feature Integration Network for Salient Object Detection in Optical Remote Sensing Images
abstract
The optical remote sensing images (RSIs) show various spatial resolutions and cluttered background, where salient objects with different scales, types, and orientations are presented in diverse RSI scenes. Therefore, it is inappropriate to directly extend cutting-edge saliency detection methods for conventional RGB images to optical RSIs. Besides, the existing saliency models targeting RSIs often render imperfect saliency maps, where some of them are with coarse boundary details. To solve this problem, this article attempts to introduce the edge information to precisely detect salient objects in RSIs. Accordingly, we propose an edge-aware multiscale feature integration network (EMFI-Net) for salient object detection by conducting multiscale feature integration under the explicit and implicit assistance of salient edge cues. Specifically, our network contains two parts including the encoder and decoder. First, the encoder extracts multiscale deep features from three RSIs with different resolutions, where the high-level deep semantic features from three RSIs are integrated using a cascaded feature fusion module. Second, the encoder explicitly enriches the multiscale deep features by integrating the salient edge cues extracted by a salient edge extraction module. Meanwhile, we also implicitly deploy an edge-aware constraint to the supervision of the saliency map prediction by introducing a hybrid loss function. Finally, the decoder integrates the enriched multiscale deep features in a coarse-to-fine way, yielding a high-quality saliency map. The experiments conducted on two public optical RSI datasets clearly prove the effectiveness and superiority of the proposed EMFI-Net against the state-of-the-art saliency models.
Xiaofei Zhou 0003, Kunye Shen, Zhi Liu 0003, Chen Gong 0002, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 I2Transformer: Intra- and Inter-Relation Embedding Transformer for TV Show Captioning
abstract
TV show captioning aims to generate a linguistic sentence based on the video and its associated subtitle. Compared to purely video-based captioning, the subtitle can provide the captioning model with useful semantic clues such as actors’ sentiments and intentions. However, the effective use of subtitle is also very challenging, because it is the pieces of scrappy information and has semantic gap with visual modality. To organize the scrappy information together and yield a powerful omni-representation for all the modalities, an efficient captioning model requires understanding video contents, subtitle semantics, and the relations in between. In this paper, we propose an Intra- and Inter-relation Embedding Transformer (I2Transformer), consisting of an Intra-relation Embedding Block (IAE) and an Inter-relation Embedding Block (IEE) under the framework of a Transformer. First, the IAE captures the intra-relation in each modality via constructing the learnable graphs. Then, IEE learns the cross attention gates, and selects useful information from each modality based on their inter-relations, so as to derive the omni-representation as the input to the Transformer. Experimental results on the public dataset show that the I2Transformer achieves the state-of-the-art performance. We also evaluate the effectiveness of the IAE and IEE on two other relevant tasks of video with text inputs,i.e., TV show retrieval and video-guided machine translation. The encouraging performance further validates that the IAE and IEE blocks have a good generalization ability. The code is available athttps://github.com/tuyunbin/I2Transformer.
Yunbin Tu, Liang Li 0003, Li Su 0003, Shengxiang Gao, Chenggang Yan 0001, Zhengjun Zha, Zhengtao Yu 0001, Qingming Huang
IEEE Trans. Image Process.5
2022 FPGA-based accelerator for object detection: a comprehensive survey
Kai Zeng 0005, Tao Shen 0004, Chenggang Yan 0001
J. Supercomput.6
2022 Rich Embedding Features for One-Shot Semantic Segmentation
abstract
One-shot semantic segmentation poses the challenging task of segmenting object regions from unseen categories with only one annotated example as guidance. Thus, how to effectively construct robust feature representations from the guidance image is crucial to the success of one-shot semantic segmentation. To this end, we propose in this article a simple, yet effective approach named rich embedding features (REFs). Given a reference image accompanied with its annotated mask, our REF constructs rich embedding features of the support object from three perspectives: 1) global embedding to capture the general characteristics; 2) peak embedding to capture the most discriminative information; 3) adaptive embedding to capture the internal long-range dependencies. By combining these informative features, we can easily harvest sufficient and rich guidance even from a single reference image. In addition to REF, we further propose a simple depth-priority context module to obtain useful contextual cues from the query image. This successfully raises the performance of one-shot semantic segmentation to a new level. We conduct experiments on pattern analysis, statical modeling and computational learning (Pascal) visual object classes (VOC) 2012 and common object in context (COCO) to demonstrate the effectiveness of our approach.
Xiaolin Zhang 0006, Yunchao Wei, Chenggang Yan 0001, Yi Yang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Age-Invariant Face Recognition by Multi-Feature Fusionand Decomposition with Self-attention
abstract
Different from general face recognition, age-invariant face recognition (AIFR) aims at matching faces with a big age gap. Previous discriminative methods usually focus on decomposing facial feature into age-related and age-invariant components, which suffer from the loss of facial identity information. In this article, we propose a novel Multi-feature Fusion and Decomposition (MFD) framework for age-invariant face recognition, which learns more discriminative and robust features and reduces the intra-class variants. Specifically, we first sample multiple face images of different ages with the same identity as a face time sequence. Then, the multi-head attention is employed to capture contextual information from facial feature series, extracted by the backbone network. Next, we combine feature decomposition with fusion based on the face time sequence to ensure that the final age-independent features effectively represent the identity information of the face and have stronger robustness against the aging process. Besides, we also mitigate imbalanced age distribution in the training data by a re-weighted age loss. We experimented with the proposed MFD over the popular CACD and CACD-VS datasets, where we show that our approach improves the AIFR performance than previous state-of-the-art methods. We simultaneously show the performance of MFD on LFW dataset.
Chenggang Yan 0001, Lixuan Meng, Liang Li 0003, Jian Yin 0003, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng
ACM Trans. Multim. Comput. Commun. Appl.1
2022 PlaneFusion: Real-Time Indoor Scene Reconstruction With Planar Prior
abstract
Real-time dense SLAM techniques aim to reconstruct the dense three-dimensional geometry of a scene in real time with an RGB or RGB-D sensor. An indoor scene is an important type of working environment for these techniques. The planar prior can be used in this scenario to improve the reconstruction quality, especially for large low-texture regions that commonly occur in an indoor scene. This article fully explores the planar prior in a dense SLAM pipeline. First, we propose a novel plane detection and segmentation method that runs at 200 Hz on a modern graphics processing unit. Our algorithm for constructing global plane constraints is very efficient; hence, we use it in the process of each input frame for the camera pose estimation while maintaining the real-time performance. Second, we propose herein a plane-based map representation that greatly reduces the memory footprint of plane regions while keeping the geometric details on planes. The experiments reveal that our system yields superior reconstruction results with planar information running at more than 30 fps. Aside from speed and storage improvements, our technique also handles the low-texture problem in plane regions.
Bingjian Gong, Zunjie Zhu, Chenggang Yan 0001, Zhiguo Shi 0001, Feng Xu 0005
IEEE Trans. Vis. Comput. Graph.3
2021 Automated Model Design and Benchmarking of Deep Learning Models for COVID-19 Detection with Chest CT Scans
abstract
The COVID-19 pandemic has spread globally for several months. Because its transmissibility and high pathogenicity seriously threaten people's lives, it is crucial to accurately and quickly detect COVID-19 infection. Many recent studies have shown that deep learning (DL) based solutions can help detect COVID-19 based on chest CT scans. However, most existing work focuses on 2D datasets, which may result in low quality models as the real CT scans are 3D images. Besides, the reported results span a broad spectrum on different datasets with a relatively unfair comparison. In this paper, we first use three state-of-the-art 3D models (ResNet3D101, DenseNet3D121, and MC3\_18) to establish the baseline performance on three publicly available chest CT scan datasets. Then we propose a differentiable neural architecture search (DNAS) framework to automatically search the 3D DL models for 3D chest CT scans classification and use the Gumbel Softmax technique to improve the search efficiency. We further exploit the Class Activation Mapping (CAM) technique on our models to provide the interpretability of the results. The experimental results show that our searched models (CovidNet3D) outperform the baseline human-designed models on three datasets with tens of times smaller model size and higher accuracy. Furthermore, the results also verify that CAM can be well applied in CovidNet3D for COVID-19 datasets to provide interpretability for medical diagnosis. Code: https://github.com/HKBU-HPML/CovidNet3D.
Xin He 0019, Xiaowen Chu 0001, Shaohuai Shi, Jiangping Tang, Xin Liu 0027, Chenggang Yan 0001, Jiyong Zhang 0001, Guiguang Ding
AAAI7
2021 R\^3Net: Relation-embedded Representation Reconstruction Network for Change Captioning
abstract
Change captioning is to use a natural language sentence to describe the fine-grained disagreement between two similar images.Viewpoint change is the most typical distractor in this task, because it changes the scale and location of the objects and overwhelms the representation of real change.In this paper, we propose a Relation-embedded Representation Reconstruction Network (R 3 Net) to explicitly distinguish the real change from the large amount of clutter and irrelevant changes.Specifically, a relation-embedded module is first devised to explore potential changed objects in the large amount of clutter.Then, based on the semantic similarities of corresponding locations in the two images, a representation reconstruction module (RRM) is designed to learn the reconstruction representation and further model the difference representation.Besides, we introduce a syntactic skeleton predictor (SSP) to enhance the semantic interaction between change localization and caption generation.Extensive experiments show that the proposed method achieves the state-of-the-art results on two public datasets 1 .
Yunbin Tu, Liang Li 0003, Chenggang Yan 0001, Shengxiang Gao, Zhengtao Yu 0001
EMNLP (1)3
2021 TraND: Transferable Neighborhood Discovery for Unsupervised Cross-Domain Gait Recognition
abstract
Gait, i.e., the movement pattern of human limbs during locomotion, is a promising biometrie for identification of persons. Despite significant improvement in gait recognition with deep learning, existing studies still neglect a more practical but challenging scenario - unsupervised cross-domain gait recognition which aims to learn a model on a labeled dataset then adapt it to an unlabeled dataset. Due to the domain shift and class gap, directly applying a model trained on one source dataset to other target datasets usually obtains very poor results. Therefore, this paper proposes a Transferable Neighborhood Discovery (TraND) framework to bridge the domain gap for unsupervised cross-domain gait recognition. To learn effective prior knowledge for gait representation, we first adopt a backbone network pre- trained on the labeled source data in a supervised manner. Then we design an end-to-end trainable approach to automatically discover the confident neighborhoods of unlabeled samples in the latent space. During training, the class consistency indicator is adopted to select confident neighborhoods of samples based on their entropy measurements. Moreover, we explore a high- entropy-first neighbor selection strategy, which can effectively transfer prior knowledge to the target domain. Our method achieves the state-of-the-art results on two public datasets, i.e., CASIA-B and OU-LP.
Jinkai Zheng, Xinchen Liu, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Xiao-Ping Zhang 0002, Tao Mei 0001
ISCAS3
2021 Mask and Predict: Multi-step Reasoning for Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to parse the image as a set of semantics, containing objects and their relations. Currently, the SGG methods only stay at presenting the intuitive detection in the image, such as the triplet "logo on board". Intuitively, we humans can further refine these intuitive detections as rational descriptions like "flower painted on surfboard". However, most of existing methods always formulate SGG as a straightforward task, only limited by the manner of one-time prediction, which focuses on a single-pass pipeline and predicts all the semantic. Therefore, to handle this problem, we propose a novel multi-step reasoning manner for SGG. Concretely, we break SGG into two explicit learning stages, including intuitive training stage (ITS) and rational training stage (RTS). In the first stage, we follow the traditional SGG processing to detect objects and relationships, yielding an intuitive scene graph. In the second stage, we perform multi-step reasoning to refine the intuitive scene graph. For each step of reasoning, it consists of two kinds of operations: mask and predict. According to primary predictions and their confidences, we constantly select and mask the low-confidence predictions, which features are optimized and predicted again. After several iterations, all of intuitive semantics will gradually tend to be revised with high confidences, yielding a rational scene graph. Extensive experiments on Visual Genome prove the superiority of the proposed method. Additional ablation studies and visualization cases further validate its effectiveness.
Hongshuo Tian, Ning Xu 0003, Anan Liu, Chenggang Yan 0001, Zhendong Mao 0001, Yongdong Zhang 0001
ACM Multimedia4
2021 Heuristic Depth Estimation with Progressive Depth Reconstruction and Confidence-Aware Loss
abstract
Recently deep learning-based depth estimation has shown the promising result, especially with the help of sparse depth reference samples. Existing works focus on directly inferring the depth information from sparse samples with high confidence. In this paper, we propose a Heuristic Depth Estimation Network (HDEN) with progressive depth reconstruction and confidence-aware loss. The HDEN leverages the reference samples with low confidence to distill the spatial geometric and local semantic information for dense depth prediction. Specifically, we first train a U-NET network to generate a coarse-level dense reference map. Second, the progressive depth reconstruction module successively reconstructs the fine-level dense depth map from different scales, where a multi-level upsampling block is designed to recover the local structure of object. Finally, the confidence-aware loss is proposed to trigger the reference samples with low confidence, which enforces the model focusing on estimating the depth of the tiny structure. Extensive experiments on the NYU-Depth-v2 and KITTI-Odometry dataset show the effectiveness of our method. Visualization results demonstrate that the dense depth maps generated by HDEN have better consistency at the entity edge with RGB image.
Liang Li 0003, Chenggang Yan 0001, Yaoqi Sun, Tao Shen 0004, Jiyong Zhang 0001
ACM Multimedia3
2021 Cross-modal semantic correlation learning by Bi-CNN network
abstract
Abstract Cross modal retrieval can retrieve images through a text query and vice versa. In recent years, cross modal retrieval has attracted extensive attention. The purpose of most now available cross modal retrieval methods is to find a common subspace and maximize the different modal correlation. To generate specific representations consistent with cross modal tasks, this paper proposes a novel cross modal retrieval framework, which integrates feature learning and latent space embedding. In detail, we proposed a deep CNN and a shallow CNN to extract the feature of the samples. The deep CNN is used to extract the representation of images, and the shallow CNN uses a multi‐dimensional kernel to extract multi‐level semantic representation of text. Meanwhile, we enhance the semantic manifold by constructing cross modal ranking and within‐modal discriminant loss to improve the division of semantic representation. Moreover, the most representative samples are selected by using online sampling strategy, so that the approach can be implemented on a large‐scale data. This approach not only increases the discriminative ability among different categories, but also maximizes the relativity between different modalities. Experiments on three real word datasets show that the proposed method is superior to the popular methods.
Liang Li 0003, Chenggang Yan 0001, Yaoqi Sun, Jiyong Zhang 0001
IET Image Process.3
2021 Joint hub identification for brain networks by multivariate graph inference
Defu Yang, Xiaofeng Zhu 0001, Chenggang Yan 0001, Zi-Wen Peng, Maria Bagonis, Paul J. Laurienti, Martin Styner, Guorong Wu 0001
Medical Image Anal.3
2021 An integrated classification model for incremental learning
Ji Hu 0002, Chenggang Yan 0001, Xin Liu 0027, Chengwei Ren, Jiyong Zhang 0001, Dongliang Peng 0001, Yi Yang 0001
Multim. Tools Appl.2
2021 Stroke prediction from electrocardiograms by deep neural network
Yifeng Xie, Hongnan Yang, Xi Yuan, Ruitao Zhang, Qianyun Zhu, Zhenhai Chu, Chengming Yang, Peiwu Qin, Chenggang Yan 0001
Multim. Tools Appl.10
2021 Deep Multi-View Enhancement Hashing for Image Retrieval
abstract
Hashing is an efficient method for nearest neighbor search in large-scale data space by embedding high-dimensional feature descriptors into a similarity preserving Hamming space with a low dimension. However, large-scale high-speed retrieval through binary code has a certain degree of reduction in retrieval accuracy compared to traditional retrieval methods. We have noticed that multi-view methods can well preserve the diverse characteristics of data. Therefore, we try to introduce the multi-view deep neural network into the hash learning field, and design an efficient and innovative retrieval model, which has achieved a significant improvement in retrieval performance. In this paper, we propose a supervised multi-view hash model which can enhance the multi-view information through neural networks. This is a completely new hash learning method that combines multi-view and deep learning methods. The proposed method utilizes an effective view stability evaluation method to actively explore the relationship among views, which will affect the optimization direction of the entire network. We have also designed a variety of multi-data fusion methods in the Hamming space to preserve the advantages of both convolution and multi-view. In order to avoid excessive computing resources on the enhancement procedure during retrieval, we set up a separate structure called memory network which participates in training together. The proposed method is systematically evaluated on the CIFAR-10, NUS-WIDE and MS-COCO datasets, and the results show that our method significantly outperforms the state-of-the-art single-view and multi-view hashing methods.
Chenggang Yan 0001, Biao Gong, Yuxuan Wei, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Evolution of ICTs-empowered-identification: A general re-ranking method for person re-identification
Tongkun Xu, Bolun Zheng, Yaoqi Sun, Anan Liu, Zhendong Mao 0001, Chenggang Yan 0001
Pattern Recognit. Lett.8
2021 Sparse intrinsic decomposition and applications
Kun Li 0001, Xinchen Ye, Chenggang Yan 0001, Jing-Yu Yang 0002
Signal Process. Image Commun.4
2021 Multi-Scale Representation Learning on Hypergraph for 3D Shape Retrieval and Recognition
abstract
Effective 3D shape retrieval and recognition are challenging but important tasks in computer vision research field, which have attracted much attention in recent decades. Although recent progress has shown significant improvement of deep learning methods on 3D shape retrieval and recognition performance, it is still under investigated of how to jointly learn an optimal representation of 3D shapes considering their relationships. To tackle this issue, we propose a multi-scale representation learning method on hypergraph for 3D shape retrieval and recognition, called multi-scale hypergraph neural network (MHGNN). In this method, the correlation among 3D shapes is formulated in a hypergraph and a hypergraph convolution process is conducted to learn the representations. Here, multiple representations can be obtained through different convolution layers, leading to multi-scale representations of 3D shapes. A fusion module is then introduced to combine these representations for 3D shape retrieval and recognition. The main advantages of our method lie in 1) the high-order correlation among 3D shapes can be investigated in the framework and 2) the joint multi-scale representation can be more robust for comparison. Comparisons with state-of-the-art methods on the public ModelNet40 dataset demonstrate remarkable performance improvement of our proposed method on the 3D shape retrieval task. Meanwhile, experiments on recognition tasks also show better results of our proposed method, which indicate the superiority of our method on learning better representation for retrieval and recognition.
Biao Gong, Fuqiang Lei, Chenggang Yan 0001, Yue Gao 0002
IEEE Trans. Image Process.5
2021 Dynamic Selective Network for RGB-D Salient Object Detection
abstract
RGB-D saliency detection is receiving more and more attention in recent years. There are many efforts have been devoted to this area, where most of them try to integrate the multi-modal information, i.e. RGB images and depth maps, via various fusion strategies. However, some of them ignore the inherent difference between the two modalities, which leads to the performance degradation when handling some challenging scenes. Therefore, in this paper, we propose a novel RGB-D saliency model, namely Dynamic Selective Network (DSNet), to perform salient object detection (SOD) in RGB-D images by taking full advantage of the complementarity between the two modalities. Specifically, we first deploy a cross-modal global context module (CGCM) to acquire the high-level semantic information, which can be used to roughly locate salient objects. Then, we design a dynamic selective module (DSM) to dynamically mine the cross-modal complementary information between RGB images and depth maps, and to further optimize the multi-level and multi-scale information by executing the gated and pooling based selection, respectively. Moreover, we conduct the boundary refinement to obtain high-quality saliency maps with clear boundary details. Extensive experiments on eight public RGB-D datasets show that the proposed DSNet achieves a competitive and excellent performance against the current 17 state-of-the-art RGB-D SOD models.
Hongfa Wen, Chenggang Yan 0001, Xiaofei Zhou 0003, Runmin Cong, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Yongjun Bao, Guiguang Ding
IEEE Trans. Image Process.2
2021 Precise No-Reference Image Quality Evaluation Based on Distortion Identification
abstract
The difficulty of no-reference image quality assessment (NR IQA) often lies in the lack of knowledge about the distortion in the image, which makes quality assessment blind and thus inefficient. To tackle such issue, in this article, we propose a novel scheme for precise NR IQA, which includes two successive steps, i.e., distortion identification and targeted quality evaluation. In the first step, we employ the well-known Inception-ResNet-v2 neural network to train a classifier that classifies the possible distortion in the image into the four most common distortion types, i.e., Gaussian white noise (WN), Gaussian blur (GB), jpeg compression (JPEG), and jpeg2000 compression (JP2K). Specifically, the deep neural network is trained on the large-scale Waterloo Exploration database, which ensures the robustness and high performance of distortion classification. In the second step, after determining the distortion type of the image, we then design a specific approach to quantify the image distortion level, which can estimate the image quality specially and more precisely. Extensive experiments performed on LIVE, TID2013, CSIQ, and Waterloo Exploration databases demonstrate that (1) the accuracy of our distortion classification is higher than that of the state-of-the-art distortion classification methods, and (2) the proposed NR IQA method outperforms the state-of-the-art NR IQA methods in quantifying the image quality.
Chenggang Yan 0001, Tong Teng, Yutao Liu 0002, Yongbing Zhang 0002, Haoqian Wang, Xiangyang Ji
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Depth Image Denoising Using Nuclear Norm and Learning Graph Model
abstract
Depth image denoising is increasingly becoming the hot research topic nowadays, because it reflects the three-dimensional scene and can be applied in various fields of computer vision. But the depth images obtained from depth camera usually contain stains such as noise, which greatly impairs the performance of depth-related applications. In this article, considering that group-based image restoration methods are more effective in gathering the similarity among patches, a group-based nuclear norm and learning graph (GNNLG) model was proposed. For each patch, we find and group the most similar patches within a searching window. The intrinsic low-rank property of the grouped patches is exploited in our model. In addition, we studied the manifold learning method and devised an effective optimized learning strategy to obtain the graph Laplacian matrix, which reflects the topological structure of image, to further impose the smoothing priors to the denoised depth image. To achieve fast speed and high convergence, the alternating direction method of multipliers is proposed to solve our GNNLG. The experimental results show that the proposed method is superior to other current state-of-the-art denoising methods in both subjective and objective criterion.
Chenggang Yan 0001, Zhisheng Li, Yongbing Zhang 0002, Yutao Liu 0002, Xiangyang Ji, Yongdong Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Unsupervised Person Re-Identification via Softened Similarity Learning
abstract
Person re-identification (re-ID) is an important topic in computer vision. This paper studies the unsupervised setting of re-ID, which does not require any labeled information and thus is freely deployed to new scenarios. There are very few studies under this setting, and one of the best approach till now used iterative clustering and classification, so that unlabeled images are clustered into pseudo classes for a classifier to get trained, and the updated features are used for clustering and so on. This approach suffers two problems, namely, the difficulty of determining the number of clusters, and the hard quantization loss in clustering. In this paper, we follow the iterative training mechanism but discard clustering, since it incurs loss from hard quantization, yet its only product, image-level similarity, can be easily replaced by pairwise computation and a softened classification task. With these improvements, our approach becomes more elegant and is more robust to hyper-parameter changes. Experiments on two image-based and video-based datasets demonstrate state-of-the-art performance under the unsupervised re-ID setting.
Yutian Lin, Lingxi Xie, Yu Wu 0011, Chenggang Yan 0001, Qi Tian 0001
CVPR4
2020 Real-World Automatic Makeup via Identity Preservation Makeup Net
abstract
This paper focuses on the real-world automatic makeup problem. Given one non-makeup target image and one reference image, the automatic makeup is to generate one face image, which maintains the original identity with the makeup style in the reference image. In the real-world scenario, face makeup task demands a robust system against the environmental variants. The two main challenges in real-world face makeup could be summarized as follow: first, the background in real-world images is complicated. The previous methods are prone to change the style of background as well; second, the foreground faces are also easy to be affected. For instance, the ``heavy'' makeup may lose the discriminative information of the original identity. To address these two challenges, we introduce a new makeup model, called Identity Preservation Makeup Net (IPM-Net), which preserves not only the background but the critical patterns of the original identity. Specifically, we disentangle the face images to two different information codes, i.e., identity content code and makeup style code. When inference, we only need to change the makeup style code to generate various makeup images of the target person. In the experiment, we show the proposed method achieves not only better accuracy in both realism (FID) and diversity (LPIPS) in the test set, but also works well on the real-world images collected from the Internet.
Zhikun Huang, Zhedong Zheng, Chenggang Yan 0001, Hongtao Xie 0001, Yaoqi Sun, Jiyong Zhang 0001
IJCAI3
2020 Diverter-Guider Recurrent Network for Diverse Poems Generation from Image
abstract
Poem generation from image aims to automatically generate the poetic sentences for presenting the image content or overtone. Previous works focused on 1-to-1 image-poem generation with the demands of poeticness and content relevance. This paper proposes the paradigm of multiple poems generation from one image, which is closer to human poetizing but more challenging. Its key problem is to simultaneously guarantee the diversity of multiple poems with poeticness and relevance. To this end, we propose an end-to-end probabilistic Diverter-Guider Recurrent Network (DG-Net), which is a context-based encoder-decoder generative model with the hierarchical stochastic variables. Specifically, the diverter-variable represents the decoding-context inferred from the input image to diversify the poem themes; the guider-variable is introduced as an attribute decoder to restricts the word-choice with supervised information. Extensive experiments on automatic evaluations and human judgments demonstrate the superior performance of DG-Net than existing poem generation methods. Qualitative study show that our model can generate diverse poems with the poeticness and relevance.
Liang Li 0003, Li Su 0003, Shuhui Wang, Chenggang Yan 0001, Zhengjun Zha, Qingming Huang
ACM Multimedia5
2020 Beyond the Parts: Learning Multi-view Cross-part Correlation for Vehicle Re-identification
abstract
Vehicle re-identification (Re-Id) is a challenging task due to the inter-class similarity, the intra-class difference, and the cross-view misalignment of vehicle parts. Although recent methods achieve great improvement by learning detailed features from keypoints or bounding boxes of parts, vehicle Re-Id is still far from being solved. Different from existing methods, we propose a Parsing-guided Cross-part Reasoning Network, named as PCRNet, for vehicle Re-Id. The PCRNet explores vehicle parsing to learn discriminative part-level features, model the correlation among vehicle parts, and achieve precise part alignment for vehicle Re-Id. To accurately segment vehicle parts, we first build a large-scale Multi-grained Vehicle Parsing (MVP) dataset from surveillance images. With the parsed parts, we extract regional features for each part and build a part-neighboring graph to explicitly model the correlation among parts. Then, the graph convolutional networks (GCNs) are adopted to propagate local information among parts, which can discover the most effective local features of varied viewpoints. Moreover, we propose a self-supervised part prediction loss to make the GCNs generate features of invisible parts from visible parts under different viewpoints. By this means, the same vehicle from different viewpoints can be matched with the well-aligned and robust feature representations. Through extensive experiments, our PCRNet significantly outperforms the state-of-the-art methods on three large-scale vehicle Re-Id datasets.
Xinchen Liu, Wu Liu 0005, Jinkai Zheng, Chenggang Yan 0001, Tao Mei 0001
ACM Multimedia4
2020 Multi-Features Fusion and Decomposition for Age-Invariant Face Recognition
abstract
Although the General Face Recognition (GFR) research achieves great success, Age-Invariant Face Recognition (AIFR) is still a challenging problem since facial appearance changing over time brings significant intra-class variations. The existing discriminative methods for the AIFR task mostly focus on decomposing the facial feature from a sigle image into age-related feature and age-independent feature for recognition, which suffer from the loss of facial identity information. To address this issue, in this work we propose a novel Multi-Features Fusion and Decomposition (MFFD) framework to learn more discriminative feature representations and alleviate the intra-class variations for AIFR. Specifically, we first sample multiple face images of different ages with the same identity as a face time series. Next, we combine feature decomposition with fusion based on the face time series to ensure that the final age-independent features effectively represent the identity information of the face and have stronger robustness against aging. Moreover, we also present two feature fusion methods and several different training strategies to explore the impact on the model. Extensive experiments on several cross-age datasets (CACD, CACD-VS) demonstrate the effectiveness of our proposed method. Besides, our method also shows comparable generalization performance on the well-known LFW dataset.
Lixuan Meng, Chenggang Yan 0001, Jian Yin 0003, Wu Liu 0005, Hongtao Xie 0001, Liang Li 0003
ACM Multimedia2
2020 Improving Just Noticeable Difference Model by Leveraging Temporal HVS Perception Characteristics
Haibing Yin, Yafen Xing, Guangjing Xia, Xiaofeng Huang, Chenggang Yan 0001
MMM (1)5
2020 Learning salient features to prevent model drift for correlation tracking
Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong, Liang Li 0003, Chenggang Yan 0001, Tao Shen 0004
Neurocomputing6
2020 Parameters Analysis of Sample Entropy, Permutation Entropy and Permutation Ratio Entropy for RR Interval Time Series
Jian Yin 0003, Pengxiang Xiao, Yungang Liu, Chenggang Yan 0001, Yatao Zhang
Inf. Process. Manag.5
2020 Cross-modal feature extraction and integration based RGBD saliency detection
Liang Pan, Xiaofei Zhou 0003, Jiyong Zhang 0001, Chenggang Yan 0001
Image Vis. Comput.5
2020 Depth-guided saliency detection via boundary information
Xiaofei Zhou 0003, Hongfa Wen, Haibing Yin, Chenggang Yan 0001
Image Vis. Comput.5
2020 Multi-stage all-zero block detection for HEVC coding using machine learning
Haibing Yin, Haoyun Yang, Xiaofeng Huang, Hongkui Wang, Chenggang Yan 0001
J. Vis. Commun. Image Represent.5
2020 Towards context-aware collaborative filtering by learning context-aware latent representations
Xin Liu 0027, Jiyong Zhang 0001, Chenggang Yan 0001
Knowl. Based Syst.3
2020 Cascaded Revision Network for Novel Object Captioning
abstract
Image captioning, a challenging task where the machine automatically describes an image with natural language, has drawn significant attention in recent years. Despite the remarkable improvements of recent approaches, however, these methods are built upon a large set of training image-sentence pairs. The expensive labor efforts hence limit the captioning model to describe the wider world. In this paper, we present a novel network structure, Cascaded Revision Network, which aims at relieving the problem by equipping the model with out-of-domain knowledge. CRN first tries its best to describe an image using the existing vocabulary from in-domain knowledge. Due to the lack of out-of-domain knowledge, the caption may be inaccurate or include ambiguous words for the image with unknown (novel) objects. We propose to re-edit the primary captioning sentence by a series of cascaded operations. We introduce a perplexity predictor to find out which words are most likely to be inaccurate given the input image. Thereafter, we utilize external knowledge from a pretrained object detection model and select more accurate words from detection results by the visual matching module. In the last step, we design a semantic matching module to ensure that the novel object is fit in the right position. By this novel cascaded captioning-revising mechanism, CRN can accurately describe images with unseen objects. We validate the proposed method with state-of-the-art performance on the held-out MSCOCO dataset as well as scale to ImageNet, demonstrating the effectiveness of our method.
Qianyu Feng, Yu Wu 0011, Hehe Fan, Chenggang Yan 0001, Mingliang Xu 0001, Yi Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 Weighted Convolutional Motion-Compensated Frame Rate Up-Conversion Using Deep Residual Network
abstract
Frame rate up-conversion (FRUC) usually suffers from unreliable motion vectors due to the absence of the current frame to be interpolated. In addition, since the majority of video sequences are usually compressed by various coding standards to reduce the data volume, the quality of the generated frames in the FRUC will be further impaired. To address this problem, we proposed two FRUC algorithms based on deep residual network. We first present a deep residual network for the FRUC (DRNFRUC), which consists of feature extraction, feature recursive analysis, and image restoration parts with a skip connection between the input and the output of the network. The proposed DRNFRUC takes the result of an arbitrary existing FRUC method as the input and is able to significantly reduce the edge blurring and blocking artifacts when the motion of the block is violent. In addition, we proposed a deep residual network with weighted convolutional motion compensation (DRNWCMC) for the FRUC, where the convolution operations can be embedded into the motion compensation interpolation (MCI) in any existing MCI-based FRUC method. In DRNWCMC, we first devise two convolutional neural networks corresponding to the forward and backward motion compensated frames, respectively. And then, the adaptive interpolation coefficients for motion compensation are designed as two$1\times1$convolutional kernels. Finally, the interpolation result of WCMC is fed into another convolutional neural network to further improve the performance. All the parameters involved in the DRNWCMC are trained simultaneously under the same cost function. The experimental results show that the two proposed algorithms can remarkably improve both the objective and subjective quality of the interpolated frames.
Yongbing Zhang 0002, Lixin Chen, Chenggang Yan 0001, Peiwu Qin, Xiangyang Ji, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.3
2020 Asymptotic Soft Filter Pruning for Deep Convolutional Neural Networks
abstract
Deeper and wider convolutional neural networks (CNNs) achieve superior performance but bring expensive computation cost. Accelerating such overparameterized neural network has received increased attention. A typical pruning algorithm is a three-stage pipeline, i.e., training, pruning, and retraining. Prevailing approaches fix the pruned filters to zero during retraining and, thus, significantly reduce the optimization space. Besides, they directly prune a large number of filters at first, which would cause unrecoverable information loss. To solve these problems, we propose an asymptotic soft filter pruning (ASFP) method to accelerate the inference procedure of the deep neural networks. First, we update the pruned filters during the retraining stage. As a result, the optimization space of the pruned model would not be reduced but be the same as that of the original model. In this way, the model has enough capacity to learn from the training data. Second, we prune the network asymptotically. We prune few filters at first and asymptotically prune more filters during the training procedure. With asymptotic pruning, the information of the training set would be gradually concentrated in the remaining filters, so the subsequent training and pruning process would be stable. The experiments show the effectiveness of our ASFP on image classification benchmarks. Notably, on ILSVRC-2012, our ASFP reduces more than 40% FLOPs on ResNet-50 with only 0.14% top-5 accuracy degradation, which is higher than the soft filter pruning by 8%.
Yang He 0002, Xuanyi Dong, Guoliang Kang, Yanwei Fu 0001, Chenggang Yan 0001, Yi Yang 0001
IEEE Trans. Cybern.5
2020 Hamming Embedding Sensitivity Guided Fusion Network for 3D Shape Representation
abstract
Three-dimensional multi-modal data are used to represent 3D objects in the real world in different ways. Features separately extracted from multimodality data are often poorly correlated. Recent solutions leveraging the attention mechanism to learn a joint-network for the fusion of multimodality features have weak generalization capability. In this paper, we propose a hamming embedding sensitivity network to address the problem of effectively fusing multimodality features. The proposed network called HamNet is the first end-to-end framework with the capacity to theoretically integrate data from all modalities with a unified architecture for 3D shape representation, which can be used for 3D shape retrieval and recognition. HamNet uses the feature concealment module to achieve effective deep feature fusion. The basic idea of the concealment module is to re-weight the features from each modality at an early stage with the hamming embedding of these modalities. The hamming embedding also provides an effective solution for fast retrieval tasks on a large scale dataset. We have evaluated the proposed method on the large-scale ModelNet40 dataset for the tasks of 3D shape classification, single modality and cross-modality retrieval. Comprehensive experiments and comparisons with state-of-the-art methods demonstrate that the proposed approach can achieve superior performance.
Biao Gong, Chenggang Yan 0001, Changqing Zou, Yue Gao 0002
IEEE Trans. Image Process.2
2020 Unsupervised Person Re-identification via Cross-Camera Similarity Exploration
abstract
Most person re-identification (re-ID) approaches are based on supervised learning, which requires manually annotated data. However, it is not only resource-intensive to acquire identity annotation but also impractical for large-scale data. To relieve this problem, we propose a cross-camera unsupervised approach that makes use of unsupervised style-transferred images to jointly optimize a convolutional neural network (CNN) and the relationship among the individual samples for person re-ID. Our algorithm considers two fundamental facts in the re- ID task, i.e., variance across diverse cameras and similarity within the same identity. In this paper, we propose an iterative framework which overcomes the camera variance and achieves across-camera similarity exploration. Specifically, we apply an unsupervised style transfer model to generate style-transferred training images with different camera styles. Then we iteratively exploit the similarity within the same identity from both the original and the style-transferred data. We start with considering each training image as a different class to initialize the Convolutional Neural Network (CNN) model. Then we measure the similarity and gradually group similar samples into one class, which increases similarity within each identity. We also introduce a diversity regularization term in the clustering to balance the cluster distribution. The experimental results demonstrate that our algorithm is not only superior to state-of-the-art unsupervised re-ID approaches, but also performs favorably compared with other competing unsupervised domain adaptation methods (UDA) and semi-supervised learning methods.
Yutian Lin, Yu Wu 0011, Chenggang Yan 0001, Mingliang Xu 0001, Yi Yang 0001
IEEE Trans. Image Process.3
2020 Mining Spatial-Temporal Similarity for Visual Tracking
abstract
Correlation filter (CF) is a critical technique to improve accuracy and speed in the field of visual object tracking. Despite being studied extensively, most existing CF methods suffer from failing to make the most of the inherent spatial-temporal prior of videos. To address this limitation, as consecutive frames are eminently resemble in most videos, we investigate a novel scheme to predict targets' future state by exploiting previous observations. Specifically, in this paper, we propose a prediction based CF tracking framework by learning the spatial-temporal similarity of consecutive frames for sample managing, template regularization, and training response pre-weighting. We model the learning problem theoretically as a novel objective and provide effective optimization algorithms to solve the learning task. In addition, we implement two CF trackers with different features. Extensive experiments are conducted on three popular benchmarks to validate our scheme. The encouraging results demonstrate that the proposed scheme can significantly boost the accuracy of CF tracking, and the two trackers achieve competitive performances against state-of-the-art trackers. We finally present a comprehensive analysis on the efficacy of our proposed method and the efficiency of our trackers to facilitate real-world visual tracking applications.
Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong, Hongtao Xie 0001, Chenggang Yan 0001
IEEE Trans. Image Process.6
2020 3D Room Layout Estimation From a Single RGB Image
abstract
3D layout is crucial for scene understanding and reconstruction, and very useful in applications like real estate and furniture design. In this paper, we propose a fully automatic solution to estimate 3D layout of an indoor scene from a single 2D image. Our technique contains two key components. Firstly, we train a neural network that directly estimates room structure lines from the input image. Secondly, we propose a novel technique to automatically identify the layout topology of an input image, followed by a nonlinear optimization with equality constraints to estimate the final 3D layout of a scene. Based on our knowledge, this is the first fully automatic technique to achieve single image-based 3D layout estimation of an indoor scene. We evaluate our method on the public datasets LSUN, Hedau and 3DGP and the results show that the proposed method achieves accurate 3D layout reconstruction on various images with different layout topologies.
Chenggang Yan 0001, Biyao Shao, Hao Zhao 0002, Ruixin Ning, Yongdong Zhang 0001, Feng Xu 0005
IEEE Trans. Multim.1
2020 STAT: Spatial-Temporal Attention Mechanism for Video Captioning
abstract
Video captioning refers to automatic generate natural language sentences, which summarize the video contents. Inspired by the visual attention mechanism of human beings, temporal attention mechanism has been widely used in video description to selectively focus on important frames. However, most existing methods based on temporal attention mechanism suffer from the problems of recognition error and detail missing, because temporal attention mechanism cannot further catch significant regions in frames. In order to address above problems, we propose the use of a novel spatial-temporal attention mechanism (STAT) within an encoder-decoder neural network for video captioning. The proposed STAT successfully takes into account both the spatial and temporal structures in a video, so it makes the decoder to automatically select the significant regions in the most relevant temporal segments for word prediction. We evaluate our STAT on two well-known benchmarks: MSVD and MSR-VTT-10K. Experimental results show that our proposed STAT achieves the state-of-the-art performance with several popular evaluation metrics: BLEU-4, METEOR, and CIDEr.
Chenggang Yan 0001, Yunbin Tu, Xingzheng Wang, Yongbing Zhang 0002, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.1
2020 Corrections to "STAT: Spatial-Temporal Attention Mechanism for Video Captioning"
abstract
Presents corrections to affiliations in the above named paper.
Chenggang Yan 0001, Yunbin Tu, Xingzheng Wang, Yongbing Zhang 0002, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.1
2020 Listen, Look, and Find the One: Robust Person Search with Multimodality Index
abstract
Person search with one portrait, which attempts to search the targets in arbitrary scenes using one portrait image at a time, is an essential yet unexplored problem in the multimedia field. Existing approaches, which predominantly depend on the visual information of persons, cannot solve problems when there are variations in the person’s appearance caused by complex environments and changes in pose, makeup, and clothing. In contrast to existing methods, in this article, we propose an associative multimodality index for person search with face, body, and voice information. In the offline stage, an associative network is proposed to learn the relationships among face, body, and voice information. It can adaptively estimate the weights of each embedding to construct an appropriate representation. The multimodality index can be built by using these representations, which exploit the face and voice as long-term keys and the body appearance as a short-term connection. In the online stage, through the multimodality association in the index, we can retrieve all targets depending only on the facial features of the query portrait. Furthermore, to evaluate our multimodality search framework and facilitate related research, we construct the Cast Search in Movies with Voice (CSM-V) dataset, a large-scale benchmark that contains 127K annotated voices corresponding to tracklets from 192 movies. According to extensive experiments on the CSM-V dataset, the proposed multimodality person search framework outperforms the state-of-the-art methods.
Xiao Wang 0029, Wu Liu 0005, Jun Chen 0001, Xiaobo Wang 0001, Chenggang Yan 0001, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2020 Enabling 5G: sentimental image dominant graph topic model for cross-modality topic detection
Liang Li 0003, Wenchao Li 0004, Jiyong Zhang 0001, Chenggang Yan 0001
Wirel. Networks5
2019 Dual-View Ranking with Hardness Assessment for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is to build recognition models for previously unseen target classes which have no labeled data for training by transferring knowledge from some other related auxiliary source classes with abundant labeled samples to the target ones with class attributes as the bridge. The key is to learn a similarity based ranking function between samples and class labels using the labeled source classes so that the proper (unseen) class label for a test sample can be identified by the function. In order to learn the function, single-view ranking based loss is widely used which aims to rank the true label prior to the other labels for a training sample. However, we argue that the ranking can be performed from the other view, which aims to place the images belonging to a label before the images from the other classes. Motivated by it, we propose a novel DuAl-view RanKing (DARK) loss for zeroshot learning simultaneously ranking labels for an image by point-to-point metric and ranking images for a label by pointto-set metric, which is capable of better modeling the relationship between images and classes. In addition, we also notice that previous ZSL approaches mostly fail to well exploit the hardness of training samples, either using only very hard ones or using all samples indiscriminately. In this work, we also introduce a sample hardness assessment method to ZSL which assigns different weights to training samples based on their hardness, which leads to a more accurate and robust ZSL model. Experiments on benchmarks demonstrate that DARK outperforms the state-of-the-arts for (generalized) ZSL.
Guiguang Ding, Jungong Han, Xiaohan Ding, Sicheng Zhao, Zheng Wang 0001, Chenggang Yan 0001, Qionghai Dai
AAAI7
2019 Recurrent Attention Model for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition is to predict attribute labels of pedestrian from surveillance images, which is a very challenging task for computer vision due to poor imaging quality and small training dataset. It is observed that many semantic pedestrian attributes to be recognised tend to show spatial locality and semantic correlations by which they can be grouped while previous works mostly ignore this phenomenon. Inspired by Recurrent Neural Network (RNN)’s super capability of learning context correlations and Attention Model’s capability of highlighting the region of interest on feature map, this paper proposes end-to-end Recurrent Convolutional (RC) and Recurrent Attention (RA) models, which are complementary to each other. RC model mines the correlations among different attribute groups with convolutional LSTM unit, while RA model takes advantage of the intra-group spatial locality and inter-group attention correlation to improve the performance of pedestrian attribute recognition. Our RA method combines the Recurrent Learning and Attention Model to highlight the spatial position on feature map and mine the attention correlations among different attribute groups to obtain more precise attention. Extensive empirical evidence shows that our recurrent model frameworks achieve state-of-the-art results, based on pedestrian attribute datasets, i.e. standard PETA and RAP datasets.
Xin Zhao 0020, Liufang Sang, Guiguang Ding, Jungong Han, Na Di, Chenggang Yan 0001
AAAI6
2019 Social Relation Recognition From Videos via Multi-Scale Spatial-Temporal Reasoning
abstract
Discovering social relations, e.g., kinship, friendship, etc., from visual contents can make machines better interpret the behaviors and emotions of human beings. Existing studies mainly focus on recognizing social relations from still images while neglecting another important media--video. On one hand, the actions and storylines in videos provide more important cues for social relation recognition. On the other hand, the key persons may appear at arbitrary spatial-temporal locations, even not in one same image from beginning to the end. To overcome these challenges, we propose a Multi-scale Spatial-Temporal Reasoning (MSTR) framework to recognize social relations from videos. For the spatial representation, we not only adopt a temporal segment network to learn global action and scene information, but also design a Triple Graphs model to capture visual relations between persons and objects. For the temporal domain, we propose a Pyramid Graph Convolutional Network to perform temporal reasoning with multi-scale receptive fields, which can obtain both long-term and short-term storylines in videos. By this means, MSTR can comprehensively explore the multi-scale actions and storylines in spatial-temporal dimensions for social relation reasoning in videos. Extensive experiments on a new large-scale Video Social Relation dataset demonstrate the effectiveness of the proposed framework.
Xinchen Liu, Wu Liu 0005, Jingwen Chen 0001, Lianli Gao, Chenggang Yan 0001, Tao Mei 0001
CVPR6
2019 Towards Better Uncertainty Sampling: Active Learning with Multiple Views for Deep Convolutional Neural Network
abstract
Convolutional neural network (CNN) has been successfully applied to many fields, such as image classification and object detection. It relies on huge amount of data. However, labelling a large amount of data is expensive. Active learning is one of the approaches to alleviate the labelling effort. We propose a new active learning approach for CNN. Different from existing active learning algorithms for CNN, first, the active query strategy is measured from multiple views, not only the last output of CNN; second, multiple views are obtained from multiple hidden layers in CNN, not from other related data or models. We evaluate our approach on three widely used datasets: Fashion-MNIST, SVHN and CIFAR-10. Experimental results show that the proposed method outperforms baseline methods in image classification.
Xiaoming Jin, Guiguang Ding, Lan Yi, Chenggang Yan 0001
ICME5
2019 Truncated Gradient Confidence-Weighted Based Online Learning for Imbalance Streaming Data
abstract
Online learning for imbalanced streaming data is an important and challenging problem for many classification tasks in the machine learning research field. Traditional online learning algorithms are mainly focused on classification tasks with balanced data, and with little consideration about the characteristics of imbalanced streaming data. In this paper, we propose a novel online learning algorithm called Truncated Gradient Confidence-Weighted (TGCW), which integrate the truncated gradient algorithm with the confidence weighted algorithm together to improve the feature selection ability while reducing the dimensions of imbalanced streaming data effectively. We study a number of classification tasks with various imbalance data ratio including the pedestrian detection application and compare the performance of the TGCW algorithm with traditional online learning algorithms, and empirical results show that the TGCW algorithm can achieve better performance consistently than other baseline approaches.
Ji Hu 0002, Chenggang Yan 0001, Xin Liu 0027, Jiyong Zhang 0001, Dongliang Peng 0001, Yi Yang 0001
ICME2
2019 Online Learning to Rank in a Listwise Approach for Information Retrieval
abstract
A common approach to learning to rank is to minimize the pair-wise loss. However, established analysis shows that pair-wise loss does not necessarily lead to an optimal list-wise ranking measures, e.g., average precision (AP) or area under precision-recall curve (AUPRC). It becomes more difficult in the online learning setting, where the data arrives sequentially and is scanned only once. This paper proposes an online learning-to-rank algorithm by minimizing the list-wise ranking error, which achieves a vanishing gap between the list-wise loss and the ranking measures. Experiments also testify the effectiveness and robustness of the proposed online List-wise algorithm.
Fan Ma, Haoyun Yang, Haibing Yin, Xiaofeng Huang, Chenggang Yan 0001
ICME5
2019 Real-time Indoor Scene Reconstruction with RGBD and Inertial Input
abstract
Camera motion estimation is a key technique for 3D scene reconstruction. Previous works usually assume slow camera motions, which limit the usage in many real cases. We propose an end-to-end 3D reconstruction system which combines color, depth and inertial measurements to achieve robust reconstruction with fast sensor motions. Our framework utilizes extended Kalman filter to fuse the three kinds of information and involve an iterative method to jointly optimize feature correspondences, camera poses and scene geometry. We also propose a novel geometry-aware patch deformation technique to adapt the feature appearance in image domain, leading to a more accurate feature matching under fast camera motions. Experiments show that our patch deformation method improves the accuracy of feature tracking, and our 3D reconstruction framework outperforms the state-of-the-art solutions under fast camera motions.
Zunjie Zhu, Feng Xu 0005, Chenggang Yan 0001, Xinhong Hao, Xiangyang Ji, Yongdong Zhang 0001, Qionghai Dai
ICME3
2019 Approximated Oracle Filter Pruning for Destructive CNN Width Optimization
abstract
It is not easy to design and run Convolutional Neural Networks (CNNs) due to: 1) finding the optimal number of filters (i.e., the width) at each layer is tricky, given an architecture; and 2) the computational intensity of CNNs impedes the deployment on computationally limited devices. Oracle Pruning is designed to remove the unimportant filters from a well-trained CNN, which estimates the filters’ importance by ablating them in turn and evaluating the model, thus delivers high accuracy but suffers from intolerable time complexity, and requires a given resulting width but cannot automatically find it. To address these problems, we propose Approximated Oracle Filter Pruning (AOFP), which keeps searching for the least important filters in a binary search manner, makes pruning attempts by masking out filters randomly, accumulates the resulting errors, and finetunes the model via a multi-path framework. As AOFP enables simultaneous pruning on multiple layers, we can prune an existing very deep CNN with acceptable time cost, negligible accuracy drop, and no heuristic knowledge, or re-design a model which exerts higher accuracy and faster inference.
Xiaohan Ding, Guiguang Ding, Jungong Han, Chenggang Yan 0001
ICML5
2019 Landmark Selection for Zero-shot Learning
abstract
Zero-shot learning (ZSL) is an emerging research topic whose goal is to build recognition models for previously unseen classes. The basic idea of ZSL is based on heterogeneous feature matching which learns a compatibility function between image and class features using seen classes. The function is constructed based on one-vs-all training in which each class has only one class feature and many image features. Existing ZSL works mostly treat all image features equivalently. However, in this paper we argue that it is more reasonable to use some representative cross-domain data instead of all. Motivated by this idea, we propose a novel approach, termed as Landmark Selection(LAST) for ZSL. LAST is able to identify representative cross-domain features which further lead to better image-class compatibility function. Experiments on several ZSL datasets including ImageNet demonstrate the superiority of LAST to the state-of-the-arts.
Guiguang Ding, Jungong Han, Chenggang Yan 0001, Jiyong Zhang 0001, Qionghai Dai
IJCAI4
2019 Joint Identification of Network Hub Nodes by Multivariate Graph Inference
Defu Yang, Chenggang Yan 0001, Feiping Nie 0001, Xiaofeng Zhu 0001, Md Asadullah Turja, Leo Zsembik, Martin Styner, Guorong Wu 0001
MICCAI (3)2
2019 Optimizing sparse tensor times matrix on GPUs
Yuchen Ma 0001, Jiajia Li 0001, Chenggang Yan 0001, Jimeng Sun 0001, Richard W. Vuduc
J. Parallel Distributed Comput.4
2019 Image classification base on PCA of multi-view deep representation
Yaoqi Sun, Liang Li 0003, Liang Zheng 0007, Ji Hu 0002, Wenchao Li 0004, Yatong Jiang, Chenggang Yan 0001
J. Vis. Commun. Image Represent.7
2019 Deep fusion based video saliency detection
Hongfa Wen, Xiaofei Zhou 0003, Yaoqi Sun, Jiyong Zhang 0001, Chenggang Yan 0001
J. Vis. Commun. Image Represent.5
2019 CFMDA: collaborative filtering-based MiRNA-disease association prediction
Zhisheng Li, Bingtao Liu, Chenggang Yan 0001
Multim. Tools Appl.3
2019 Real-time indoor scene reconstruction with Manhattan assumption
Zunjie Zhu, Feng Xu 0005, Chenggang Yan 0001, Bingjian Gong, Yongdong Zhang 0001, Qionghai Dai
Multim. Tools Appl.3
2019 Improving person re-identification by attribute and identity learning
Yutian Lin, Liang Zheng 0001, Zhedong Zheng, Yu Wu 0011, Zhilan Hu, Chenggang Yan 0001, Yi Yang 0001
Pattern Recognit.6
2019 Double-Bit Quantization and Index Hashing for Nearest Neighbor Search
abstract
As binary code is storage efficient and fast to compute, it has become a trend to compact real-valued data to binary codes for the nearest neighbors (NN) search in a large-scale database. However, the use of binary code for the NN search leads to low retrieval accuracy. To increase the discriminability of the binary codes of existing hash functions, in this paper, we propose a framework of double-bit quantization and index hashing for an effective NN search. The main contributions of our framework are: first, a novel double-bit quantization (DBQ) is designed to assign more bits to each dimension for higher retrieval accuracy; second, a double-bit index hashing (DBIH) is presented to efficiently index binary codes generated by DBQ; and third, a weighted distance measurement for DBQ binary codes is put forward to re-rank the search results from DBIH. The empirical results on three benchmark databases demonstrate the superiority of our framework over existing approaches in terms of both retrieval accuracy and query efficiency. Specifically, we observe an absolute improvement on precision of 10%-25% in most cases and the query speed increases over 30 times compared to traditional binary embedding methods and linear scan, respectively.
Hongtao Xie 0001, Zhendong Mao 0001, Yongdong Zhang 0001, Chenggang Yan 0001, Zhineng Chen
IEEE Trans. Multim.5
2019 Cross-Modality Bridging and Knowledge Transferring for Image Understanding
abstract
The understanding of web images has been a hot research topic in both artificial intelligence and multimedia content analysis domains. The web images are composed of various complex foregrounds and backgrounds, which makes the design of an accurate and robust learning algorithm a challenging task. To solve the above significant problem, first, we learn a cross-modality bridging dictionary for the deep and complete understanding of a vast quantity of web images. The proposed algorithm leverages the visual features into the semantic concept probability distribution, which can construct a global semantic description for images while preserving the local geometric structure. To discover and model the occurrence patterns between intra- and inter-categories, multi-task learning is introduced for formulating the objective formulation with Capped-ℓ1penalty, which can obtain the optimal solution with a higher probability and outperform the traditional convex function-based methods. Second, we propose a knowledge-based concept transferring algorithm to discover the underlying relations of different categories. This distribution probability transferring among categories can bring the more robust global feature representation, and enable the image semantic representation to generalize better as the scenario becomes larger. Experimental comparisons and performance discussion with classical methods on the ImageNet, Caltech-256, SUN397, and Scene15 datasets show the effectiveness of our proposed method at three traditional image understanding tasks.
Chenggang Yan 0001, Liang Li 0003, Chunjie Zhang 0001, Bingtao Liu, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.1
2019 Hierarchical Feature Selection for Random Projection
abstract
Random projection is a popular machine learning algorithm, which can be implemented by neural networks and trained in a very efficient manner. However, the number of features should be large enough when applied to a rather large-scale data set, which results in slow speed in testing procedure and more storage space under some circumstances. Furthermore, some of the features are redundant and even noisy since they are randomly generated, so the performance may be affected by these features. To remedy these problems, an effective feature selection method is introduced to select useful features hierarchically. Specifically, a novel criterion is proposed to select useful neurons for neural networks, which establishes a new way for network architecture design. The testing time and accuracy of the proposed method are improved compared with traditional methods and some variations on both classification and regression tasks. Extensive experiments confirm the effectiveness of the proposed method.
Qi Wang 0009, Jia Wan 0001, Feiping Nie 0001, Bo Liu 0006, Chenggang Yan 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2018 Memory Matching Networks for One-Shot Image Recognition
abstract
In this paper, we introduce the new ideas of augmenting Convolutional Neural Networks (CNNs) with Memory and learning to learn the network parameters for the unlabelled images on the fly in one-shot learning. Specifically, we present Memory Matching Networks (MM-Net) - a novel deep architecture that explores the training procedure, following the philosophy that training and test conditions must match. Technically, MM-Net writes the features of a set of labelled images (support set) into memory and reads from memory when performing inference to holistically leverage the knowledge in the set. Meanwhile, a Contextual Learner employs the memory slots in a sequential manner to predict the parameters of CNNs for unlabelled images. The whole architecture is trained by once showing only a few examples per class and switching the learning from minibatch to minibatch, which is tailored for one-shot learning when presented with a few examples of new categories at test time. Unlike the conventional one-shot learning approaches, our MM-Net could output one unified model irrespective of the number of shots and categories. Extensive experiments are conducted on two public datasets, i.e., Omniglot and miniImageNet, and superior results are reported when compared to state-of-the-art approaches. More remarkably, our MM-Net improves one-shot accuracy on Omniglot from 98.95% to 99.28% and from 49.21% to 53.37% on miniImageNet.
Yingwei Pan, Ting Yao 0003, Chenggang Yan 0001, Tao Mei 0001
CVPR4
2018 Watching a Small Portion could be as Good as Watching All: Towards Efficient Video Classification
abstract
We aim to significantly reduce the computational cost for classification of temporally untrimmed videos while retaining similar accuracy. Existing video classification methods sample frames with a predefined frequency over entire video. Differently, we propose an end-to-end deep reinforcement approach which enables an agent to classify videos by watching a very small portion of frames like what we do. We make two main contributions. First, information is not equally distributed in video frames along time. An agent needs to watch more carefully when a clip is informative and skip the frames if they are redundant or irrelevant. The proposed approach enables the agent to adapt sampling rate to video content and skip most of the frames without the loss of information. Second, in order to have a confident decision, the number of frames that should be watched by an agent varies greatly from one video to another. We incorporate an adaptive stop network to measure confidence score and generate timely trigger to stop the agent watching videos, which improves efficiency without loss of accuracy. Our approach reduces the computational cost significantly for the large-scale YouTube-8M dataset, while the accuracy remains the same.
Hehe Fan, Zhongwen Xu, Linchao Zhu, Chenggang Yan 0001, Jianjun Ge, Yi Yang 0001
IJCAI4
2018 Uyghur Text Localization with Fast Component Detection
Hongtao Xie 0001, Yue Hu 0002, Chenggang Yan 0001
MMM (1)4
2018 DPFMDA: Distributed and privatized framework for miRNA-Disease association prediction
Lixin Chen, Bingtao Liu, Chenggang Yan 0001
Pattern Recognit. Lett.3
2018 CNNs-Based RGB-D Saliency Detection via Cross-View Transfer and Multiview Fusion
abstract
Salient object detection from RGB-D images aims to utilize both the depth view and RGB view to automatically localize objects of human interest in the scene. Although a few earlier efforts have been devoted to the study of this paper in recent years, two major challenges still remain: 1) how to leverage the depth view effectively to model the depth-induced saliency and 2) how to implement an optimal combination of the RGB view and depth view, which can make full use of complementary information among them. To address these two challenges, this paper proposes a novel framework based on convolutional neural networks (CNNs), which transfers the structure of the RGB-based deep neural network to be applicable for depth view and fuses the deep representations of both views automatically to obtain the final saliency map. In the proposed framework, the first challenge is modeled as a cross-view transfer problem and addressed by using the task-relevant initialization and adding deep supervision in hidden layer. The second challenge is addressed by a multiview CNN fusion model through a combination layer connecting the representation layers of RGB view and depth view. Comprehensive experiments on four benchmark datasets demonstrate the significant and consistent improvements of the proposed approach over other state-of-the-art methods.
Junwei Han 0001, Hao Chen 0011, Nian Liu 0002, Chenggang Yan 0001, Xuelong Li 0001
IEEE Trans. Cybern.4
2018 AutoBD: Automated Bi-Level Description for Scalable Fine-Grained Visual Categorization
abstract
Compared with traditional image classification, fine-grained visual categorization is a more challenging task, because it targets to classify objects belonging to the same species, e.g., classify hundreds of birds or cars. In the past several years, researchers have made many achievements on this topic. However, most of them are heavily dependent on the artificial annotations, e.g., bounding boxes, part annotations, and so on. The requirement of artificial annotations largely hinders the scalability and application. Motivated to release such dependence, this paper proposes a robust and discriminative visual description named Automated Bi-level Description (AutoBD). “Bi-level” denotes two complementary part-level and object-level visual descriptions, respectively. AutoBD is “automated,” because it only requires the image-level labels of training images and does not need any annotations for testing images. Compared with the part annotations labeled by the human, the image-level labels can be easily acquired, which thus makes AutoBD suitable for large-scale visual categorization. Specifically, the part-level description is extracted by identifying the local region saliently representing the visual distinctiveness. The object-level description is extracted from object bounding boxes generated with a co-localization algorithm. Although only using the image-level labels, AutoBD outperforms the recent studies on two public benchmark, i.e., classification accuracy achieves 81.6% on CUB-200-2011 and 88.9% on Car-196, respectively. On the large-scale Birdsnap data set, AutoBD achieves the accuracy of 68%, which is currently the best performance to the best of our knowledge.
Hantao Yao, Shiliang Zhang, Chenggang Yan 0001, Yongdong Zhang 0001, Jintao Li 0001, Qi Tian 0001
IEEE Trans. Image Process.3
2018 Adaptive Residual Networks for High-Quality Image Restoration
abstract
Image restoration methods based on convolutional neural networks have shown great success in the literature. However, since most of networks are not deep enough, there is still some room for the performance improvement. On the other hand, though some models are deep and introduce shortcuts for easy training, they ignore the importance of location and scaling of different inputs within the shortcuts. As a result, existing networks can only handle one specific image restoration application. To address such problems, we propose a novel adaptive residual network (ARN) for high-quality image restoration in this paper. Our ARN is a deep residual network, which is composed of convolutional layers, parametric rectified linear unit layers, and some adaptive shortcuts. We assign different scaling parameters to different inputs of the shortcuts, where the scaling is considered as part parameters of the ARN and trained adaptively according to different applications. Due to the special construction of ARN, it can solve many image restoration problems and have superior performance. We demonstrate its capabilities with three representative applications, including Gaussian image denoising, single image super resolution, and JPEG image deblocking. Experimental results prove that our model greatly outperforms numerous state-of-the-art restoration methods in terms of both peak signal-to-noise ratio and structure similarity index metrics, e.g., it achieves 0.2-0.3 dB gain in average compared with the second best method at a wide range of situations.
Yongbing Zhang 0002, Chenggang Yan 0001, Xiangyang Ji, Qionghai Dai
IEEE Trans. Image Process.3
2018 Effective Uyghur Language Text Detection in Complex Background Images for Traffic Prompt Identification
abstract
Text detection in complex background images is a challenging task for intelligent vehicles. Actually, almost all the widely-used systems focus on commonly used languages while for some minority languages, such as the Uyghur language, text detection is paid less attention. In this paper, we propose an effective Uyghur language text detection system in complex background images. First, a new channel-enhanced maximally stable extremal regions (MSERs) algorithm is put forward to detect component candidates. Second, a two-layer filtering mechanism is designed to remove most non-character regions. Third, the remaining component regions are connected into short chains, and the short chains are extended by a novel extension algorithm to connect the missed MSERs. Finally, a two-layer chain elimination filter is proposed to prune the non-text chains. To evaluate the system, we build a new data set by various Uyghur texts with complex backgrounds. Extensive experimental comparisons show that our system is obviously effective for Uyghur language text detection in complex background images. The F-measure is 85%, which is much better than the state-of-the-art performance of 75.5%.
Chenggang Yan 0001, Hongtao Xie 0001, Jian Yin 0003, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Intell. Transp. Syst.1
2018 Supervised Hash Coding With Deep Neural Network for Environment Perception of Intelligent Vehicles
abstract
Image content analysis is an important surround perception modality of intelligent vehicles. In order to efficiently recognize the on-road environment based on image content analysis from the large-scale scene database, relevant images retrieval becomes one of the fundamental problems. To improve the efficiency of calculating similarities between images, hashing techniques have received increasing attentions. For most existing hash methods, the suboptimal binary codes are generated, as the hand-crafted feature representation is not optimally compatible with the binary codes. In this paper, a one-stage supervised deep hashing framework (SDHP) is proposed to learn high-quality binary codes. A deep convolutional neural network is implemented, and we enforce the learned codes to meet the following criterions: 1) similar images should be encoded into similar binary codes, and vice versa; 2) the quantization loss from Euclidean space to Hamming space should be minimized; and 3) the learned codes should be evenly distributed. The method is further extended into SDHP+ to improve the discriminative power of binary codes. Extensive experimental comparisons with state-of-the-art hashing algorithms are conducted on CIFAR-10 and NUS-WIDE, the MAP of SDHP reaches to 87.67% and 77.48% with 48 b, respectively, and the MAP of SDHP+ reaches to 91.16%, 81.08% with 12 b, 48 b on CIFAR-10 and NUS-WIDE, respectively. It illustrates that the proposed method can obviously improve the search accuracy.
Chenggang Yan 0001, Hongtao Xie 0001, Dongbao Yang, Jian Yin 0003, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Intell. Transp. Syst.1
2018 A Fast Uyghur Text Detector for Complex Background Images
abstract
Uyghur text localization in images with complex backgrounds is a challenging yet important task for many applications. Generally, Uyghur characters in images consist of strokes with uniform features, and they are distinct from backgrounds in color, intensity, and texture. Based on these differences, we propose a FASTroke keypoint extractor, which is fast and stroke-specific. Compared with the commonly used MSER detector, FASTroke produces less than twice the amount of components and recognizes at least 10% more characters. While the characters in a line usually have uniform features such as size, color, and stroke width, a component similarity based clustering is presented without component-level classification. It incurs no extra errors by incorporating a component-level classifier while the computing cost is drastically reduced. The experiments show that the proposed method can achieve the best performance on the UICBI-500 benchmark dataset.
Chenggang Yan 0001, Hongtao Xie 0001, Zhengjun Zha, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.1
2018 Unsupervised Person Re-identification: Clustering and Fine-tuning
abstract
The superiority of deeply learned pedestrian representations has been reported in very recent literature of person re-identification (re-ID). In this article, we consider the more pragmatic issue of learning a deep feature with no or only a few labels. We propose a progressive unsupervised learning (PUL) method to transfer pretrained deep representations to unseen domains. Our method is easy to implement and can be viewed as an effective baseline for unsupervised re-ID feature learning. Specifically, PUL iterates between (1) pedestrian clustering and (2) fine-tuning of the convolutional neural network (CNN) to improve the initialization model trained on the irrelevant labeled dataset. Since the clustering results can be very noisy, we add a selection operation between the clustering and fine-tuning. At the beginning, when the model is weak, CNN is fine-tuned on a small amount of reliable examples that locate near to cluster centroids in the feature space. As the model becomes stronger, in subsequent iterations, more images are being adaptively selected as CNN training samples. Progressively, pedestrian clustering and the CNN model are improved simultaneously until algorithm convergence. This process is naturally formulated as self-paced learning. We then point out promising directions that may lead to further improvement. Extensive experiments on three large-scale re-ID datasets demonstrate that PUL outputs discriminative features that improve the re-ID accuracy. Our code has been released at https://github.com/hehefan/Unsupervised-Person-Re-identification-Clustering-and-Fine-tuning.
Hehe Fan, Liang Zheng 0001, Chenggang Yan 0001, Yi Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2017 Video Description with Spatial-Temporal Attention
abstract
Temporal attention has been widely used in video description to adaptively focus on important frames. However, most existing methods based on temporal attention suffer from the problems of recognition error and detail missing, because only coarse frame-level global features are employed. Inspired by recent successful work in image description using spatial attention, we propose a spatial-temporal attention (STAT) method to address such problems. In particular, first, we take advantage of object-level local features to address the problem of detail missing. Second, the STAT method further selects relevant local features by spatial attention and then attend to important frames by temporal attention to recognize related semantics. The proposed two-stage attention mechanism can recognize the salient objects more precisely with high recall and automatically focus on the most relevant spatial-temporal segments given the sentence context. Extensive experiments on two well-known benchmarks suggest that STAT method outperforms the state-of-the-art methods on MSVD with BLEU4 score 0.511, and achieves superior BLEU4 score 0.374 on MSR-VTT-10K. Compared to the method without local features, the relative improvements derived from our STAT method are 10.1% and 0.8% respectively on two benchmarks. Compared to the method using only temporal attention, the relative improvements derived from our STAT method are 18.3% and 9.0% respectively on two benchmarks.
Yunbin Tu, Xishan Zhang, Bingtao Liu, Chenggang Yan 0001
ACM Multimedia4
2017 Long non-coding RNAs and complex diseases: from experimental results to computational models
abstract
LncRNAs have attracted lots of attentions from researchers worldwide in recent decades. With the rapid advances in both experimental technology and computational prediction algorithm, thousands of lncRNA have been identified in eukaryotic organisms ranging from nematodes to humans in the past few years. More and more research evidences have indicated that lncRNAs are involved in almost the whole life cycle of cells through different mechanisms and play important roles in many critical biological processes. Therefore, it is not surprising that the mutations and dysregulations of lncRNAs would contribute to the development of various human complex diseases. In this review, we first made a brief introduction about the functions of lncRNAs, five important lncRNA-related diseases, five critical disease-related lncRNAs and some important publicly available lncRNA-related databases about sequence, expression, function, etc. Nowadays, only a limited number of lncRNAs have been experimentally reported to be related to human diseases. Therefore, analyzing available lncRNA-disease associations and predicting potential human lncRNA-disease associations have become important tasks of bioinformatics, which would benefit human complex diseases mechanism understanding at lncRNA level, disease biomarker detection and disease diagnosis, treatment, prognosis and prevention. Furthermore, we introduced some state-of-the-art computational models, which could be effectively used to identify disease-related lncRNAs on a large scale and select the most promising disease-related lncRNAs for experimental validation. We also analyzed the limitations of these models and discussed the future directions of developing computational models for lncRNA research.
Xing Chen 0001, Chenggang Yan 0001, Xu Zhang 0028, Zhu-Hong You
Briefings Bioinform.2
2017 Deep learning based basketball video analysis for intelligent arena application
Wu Liu 0005, Chenggang Yan 0001, Jiangyu Liu, Huadong Ma
Multim. Tools Appl.2
2017 A survey of memory deduplication approaches for intelligent urban computing
Chenggang Yan 0001, Bingtao Liu, Licheng Chen
Mach. Vis. Appl.2
2017 Three-dimensional laser scanning under the pinhole camera with lens distortion
Binbin Lv, Liang Li 0003, Chenggang Yan 0001
Mach. Vis. Appl.3
2017 Triple-Bit Quantization with Asymmetric Distance for Image Content Security
Hongtao Xie 0001, Chenggang Yan 0001
Mach. Vis. Appl.3
2016 Drug-target interaction prediction: databases, web servers and computational models
abstract
Identification of drug-target interactions is an important process in drug discovery. Although high-throughput screening and other biological assays are becoming available, experimental methods for drug-target interaction identification remain to be extremely costly, time-consuming and challenging even nowadays. Therefore, various computational models have been developed to predict potential drug-target associations on a large scale. In this review, databases and web servers involved in drug-target identification and drug discovery are summarized. In addition, we mainly introduced some state-of-the-art computational models for drug-target interactions prediction, including network-based method, machine learning-based method and so on. Specially, for the machine learning-based method, much attention was paid to supervised and semi-supervised models, which have essential difference in the adoption of negative samples. Although significant improvements for drug-target interaction prediction have been obtained by many effective computational models, both network-based and machine learning-based methods have their disadvantages, respectively. Furthermore, we discuss the future directions of the network-based drug discovery and network approach for personalized drug discovery based on personalized medicine, genome sequencing, tumor clone-based network and cancer hallmark-based network. Finally, we discussed the new evaluation validation framework and the formulation of drug-target interactions prediction problem by more realistic regression formulation based on quantitative bioactivity data.
Xing Chen 0001, Chenggang Yan 0001, Xu Zhang 0028, Jian Yin 0003, Yongdong Zhang 0001
Briefings Bioinform.2
2016 Distributed image understanding with semantic dictionary and semantic expansion
Liang Li 0003, Chenggang Yan 0001, Xing Chen 0001, Chunjie Zhang 0001, Jian Yin 0003, Baochen Jiang, Qingming Huang
Neurocomputing2
2015 Unobtrusive Sensing Incremental Social Contexts Using Fuzzy Class Incremental Learning
abstract
By utilizing captured characteristics of surrounding contexts through widely used Bluetooth sensor, user-centric social contexts can be effectively sensed and discovered by dynamic Bluetooth information. At present, state-of-the-art approaches for building classifiers can basically recognize limited classes trained in the learning phase; however, due to the complex diversity of social contextual behavior, the built classifier seldom deals with newly appeared contexts, which results in degrading the recognition performance greatly. To address this problem, we propose, an OSELM (online sequential extreme learning machine) based class incremental learning method for continuous and unobtrusive sensing new classes of social contexts from dynamic Bluetooth data alone. We integrate fuzzy clustering technique and OSELM to discover and recognize social contextual behaviors by real-world Bluetooth sensor data. Experimental results show that our method can automatically cope with incremental classes of social contexts that appear unpredictably in the real-world. Further, our proposed method have the effective recognition capability for both original known classes and newly appeared unknown classes, respectively.
Zhenyu Chen 0003, Yiqiang Chen 0001, Xingyu Gao 0001, Shuangquan Wang, Lisha Hu, Chenggang Yan 0001, Nicholas D. Lane, Chunyan Miao
ICDM6
2015 Fast approximate matching of binary codes with distinctive bits
Chenggang Yan 0001, Hongtao Xie 0001, Yanping Ma, Qiong Dai
Frontiers Comput. Sci.1
2015 Memory bandwidth optimization of SpMV on GPGPUs
Chenggang Yan 0001, Hui Yu 0010, Weizhi Xu 0001, Yingping Zhang, Bochuan Chen, Zhu Tian, Jian Yin 0003
Frontiers Comput. Sci.1
2015 Robust skin detection in real-world images
Lei Huang 0010, Zhiqiang Wei 0002, Bo-Wei Chen, Chenggang Yan 0001, Jie Nie, Jian Yin 0003, Baochen Jiang
J. Vis. Commun. Image Represent.5
2015 LSH-based semantic dictionary learning for large scale image understanding
Liang Li 0003, Chenggang Yan 0001, Bo-Wei Chen, Shuqiang Jiang, Qingming Huang
J. Vis. Commun. Image Represent.2
2014 Finding suits in images of people in unconstrained environments
Chenggang Yan 0001, Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie, Bochuan Chen, Yingping Zhang
J. Vis. Commun. Image Represent.1
2014 Fusing multi-cues description for partial-duplicate image retrieval
Chenggang Yan 0001, Liang Li 0003, Jian Yin 0003, Hailong Shi, Shuqiang Jiang, Qingming Huang
J. Vis. Commun. Image Represent.1
2014 Extracting salient region for pornographic image detection
Chenggang Yan 0001, Hongtao Xie 0001, Zhuhua Liao, Jian Yin 0003
J. Vis. Commun. Image Represent.1
2014 A Highly Parallel Framework for HEVC Coding Unit Partitioning Tree Decision on Many-core Processors
abstract
High Efficiency Video Coding (HEVC) uses a very flexible tree structure to organize coding units, which leads to a superior coding efficiency compared with previous video coding standards. However, such a flexible coding unit tree structure also places a great challenge for encoders. In order to fully exploit the coding efficiency brought by this structure, huge amount of computational complexity is needed for an encoder to decide the optimal coding unit tree for each image block. One way to achieve this is to use parallel computing enabled by many-core processors. In this paper, we analyze the challenge to use many-core processors to make coding unit tree decision. Through in-depth understanding of the dependency among different coding units, we propose a parallel framework to decide coding unit trees. Experimental results show that, on the Tile64 platform, our proposed method achieves averagely more than 11 and 16 times speedup for 1920x1080 and 2560x1600 video sequences, respectively, without any coding efficiency degradation.
Chenggang Yan 0001, Yongdong Zhang 0001, Jizheng Xu, Liang Li 0003, Qionghai Dai, Feng Wu 0001
IEEE Signal Process. Lett.1
2014 Efficient Parallel Framework for HEVC Motion Estimation on Many-Core Processors
abstract
High Efficiency Video Coding (HEVC) provides superior coding efficiency than previous video coding standards at the cost of increasing encoding complexity. The complexity increase of motion estimation (ME) procedure is rather significant, especially when considering the complicated partitioning structure of HEVC. To fully exploit the coding efficiency brought by HEVC requires a huge amount of computations. In this paper, we analyze the ME structure in HEVC and propose a parallel framework to decouple ME for different partitions on many-core processors. Based on local parallel method (LPM), we first use the directed acyclic graph (DAG)-based order to parallelize coding tree units (CTUs) and adopt improved LPM (ILPM) within each CTU (DAGILPM), which exploits the CTU-level and prediction unit (PU)-level parallelism. Then, we find that there exist completely independent PUs (CIPUs) and partially independent PUs (PIPUs). When the degree of parallelism (DP) is smaller than the maximum DP of DAGILPM, we process the CIPUs and PIPUs, which further increases the DP. The data dependencies and coding efficiency stay the same as LPM. Experiments show that on a 64-core system, compared with serial execution, our proposed scheme achieves more than 30 and 40 times speedup for 1920 × 1080 and 2560 × 1600 video sequences, respectively.
Chenggang Yan 0001, Yongdong Zhang 0001, Jizheng Xu, Jun Zhang 0007, Qionghai Dai, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 Highly Parallel Framework for HEVC Motion Estimation on Many-Core Platform
abstract
As the next generation standard of video coding, High Efficiency Video Coding (HEVC) is expected to be more complex than H.264/AVC. Many-core platforms are good candidates for speeding up HEVC in the case that HEVC can provide sufficient parallelism. The local parallel method (LPM) is the most promising parallel proposal for HEVC motion estimation (ME), but it can't provide sufficient parallelism for many-core platforms. On the premise of keeping the data dependencies and coding efficiency the same as the LPM, we propose a highly parallel framework to exploit the implicit parallelism. Compared with the well-known LPM, experiments conducted on a 64-core system show that our proposed method achieves averagely more than 10 and 13 times speedup for 1920×1080 and 2560×1600 video sequences, respectively.
Chenggang Yan 0001, Yongdong Zhang 0001, Liang Li 0003
DCC1
2013 Efficient Parallel Framework for HEVC Deblocking Filter on Many-Core Platform
abstract
Summary form only given. Many-core platforms are good candidates for speeding up High Efficiency Video Coding (HEVC) in the case that HEVC can provide sufficient parallelism. As the most promising proposal for parallelizing HEVC deblocking filter (DF), the order-changed parallel method (OCPM) changes the order of filtering and incurs considerable loss in coding efficiency. Meanwhile, the parallelism of OCPM still has some room for improvement. In this paper, we propose an efficient parallel framework for HEVC DF, which exploits the implicit parallelism and keeps the filtering order of DF unchanged. Compared with the well-known OCPM, experiments conducted on a 64-core system show that our proposed method saves averagely 37.18% and 37.93% DF time with different quantization parameters (QPs). Meanwhile, our proposed method improves coding efficiency, which achieves an average BD-rate reduction of 0.09%, 0.11% and 0.12% for Y, U and V components, respectively.
Chenggang Yan 0001, Yongdong Zhang 0001, Liang Li 0003
DCC1
2013 Efficient HEVC to H.264/AVC Transcoding with Fast Intra Mode Decision
Jun Zhang 0007, Yongdong Zhang 0001, Chenggang Yan 0001
MMM (1)4
2012 Efficient Parallel Framework for H.264/AVC Deblocking Filter on Many-Core Platform
abstract
The H.264/AVC deblocking filter is becoming the performance bottleneck of H.264/AVC parallelization on many-core platform. Efficient parallelization of the deblocking filter on a many-core platform is challenging, because the deblocking filter has complicated data dependencies, which provide insufficient parallelism for so many cores. Furthermore, parallelization may have significant synchronization and load imbalance overhead. At present, research on the parallelizing deblocking filter on a many-core platform is rare and focuses on data-level parallelization. In this paper, we propose a three-step framework considering task-level segmentation and data-level parallelization to efficiently parallelize the deblocking filter. First, we review the entire deblocking filter process in 4 × 4 block edge-level and divide it into two parts: 1) boundary strength computation (BSC) and 2) edge discrimination and filtering (EDF), which increases the parallelism. Then, we apply the Markov empirical transition probability matrix and Huffman tree (METPMHT) to the BSC, which alleviate the load imbalance problem. Finally, we use an independent pixel connected area parallelization (IPCAP) for the EDF, which increases the parallelism and reduces the synchronization. In experiments, we apply our parallel method to the deblocking filter of the H.264/AVC reference software JM15.1 on the Tile64 platform without any Tile64 platform-based optimizations. Compared to the well-known 2D-wavefront method, the proposed method achieves on average 14.85, 17.83, and 10.60 times speed-up for QCIF, CIF, and HD videos using 62 cores, respectively.
Yongdong Zhang 0001, Chenggang Yan 0001, Yike Ma
IEEE Trans. Multim.2
2011 Parallel deblocking filter for H.264/AVC implemented on Tile64 platform
abstract
For the purpose of accelerating deblocking filter, which accounts for a significant percentage of H.264/AVC decoding time, some researchers use multi-core platforms to achieve the required performance. We study the problem under the context of many-core systems. Parallelization of deblocking filter on many-core platform is challenging not only because deblocking filter has complicated data dependencies which provides insufficient parallelism for so many cores but also because parallelization may have significant synchronization overhead. We present a new method to exploit the implicit parallelism and reduce the synchronization overhead. We apply our implementation to the deblocking filter of the H.264/AVC reference software JM15.1 on Tile64 platform. The proposed method achieves up to 817%, 604% and 532% speedup for CIF, SD and HD videos compared to the well-known wavefront method using 62 cores, respectively.
Chenggang Yan 0001, Yongdong Zhang 0001, Yike Ma, Licheng Chen, Lingjun Fan, Yasong Zheng
ICME1
2011 Parallel Deblocking Filter for H.264/AVC on the TILERA Many-Core Systems
Chenggang Yan 0001, Yongdong Zhang 0001
MMM (1)1