VLDB 2026 Research / reviewers in the wild / expert
Chengdong Wu 0001
dblp:26/1959
· DBLP profile ↗
63ranked-venue papers
1as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 21 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 8 since 2021Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BDTSNet: A novel bidirectional two-stream network for video-based human action recognition
Chuanjiang Leng, Chengdong Wu 0001, Ange Chen, Hexiao Li, Hao Wu 0064 |
Signal Process. Image Commun. | 2 |
| 2026 | SeaAnchor-GS: Identity-Anchored Gaussian Splatting for High-Fidelity Dynamic Underwater Scene ReconstructionabstractDynamic underwater novel-view synthesis remains challenging because refraction, scattering, and particle interference undermine correspondence reliability, geometric consistency, and appearance stability. We propose SeaAnchor-GS, a robust dynamic 3D Gaussian splatting framework for underwater scene reconstruction. Each Gaussian is associated with a persistent identity embedding and deformed via identity–time conditioning, which improves deformation estimation under unstable underwater observations. A dual-branch residual dynamics module captures both dominant motion and fine-scale variations, while confidence-guided sampling, progressive deformation activation, and neighborhood-consistency regularization enhance optimization robustness and model compactness. Extensive experiments on dynamic underwater benchmarks show consistent improvements in reconstruction fidelity and perceptual quality, with a favorable balance between quality, efficiency, and representation compactness. Additional results on a static underwater benchmark suggest that the proposed representation remains competitive beyond the dynamic setting. Yaoming Zhuang, Tongrui Liu, Yifan Chao, Hao Wu 0064, Chengdong Wu 0001, Zhanlin Liu |
IEEE Signal Process. Lett. | 6 |
| 2026 | Deep Fourier-Embedded Network for RGB and Thermal Salient Object DetectionabstractThe rapid development of deep learning has significantly improved salient object detection (SOD) combining both RGB and thermal (RGB-T) images. However, existing Transformer-based RGB-T SOD models with quadratic complexity are memory-intensive, limiting their application in high-resolution bimodal feature fusion. To overcome this limitation, we propose a purely Fourier Transform-based model, namely Deep Fourier-embedded Network (FreqSal), for accurate RGB-T SOD. Specifically, we leverage the efficiency of Fast Fourier Transform with linear complexity to design three key components: (1) To fuse RGB and thermal modalities, we propose Modal-coordinated Perception Attention, which aligns and enhances bimodal Fourier representation in multiple dimensions; (2) To clarify object edges and suppress noise, we design Frequency-decomposed Edge-aware Block, which deeply decomposes and filters Fourier components of low-level features; (3) To accurately decode features, we propose Fourier Residual Channel Attention Block, which prioritizes high-frequency information while aligning channel-wise global relationships. Additionally, even when converged, existing deep learning-based SOD models’ predictions still exhibit frequency gaps relative to ground-truth. To address this problem, we propose Co-focus Frequency Loss, which dynamically weights hard frequencies during edge frequency reconstruction by cross-referencing bimodal edge information in the Fourier domain. Extensive experiments on ten bimodal SOD benchmark datasets demonstrate that FreqSal outperforms twenty-nine existing state-of-the-art bimodal SOD models. Comprehensive ablation studies further validate the value and effectiveness of our newly proposed components. The code is available at https://github.com/JoshuaLPF/FreqSal. Pengfei Lyu, Xiaosheng Yu 0001, Pak-Hei Yeung, Chengdong Wu 0001, Jagath C. Rajapakse |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | EISegNet: Enhancing Instrument Segmentation Network via Dual-View Disparity EstimationabstractAccurate segmentation of endoscopic instruments is essential in robot-assisted surgery, supporting precis enavigation, enhancing safety, and advancing surgical automation. However, this task is challenging due to factors like complex environments, instrument-tissue similarity, and lighting variations. Instruments, due to their material properties, have distinct depth distributions compared to surrounding tissues. This aspect is often overlooked in monocular video segmentation methods.To address this issue, we propose EISegNet, a multi-task framework that prioritizes instrument segmentation with an auxiliary disparity estimation task. The framework integrates an asymmetric cross-attention mechanism to enhance segmentation performance by fusing features from both tasks. Moreover, by leveraging the geometric properties of motion, EISegNet adapts the stereo disparity estimation strategy for dual-view depth estimation, broadening its applicability to various endoscopic surgeries beyond laparoscopic procedures. Furthermore, EISegNet incorporates a Gaussian-weighted loss function to emphasize edge features, which are particularly challenging for disparity estimation. This function reduces overall loss and improves segmentation accuracy. Extensive cross-dataset experiments demonstrate the superior accuracy and generalization of our method, achieving a 5.97% increase in IoU (Intersection over Union). Qualitative evaluations on clinical datasets further demonstrate the promising performance in real-world scenarios. Yongming Yang, Zhaoshuo Diao, Ziliang Song, Shenglin Zhang, Tiancong Liu, Chengdong Wu 0001, Weiliang Bai, Hao Liu 0008 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Structure-semantic co-alignment network for RGB-T salient object detection
Chuanyue Bai, Lingying Shen, Chengdong Wu 0001 |
Vis. Comput. | 4 |
| 2025 | PONet: Prototype optimization network for few-shot medical image segmentation
Xiaosheng Yu 0001, Jianning Chi, Chengdong Wu 0001, Xiujing Gao |
Neurocomputing | 4 |
| 2025 | Conditional variational underwater image enhancement with kernel decomposition and adaptive hybrid normalization
Haopeng Zhang 0018, Hongli Xu 0003, Hao Liu 0008, Xiaosheng Yu 0001, Xiangyue Zhang, Chengdong Wu 0001 |
Neurocomputing | 6 |
| 2025 | Efficient Fourier Filtering Network With Contrastive Learning for AAV-Based Unaligned Bimodal Salient Object DetectionabstractUnmanned aerial vehicle (UAV)-based bi-modal salient object detection (BSOD) aims to segment salient objects in a scene utilizing complementary cues in unaligned RGB and thermal image pairs. However, the high computational expense of existing UAV-based BSOD models limits their applicability to real-world UAV devices. To address this problem, we propose an efficient Fourier filter network with contrastive learning that achieves both real-time and accurate performance. Specifically, we first design a semantic contrastive alignment loss to align the two modalities at the semantic level, which facilitates mutual refinement in a parameter-free way. Second, inspired by the fast Fourier transform that obtains global relevance in linear complexity, we propose synchronized alignment fusion, which aligns and fuses bi-modal features in the channel and spatial dimensions by a hierarchical filtering mechanism. Our proposed model, AlignSal, reduces the number of parameters by 70.0%, decreases the floating point operations by 49.4%, and increases the inference speed by 152.5% compared to the cutting-edge BSOD model (i.e., MROS). Extensive experiments on the UAV RGB-T 2400 and seven bi-modal dense prediction datasets demonstrate that AlignSal achieves both real-time inference speed and better performance and generalizability compared to nineteen state-of-the-art models across most evaluation metrics. In addition, our ablation studies further verify AlignSal’s potential in boosting the performance of existing aligned BSOD models on UAV-based unaligned data. The code is available at: https://github.com/JoshuaLPF/AlignSal. Pengfei Lyu, Pak-Hei Yeung, Xiaosheng Yu 0001, Xiufei Cheng, Chengdong Wu 0001, Jagath C. Rajapakse |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | CDF-UIE: Leveraging Cross-Domain Fusion for Underwater Image EnhancementabstractUnderwater image enhancement (UIE) aims to restore image quality by mitigating inherent degradations in underwater imaging systems. While existing learning-based methods show promise, they face limitations in separating and processing frequency components, effectively fusing domain information, and balancing the enhancement of structures and details. To resolve these limitations, we propose cross-domain fusion (CDF)-UIE, a novel network that leverages and fuses cross-domain information for mitigating the degradation in underwater images. CDF-UIE first performs domain decoupling of input features using the proposed spatial-frequency decoupling (SFD) block. Then, we design an innovative CDF block, which effectively bridges the spatial- and frequency-domain features through the cross-domain attention mechanism. To produce stable and detailed enhanced outputs, we exploit the coarse and fine-scale information in the image reconstruction stage. In addition, we introduce a multiscale objective function that incorporates pixel-level, structural, and perceptual constraints to guide the enhancement process. We conduct extensive experiments on six diverse real-world underwater image datasets. Comprehensive experiments and real-world application tests demonstrate that CDF-UIE significantly outperforms existing methods, offering promising future applications in various underwater scenarios. The source code is available athttps://github.com/hpzhan66/CDF-UIE. Haopeng Zhang 0018, Hongli Xu 0003, Xiaosheng Yu 0001, Xiangyue Zhang, Xiujing Gao, Chengdong Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | TwinsTNet: Broad-View Twins Transformer Network for Bi-Modal Salient Object DetectionabstractExploring complementary information between RGB and thermal/depth modalities is crucial for bi-modal salient object detection (BSOD). However, the distinct characteristics of different modalities often lead to large differences in information distributions. Existing models, which rely on convolutional operations or plug-and-play attention mechanisms, struggle to address this issue. To overcome this challenge, we rethink the relationship between information complementarity and long-range relevance, and propose a uniform broad-view Twins Transformer Network (TwinsTNet) for accurate BSOD. Specifically, to efficiently fuse bi-modal information, we first design the Cross-Modal Federated Attention (CMFA), which mines complementary cues across modalities through element-wise global dependency. Second, to ensure accurate modality fusion, we propose the Semantic Consistency Attention Loss, which supervises the co-attention feature in CMFA using the ground-truth-generated attention map. Additionally, existing BSOD models lack the exploration of inter-layer interactions, for which we propose the Cross-Scale Retracing Attention (CSRA), which retrieves query-relevant information from stacked features of all previous layers, enabling flexible cross-layer interactions. The cooperation between CMFA and CSRA mitigates inductive bias in both modality and layer dimensions, enhancing TwinsTNet's representational capability. Extensive experiments demonstrate that TwinsTNet outperforms twenty-two existing state-of-the-art models on ten BSOD benchmark datasets. The code is available at: https://github.com/JoshuaLPF/TwinsTNet. Pengfei Lyu, Xiaosheng Yu 0001, Jianning Chi, Hao Wu 0064, Chengdong Wu 0001, Jagath C. Rajapakse |
IEEE Trans. Image Process. | 5 |
| 2025 | CMT-6D: a lightweight iterative 6DoF pose estimation network based on cross-modal Transformer
Suyi Liu, Chengdong Wu 0001, Jianning Chi, Xiaosheng Yu 0001, Longxing Wei, Chuanjiang Leng |
Vis. Comput. | 3 |
| 2024 | Geometry-aided Underwater 3D Mapping Using Side-scan SonarabstractIn recent years, the interest in underwater exploration with Autonomous Underwater Vehicles (AUVs) equipped with side-scan sonars (SSS) has grown considerably. However, state-of-the-art SSS Simultaneous Localization and Mapping (SLAM) systems encounter challenges in data association across large viewpoint changes. Additionally, these systems assume that the seabed is a flat surface, leading to significant mapping error in uneven underwater terrains. To address these challenges, we propose a framework that leverages the side-scan sonar geometry to facilitate data association and improve mapping accuracy. The framework begins with a preprocessing module that extracts feature points and provides initial estimates of the elevation angles of the landmarks. Then, a non-consecutive data association module applies epipolar line search to establish correspondences between the current and historical frames. Finally, the mapping module uses side-scan sonar bundle adjustment to recover the positions of the landmarks. The proposed method is evaluated using an underwater terraced fields dataset. Our method achieves over 90% matching rate and reduces the average mapping error from 3.799 to 0.134. Yiqiao Yang, Chenglin Pang, Chengdong Wu 0001, Zheng Fang 0001 |
IROS | 3 |
| 2024 | A Gradient Vector Self-Learning Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) is challenging due to low target-background contrast and small target size, leading to missed detections and false alarms (Fas). To address these problems, a novel gradient vector self-learning network (GVSLNet) is proposed. First, the self-learning gradient vector (SLGV) module is designed based on the unique high correlation of infrared small target gradient. Traditional gradient vector field operators cannot update and learn the deep features. SLGV overcomes these limitations by using CNN to adaptively learn and calculate the gradient vector of infrared images, improving the ability to distinguish between targets and backgrounds under complex environments. Then, edge features are encoded into the global-local attention fusion (GLAF) module, which is based on Transformer and dilated convolution. Infrared small target images often display nonlocal self-similarity, where background signals tend to share similar structures. The GLAF module leverages this characteristic to further enhance target intensity while effectively suppressing background noise. The proposed GVSLNet can dynamically calculate gradients based on scene information, improving the detection ability of small targets in complex environments. Experimental results prove that the proposed GVSLNet outperforms state-of-the-art methods on public datasets while maintaining high inference speed. Xiangyue Zhang, Xinhao Zheng, Chengdong Wu 0001, Jingyu Ru |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Generative facial prior embedded degradation adaption network for heterogeneous face hallucination
Jianning Chi, Chengdong Wu 0001, Hao Wu 0064 |
Multim. Tools Appl. | 4 |
| 2024 | Progressive local-to-global vision transformer for occluded face hallucination
Jianning Chi, Chengdong Wu 0001, Xiaosheng Yu 0001, Hao Wu 0064 |
Multim. Tools Appl. | 3 |
| 2024 | BDNet: a method based on forward and backward convolutional networks for action recognition in videos
Chuanjiang Leng, Qichuan Ding, Chengdong Wu 0001, Ange Chen, Hao Wu 0064 |
Vis. Comput. | 3 |
| 2023 | Low-Dose CT Image Super-Resolution Network with Dual-Guidance Feature Distillation and Dual-Path Content Communication
Jianning Chi, Zhiyi Sun, Tianli Zhao, Xiaosheng Yu 0001, Chengdong Wu 0001 |
MICCAI (10) | 6 |
| 2023 | Automatic video clip and mixing based on semantic sentence matching
Zixi Jia, Zhengjun Du, Jingyu Ru, Chengdong Wu 0001, Shuangjiang Yu, Changsheng Sun, Ao Lyu |
Appl. Intell. | 6 |
| 2023 | Neural network equivalent model for highly efficient massive data classification
Siquan Yu, Zhi Han, Yandong Tang, Chengdong Wu 0001 |
Sci. China Inf. Sci. | 4 |
| 2023 | A geometry-aware deep network for depth estimation in monocular endoscopy
Yongming Yang, Shuwei Shao, Chengdong Wu 0001, Hao Liu 0008 |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Cross-view information interaction and feedback network for face hallucination
Jianning Chi, Chengdong Wu 0001, Xiaosheng Yu 0001, Hao Wu 0064 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Cross-modal co-feedback cellular automata for RGB-T saliency detection
Hao Wu 0064, Chengdong Wu 0001 |
Pattern Recognit. | 3 |
| 2023 | Infrared Small Target Detection Based on Gradient Correlation Filtering and Contrast MeasurementabstractInfrared small target detection under complex backgrounds, especially in dense cloud and changeable clutter scenes, has always been a challenging research task. In order to improve the detection ability of small targets under complex backgrounds, an infrared small target detection method based on gradient correlation filtering and gradient contrast measurement (GCF-CM) is proposed in this article. The infrared gradient vector field (IGVF) of the original image is first constructed through the facet model. Then, considering the unique gradient characteristics of small targets, a gradient correlation filtering (GCF) method is proposed to filter small targets and background clutters. Meanwhile, a gradient contrast measurement (GCM) method is designed to further enhance the intensity of the small target. Finally, after fusing the two response maps, an adaptive threshold is adopted to extract small targets. Experimental results demonstrate that the proposed method can improve the intensity of the small target and suppress clutter sufficiently. In comparison with other excellent methods, the proposed method exhibits a robust detection performance. Xiangyue Zhang, Jingyu Ru, Chengdong Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Unsupervised Multi-Subclass Saliency Classification for Salient Object DetectionabstractNumerous bottom-up salient object detection algorithms formulate the problem as a classification task. For an input image, these methods usually utilize prior cues to select some regions as training set, and learn a classifier to classify all regions into foreground/background. However, such binary classification based approaches suffer from accuracy problems in some complex scenes. To this end, we propose a novel framework, namely Multi-Subclass Classification with Label Distribution Learning (MSCLDL). Specifically, prior knowledge is firstly employed to build a training set from input image, in which each sample is associated with one of two class labels. Previous works usually learn directly a binary classification model from training set. Different with them, we further decompose two classes into a certain number of subclasses, each sample is thus described by one of multiple subclass labels. Based on the multi-subclass training set, we learn a label distribution model to predict the subclass label of each image region. Furthermore, the saliency value of each image region could be computed via exploring the relationship class and subclass labels. The MSCLDL could overcome the limitation of existing classification-based algorithms in some challenging scenes. Finally, a novel refinement technology is presented to further refine the saliency map obtained by MSCLDL. We compare the proposed method and other state-of-the-art methods on four benchmark datasets, the superiority of our model is adequately demonstrated via the experimental results analysis. Chengdong Wu 0001, Hao Wu 0064, Xiaosheng Yu 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Over-sampling strategy-based class-imbalanced salient object detection and its application in underwater scene
Chengdong Wu 0001, Hao Wu 0064, Xiaosheng Yu 0001 |
Vis. Comput. | 2 |
| 2022 | Facial Landmarks and Generative Priors Guided Blind Face RestorationabstractBlind face restoration (BFR) from severely degraded face images is important in face image processing, and has attracted increasing attention due to its wide applications. How-ever, due to the complex unknown degradations in real-world scenarios, existing priors-based methods tend to restore faces with unstable quality. In this paper, we propose a Facial Landmarks and Generative Priors Guided Blind Face Restoration Network (FGPNet) to seamlessly integrate the advantages of generative priors and face-specific geometry priors. Specifically, we pretrain a high-quality (HQ) face synthesis generative adversarial network (GAN) and a landmarks prediction network, and then embed them into a U-shaped deep neural network (DNN) as decoder priors to guide face restoration, during which the generative priors can provide adequate details and the landmarks priors provide geometry and semantic information. Furthermore, we design facial priors fusion (FPF) blocks to incorporate the prior features from pretrained face synthesis GAN and landmarks prediction network in an adaptive and progressive manner, making our FGPNet exhibits good generalization in real-world application. Experiments demonstrate the superiority of our FGPNet in comparison to state-of-the-arts, and also show its potential in handling real-world low-quality images from several practical applications. Zi Teng, Chengdong Wu 0001, Sonya A. Coleman |
INDIN | 3 |
| 2022 | An Infrared Small Target Detection Method Based on Gradient Correlation MeasureabstractTo overcome the interference of complex background and improve the detection ability of infrared small target under low signal-to-clutter ratio (SCR) scenes, a novel detection method based on gradient correlation measure (GCM) is proposed in this letter. Initially, the infrared gradient vector field (IGVF) of the original image is constructed based on the facet model. Then, a gradient correlation template is designed to distinguish the difference of local gradient between small targets and background. Finally, an adaptive threshold is adopted to extract small targets from background clutter. The proposed GCM method can identify the unique gradient characteristics of small targets. Experimental evaluations prove that the proposed method can achieve higher SCR scores in complex backgrounds. Especially in the scene where the gray contrast of small targets is low, the proposed GCM method shows a more robust detection performance. Xiangyue Zhang, Jingyu Ru, Chengdong Wu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Optic disc detection based on fully convolutional neural network and structured matrix decomposition
Ying Wang 0152, Xiaosheng Yu 0001, Chengdong Wu 0001 |
Multim. Tools Appl. | 3 |
| 2022 | MID-UNet: Multi-input directional UNet for COVID-19 lung infection segmentation from CT images
Jianning Chi, Xiaoying Han, Chengdong Wu 0001, Xiaosheng Yu 0001 |
Signal Process. Image Commun. | 5 |
| 2021 | Brain tumor segmentation in MR images using a sparse constrained level set algorithm
Xiaoliang Lei, Xiaosheng Yu 0001, Jianning Chi, Ying Wang 0152, Jingsi Zhang, Chengdong Wu 0001 |
Expert Syst. Appl. | 6 |
| 2021 | X-Net: Multi-branch UNet-like network for liver and tumor segmentation from 3D abdominal CT scans
Jianning Chi, Xiaoying Han, Chengdong Wu 0001 |
Neurocomputing | 3 |
| 2021 | Image super-resolution using multi-granularity perception and pyramid attention networks
Chengdong Wu 0001, Jianning Chi, Xiaosheng Yu 0001, Hao Wu 0064 |
Neurocomputing | 2 |
| 2021 | Unknown hostile environment-oriented autonomous WSN deployment using a mobile robot
Sheng Feng, Haiyan Shi, Longjun Huang, Shigen Shen, Shui Yu 0001, Hua Peng, Chengdong Wu 0001 |
J. Netw. Comput. Appl. | 7 |
| 2021 | Augmented two stream network for robust action recognition adaptive to various action videos
Chuanjiang Leng, Qichuan Ding, Chengdong Wu 0001, Ange Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Underwater image super-resolution using multi-stage information distillation networks
Hao Wu 0064, Jianning Chi, Xiaosheng Yu 0001, Chengdong Wu 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2021 | DCLNet: Dual Closed-loop Networks for face super-resolution
Chengdong Wu 0001, Jianning Chi, Xiaosheng Yu 0001, Hao Wu 0064 |
Knowl. Based Syst. | 3 |
| 2021 | Efficient and robust unsupervised inverse intensity compensation for stereo image registration under radiometric changes
Chengdong Wu 0001, Daokui Qu, Haibo Sun, Jilai Song |
Signal Process. Image Commun. | 2 |
| 2021 | Accurate and Efficient Stereo Matching by Log-Angle and Pyramid-TreeabstractEfficient and fast stereo matching is a challenging task due to the presence of occlusion and low texture areas. In stereo matching, the correspondence between left and right images may be difficult owing to the lack of matching information. The cost metrics proposed before are not robust enough or are computationally expensive. In this work, we propose a novel generic tree structure, Pyramid-tree, which improves the single mode of traditional tree, and can achieve cross-regional connection between different regions with similar colors and similar depths. This unique structure can achieve cross-regional cost smoothing, significantly reducing the possibility of mismatch due to occlusion or lack of corresponding matching information, and has stronger robustness to occlusion and low texture regions. In addition, we also propose a new bearings-only cost metric, Log-angle, which is not affected by occlusion, low texture, illumination and other factors. Log-angle combines with traditional metrics can show better performance. We show that the Pyramid-tree structure and Log-angle are very important as it efficiently expands the state-of-the-art stereo matching methods and leads to significant improvements. Qualitative and quantitative experiments on Middlebury data sets verify the superior performance of the algorithm, and very effective compromise between the accuracy and computation load is achieved. Chengdong Wu 0001, Daokui Qu, Haibo Sun, Jilai Song |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Kalman Filter for Spatial-Temporal Regularized Correlation FiltersabstractWe consider visual tracking in numerous applications of computer vision and seek to achieve optimal tracking accuracy and robustness based on various evaluation criteria for applications in intelligent monitoring during disaster recovery activities. We propose a novel framework to integrate a Kalman filter (KF) with spatial-temporal regularized correlation filters (STRCF) for visual tracking to overcome the instability problem due to large-scale application variation. To solve the problem of target loss caused by sudden acceleration and steering, we present a stride length control method to limit the maximum amplitude of the output state of the framework, which provides a reasonable constraint based on the laws of motion of objects in real-world scenarios. Moreover, we analyze the attributes influencing the performance of the proposed framework in large-scale experiments. The experimental results illustrate that the proposed framework outperforms STRCF on OTB-2013, OTB-2015 and Temple-Color datasets for some specific attributes and achieves optimal visual tracking for computer vision. Compared with STRCF, our framework achieves AUC gains of 2.8%, 2%, 1.8%, 1.3%, and 2.4% for the background clutter, illumination variation, occlusion, out-of-plane rotation, and out-of-view attributes on the OTB-2015 datasets, respectively. For sporting events, our framework presents much better performance and greater robustness than its competitors. Sheng Feng, Keli Hu, En Fan, Liping Zhao 0005, Chengdong Wu 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Deep Retinex Network for Single Image DehazingabstractIn this paper, we propose a retinex-based decomposition model for a hazy image and a novel end-to-end image dehazing network. In the model, the illumination of the hazy image is decomposed into natural illumination for the haze-free image and residual illumination caused by haze. Based on this model, we design a deep retinex dehazing network (RDN) to jointly estimate the residual illumination map and the haze-free image. Our RDN consists of a multiscale residual dense network for estimating the residual illumination map and a U-Net with channel and spatial attention mechanisms for image dehazing. The multiscale residual dense network can simultaneously capture global contextual information from small-scale receptive fields and local detailed information from large-scale receptive fields to precisely estimate the residual illumination map caused by haze. In the dehazing U-Net, we apply the channel and spatial attention mechanisms in the skip connection of the U-Net to achieve a trade-off between overdehazing and underdehazing by automatically adjusting the channel-wise and pixel-wise attention weights. Compared with scattering model-based networks, fully data-driven networks, and prior-based dehazing methods, our RDN can avoid the errors associated with the simplified scattering model and provide better generalization ability with no dependence on prior information. Extensive experiments show the superiority of the RDN to various state-of-the-art methods. Pengyue Li, Jiandong Tian, Yandong Tang, Guolin Wang, Chengdong Wu 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Neural Coding Strategies for Event-Based Vision DataabstractNeural coding schemes are powerful tools used within neuroscience. This paper introduces three different neural coding scheme formations for event-based vision data which are designed to emulate the neural behaviour exhibited by neurons under stimuli. Presented are phase-of-firing and two sparse neural coding schemes. It is determined that machine learning approaches, i.e. Convolutional Neural Network combined with a Stacked Autoencoder network, produce powerful descriptors of the patterns within events. These coding schemes are deployed in an existing action recognition template and evaluated using two popular event-based data sets. Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001 |
ICASSP | 6 |
| 2020 | Post-Stimulus Time-Dependent Event DescriptorabstractEvent-based image processing is a relatively new domain in the field of computer vision. Much research has been carried out on adapting event-based data to comply with established techniques from frame-based computer vision. On the contrary, this paper presents a descriptor which is designed specifically for direct use with event-based data and therefore can be considered to be a pure event-based vision descriptor as it only uses events emitted from event-based vision devices without transforming the data to accommodate frame-based vision techniques. This novel descriptor is known as the Post-stimulus Time-dependent Event Descriptor (P-TED). P-TED is comprised of two features extracted from event data which describe motion and the underlying pattern of transmission respectively. Furthermore a framework is presented which leverages the P-TED descriptor to classify motions within event data. This framework is compared against another state-of-the-art event-based vision descriptor as well as an established frame-based approach. Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001 |
ICIP | 6 |
| 2020 | Reducing-Over-Time Tree for Event-based DataabstractThis paper presents a novel Reducing-Over-Time (ROT) binary tree structure for event-based vision data and subtypes of the tree structure. A framework is presented using ROT, that takes advantage of the self-balancing and self-pruning nature of the tree structure to extract spatial-temporal information. The ROT framework is paired with an established motion classification technique and performance is evaluated against other state-of-the-art techniques using four datasets. Additionally, the ROT framework as a processing platform is compared with other event-based vision processing platforms in terms of memory usage and is found to be one of the most memory efficient platforms available. Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001 |
ICPR | 6 |
| 2020 | Salient object detection via effective background prior and novel graph
Yunhe Wu, Chengdong Wu 0001 |
Multim. Tools Appl. | 3 |
| 2020 | Bagging-based saliency distribution learning for visual saliency detection
Xiaosheng Yu 0001, Yunhe Wu, Chengdong Wu 0001 |
Signal Process. Image Commun. | 4 |
| 2019 | Saliency detection via integrating deep learning architecture and low-level features
Jianning Chi, Chengdong Wu 0001, Xiaosheng Yu 0001, Hao Chu |
Neurocomputing | 2 |
| 2019 | Stacked dense networks for single-image snow removal
Pengyue Li, Mengshen Yun, Jiandong Tian, Yandong Tang, Guolin Wang, Chengdong Wu 0001 |
Neurocomputing | 6 |
| 2019 | Three-dimensional robot localization using cameras in wireless multimedia sensor networks
Sheng Feng, Shigen Shen, Longjun Huang, Adam C. Champion, Shui Yu 0001, Chengdong Wu 0001, Yunzhou Zhang |
J. Netw. Comput. Appl. | 6 |
| 2019 | Salient object detection based on novel graph model
Xiaosheng Yu 0001, Ying Wang 0152, Chengdong Wu 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Probabilistic coverage in directional sensor networks
Pengju Si, Chengdong Wu 0001, Yunzhou Zhang, Hao Chu, He Teng |
Wirel. Networks | 2 |
| 2018 | Fractional Order Flight Control of Quadrotor UAS: an OS4 Benchmark Environment and a Case StudyabstractThe OS4 quadrotor is a classic quadrotor simulation platform. So far, many different kinds of controllers have been designed based on its plant model. Most of the research only provided a numerical simulation to verify their designed controllers. Only a few researches have put the proposed controllers back to OS4 quadrotor to verify, but they didn't share the project folder to let others continue their work. More open-source and well-documented codes are needed to accelerate the application of fractional order controllers in industry. This paper updated the OS4 folder for the latest MATLAB version. A case of study demonstrated the workflow to design a fractional order proportional derivative controller for the simulated drone. Comparisons showed that fractional order controllers perform better in a nonlinear system like OS4 than integer order PID controllers. An impulse disturbance scenario is also used as a testbed. Project folder can be accessed from: https://ww2.mathworks.cn/matlabcentral/fileexchange/67882-os4-foc. Related videos can be found from this link: https://youtu.be/heuz4tFqf64. Bo Shang, Yunzhou Zhang, Chengdong Wu 0001, YangQuan Chen |
ICARCV | 3 |
| 2018 | Sensor-based Vital Sign Monitoring, Analysis and Visualisation for Ageing in PlaceabstractWith the ever-increasing global population and average life expectancy, care homes and care at home services are continuously being stretched beyond capacity. Recent developments in tactile sensing have enabled robot systems to measure human vital signs such as beats per minute (BPM), Respiratory Rate (RR) and Capillary Refill Time (CRT). Using robotic systems to measure vital sign data in the home of an elderly or disabled person would greatly assist medical and health services. This paper proposes the use of a vital sign measuring robotic system together with Cloud computing to intelligently process big data and ascertain the current health status of the service user without the need to expose their identity or burden health professionals. Furthermore, a method that enables medical professionals to visualise the data for a complete geographical region as well as for individual patients is presented and hence we provide details of a closed loop system to support ageing-in-place. E. P. Kerr, Sonya A. Coleman, Dermot Kerr, Philip J. Vance, Bryan Gardiner, Chengdong Wu 0001 |
IJCNN | 8 |
| 2016 | Visual saliency detection: From space to frequency
Dongyue Chen 0001, Tong Jia 0001, Chengdong Wu 0001 |
Signal Process. Image Commun. | 3 |
| 2015 | Conjugate gradient algorithm for efficient covariance tracking with Jensen-Bregman LogDet metricabstractRegion covariance descriptor that fuses multiple features compactly has proven to be very effective for visual tracking. While working effectively, the exhaustive global search strategy of covariance tracking is still inefficient, and there is much room for improvement. It may cause inconsecutive tracking trajectory and distraction. A suitable region similarity metric for covariance matching between the candidate object region and a given appearance template is of much importance. However, the computational burden of the metric, especially for large matrices under Riemannian space, may hinder its application in gradient‐based algorithms. In this study, the authors propose an algorithm which, by minimising the metric function, exploits an efficient conjugate gradient method to iteratively search the best matched candidate, and determines the search step size by non‐monotonic liner strategy. Then, an inferential reasoning in view of new efficient metric is derived for the gradient‐based algorithm. The authors test the proposed tracking method on test baseline dataset. Both quantitative and qualitative results demonstrate the effectiveness of the proposed algorithm compared with other state‐of‐the‐art methods. Qiang Guo 0003, Chengdong Wu 0001, Xiaohong Lu |
IET Comput. Vis. | 2 |
| 2014 | Active tension optimal control for WT wheelchair robot by using a novel control law for holonomic or nonholonomic systems
Xiaofan Li 0005, Chengdong Wu 0001 |
Sci. China Inf. Sci. | 5 |
| 2013 | AFM-Based Robotic Nano-Hand for Stable Manipulation at NanoscaleabstractOne of the major limitations for Atomic Force Microscopy (AFM)-based nanomanipulation is that AFM only has one sharp tip as the end-effector, and can only apply a point force to the nanoobject, which makes it extremely difficult to achieve a stable manipulation. For example, the AFM tip tends to slip-away during nanoparticle manipulation due to its small touch area, and there is no available strategy to manipulate a nanorod in a constant posture with a single tip since the applied point force can make the nanorod rotate more easily. In this paper, a robotic nano-hand method is proposed to solve these problems. The basic idea is using a single tip to mimic the manipulation effect that multi-AFM tip can achieve through the planned high speed sequential tip pushing. The theoretical behavior models of nanoparticle and nanorod are developed, based on which the moving speed and trajectory of the AFM tip are planned artfully to form a nano-hand. In this way, the slip-away problem during nanoparticle manipulation can be get rid of efficiently, and a posture constant manipulation for nanorod can be achieved. The simulation and experimental results demonstrate the effectiveness and advantages of the proposed method. Lianqing Liu, Ning Xi 0001, Yuechao Wang, Chengdong Wu 0001, Zaili Dong |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2012 | Classification of multi-channels SEMG signals using wavelet and neural networks on assistive robotabstractRecently, the robot technology research is changing from manufacturing industry to non-manufacturing industry, especially the service industry related to the human life. Assistive robot is a kind of novel service robot. It can not only help the elder and disabled people to rehabilitate their impaired musculoskeletal functions, but also help healthy people to perform tasks requiring large forces. This kind of robot has a broad application prospect in many areas, such as medical rehabilitation, special military operations, special/high intensity physical labour, space, sports, and entertainment. SEMG (Surface Electromyography) of Palmaris longus, brachioradialis, flexor carpiulnaris and biceps brachii are analysed with a wavelet transform method. The absolute variance of 3-layer wavelet coefficients is distilled and regarded as signal characteristics to compose eigenvectors. The eigenvectors are input data of a neural network classifier used to identify 5 different kinds of movement patterns including wrist flexor, wrist extensor, elbow flexion, forearm pronation and forearm rotation. Experiments verify the effectiveness of the proposed method. Yong Yue 0001, Carsten Maple, Beisheng Liu, Chengdong Wu 0001 |
INDIN | 5 |
| 2012 | A Novel Method of River Detection for High Resolution Remote Sensing Image Based on Corner Feature and SVM
Ziheng Tian, Chengdong Wu 0001, Dongyue Chen 0001, Xiaosheng Yu 0001 |
ISNN (2) | 2 |
| 2012 | A Remote Sensing Image Matching Algorithm Based on the Feature Extraction
Chengdong Wu 0001, Dongyue Chen 0001, Xiaosheng Yu 0001 |
ISNN (2) | 1 |
| 2012 | Gradient Vector Flow Based on Anisotropic Diffusion
Xiaosheng Yu 0001, Chengdong Wu 0001, Dongyue Chen 0001, Tong Jia 0001 |
ISNN (2) | 2 |
| 2010 | Shape-shifting robot path planning method based on reconfiguration performanceabstractA shape-shifting robot “AMOEBA-I” has diverse configurations, and the accessibility of the robot can be reinforced in the narrow space by changing the configurations. In this paper, a path planning method is presented corresponding to the unique reconfiguration ability of this robot. This method can automatically adjust the relation between the rapid movement and the secure mobile position of the robot based on the distribution of obstacles that are around the robot, and this method can also automatically select an appropriate configuration to adapt the robot to the environmental variation according to the information of the current conditions. The problem of deadlock is avoided by using the Boundary Following. Further, a Reconfiguration Memory is provided to optimize the trace of the Boundary Following, and it can help the robot to search a new path. Simulation results validated the advantage of the proposed method which can get the best out of the unique accessibility of the shape-shifting robot and reduced effectively the length of the robot traveling path. Tonglin Liu, Chengdong Wu 0001, Bin Li 0001 |
IROS | 2 |
| 2010 | Frequency Spectrum Modification: A New Model for Visual Saliency Detection
Dongyue Chen 0001, Chengdong Wu 0001 |
ISNN (2) | 3 |
| 2008 | Distributed energy-based multi-source localization in wireless sensor networkabstractMulti-source localization is an open and challenging research problem in the energy-based wireless sensor network (WSN) of acoustic sensors. Classic maximum likelihood (ML) algorithm can not work well due to the high computation demand. An expectation maximization (EM) algorithm was proposed in our previous paper to approximate the optimal solution with lower computational complexity. However that algorithm is centralized in nature in which all the observed data from each sensor node must be sent to the fusion centre. It has several drawbacks including the relying on the centre which may be damaged or shut down, poor scalability with the increasing of network size, and the high communication overhead requirement when a sensor is far away from the centre. In this paper we propose a distributed EM algorithm for multi-source localization in energy-based WSN. It can achieve satisfactory localization accuracy with significantly low communication and computation cost. Wei Meng 0002, Wendong Xiao, Chengdong Wu 0001 |
SMC | 3 |