Dongchen Zhu

dblp:205/7683 · DBLP profile ↗
← Back
39ranked-venue papers
2as first author
37since 2021 · last 2026
0000-0002-1579-3942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 18 since 2021Systems, architecture and hardware · 13 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Energy-guided Dual Domain-invariant Prompting Framework with Fourier Regularization for Generalized Few-Shot Medical Segmentation
abstract
Precise segmentation of organ and tissue lesions is essential for clinical diagnosis and treatment. Despite the progress of deep learning and foundation segmentation models, their domain generalization capability remains limited particularly when dealing with cross-domain scenarios or unseen data, leading to significant performance degradation. Current medical SAM-based generalization methods face two primary challenges: First, existing prompt-tuning strategies inadequately capture key domain-invariant features; Second, the reliance on fully labeled source domain data is unrealistic in clinical practice. To address these challenges, we propose a novel Dual domain-Invariant Prompt Optimization (DIPO) enhanced by energy-guided augmentation and frequency consistency regularization for few-shot medical image segmentation generalization. Our approach introduces a multi-band momentum enhancement strategy to dynamically augment source data by leveraging diverse frequency bands of the Fourier amplitude spectrum. Furthermore, we integrate multiscale geometric representation-based non-subsampled shearlet transform and text prompts to strengthen the extraction of shape- and texture-related domain-invariant features. Finally, we employ frequency consistency regularization to refine model robustness using predictions from unlabeled data. Experimental results in prostate and fundus datasets demonstrate that our method significantly outperforms current state-of-the-art methods.
Shaolei Liu, Dongchen Zhu, Jiamao Li
AAAI3
2026 DAWDet: A dynamic content-aware multi-branch framework with adaptive wavelet boosting for small object detection
Shaolei Liu, Dongchen Zhu, Lei Wang 0202, Jiamao Li
Pattern Recognit.3
2026 Learning to chase: Adaptive audio-visual navigation for moving sounds in complex environments
Yuanzheng He, Yuyi Liu, Chenfan Zhang, Dongchen Zhu, Lei Wang 0202
Pattern Recognit. Lett.4
2026 OMFlow: Optimizing optical flow via occlusion motion estimation
Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
Pattern Recognit. Lett.3
2026 A Unified Task Trajectory Planner Based on High-Order Motion Information and Virtual Impedance: Experimental Validation on a Humanoid Upper-Limb Robot
abstract
This article proposes a unified task trajectory planner (UTTP) that integrates motion observation, trajectory planning and obstacle avoidance function, which is a key technology for dynamic target operations. The high-order motion observer (HOMO) for capturing target motion information, and the virtual impedance model for noise reduction and obstacle avoidance are two cores of UTTP. First, in HOMO, an exploratory model and optimal policy iteration method are established based on dynamics principles. Guided by the optimal policy, the exploratory model simulates the target's dynamic behavior and outputs pose, velocity, and acceleration in real-time. Second, a controller is designed based on the high-order motion information observed by HOMO and the positions of obstacles, driving the virtual impedance model to produce safe and smooth task trajectories. The introduction of UTTP can effectively enhance the robot's autonomous decision-making capabilities in dynamic environments. Finally, simulation analyses and experiments on a hyperredundant humanoid upper-limb robot support the effectiveness and superiority of the proposed algorithm in terms of accuracy and smoothness.
Jiaxiu Liu, HongZhe Jin, Fengjia Ju, Dongchen Zhu, Jie Zhao 0003
IEEE Trans. Ind. Informatics4
2025 Mutual Semantic Bridged Tri-Tower Fusion for Audio-Visual Segmentation
abstract
Community researchers have developed various advanced audio-visual segmentation (AVS) models to accurately segment sound-producing objects. However, existing methods face two key limitations: 1) they lack effective extraction of audio semantics, resulting in an insufficiently accurate reference guidance for segmentation, and 2) they directly align cross-modal information without adequately considering the significant semantic gaps and cross-modal noise between audio and visual features, resulting in catastrophic forgetting of audio information. To address these challenges, we propose a novel Mutual Semantic-Bridged Tri-Tower Fusion Network. First, we introduce a Mutual Semantic Encoder (MSE) to extract global mutual semantic embeddings, enabling more robust semantic representations of sound-producing objects. Second, we design the Tri-Tower Semantic Fusion (TTSF) mechanism, which leverages mutual semantics as an anchor to integrate visual, semantic, and audio information, facilitating deeper interactions of sound-producing object features across modalities. Experimental results demonstrate that our approach significantly improves performance on AVS benchmarks.
Jingqi Qu, Dongchen Zhu, Jiamao Li
ICME3
2025 The Motion in the Details: Adapting CLIP for Action Recognition via Dual-prompt Guidance
abstract
Recently, large-scale pre-trained language-image models like CLIP have shown extraordinary capabilities for understanding image-level objects, but naively transferring such models to video recognition is still unsatisfactory. Existing methods design plugged temporal modules into the pre-trained model or explore the vision-text relation to improve the performance, which either demonstrated insufficient attention to the frame-wise action subject or is limited by the unreliable prompt guidance, struggling to achieve better adaptation. In this work, we present DP-CLIP, a novel dual-prompt guidance mechanism to disentangle the adaptation in textual and temporal aspects. Specifically, DP-CLIP consists of an explicit instruction-filtered caption prompt guidance (ECPG) module and an implicit action subject prompt mining (ISPM) module to maximize the textual-visual alignment and enhance the temporal reasoning ability. The former designs an instruction-filtered strategy to generate reliable and reasonable semantic captions matching with the video details, narrowing the gap between videos and labels. Further, the ISPM emphasizes action-related discriminative clues in temporal via highlighting the action subject and refining the motion cues across frames, immune to irrelevant interference. Extensive experiments on Kinetics-400, HMDB-51 and UCF-101 demonstrate that our method achieves state-of-the-art performance across fully-supervised and zero-shot settings.
Longjuan Sun, Xixia Xu, Dongchen Zhu, Jiamao Li
ICME3
2025 $\mathbf{F}^{2} \mathbf{R}^{2}$: Frequency Filtering-Based Rectification Robustness Method for Stereo Matching
abstract
Most stereo matching networks assume that the stereo images are perfectly rectified, ignoring the perturbation of extrinsic parameters due to collisions, mechanical vibrations, and thermal expansion. This leads to poor rectification robustness in real-world stereo systems. That is, even minor rectification errors can lead to failure, making stereo systems unreliable for long-term autonomous operation in complex environments. In this paper, we are the first to propose a frequency filtering-based rectification robustness ($\mathbf{F}^{2} \mathbf{R}^{2}$) method for stereo matching, which aims to enhance the robustness of existing stereo networks to rectification errors. Specifically, we propose a sensitive frequency filter (SFF) to remove components susceptible to rectification errors within the frequency domain. SFF achieves the filtering through the learning-based adaptive filtering mask (AFM) guided by the spatial-frequency mapping modulation mask (SFM). Moreover, we build the matching feature reconstruction module (MFRM) to recover the features lost during filtering to benefit cost aggregation. Comprehensive experiments on simulated datasets and self-collected data validate that our method can significantly enhance the rectification robustness of stereo matching networks.
Haolong Zhou, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICRA2
2025 A Fast and Accurate ANN-SNN Conversion Algorithm with Negative Spikes
abstract
Spiking neural network (SNN) is an event-driven neural network that can greatly reduce the power consumption of the conventional artificial neural networks (ANN). Many ANN models can be converted to SNN models when the activation function is ReLU. For ANN models with other activation functions, such as the Leaky ReLU function, the converted SNN models either suffer from serious accuracy degradation or require a long time step. In this paper, we propose a fast and accurate ANN-SNN conversion algorithm for models with the Leaky ReLU function. We design a novel neuron model that supports negative spikes. To address the problem of long tail distribution in the activation values, we propose a threshold optimization algorithm based on the variance of the activation values. To avoid the problem of error accumulation, we jointly calibrate all layers in the SNN model with adaptive weighting. Experiment results verify the effectiveness of the proposed algorithm.
Xu Wang 0035, Dongchen Zhu, Jiamao Li
IJCAI2
2025 AKP: Actionable Knowledge-Augmented Agent Planning with Large Language Models
abstract
Large Language Models (LLMs) have demonstrated remarkable planning capabilities in complex language reasoning tasks. However, existing multi-step reasoning technologies easily introduces potential unreliable and inaccurate error accumulations over long-horizon action steps, and thereby making it exceedingly difficult to accurately explore the exponentially large search space. This deficiency primarily stems from the short-sighted greedy decoding of the next admissible action, and the lack of build-in actionable knowledge, which exacerbates the generation of hallucinatory or conflicting action sequences. To address these limitations, we propose Actionable Knowledge-augmented agent Planning (AKP), a novel LLM-based framework designed to enhance task planning of language agents by incorporating external actional knowledge. Specifically, we introduce a new knowledgeable policy called GVR by exploring future longer-term action paths, which leverages multiple domain foundation models (Verifier, Rewarder) trained on actional knowledge to jointly guide and calibrate more reasonable action generation. Additionally, AKP integrates the learned GVR into Monte Carlo Tree Search (MCTS) for deliberate planning, predicting various potential actions and iteratively refining alternative plans through lookahead and backtracking. Experimental results demonstrate that AKP significantly outperforms existing baselines, showcasing superior planning performance in tackling complex goal tasks across diverse embodied environments.
Dongchen Zhu, Lei Wang 0202, Jiamao Li
INDIN2
2025 WKAgent: World Knowledge-Guided Agent for Task Planning with Large Language Models
abstract
Recent advancements in using large language models (LLMs) as agent policies have demonstrated impressive planning capabilities for multi-step reasoning and planning tasks. Despite their achievements, LLM-based agents are prone to trial-and-error and planning hallucinations when generating action plans in embodied environments. This limitation stems from their lack of intrinsic world knowledge, resulting in a poor understanding of the physical world and ineffective decision-making. Imitating humans’ mental world model which provides commonsense prior knowledge when planning task, in this paper, we propose World Knowledge-Guided Agent Planning (WKAgent), a novel LLM planning approach that repurposes the LLM both a world knowledge model and an action reasoning model, and incorporates a search algorithm, such as Monte Carlo Tree Search (MCTS) to enhance the planning capabilities of language agents. Specifically, WKAgent empowers LLMs to self-synthesize dynamic world knowledge from expert trajectories to guide deliberate planning akin to human brains, which involves constantly tracking the goal instruction, summarizing world state changes, and anticipating the next course of actions during planning. Furthermore, WKAgent employs a lightweight knowledgeable policy model learned from contrastive correct-incorrect world knowledge-action pairs, aiming to constrain the next actions and calibrate more reasonable action paths. Experimental results on three complex tasks demonstrate that WKAgent can achieve superior task planning performance compared to various baselines. Further analysis indicates the explicit world knowledge from our WKAgent can improve the essential understanding capabilities of the real physical world and effectively alleviate the blind trial-and-error. Moreover, the knowledgeable learning policy model can steer the reasoning process towards generating more feasible action plans and mitigate the planning hallucinations.
Dongchen Zhu, Lei Wang 0202, Jiamao Li
INDIN2
2025 PCGE: Boosting 3D Visual Grounding via Progressive Comprehension and Geometric-topology Perception Enhancement
abstract
The 3D visual grounding task aims to establish correspondences between the 3D physical world and textual descriptions. Despite significant progress having been made, it still suffers from some challenges that need to be solved. a) Scene-agnostic text reasoning causes misaligned target region concentration. b) The regional pseudo-center interferences result in an inaccurate geometric center. c) Multi-modal features overemphasize semantics, leading to degradation in geometric topological perception for size regression. To address these issues, we creatively propose a Progressive Comprehension and Geometric-topology Perception Enhancement (PCGE) one-stage framework, which decouples the task into keypoint estimation and size regression under textual constraints. Specifically, to enable coarse-to-fine keypoint estimation, we propose the STAR module to focus the target region approximately with a scene-specific reasoning mechanism, while the K2C module performs geometric calibration to alleviate pseudo-center bias. For size regression, we propose GTE to enhance the geometric boundary perception during the decoding process, improving size regression via establishing topological matrices. Compared with previous methods, our approach achieves state-of-the-art performance on ScanRefer and Sr3D, with 3.94% leads of [email protected] on ScanRefer, and 3.7% leads on Sr3D.
Zeyue Wang 0002, Xixia Xu, Dongchen Zhu, Jiamao Li
IROS4
2024 Rotated Orthographic Projection for Self-supervised 3D Human Pose Estimation
Yixuan Pan, Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ECCV (69)4
2024 IPHGaze: Image Pyramid Gaze Estimation with Head Pose Guidance
Hekuangyi Che, Dongchen Zhu, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR (28)2
2024 MemoFlow: Modifying Explicit Motion of Inconsistency in Optical Flow
Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICPR (30)3
2024 BCNet: Binocular Cooperative Network for Gaze Estimation
Dongchen Zhu, Minjing Lin, Hekuangyi Che, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR (28)1
2024 CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers
abstract
With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we present a novel transformer-based framework named CVFormer, which leverages two-dimensional circum-views from the ego to excavate three-dimensional features of the surrounding environment. Circum-views provide a novel solution for effectively addressing the representation of dense and fine-grained scenes. Specifically, a multi-attention module CTMA is designed for fusing temporal features from circum-views to fully exploit the spatiotemporal correlations between frames and capture more comprehensive clues. Furthermore, a novel 2D projection constraint is established by observing objects from different perspective directions, and multiple 3D constraints based on object invariance and semantic consistency are also conducted for supervising the network, which enhances its performance of understanding the scene. Experimental results on nuScenes dataset demonstrate that the proposed CVFormer obviously outperforms existing methods for occupancy prediction.
Zhengqi Bai, Wenjun Shi, Dongchen Zhu, Hanlong Kang, Gang Ye, Lei Wang 0202, Jiamao Li
ICRA3
2024 BEE-Net: Bridging Semantic and Instance with Gated Encoding and Edge Constraint for Efficient Panoptic Segmentation
abstract
Panoptic segmentation is a challenging perception task, which can help robots to comprehensively perceive the surrounding environment. In the task, we notice that semantic, instance, and panoptic have rich relations, however, which are rarely explored. In this work, we propose a novel panoptic, instance, and semantic bridged network to delve into the reciprocal relation. To make semantic and instance benefit from each other, we design a novel Gated Encoding (GE) module, incorporating complementary cues between semantic and instance heads through the gated mechanism. In addition, a novel edge-aware consistency constraint among edges of each task is presented, which exhaustedly exploits geometric constraints, to boost the segmentation quality of challenging edges. Experimental results on the Cityscapes and MS-COCO datasets demonstrate that our approach achieves state-of-the-art performance in an efficient CNN-based paradigm, attaining a balance between accuracy and efficiency.
Dongchen Zhu, Wenjun Shi, Gang Ye, Lei Wang 0202, Jiamao Li
ICRA3
2024 ESD-Pose: Enhanced Semantic Discrimination for Generalizable 6D Pose Estimation
Xingyuan Deng, Kangru Wang, Lei Wang 0202, Dongchen Zhu, Jiamao Li
PRCV (6)4
2024 Discriminative-Guided Diffusion-Based Self-supervised Monocular Depth Estimation
Dongchen Zhu, Lei Wang 0202, Jiamao Li
PRCV (6)3
2023 CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing
abstract
The weakly-supervised audio-visual video parsing (AVVP) task aims to parse duration and categories of each snippet when only the video-level event labels are provided. Most methods either leverage attention mechanisms to explore cross-modal and cross-video event semantics or alleviate label noise to improve performance. However, the distributional modality discrepancy caused by the heterogeneity of signals remains a significant challenge. To this end, we propose a novel cross-modal common-specific feature learning method (cm-CS) to map the modal features into modality-common and modality-specific subspaces. The former aims to capture similar high-level scene cue across different modalities, while the later attempts to capture specific cue. The proposed method is applied among and across in-visual 2D-3D modalities, audio-visual modalities, respectively. In addition, we design a training strategy to strengthen the learning of similarity and differences across modalities. Experiments show a large improvement of our method against existing works on the Look, Listen, and Parse (LLP) dataset (e.g. from 58.9% to 62.9% in video-level visual metric).
Dongchen Zhu, Wenjun Shi, Jiamao Li
ICASSP2
2023 Fast Extrinsic Calibration for Multiple Inertial Measurement Units in Visual-Inertial System
abstract
In this paper, we propose a fast extrinsic calibration method for fusing multiple inertial measurement units (MIMU) to improve visual-inertial odometry (VIO) localization accuracy. Currently, data fusion algorithms for MIMU highly depend on the number of inertial sensors. Based on the assumption that extrinsic parameters between inertial sensors are perfectly calibrated, the fusion algorithm provides better localization accuracy with more IMUs, while neglecting the effect of extrinsic calibration error. Our method builds two non-linear least-squares problems to estimate the MIMU relative position and orientation separately, independent of external sensors and inertial noises online estimation. Then we give the general form of the virtual IMU (VIMU) method and propose its propagation on manifold. We perform our method on datasets, our self-made sensor board, and board with different IMUs, validating the superiority of our method over competing methods concerning speed, accuracy, and robustness. In the simulation experiment, we show that only fusing two IMUs with our calibration method to predict motion can rival nine IMUs. Real-world experiments demonstrate better localization accuracy of the VIO integrated with our calibration method and VIMU propagation on manifold.
Youwei Yu, Fengjie Fu, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICRA5
2023 FeatDANet: Feature-level Domain Adaptation Network for Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) is proposed to better adapt the network trained on labeled synthetic data to unlabeled real-world data for addressing the annotation cost. However, most of these methods pay more attention to domain distributions in input and output stages while ignoring the important differences in semantic expressions and local details in middle feature stages. Therefore, a novel UDA network named FeatDANet is presented to align feature-level domain distributions at each encoder layer. Specifically, two attention-based modules abbreviated as IFAM and DFLM are designed and implemented by mixing queries and keys between domains for advisable domain adaptation. The former realizes Inter-domain Features Alignment by transferring feature style, and the latter achieves Domain-invariant Features Learning robustly for the domain shift. Furthermore, FeatDANet is constructed as a self-training network with three weight-sharing branches, and an improved pseudo-labels learning strategy is suggested by identifying more confident pseudolabels and maximizing the use of pseudo-labels. It increases the participation of unlabeled data and also ensures stability in training. Extensive experiments show that FeatDANet achieves state-of-the-art performances on the tasks of GTA→Cityscapes and Synthia→Cityscapes.
Wenjun Shi, Dongchen Zhu, Jiamao Li
IROS3
2023 MFCFlow: A Motion Feature Compensated Multi-Frame Recurrent Network for Optical Flow Estimation
abstract
Occlusions have long been a hard nut to crack in optical flow estimation due to ambiguous pixels matching between abutting images. Current methods only take two consecutive images as input, which is challenging to capture temporal coherence and reason about occluded regions. In this paper, we propose a novel optical flow estimation framework, namely MFCFlow, which attempts to compensate for the information of occlusions by mining and transferring motion features between multiple frames. Specifically, we construct a Motion-guided Feature Compensation cell (MFC cell) to enhance the ambiguous motion features according to the correlation of previous features obtained by attention-based structure. Furthermore, a TopK attention strategy is developed and embedded into the MFC cell to improve the subsequent matching quality. Extensive experiments demonstrate that our MFCFlow achieves significant improvements in occluded regions and attains state-of-the-art performances on both Sintel and KITTI benchmarks among other multi-frame optical flow methods.
Yonghu Chen, Dongchen Zhu, Wenjun Shi, Jiamao Li
WACV2
2023 SD-Pose: Structural Discrepancy Aware Category-Level 6D Object Pose Estimation
abstract
Category-level 6D object pose estimation aims to predict the full pose and size information for previously unseen instances from known categories, which is an essential portion of robot grasping and augmented reality. However, the core challenge of this task still is the enormous shape variation within each category. With regard to the challenge, we propose a novel framework SD-Pose, which utilizes the instance-category structural discrepancy and the potential geometric-semantic association to enhance the exploration of the intra-class shape information. Specifically, an information exchange augmentation (IEA) module is introduced to supplement the instance-category structural information by their structural discrepancy, thus facilitating the enhanced geometric information to contain both the character of instance shape and the commonality of category structure. For complementing the deficiencies of structural information adaptively, a semantic dynamic fusion (SDF) module is further designed to fuse semantic and geometric features. Finally, the proposed SD-Pose framework equipped with the IEA and SDF modules hierarchically supplements instance-category structural information in a stacked manner and achieves state-of-the-art performance on the CAMERA25 and REAL275 datasets.
Dongchen Zhu, Wenjun Shi, Jiamao Li
WACV2
2022 Traffic Sign Instances Segmentation Using Aliased Residual Structure and Adaptive Focus Localizer
abstract
Traffic sign recognition plays a crucial role in both unmanned vehicles and advanced driver assistance systems. Although many recent deep-learning-based approaches have made some progress on this task, it still suffers from the significant changes in scale and rotation, the presence of objects of the same color appearance, as well as the high similarity between classes. In this paper, we first propose a traffic sign detection network, named AAMNet, by adopting the generally framework of one-stage anchor-free instance segmentation models. For the feature extraction, a novel Aliased residual block is designed to support the encoder to retain more detailed information from the shallow low-level features. For decoder, we introduce a channel attention module into the mask generator to implement an Adaptive focus localizer head, which can filter out the irrelevant prototype masks. A Mask guided center loss is further constructed to improve the localization accuracy. Then, considering the difficulty of text based traffic signs recognition, for the first time, we developed a text-traffic-signs dataset TextTSD covering rich scenes and multiple languages. Extensive experiments both on TT100k and TextTSD show that our AAMNet gains a competitive performance compared with several state-of-the-art methods.
Wenjun Shi, Yingjun Shi, Dongchen Zhu, Jiamao Li
ICPR3
2022 SRNet: Structural Relation-aware Network for Head Pose Estimation
abstract
Estimating head pose from a single RGB image has recently attracted considerable research attention. Prior arts employ a CNN backbone to process face images and then directly output Euler angles. We argue that they may ignore essential features that are highly correlated to head pose due to the non-global perspective, and the ambiguity and discontinuity issues of Euler angles representation could interfere with the performance of challenging samples. In this paper, we formulate the head pose estimation problem into quaternion representation space and propose a novel framework named Structural Relation-aware Network (SRNet). Different from previous methods, our SRNet explicitly explores the correlation among different regions of the face for mining global facial structure information. Furthermore, in order to boost robustness and generalization of the model, a hard example mining (HEM) strategy is designed to mitigate the data imbalance issue by adjusting the contributions of examples in different states to loss. Extensive experiments demonstrate that our method outperforms the current state-of-the-art alternatives on the public benchmark datasets: AFLW2000 and BIWI.
Zhaoxiang Zeng, Dongchen Zhu, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR2
2022 Dual-Neighborhood Feature Aggregation Network for Point Cloud Semantic Segmentation
abstract
Neighborhood construction plays a key role in point cloud processing. However, existing models only use a single neighborhood construction method to extract neighborhood features, which limits their scene understanding ability. In this paper, we propose a learnable Dual-Neighborhood Feature Aggregation (DNFA) module embedded in the encoder that builds and aggregates comprehensive surrounding knowledge of point clouds. In this module, we first construct two kinds of neighborhoods and design corresponding feature enhancement blocks, including a Basic Local Structure Encoding (BLSE) block and an Extended Context Encoding (ECE) block. The two blocks mine structural and contextual cues for enhancing neighborhood features, respectively. Second, we propose a Geometry-Aware Compound Aggregation (GACA) block, which introduces a functionally complementary compound pooling strategy to aggregate richer neighborhood features. To fully learn the neighborhood distribution, we absorb the geometric location information during the aggregation process. The proposed module is integrated into an MLP-based large-scale 3D processing architecture, which constitutes a 3D semantic segmentation network called DNFA-Net. Extensive experiments on public datasets containing indoor and outdoor scenes validate the superiority of DNFA-Net.
Minghong Chen, Wenjun Shi, Dongchen Zhu, Jiamao Li
ICTAI4
2022 J-RR: Joint Monocular Depth Estimation and Semantic Edge Detection Exploiting Reciprocal Relations
abstract
Depth estimation and semantic edge detection are two key tasks in computer vision, which have made great progress. To date, how to associatively predict the depth and the semantic edge is rarely explored. In this work, we first propose a flexible two-branch framework that can make the two tasks take advantage of each other, achieving a win-win situation. Specifically, for the semantic edge detection branch, an Enhanced Edge Weighting strategy (EEW) is designed, which learns weight information from the by-product of depth branch, depth edge, to enhance edge perception in features. Meanwhile, we make depth estimation benefit from semantic edge detection through introducing Depth Edge Semantic Classification module (DESC). Furthermore, a double reconstruction (D-reconstruction) approach is presented, together with semantic edge-guided disparity smoothing loss to mitigate the ambiguities of the self-supervised manner for depth estimation. Experiments on the Cityscapes dataset demonstrate that our framework outperforms the state-of-the-art method in depth estimation along with a significant improvement in semantic edge detection.
Deming Wu, Dongchen Zhu, Wenjun Shi, Jiamao Li
IROS2
2022 Spatiotemporally Enhanced Photometric Loss for Self-Supervised Monocular Depth Estimation
abstract
Recovering depth information from a single image is a long-standing challenge, and self-supervised depth estimation methods have gradually attracted attention due to not relying on high-cost ground truth. Constructing an accurate photometric loss based on photometric consistency is crucial for these self-supervised methods to obtain high-quality depth maps. However, the photometric loss in most studies treats all pixels indiscriminately, resulting in poor performance. In this paper, we propose two modules based on the spatial and temporal cues to refine the photometric loss. Delving into the geometric model of photometric consistency, we introduce a depth-aware pixel correspondence module (DPC) inside the monocular depth estimation pipeline. It reduces the uncertainty of photometric errors by applying the homography matrix to the projection of corresponding pixels in far regions instead of the fundamental matrix. Furthermore, we design an omnidirectional auto-masking module (OA) to boost the robustness of our model, which utilizes temporal sequences to generate disturbance poses and hypothetical views to distin-guish dynamic objects with different directions that violate the photometric consistency. Experiments on the KITTI and the Make3d datasets reveal that our framework achieves state-of-the-art performance.
Dongchen Zhu, Wenjun Shi, Jiamao Li
IROS2
2022 EFG-Net: A Unified Framework for Estimating Eye Gaze and Face Gaze Simultaneously
Hekuangyi Che, Dongchen Zhu, Minjing Lin, Wenjun Shi, Jiamao Li
PRCV (1)2
2022 SemRegionNet: Region ensemble 3D semantic instance segmentation network with semantic spatial aware discriminative loss
Dongchen Zhu, Wenjun Shi, Jiamao Li
Neurocomputing2
2022 CSF: Closed-mask-guided semantic fusion method for semantic perception of unknown scenes
Minghong Chen, Ruijun Shu, Dongchen Zhu, Jiamao Li
Pattern Recognit. Lett.3
2022 RGB-D Semantic Segmentation and Label-Oriented Voxelgrid Fusion for Accurate 3D Semantic Mapping
abstract
The 3D semantic map plays an increasingly important role in a wide variety of applications, especially for many kinds of task-driven robots. In this paper, we present a semantic mapping methodology for 3D semantic map obtaining from RGB-D scans. In contrast to existing methods that use 3D annotated information as supervisory, we focus on accurate 2D frame labeling and combine labels in 3D space using semantic fusion mechanism. For scene parsing, a two-stream network with a novel discriminatory mask loss is proposed to explore sufficient extraction and fusion of RGB and depth information achieving steadily semantic segmentation. The discriminatory mask guides the cross-entropy loss function and interprets the influence of different pixels on back-propagation, which reduces the harmful effects of the depth noise or the fallible annotation at the edges of objects. After the correspondences between frames are provided, these semantic frames are fused in unified 3D coordinates using the novel label-oriented voxelgrid filter. It can ensure the intra-frame spatial continuity and the inter-frame spatiotemporal consistency through introducing the label-oriented statistical principle into labeled point clouds. In order to avoid the unfavorable interference between uncorrelated frames, we further propose an adaptive grouping algorithm by applying the view frustum filter to group frames with sufficient overlap as a segment. To this end, we demonstrate the effectiveness of the proposed method on the 2D/3D semantic label benchmark of ScanNetv2 and Cityscapes datasets.
Wenjun Shi, Dongchen Zhu, Xianshun Wang, Jiamao Li
IEEE Trans. Circuits Syst. Video Technol.3
2021 Camera Parameters Aware Motion Segmentation Network with Compensated Optical Flow
abstract
Learning to distinguish independent moving objects from the observed optical flow with a moving camera remains challenging. In this work, we first present a novel camera pose compensation (CPC) scheme. With the help of ingenious geometric analysis, it breaks the observed optical flow into patterns that are easier to interpret for the motion segmentation network. Secondly, we further refine such compensation with a camera parameter aware (CPA) module to account for poses’ errors in the CPC processing and enhance the entire network’s tolerance to noises. Additionally, an MMPNet is developed to intensify the identification ability of overall motion patterns. It reaches a larger receptive field with a bottom-up information transmission structure and integrates motion information at different granularities. We demonstrate the benefits of our framework on FlyingThings3D and Monkaa datasets. Without the complement of semantic information, our approach outperforms the top methods for moving objects segmentation.
Xianshun Wang, Dongchen Zhu, Shaojie Xu, Wenjun Shi, Jiamao Li
IROS2
2021 Contour-Aware Panoptic Segmentation Network
Dongchen Zhu, Wenjun Shi, Jiamao Li
PRCV (2)2
2021 A Less-constrained Sclera Recognition Method based on Stem-and-leaf Branches Network
Dongchen Zhu, Jiamao Li, Jingquan Peng, Xianshun Wang
Pattern Recognit. Lett.1
2020 Richer Aggregated Features for Optical Flow Estimation with Edge-aware Refinement
abstract
Recent CNN-based optical flow approaches have a separated structure of feature extraction and flow estimation. The core task of optical flow is finding the corresponding points while rich representation is just the key part of such matching problems. However, the prior work usually pays more attention to the design of flow decoder than the feature extraction. In this paper, we present a novel optical flow estimation network to enrich the feature representation of each pyramid level, with a hierarchical dilated architecture and a bottom-up aggregation scheme. In addition, inspired by edge guided classical methods, we bring the edge-aware idea into our approach and propose an edge-aware refinement (EAR) subnetwork to handle motion boundaries. Using the same decoding structure as PWC-Net, our network outperforms it by a large margin and leads all its derivatives both on KITTI-2012 and KITTI-2015. Further performance analysis proves the effectiveness of proposed ideas.
Xianshun Wang, Dongchen Zhu, Jiafei Song, Jiamao Li
IROS2
2020 RegionNet: Region-feature-enhanced 3D Scene Understanding Network with Dual Spatial-aware Discriminative Loss
abstract
Neural networks have recently achieved impressive success in semantic and instance segmentation on 2D images. However, their capabilities have not been fully explored to address semantic instance segmentation on unstructured 3D point cloud data. Digging into the regional feature representation to boost point cloud comprehension, we propose a region-feature-enhanced structure consisting of adaptive regional feature complementary (ARFC) module and affinity-based regional relational reasoning (AR3) module. The ARFC module aims to complement low-level features of sparse regions adaptively. The AR3module emphasizes on mining the potential reasoning relationships between high-level features based on affinity. Both the ARFC and AR3modules are plug-and-play. Besides, a novel dual spatial-aware discriminative loss is proposed to improve the discrimination of instance embedding. Our proposal-free point cloud instance segmentation network (RegionNet) equipped with the region-feature-enhanced structure and dual spatial-aware discriminative loss achieves state-of-the-art performance on S3DIS dataset and ScanNet-v2 dataset.
Dongchen Zhu, Xiaoqing Ye, Wenjun Shi, Minghong Chen, Jiamao Li
IROS2