Jiamao Li

dblp:205/7918 · DBLP profile ↗
← Back
53ranked-venue papers
0as first author
48since 2021 · last 2026
0000-0002-7478-4544ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 24 since 2021Systems, architecture and hardware · 14 · 11 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Energy-guided Dual Domain-invariant Prompting Framework with Fourier Regularization for Generalized Few-Shot Medical Segmentation
abstract
Precise segmentation of organ and tissue lesions is essential for clinical diagnosis and treatment. Despite the progress of deep learning and foundation segmentation models, their domain generalization capability remains limited particularly when dealing with cross-domain scenarios or unseen data, leading to significant performance degradation. Current medical SAM-based generalization methods face two primary challenges: First, existing prompt-tuning strategies inadequately capture key domain-invariant features; Second, the reliance on fully labeled source domain data is unrealistic in clinical practice. To address these challenges, we propose a novel Dual domain-Invariant Prompt Optimization (DIPO) enhanced by energy-guided augmentation and frequency consistency regularization for few-shot medical image segmentation generalization. Our approach introduces a multi-band momentum enhancement strategy to dynamically augment source data by leveraging diverse frequency bands of the Fourier amplitude spectrum. Furthermore, we integrate multiscale geometric representation-based non-subsampled shearlet transform and text prompts to strengthen the extraction of shape- and texture-related domain-invariant features. Finally, we employ frequency consistency regularization to refine model robustness using predictions from unlabeled data. Experimental results in prostate and fundus datasets demonstrate that our method significantly outperforms current state-of-the-art methods.
Shaolei Liu, Dongchen Zhu, Jiamao Li
AAAI4
2026 DAWDet: A dynamic content-aware multi-branch framework with adaptive wavelet boosting for small object detection
Shaolei Liu, Dongchen Zhu, Lei Wang 0202, Jiamao Li
Pattern Recognit.5
2026 OMFlow: Optimizing optical flow via occlusion motion estimation
Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
Pattern Recognit. Lett.6
2026 Hierarchical Contrastive Consistency for Human Pose Estimation in Images and Videos
abstract
Human pose estimation (HPE) is an invaluable task in computer vision with various practical applications. This paper proposes a novel Hierarchical Contrastive Consistensy constraint (HICCON) to improve the HPE in both images and videos, which describes the input into multi-granular representations at spatial and temporal domain and performs multi-level feature consistency by exploring the characteristic of human structure and time sequence. The hierarchical contrast is conducted at four levels: keypoint-level, part-level, instance-level and clip-level. In spatial, we consider keypoint-level and part-level consistency across instances within frame to enhance the fine-grained keypoint robustness. The former conducts the single keypoint feature contrast across instances to improve the category-specific keypoint features. The latter explores the specific pair-wise features for preserving the instructive relation. In temporal, we develop the instance-level and clip-level feature consistency across frames to capture more discriminative temporal representations. The former discriminates instance features across frames within the same video, whereas the clip-level constraint aims to discriminate consistent features from different videos in order to capture more distinctive temporal features. Extensive experiments on kinds of architectures across datasets i.e, PoseTrack2017, PoseTrack2018 and PoseTrack2021 show the HICCON achieves about 1.5% improvement than baseline. Besides, the proposed method unleashes the potential of the contrastive learning in HPE field.
Xixia Xu, Qi Zou 0001, Jiamao Li
IEEE Trans. Circuits Syst. Video Technol.3
2026 MMCPose: Multimodal Condition-Driven 3D Human Pose Estimation Via Diffusion Models
abstract
Nowadays, diffusion-based methods for monocular 3D human pose estimation (3D HPE) have achieved state-of-the-art performance by directly regressing the 3D joint coordinates from the 2D observations. Although some methods incorporated the human body prior to improve the denoising quality, the absense of the structural relation and pose-aware guidance make these models prone to generating unreasonable poses. The challenge is noticeable in complex conditions such as occlusions and crowded scenarios. To alleviate this, we present MMCPose, a novel Multi-modal Condition-driven 3D HPE framework via diffusion models that capitalizes on the benefits of the multi-modal conditioning input. Specifically, we propose Multi-modal Condition Learning (MCL) strategy to incorporate multi-modal conditions such as joint- wise relation, part-aware prompt and pose-aware mask to improve the generation quality. The MCL block consists of (i) Joint- wise Relation Condition Learning (JRCL) models the flexible joint- wise relation via GCN to mitigate disturbances arising from confused joints. (ii) Part-aware Prompt Condition Learning (PPCL) constructs multi-granular prompts via accessible texts and feasible knowledge of body parts with learnable prompts to model implicit textual guidance. (iii) Pose-aware Mask Condition Learning (PMCL) designs a pose-specific mask to increase the model's emphasis to the pose region, augmenting the precision in capturing intricate pose details. Furthermore, we explore a multi-modal condition-pose interaction learning (MCPI) mechanism to establish interaction between the learned multi-modal conditions and poses to maximize the power of condition effect. This method fully unleashes the perceptual capability of the multi-modal conditions in diffusion-based 3D HPE. Extensive evaluations conducted on two popular benchmarks (e.g., Human3.6 M, MPI-INF-3DHP) and achieve new state-of-the-art performance.
Xixia Xu, Jiamao Li
IEEE Trans. Multim.2
2026 Stable Kinematics for Multirobot Collaborative Transporting System With a Deformable Sheet
abstract
A deformable, flexible sheet that is held by multiple robots can be used for manipulating and transporting an object. For such multi-robot collaborative transporting system, forward kinematics provide object's potential positions where it might remain stationary on the sheet. Some of these forward kinematics solutions are however demonstrated unstable under object's position or robot formation perturbations. This paper presents stability criteria of the kinematics solutions of the object on the deformable sheet held by multi-robot system. We capture the sheet deformation and object position using the virtual variable cable model. A constrained quadratic problem is formulated to obtain the forward kinematics solution for a given multi-robot configuration. Linear dependence of active constraints and stability multipliers are used to assess the stability of the kinematics solutions. Two stability criteria are proposed and analyzed under object position or robot formation perturbations. These criteria are related to the properties and conditions of the stability multipliers. We present an efficient computational algorithm to determine stable kinematics, which are a small fraction of feasible solutions. Experimental results and case studies are presented to validate and demonstrate the effectiveness and efficiency of the analyses and algorithms.
Wenyao Ma, Jiamao Li, Jingang Yi, Zhenhua Xiong 0001
IEEE Trans. Robotics3
2025 Mutual Semantic Bridged Tri-Tower Fusion for Audio-Visual Segmentation
abstract
Community researchers have developed various advanced audio-visual segmentation (AVS) models to accurately segment sound-producing objects. However, existing methods face two key limitations: 1) they lack effective extraction of audio semantics, resulting in an insufficiently accurate reference guidance for segmentation, and 2) they directly align cross-modal information without adequately considering the significant semantic gaps and cross-modal noise between audio and visual features, resulting in catastrophic forgetting of audio information. To address these challenges, we propose a novel Mutual Semantic-Bridged Tri-Tower Fusion Network. First, we introduce a Mutual Semantic Encoder (MSE) to extract global mutual semantic embeddings, enabling more robust semantic representations of sound-producing objects. Second, we design the Tri-Tower Semantic Fusion (TTSF) mechanism, which leverages mutual semantics as an anchor to integrate visual, semantic, and audio information, facilitating deeper interactions of sound-producing object features across modalities. Experimental results demonstrate that our approach significantly improves performance on AVS benchmarks.
Jingqi Qu, Dongchen Zhu, Jiamao Li
ICME4
2025 The Motion in the Details: Adapting CLIP for Action Recognition via Dual-prompt Guidance
abstract
Recently, large-scale pre-trained language-image models like CLIP have shown extraordinary capabilities for understanding image-level objects, but naively transferring such models to video recognition is still unsatisfactory. Existing methods design plugged temporal modules into the pre-trained model or explore the vision-text relation to improve the performance, which either demonstrated insufficient attention to the frame-wise action subject or is limited by the unreliable prompt guidance, struggling to achieve better adaptation. In this work, we present DP-CLIP, a novel dual-prompt guidance mechanism to disentangle the adaptation in textual and temporal aspects. Specifically, DP-CLIP consists of an explicit instruction-filtered caption prompt guidance (ECPG) module and an implicit action subject prompt mining (ISPM) module to maximize the textual-visual alignment and enhance the temporal reasoning ability. The former designs an instruction-filtered strategy to generate reliable and reasonable semantic captions matching with the video details, narrowing the gap between videos and labels. Further, the ISPM emphasizes action-related discriminative clues in temporal via highlighting the action subject and refining the motion cues across frames, immune to irrelevant interference. Extensive experiments on Kinetics-400, HMDB-51 and UCF-101 demonstrate that our method achieves state-of-the-art performance across fully-supervised and zero-shot settings.
Longjuan Sun, Xixia Xu, Dongchen Zhu, Jiamao Li
ICME4
2025 $\mathbf{F}^{2} \mathbf{R}^{2}$: Frequency Filtering-Based Rectification Robustness Method for Stereo Matching
abstract
Most stereo matching networks assume that the stereo images are perfectly rectified, ignoring the perturbation of extrinsic parameters due to collisions, mechanical vibrations, and thermal expansion. This leads to poor rectification robustness in real-world stereo systems. That is, even minor rectification errors can lead to failure, making stereo systems unreliable for long-term autonomous operation in complex environments. In this paper, we are the first to propose a frequency filtering-based rectification robustness ($\mathbf{F}^{2} \mathbf{R}^{2}$) method for stereo matching, which aims to enhance the robustness of existing stereo networks to rectification errors. Specifically, we propose a sensitive frequency filter (SFF) to remove components susceptible to rectification errors within the frequency domain. SFF achieves the filtering through the learning-based adaptive filtering mask (AFM) guided by the spatial-frequency mapping modulation mask (SFM). Moreover, we build the matching feature reconstruction module (MFRM) to recover the features lost during filtering to benefit cost aggregation. Comprehensive experiments on simulated datasets and self-collected data validate that our method can significantly enhance the rectification robustness of stereo matching networks.
Haolong Zhou, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICRA5
2025 A Fast and Accurate ANN-SNN Conversion Algorithm with Negative Spikes
abstract
Spiking neural network (SNN) is an event-driven neural network that can greatly reduce the power consumption of the conventional artificial neural networks (ANN). Many ANN models can be converted to SNN models when the activation function is ReLU. For ANN models with other activation functions, such as the Leaky ReLU function, the converted SNN models either suffer from serious accuracy degradation or require a long time step. In this paper, we propose a fast and accurate ANN-SNN conversion algorithm for models with the Leaky ReLU function. We design a novel neuron model that supports negative spikes. To address the problem of long tail distribution in the activation values, we propose a threshold optimization algorithm based on the variance of the activation values. To avoid the problem of error accumulation, we jointly calibrate all layers in the SNN model with adaptive weighting. Experiment results verify the effectiveness of the proposed algorithm.
Xu Wang 0035, Dongchen Zhu, Jiamao Li
IJCAI3
2025 AKP: Actionable Knowledge-Augmented Agent Planning with Large Language Models
abstract
Large Language Models (LLMs) have demonstrated remarkable planning capabilities in complex language reasoning tasks. However, existing multi-step reasoning technologies easily introduces potential unreliable and inaccurate error accumulations over long-horizon action steps, and thereby making it exceedingly difficult to accurately explore the exponentially large search space. This deficiency primarily stems from the short-sighted greedy decoding of the next admissible action, and the lack of build-in actionable knowledge, which exacerbates the generation of hallucinatory or conflicting action sequences. To address these limitations, we propose Actionable Knowledge-augmented agent Planning (AKP), a novel LLM-based framework designed to enhance task planning of language agents by incorporating external actional knowledge. Specifically, we introduce a new knowledgeable policy called GVR by exploring future longer-term action paths, which leverages multiple domain foundation models (Verifier, Rewarder) trained on actional knowledge to jointly guide and calibrate more reasonable action generation. Additionally, AKP integrates the learned GVR into Monte Carlo Tree Search (MCTS) for deliberate planning, predicting various potential actions and iteratively refining alternative plans through lookahead and backtracking. Experimental results demonstrate that AKP significantly outperforms existing baselines, showcasing superior planning performance in tackling complex goal tasks across diverse embodied environments.
Dongchen Zhu, Lei Wang 0202, Jiamao Li
INDIN4
2025 WKAgent: World Knowledge-Guided Agent for Task Planning with Large Language Models
abstract
Recent advancements in using large language models (LLMs) as agent policies have demonstrated impressive planning capabilities for multi-step reasoning and planning tasks. Despite their achievements, LLM-based agents are prone to trial-and-error and planning hallucinations when generating action plans in embodied environments. This limitation stems from their lack of intrinsic world knowledge, resulting in a poor understanding of the physical world and ineffective decision-making. Imitating humans’ mental world model which provides commonsense prior knowledge when planning task, in this paper, we propose World Knowledge-Guided Agent Planning (WKAgent), a novel LLM planning approach that repurposes the LLM both a world knowledge model and an action reasoning model, and incorporates a search algorithm, such as Monte Carlo Tree Search (MCTS) to enhance the planning capabilities of language agents. Specifically, WKAgent empowers LLMs to self-synthesize dynamic world knowledge from expert trajectories to guide deliberate planning akin to human brains, which involves constantly tracking the goal instruction, summarizing world state changes, and anticipating the next course of actions during planning. Furthermore, WKAgent employs a lightweight knowledgeable policy model learned from contrastive correct-incorrect world knowledge-action pairs, aiming to constrain the next actions and calibrate more reasonable action paths. Experimental results on three complex tasks demonstrate that WKAgent can achieve superior task planning performance compared to various baselines. Further analysis indicates the explicit world knowledge from our WKAgent can improve the essential understanding capabilities of the real physical world and effectively alleviate the blind trial-and-error. Moreover, the knowledgeable learning policy model can steer the reasoning process towards generating more feasible action plans and mitigate the planning hallucinations.
Dongchen Zhu, Lei Wang 0202, Jiamao Li
INDIN4
2025 PCGE: Boosting 3D Visual Grounding via Progressive Comprehension and Geometric-topology Perception Enhancement
abstract
The 3D visual grounding task aims to establish correspondences between the 3D physical world and textual descriptions. Despite significant progress having been made, it still suffers from some challenges that need to be solved. a) Scene-agnostic text reasoning causes misaligned target region concentration. b) The regional pseudo-center interferences result in an inaccurate geometric center. c) Multi-modal features overemphasize semantics, leading to degradation in geometric topological perception for size regression. To address these issues, we creatively propose a Progressive Comprehension and Geometric-topology Perception Enhancement (PCGE) one-stage framework, which decouples the task into keypoint estimation and size regression under textual constraints. Specifically, to enable coarse-to-fine keypoint estimation, we propose the STAR module to focus the target region approximately with a scene-specific reasoning mechanism, while the K2C module performs geometric calibration to alleviate pseudo-center bias. For size regression, we propose GTE to enhance the geometric boundary perception during the decoding process, improving size regression via establishing topological matrices. Compared with previous methods, our approach achieves state-of-the-art performance on ScanRefer and Sr3D, with 3.94% leads of [email protected] on ScanRefer, and 3.7% leads on Sr3D.
Zeyue Wang 0002, Xixia Xu, Dongchen Zhu, Jiamao Li
IROS5
2025 S2AFormer: Strip Self-Attention for Efficient Vision Transformer
abstract
The Vision Transformer (ViT) has achieved remarkable success in computer vision due to its powerful token mixer, which effectively captures global dependencies among all tokens. However, the quadratic complexity of standard self-attention with respect to the number of tokens severely hampers its computational efficiency in practical deployment. Although recent hybrid approaches have sought to combine the strengths of convolutions and self-attention to improve the performance-efficiency trade-off, the costly pairwise token interactions and heavy matrix operations in conventional self-attention remain a critical bottleneck. To overcome this limitation, we introduce S2AFormer, an efficient Vision Transformer architecture built around a novel Strip Self-Attention (SSA) mechanism. Our design incorporates lightweight yet effective Hybrid Perception Blocks (HPBs) that seamlessly fuse the local inductive biases of CNNs with the global modeling capability of Transformer-style attention. The core innovation of SSA lies in simultaneously reducing the spatial resolution of the key ( $K$ ) and value ( $V$ ) tensors while compressing the channel dimension of the query ( $Q$ ) and key ( $K$ ) tensors. This joint spatial-and-channel compression dramatically lowers computational cost without sacrificing representational power, achieving an excellent balance between accuracy and efficiency. We extensively evaluate S2AFormer on a wide range of vision tasks, including image classification (ImageNet-1K), semantic segmentation (ADE20K), and object detection/instance segmentation (COCO). Experimental results consistently show that S2AFormer delivers substantial accuracy improvements together with superior inference speed and throughput across both GPU and non-GPU platforms, establishing it as a highly competitive solution in the landscape of efficient Vision Transformers.
Guoan Xu, Wenfeng Huang, Wenjing Jia, Jiamao Li, Guangwei Gao, Guo-Jun Qi
IEEE Trans. Image Process.4
2025 Rethinking the Sparse End-to-End Multiperson Pose Estimation
abstract
Current methods of multiperson pose estimation (MPPE) typically treat the human detection and association of joints separately. They introduce complex hand-crafted pose-processes like RoI cropping, NMS and grouping or rely on dense representations to preserve the spatial features. In this article, we dive a deeper thought into this task and propose a simpler and effective framework, termed SparsePose, which can directly predict multiperson joint coordinates from the full image without any post-processes and dense representations. In SparsePose, the full-body instances are decoupled by exploring spatial-aware feature learning (SFL) without box and classification supervision. For improving the quality of instance map, the instance contrastive constraint (ICC) and center correction (CC) strategy are proposed to make the instance-wise spatial feature more discriminative. Importantly, we propose a visibility-guided weighting mechanism to enable model be confident to the visible joint predictions and insensitive to the occlusions or partial bodies. In general, SparsePose is conceptually simpler and plays favorably against the existing counterparts on three benchmarks in terms of both accuracy and efficiency.
Xixia Xu, Qi Zou 0001, Jiamao Li
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Rotated Orthographic Projection for Self-supervised 3D Human Pose Estimation
Yixuan Pan, Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ECCV (69)6
2024 VTR: Bidirectional Video-Textual Transmission Rail for CLIP-based Video Recognition
abstract
There are two key issues when transferring vision-language model like CLIP for video recognition: bidirectional video-textual transmission and temporal modeling. To address the issues, we propose a novel framework named Video-Textual Transmission Rail (VTR) which enables bidirectional transmission of visual-textual representations and temporal modeling concurrently. Specifically, Message Rail (MR) tokens are proposed within VTR to realize the bidirectional transfer of not only high-level visual-language knowledge but also low-level knowledge. For temporal modeling, we introduce the Temporal Elite (TE) - a temporal modeling module within VTR, providing VTR with the capability of both short- and long-range temporal modeling under textual supervision thanks to MR tokens. Extensive experiments on popular video datasets (i.e., Kinetics-400, Something-Something-v2, UCF-101 and HMDB-51) demonstrate that VTR achieves state-of-the-art performance in fully-supervised, zero-shot and few-shot video recognition.
Shaoqi Yu, Jiamao Li
ICME4
2024 IPHGaze: Image Pyramid Gaze Estimation with Head Pose Guidance
Hekuangyi Che, Dongchen Zhu, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR (28)7
2024 MemoFlow: Modifying Explicit Motion of Inconsistency in Optical Flow
Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICPR (30)5
2024 BCNet: Binocular Cooperative Network for Gaze Estimation
Dongchen Zhu, Minjing Lin, Hekuangyi Che, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR (28)8
2024 CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers
abstract
With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we present a novel transformer-based framework named CVFormer, which leverages two-dimensional circum-views from the ego to excavate three-dimensional features of the surrounding environment. Circum-views provide a novel solution for effectively addressing the representation of dense and fine-grained scenes. Specifically, a multi-attention module CTMA is designed for fusing temporal features from circum-views to fully exploit the spatiotemporal correlations between frames and capture more comprehensive clues. Furthermore, a novel 2D projection constraint is established by observing objects from different perspective directions, and multiple 3D constraints based on object invariance and semantic consistency are also conducted for supervising the network, which enhances its performance of understanding the scene. Experimental results on nuScenes dataset demonstrate that the proposed CVFormer obviously outperforms existing methods for occupancy prediction.
Zhengqi Bai, Wenjun Shi, Dongchen Zhu, Hanlong Kang, Gang Ye, Lei Wang 0202, Jiamao Li
ICRA11
2024 BEE-Net: Bridging Semantic and Instance with Gated Encoding and Edge Constraint for Efficient Panoptic Segmentation
abstract
Panoptic segmentation is a challenging perception task, which can help robots to comprehensively perceive the surrounding environment. In the task, we notice that semantic, instance, and panoptic have rich relations, however, which are rarely explored. In this work, we propose a novel panoptic, instance, and semantic bridged network to delve into the reciprocal relation. To make semantic and instance benefit from each other, we design a novel Gated Encoding (GE) module, incorporating complementary cues between semantic and instance heads through the gated mechanism. In addition, a novel edge-aware consistency constraint among edges of each task is presented, which exhaustedly exploits geometric constraints, to boost the segmentation quality of challenging edges. Experimental results on the Cityscapes and MS-COCO datasets demonstrate that our approach achieves state-of-the-art performance in an efficient CNN-based paradigm, attaining a balance between accuracy and efficiency.
Dongchen Zhu, Wenjun Shi, Gang Ye, Lei Wang 0202, Jiamao Li
ICRA11
2024 Scalable Complex Scene Understanding for Edge Computing
abstract
This paper aims to develop flexible and scalable scene understanding capabilities usable on resource-constrained devices. We propose an edge computing framework to enable large-scale object detection within complex visual scenes. The framework attempts to emulate human intuition - leveraging commonsense reasoning and fast concept learning from small data. Structured knowledge about objects and their relationships guides efficient deduction of scene contents. Specifically, we develop an ontology-based scene deduction model to refine an initial scene parsing, which incorporates declarative constraints on objects, their attributes, and interactions. It represents such domain knowledge explicitly and reasons over ontological relationships. By representing this domain knowledge explicitly and reasoning over ontological relationships, the model refines an initial scene parsing to better reflect commonsense consistency. The evaluation compared the proposed method to other large-scale object detection methods. The results show that our ontology-driven deduction model scale-ups the data-driven scene parser’s capability while reducing the data and compute requirements.
Nanxi Chen, Xu Wang 0035, Jiamao Li
ICWS4
2024 ESD-Pose: Enhanced Semantic Discrimination for Generalizable 6D Pose Estimation
Xingyuan Deng, Kangru Wang, Lei Wang 0202, Dongchen Zhu, Jiamao Li
PRCV (6)5
2024 Discriminative-Guided Diffusion-Based Self-supervised Monocular Depth Estimation
Dongchen Zhu, Lei Wang 0202, Jiamao Li
PRCV (6)6
2024 Paired relation feature network for spatial relation recognition
Nanxi Chen, Xu Wang 0035, Jiamao Li
Pattern Recognit. Lett.4
2024 Continual Multiview Spectral Clustering via Multilevel Knowledge
abstract
Multiview clustering aims to integrate multiple features from different views to benefit the clustering task, which has attracted much attention in recent years. Most previous research has focused on exploring multiview clustering with a fixed set of tasks. However, it is still challenging to efficiently integrate with new clustering tasks, as it requires repeated access to previous data. To address the above challenges, this letter proposes a novel continual multiview spectral clustering model. The proposed model can efficiently achieve clustering in new tasks by transferring the accumulated knowledge from past tasks, while continuously refining the knowledge to ultimately improve the performance of all clustering tasks. Specifically, the knowledge sharing among different clustering tasks is considered at multiple levels, preserving both the heterogeneous distribution of different views and the relationships between multiple views. Meanwhile, our method is modelled by deep nonlinear structures, which allows to capture more hidden knowledge. In addition, an efficient alternating optimization algorithm is proposed to refine the knowledge online. The superior experimental results on several benchmark datasets show the effectiveness and efficiency of our method compared with other state-of-the-art models.
Kangru Wang, Lei Wang 0202, Jiamao Li
IEEE Signal Process. Lett.4
2023 CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing
abstract
The weakly-supervised audio-visual video parsing (AVVP) task aims to parse duration and categories of each snippet when only the video-level event labels are provided. Most methods either leverage attention mechanisms to explore cross-modal and cross-video event semantics or alleviate label noise to improve performance. However, the distributional modality discrepancy caused by the heterogeneity of signals remains a significant challenge. To this end, we propose a novel cross-modal common-specific feature learning method (cm-CS) to map the modal features into modality-common and modality-specific subspaces. The former aims to capture similar high-level scene cue across different modalities, while the later attempts to capture specific cue. The proposed method is applied among and across in-visual 2D-3D modalities, audio-visual modalities, respectively. In addition, we design a training strategy to strengthen the learning of similarity and differences across modalities. Experiments show a large improvement of our method against existing works on the Look, Listen, and Parse (LLP) dataset (e.g. from 58.9% to 62.9% in video-level visual metric).
Dongchen Zhu, Wenjun Shi, Jiamao Li
ICASSP6
2023 Fast Extrinsic Calibration for Multiple Inertial Measurement Units in Visual-Inertial System
abstract
In this paper, we propose a fast extrinsic calibration method for fusing multiple inertial measurement units (MIMU) to improve visual-inertial odometry (VIO) localization accuracy. Currently, data fusion algorithms for MIMU highly depend on the number of inertial sensors. Based on the assumption that extrinsic parameters between inertial sensors are perfectly calibrated, the fusion algorithm provides better localization accuracy with more IMUs, while neglecting the effect of extrinsic calibration error. Our method builds two non-linear least-squares problems to estimate the MIMU relative position and orientation separately, independent of external sensors and inertial noises online estimation. Then we give the general form of the virtual IMU (VIMU) method and propose its propagation on manifold. We perform our method on datasets, our self-made sensor board, and board with different IMUs, validating the superiority of our method over competing methods concerning speed, accuracy, and robustness. In the simulation experiment, we show that only fusing two IMUs with our calibration method to predict motion can rival nine IMUs. Real-world experiments demonstrate better localization accuracy of the VIO integrated with our calibration method and VIMU propagation on manifold.
Youwei Yu, Fengjie Fu, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICRA8
2023 FeatDANet: Feature-level Domain Adaptation Network for Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) is proposed to better adapt the network trained on labeled synthetic data to unlabeled real-world data for addressing the annotation cost. However, most of these methods pay more attention to domain distributions in input and output stages while ignoring the important differences in semantic expressions and local details in middle feature stages. Therefore, a novel UDA network named FeatDANet is presented to align feature-level domain distributions at each encoder layer. Specifically, two attention-based modules abbreviated as IFAM and DFLM are designed and implemented by mixing queries and keys between domains for advisable domain adaptation. The former realizes Inter-domain Features Alignment by transferring feature style, and the latter achieves Domain-invariant Features Learning robustly for the domain shift. Furthermore, FeatDANet is constructed as a self-training network with three weight-sharing branches, and an improved pseudo-labels learning strategy is suggested by identifying more confident pseudolabels and maximizing the use of pseudo-labels. It increases the participation of unlabeled data and also ensures stability in training. Extensive experiments show that FeatDANet achieves state-of-the-art performances on the tasks of GTA→Cityscapes and Synthia→Cityscapes.
Wenjun Shi, Dongchen Zhu, Jiamao Li
IROS6
2023 MFCFlow: A Motion Feature Compensated Multi-Frame Recurrent Network for Optical Flow Estimation
abstract
Occlusions have long been a hard nut to crack in optical flow estimation due to ambiguous pixels matching between abutting images. Current methods only take two consecutive images as input, which is challenging to capture temporal coherence and reason about occluded regions. In this paper, we propose a novel optical flow estimation framework, namely MFCFlow, which attempts to compensate for the information of occlusions by mining and transferring motion features between multiple frames. Specifically, we construct a Motion-guided Feature Compensation cell (MFC cell) to enhance the ambiguous motion features according to the correlation of previous features obtained by attention-based structure. Furthermore, a TopK attention strategy is developed and embedded into the MFC cell to improve the subsequent matching quality. Extensive experiments demonstrate that our MFCFlow achieves significant improvements in occluded regions and attains state-of-the-art performances on both Sintel and KITTI benchmarks among other multi-frame optical flow methods.
Yonghu Chen, Dongchen Zhu, Wenjun Shi, Jiamao Li
WACV7
2023 SD-Pose: Structural Discrepancy Aware Category-Level 6D Object Pose Estimation
abstract
Category-level 6D object pose estimation aims to predict the full pose and size information for previously unseen instances from known categories, which is an essential portion of robot grasping and augmented reality. However, the core challenge of this task still is the enormous shape variation within each category. With regard to the challenge, we propose a novel framework SD-Pose, which utilizes the instance-category structural discrepancy and the potential geometric-semantic association to enhance the exploration of the intra-class shape information. Specifically, an information exchange augmentation (IEA) module is introduced to supplement the instance-category structural information by their structural discrepancy, thus facilitating the enhanced geometric information to contain both the character of instance shape and the commonality of category structure. For complementing the deficiencies of structural information adaptively, a semantic dynamic fusion (SDF) module is further designed to fuse semantic and geometric features. Finally, the proposed SD-Pose framework equipped with the IEA and SDF modules hierarchically supplements instance-category structural information in a stacked manner and achieves state-of-the-art performance on the CAMERA25 and REAL275 datasets.
Dongchen Zhu, Wenjun Shi, Jiamao Li
WACV7
2023 HandGCNFormer: A Novel Topology-Aware Transformer Network for 3D Hand Pose Estimation
abstract
Despite the substantial progress in 3D hand pose estimation, inferring plausible and accurate poses in the presence of severe self-occlusion and high self-similarity remains an inherent challenge. To mitigate the ambiguity arising from invisible and similar joints, we propose a novel Topology-aware Transformer network named HandGCNFormer, incorporating the prior knowledge of hand kinematic topology into the network while modeling long-range context information. Specifically, we present a novel Graphformer decoder with an additional node-offset graph convolutional layer (NoffGConv) that optimizes the synergy of Transformer and GCN, capturing long-range dependencies as well as local topology connection between joints. Furthermore, we replace the standard MLP prediction head with a novel Topology-aware head to better utilize local topology constraints for more plausible and accurate poses. Our method achieves state-of-the-art performance on four challenging datasets including Hands2017, NYU, ICVL, and MSRA.
Yintong Wang, Jiamao Li
WACV3
2023 Agile Services Provisioning for Learning-Based Applications in Fog Computing Networks
abstract
Fog computing has emerged as a promising solution for service provisioning. To support the increasing number of smart end devices and their requirement for intelligent data processing, numerous learning-based applications will be deployed at the edge network to provide services with an acceptable level of latency. However, such applications assume that the training and the actual data are independent and identically distributed, which is impractical in the dynamic and heterogeneous environment of the real world. This makes service matchmaking a challenge because even if the service functionality matches user requirements, the mismatched actual data will inevitably result in the degradation of service quality. This paper targets the mismatching issue of learning-based applications and the high requirement of intelligent data processing at the network edge and proposes a novel service provisioning model. This model introduces a novel service description model to resolve data mismatch, a clustering algorithm that pre-processes requests to deal with high concurrency requirements, and a heuristic joint service searching method to reduce traffic costs. This work was evaluated through a case study and simulations. Evaluation results show that the proposed model can achieve a good quality of services and reduce response time as well as traffic costs.
Nanxi Chen, Yanbei Li, Hongfeng Shu, Jiamao Li
IEEE Trans. Serv. Comput.5
2022 Traffic Sign Instances Segmentation Using Aliased Residual Structure and Adaptive Focus Localizer
abstract
Traffic sign recognition plays a crucial role in both unmanned vehicles and advanced driver assistance systems. Although many recent deep-learning-based approaches have made some progress on this task, it still suffers from the significant changes in scale and rotation, the presence of objects of the same color appearance, as well as the high similarity between classes. In this paper, we first propose a traffic sign detection network, named AAMNet, by adopting the generally framework of one-stage anchor-free instance segmentation models. For the feature extraction, a novel Aliased residual block is designed to support the encoder to retain more detailed information from the shallow low-level features. For decoder, we introduce a channel attention module into the mask generator to implement an Adaptive focus localizer head, which can filter out the irrelevant prototype masks. A Mask guided center loss is further constructed to improve the localization accuracy. Then, considering the difficulty of text based traffic signs recognition, for the first time, we developed a text-traffic-signs dataset TextTSD covering rich scenes and multiple languages. Extensive experiments both on TT100k and TextTSD show that our AAMNet gains a competitive performance compared with several state-of-the-art methods.
Wenjun Shi, Yingjun Shi, Dongchen Zhu, Jiamao Li
ICPR5
2022 SRNet: Structural Relation-aware Network for Head Pose Estimation
abstract
Estimating head pose from a single RGB image has recently attracted considerable research attention. Prior arts employ a CNN backbone to process face images and then directly output Euler angles. We argue that they may ignore essential features that are highly correlated to head pose due to the non-global perspective, and the ambiguity and discontinuity issues of Euler angles representation could interfere with the performance of challenging samples. In this paper, we formulate the head pose estimation problem into quaternion representation space and propose a novel framework named Structural Relation-aware Network (SRNet). Different from previous methods, our SRNet explicitly explores the correlation among different regions of the face for mining global facial structure information. Furthermore, in order to boost robustness and generalization of the model, a hard example mining (HEM) strategy is designed to mitigate the data imbalance issue by adjusting the contributions of examples in different states to loss. Extensive experiments demonstrate that our method outperforms the current state-of-the-art alternatives on the public benchmark datasets: AFLW2000 and BIWI.
Zhaoxiang Zeng, Dongchen Zhu, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR7
2022 Dual-Neighborhood Feature Aggregation Network for Point Cloud Semantic Segmentation
abstract
Neighborhood construction plays a key role in point cloud processing. However, existing models only use a single neighborhood construction method to extract neighborhood features, which limits their scene understanding ability. In this paper, we propose a learnable Dual-Neighborhood Feature Aggregation (DNFA) module embedded in the encoder that builds and aggregates comprehensive surrounding knowledge of point clouds. In this module, we first construct two kinds of neighborhoods and design corresponding feature enhancement blocks, including a Basic Local Structure Encoding (BLSE) block and an Extended Context Encoding (ECE) block. The two blocks mine structural and contextual cues for enhancing neighborhood features, respectively. Second, we propose a Geometry-Aware Compound Aggregation (GACA) block, which introduces a functionally complementary compound pooling strategy to aggregate richer neighborhood features. To fully learn the neighborhood distribution, we absorb the geometric location information during the aggregation process. The proposed module is integrated into an MLP-based large-scale 3D processing architecture, which constitutes a 3D semantic segmentation network called DNFA-Net. Extensive experiments on public datasets containing indoor and outdoor scenes validate the superiority of DNFA-Net.
Minghong Chen, Wenjun Shi, Dongchen Zhu, Jiamao Li
ICTAI6
2022 J-RR: Joint Monocular Depth Estimation and Semantic Edge Detection Exploiting Reciprocal Relations
abstract
Depth estimation and semantic edge detection are two key tasks in computer vision, which have made great progress. To date, how to associatively predict the depth and the semantic edge is rarely explored. In this work, we first propose a flexible two-branch framework that can make the two tasks take advantage of each other, achieving a win-win situation. Specifically, for the semantic edge detection branch, an Enhanced Edge Weighting strategy (EEW) is designed, which learns weight information from the by-product of depth branch, depth edge, to enhance edge perception in features. Meanwhile, we make depth estimation benefit from semantic edge detection through introducing Depth Edge Semantic Classification module (DESC). Furthermore, a double reconstruction (D-reconstruction) approach is presented, together with semantic edge-guided disparity smoothing loss to mitigate the ambiguities of the self-supervised manner for depth estimation. Experiments on the Cityscapes dataset demonstrate that our framework outperforms the state-of-the-art method in depth estimation along with a significant improvement in semantic edge detection.
Deming Wu, Dongchen Zhu, Wenjun Shi, Jiamao Li
IROS6
2022 Spatiotemporally Enhanced Photometric Loss for Self-Supervised Monocular Depth Estimation
abstract
Recovering depth information from a single image is a long-standing challenge, and self-supervised depth estimation methods have gradually attracted attention due to not relying on high-cost ground truth. Constructing an accurate photometric loss based on photometric consistency is crucial for these self-supervised methods to obtain high-quality depth maps. However, the photometric loss in most studies treats all pixels indiscriminately, resulting in poor performance. In this paper, we propose two modules based on the spatial and temporal cues to refine the photometric loss. Delving into the geometric model of photometric consistency, we introduce a depth-aware pixel correspondence module (DPC) inside the monocular depth estimation pipeline. It reduces the uncertainty of photometric errors by applying the homography matrix to the projection of corresponding pixels in far regions instead of the fundamental matrix. Furthermore, we design an omnidirectional auto-masking module (OA) to boost the robustness of our model, which utilizes temporal sequences to generate disturbance poses and hypothetical views to distin-guish dynamic objects with different directions that violate the photometric consistency. Experiments on the KITTI and the Make3d datasets reveal that our framework achieves state-of-the-art performance.
Dongchen Zhu, Wenjun Shi, Jiamao Li
IROS7
2022 EFG-Net: A Unified Framework for Estimating Eye Gaze and Face Gaze Simultaneously
Hekuangyi Che, Dongchen Zhu, Minjing Lin, Wenjun Shi, Jiamao Li
PRCV (1)8
2022 SemRegionNet: Region ensemble 3D semantic instance segmentation network with semantic spatial aware discriminative loss
Dongchen Zhu, Wenjun Shi, Jiamao Li
Neurocomputing4
2022 CSF: Closed-mask-guided semantic fusion method for semantic perception of unknown scenes
Minghong Chen, Ruijun Shu, Dongchen Zhu, Jiamao Li
Pattern Recognit. Lett.4
2022 RGB-D Semantic Segmentation and Label-Oriented Voxelgrid Fusion for Accurate 3D Semantic Mapping
abstract
The 3D semantic map plays an increasingly important role in a wide variety of applications, especially for many kinds of task-driven robots. In this paper, we present a semantic mapping methodology for 3D semantic map obtaining from RGB-D scans. In contrast to existing methods that use 3D annotated information as supervisory, we focus on accurate 2D frame labeling and combine labels in 3D space using semantic fusion mechanism. For scene parsing, a two-stream network with a novel discriminatory mask loss is proposed to explore sufficient extraction and fusion of RGB and depth information achieving steadily semantic segmentation. The discriminatory mask guides the cross-entropy loss function and interprets the influence of different pixels on back-propagation, which reduces the harmful effects of the depth noise or the fallible annotation at the edges of objects. After the correspondences between frames are provided, these semantic frames are fused in unified 3D coordinates using the novel label-oriented voxelgrid filter. It can ensure the intra-frame spatial continuity and the inter-frame spatiotemporal consistency through introducing the label-oriented statistical principle into labeled point clouds. In order to avoid the unfavorable interference between uncorrelated frames, we further propose an adaptive grouping algorithm by applying the view frustum filter to group frames with sufficient overlap as a segment. To this end, we demonstrate the effectiveness of the proposed method on the 2D/3D semantic label benchmark of ScanNetv2 and Cityscapes datasets.
Wenjun Shi, Dongchen Zhu, Xianshun Wang, Jiamao Li
IEEE Trans. Circuits Syst. Video Technol.6
2021 Camera Parameters Aware Motion Segmentation Network with Compensated Optical Flow
abstract
Learning to distinguish independent moving objects from the observed optical flow with a moving camera remains challenging. In this work, we first present a novel camera pose compensation (CPC) scheme. With the help of ingenious geometric analysis, it breaks the observed optical flow into patterns that are easier to interpret for the motion segmentation network. Secondly, we further refine such compensation with a camera parameter aware (CPA) module to account for poses’ errors in the CPC processing and enhance the entire network’s tolerance to noises. Additionally, an MMPNet is developed to intensify the identification ability of overall motion patterns. It reaches a larger receptive field with a bottom-up information transmission structure and integrates motion information at different granularities. We demonstrate the benefits of our framework on FlyingThings3D and Monkaa datasets. Without the complement of semantic information, our approach outperforms the top methods for moving objects segmentation.
Xianshun Wang, Dongchen Zhu, Shaojie Xu, Wenjun Shi, Jiamao Li
IROS6
2021 Contour-Aware Panoptic Segmentation Network
Dongchen Zhu, Wenjun Shi, Jiamao Li
PRCV (2)5
2021 SASO: Joint 3D semantic-instance segmentation via multi-scale semantic association and salient point clustering optimization
abstract
Abstract Jointly performing semantic and instance segmentation of 3D point cloud remains a challenging task. In this work, a novel framework called joint 3D semantic‐instance segmentation via multi‐scale Semantic Association and Salient point clustering Optimization was proposed to tackle this problem. Inspired by the inherent correlation among objects in semantic space, a Multi‐scale Semantic Association (MSA) module to explore the constructive effect of the context information for semantic segmentation is designed. For instance, segmentation, different from previous works utilising clustering only in inference procedure, a Salient Point Clustering Optimization (SPCO) module is put forward to introduce the clustering algorithm into the training phase, which impels the network to focus on points that are difficult to be distinguished. Furthermore, affected by the inherent structure of indoor scenes, the problem of uneven distribution of categories has rarely been considered in the previous work, but it significantly limits the performance of 3D scene perception. To address the issue, an adaptive Water Filling Sampling (WFS) algorithm to balance the category distribution of training data is presented. Extensive experiments on a variety of changing datasets show that the authors’ method outperforms the state‐of‐the‐art methods in both tasks of semantic segmentation and instance segmentation.
Jingang Tan, Kangru Wang, Jiamao Li
IET Comput. Vis.4
2021 HCFS3D: Hierarchical coupled feature selection network for 3D semantic and instance segmentation
Jingang Tan, Kangru Wang, Jiamao Li
Image Vis. Comput.5
2021 A Less-constrained Sclera Recognition Method based on Stem-and-leaf Branches Network
Dongchen Zhu, Jiamao Li, Jingquan Peng, Xianshun Wang
Pattern Recognit. Lett.2
2020 3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature Selection
abstract
We propose a novel fast and robust 3D point clouds segmentation framework via coupled feature selection, named 3DCFS, that jointly performs semantic and instance segmentation. Inspired by the human scene perception process, we design a novel coupled feature selection module, named CFSM, that adaptively selects and fuses the reciprocal semantic and instance features from two tasks in a coupled manner. To further boost the performance of the instance segmentation task in our 3DCFS, we investigate a loss function that helps the model learn to balance the magnitudes of the output embedding dimensions during training, which makes calculating the Euclidean distance more reliable and enhances the generalizability of the model. Extensive experiments demonstrate that our 3DCFS outperforms state-of-the-art methods on benchmark datasets in terms of accuracy, speed and computational cost. Codes are available at: https://github.com/Biotan/3DCFS.
Liang Du 0004, Jingang Tan, Xiangyang Xue 0001, Hongkai Wen 0001, Jianfeng Feng, Jiamao Li
ICRA7
2020 Richer Aggregated Features for Optical Flow Estimation with Edge-aware Refinement
abstract
Recent CNN-based optical flow approaches have a separated structure of feature extraction and flow estimation. The core task of optical flow is finding the corresponding points while rich representation is just the key part of such matching problems. However, the prior work usually pays more attention to the design of flow decoder than the feature extraction. In this paper, we present a novel optical flow estimation network to enrich the feature representation of each pyramid level, with a hierarchical dilated architecture and a bottom-up aggregation scheme. In addition, inspired by edge guided classical methods, we bring the edge-aware idea into our approach and propose an edge-aware refinement (EAR) subnetwork to handle motion boundaries. Using the same decoding structure as PWC-Net, our network outperforms it by a large margin and leads all its derivatives both on KITTI-2012 and KITTI-2015. Further performance analysis proves the effectiveness of proposed ideas.
Xianshun Wang, Dongchen Zhu, Jiafei Song, Jiamao Li
IROS5
2020 RegionNet: Region-feature-enhanced 3D Scene Understanding Network with Dual Spatial-aware Discriminative Loss
abstract
Neural networks have recently achieved impressive success in semantic and instance segmentation on 2D images. However, their capabilities have not been fully explored to address semantic instance segmentation on unstructured 3D point cloud data. Digging into the regional feature representation to boost point cloud comprehension, we propose a region-feature-enhanced structure consisting of adaptive regional feature complementary (ARFC) module and affinity-based regional relational reasoning (AR3) module. The ARFC module aims to complement low-level features of sparse regions adaptively. The AR3module emphasizes on mining the potential reasoning relationships between high-level features based on affinity. Both the ARFC and AR3modules are plug-and-play. Besides, a novel dual spatial-aware discriminative loss is proposed to improve the discrimination of instance embedding. Our proposal-free point cloud instance segmentation network (RegionNet) equipped with the region-feature-enhanced structure and dual spatial-aware discriminative loss achieves state-of-the-art performance on S3DIS dataset and ScanNet-v2 dataset.
Dongchen Zhu, Xiaoqing Ye, Wenjun Shi, Minghong Chen, Jiamao Li
IROS6
2018 3D Recurrent Neural Networks with Context Fusion for Point Cloud Semantic Segmentation
Xiaoqing Ye, Jiamao Li, Hexiao Huang, Liang Du 0004
ECCV (7)2
2017 Order-Based Disparity Refinement Including Occlusion Handling for Stereo Matching
abstract
Accurate stereo matching is still challenging in case of weakly textured areas, discontinuities, and occlusions. Besides, occlusion recovery is often regarded as a subordinate problem and simply handled. To obtain dense high-accuracy depth maps, this letter proposes an efficient multistep disparity refinement framework with occlusion handling. The framework is implemented by classifying the outliers into leftmost occlusions, nonborder occlusions, as well as mismatches, and employing different strategies to recover them. To recover occlusions, a filling order is specially introduced to avoid error propagation and surface decision based on local image content is performed when more than one background surface exists. The evaluations on Middlebury datasets and comparisons with other refinement algorithms show the superiority and robustness of our method.
Xiaoqing Ye, Yuzhang Gu, Jiamao Li
IEEE Signal Process. Lett.4