Jianxin Pang

dblp:41/813 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-3985-5802ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Complexity-Aware Policy via Heterogeneous Experts for Robotic Manipulation
abstract
Robotic manipulation requires learning a generalizable policy that can adapt to complicated new environments. However, existing methods typically overlook the inherent task complexity and employ a policy with the same budget for tasks with varied difficulties, facing challenges in inefficient computational resource allocation and zero-shot generalization. In this work, we identify three facets of complexity imbalance issues in the current manipulation tasks at the Inter-task, Intra-task, and Noise-timesteps levels. To address this gap, we introduce the Complexity-Aware Policy (CAP), a novel approach integrating flow matching with a Transformer-based backbone and a Mixture of Heterogeneous Experts (MoHE) structure for policy learning. By leveraging Rectified Flow and dynamically adjusting model capacity based on task complexity, which is assessed through features like object counts and precision needs, our method allocates computational resources efficiently and effectively. This results in faster convergence, optimized computational resource usage, and improved precision across diverse manipulation tasks. Our proposed method achieves the state-of-the-art performance on widely-used CALVIN, LIBERO, and SimplerEnv benchmarks, and is further validated through six real-world experiments, where it consistently outperforms baseline methods across all tasks.
Yunzhe Hu, Ge Yuan, Difan Zou, Jianxin Pang, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Free Your Hands: Human-Demonstration-Free Diffusion Policy Adaptation for Robotic Manipulation
abstract
Diffusion-based policies demonstrate remarkable capabilities in generating robust and precise actions for embodied tasks. However, when adapting to a new environment with visual domain gap such as lighting variation, object appearance change, different background, and unseen distractors, these methods require a substantial amount of human demonstrations gathered through teleoperation. This data collection process incurs significant costs in terms of both time and financial resources. To this end, we introduceFreeHand, which enables cross-environment adaptation without requiring additional teleoperated demonstrations,freeing human handsfrom the time-consuming data collection process. Specifically, we find that a well-trained diffusion policy network struggles to adapt to new environments due to distribution mismatches at both the visual and policy levels. Simply aligning domains at the visual level is insufficient, as even subtle visual changes in the environment can lead to severe action failures. Therefore, we propose a two-level alignment scheme. The visual-level alignment is achieved through adversarial training between the visual features from old and new environments. For the policy-level alignment, we introduce a novel noise-aware adversarial learning strategy for the diffusion policy. We validate our approach through comprehensive cross-domain adaptation experiments under three settings: Sim2Sim (PushT, CALVIN, RoboTwin), Real2Sim (SimplerEnv), and Real2Real (real-world robotic deployment). Our method demonstrates significant improvement with minimal additional parameter overhead, showcasing its effectiveness and efficiency.
Ge Yuan, Jing Zhang 0017, Jianxin Pang, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Trustworthy Driver State Perception via Contextual Interaction-Driven Evidential Vision-Language Fusion in Vehicular Cyber-Physical Systems
abstract
A vision-driven driver monitoring system plays a vital role of vehicular cyber-physical systems (VCPS) to guarantee the driving safety. Recent advances focus on modeling a deep learning-based method to realize the driver monitoring system, which benefits from the powerful capability of data-driven feature extraction. Although the acceptable performances of driver state monitoring methods are achieved, there is still a gap between the emerged techniques and actual application scenarios. First, the human-centric visual appearances are not involved comprehensively to represent the driver states, resulting in ignoring the contextual interaction of behaviors. Second, the inherent uncertainty of driver situation is not considered, while the unreliable samples would lead to the untrustworthy results. In this paper, we focus on a vision-based driver state monitoring method, where a trustworthy driver state perception (TDSP) is proposed via human-centric contextual interaction-driven evidential vision-language fusion in VCPS. Specifically, a vision-language model-based architecture is first modified in temporal dimension to represent the visual human-centric contextual interactions, while a vision-language consistency loss is designed to mitigate the gap between visual and textual representations. Then, an evidence-based learning method is introduced to jointly conduct the classification and uncertainty estimation for driver states. Furthermore, to model the human-centric contextual interactions towards the evidence-based paradigm comprehensively, Dempster-Shafer theory-based combination rule is introduced to fuse the visual and textual representations. Extensive experiments are conducted on two public benchmarks, where the superiority of TDSP is demonstrated compared with the state-of-the-art methods. The superior performance of TDSP to recognize dangerous states are 85.41% and 83.12% in terms of accuracy and F1 score, which outperforms the state-of-the-art methods by 4.68% and 3.99%. Moreover, we validate the reliability of TDSP against the noisy data for VCPS. The code will be public at https://github.com/w64228013/TDSP.
Chuanfei Hu, Xinde Li, Jianxin Pang
IEEE Trans. Intell. Transp. Syst.3
2025 On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
abstract
Diffusion Policies have significantly advanced robotic manipulation tasks via imitation learning, but their application on resource-constrained mobile platforms remains challenging due to computational inefficiency and extensive memory footprint. In this paper, we propose LightDP, a novel framework specifically designed to accelerate Diffusion Policies for real-time deployment on mobile devices. LightDP addresses the computational bottleneck through two core strategies: network compression of the denoising modules and reduction of the required sampling steps. We first conduct an extensive computational analysis on existing Diffusion Policy architectures, identifying the denoising network as the primary contributor to latency. To overcome performance degradation typically associated with conventional pruning methods, we introduce a unified pruning and retraining pipeline, optimizing the model's post-pruning recoverability explicitly. Furthermore, we combine pruning techniques with consistency distillation to effectively reduce sampling steps while maintaining action prediction accuracy. Experimental evaluations on the standard datasets, \ie, PushT, Robomimic, CALVIN, and LIBERO, demonstrate that LightDP achieves real-time action prediction on mobile devices with competitive performance, marking an important step toward practical deployment of diffusion-based policies in resource-limited environments. Extensive real-world experiments also show the proposed LightDP can achieve performance comparable to state-of-the-art Diffusion Policies.
Yiming Wu 0005, Huan Wang 0001, Jianxin Pang, Dong Xu 0001
ICCV4
2025 The Sampling-Gaussian for Stereo Matching
abstract
The soft-argmax operation is widely adopted in neural network-based stereo matching methods to enable differentiable regression of disparity. However, networks trained with soft-argmax tend to predict multimodal probability distributions due to the absence of explicit constraints on the shape of the distribution. Previous methods leveraged Laplacian distributions and cross-entropy for training but failed to effectively improve accuracy and even increased the network’s processing time. In this paper, we propose a novel method called Sampling-Gaussian as a substitute for soft-argmax. It improves accuracy without increasing inference time. We innovatively interpret the training process as minimizing the distance in vector space and propose a combined loss of L1 loss and cosine similarity loss. We leveraged the normalized discrete Gaussian distribution for supervision. Moreover, we identified two issues in previous methods and proposed extending the disparity range and employing bilinear interpolation as solutions. We have conducted comprehensive experiments to demonstrate the superior performance of our Sampling-Gaussian method. The experimental results prove that we have achieved better accuracy on five baseline methods across four datasets. Moreover, we have achieved significant improvements on small datasets and models with weaker generalization capabilities. Our method is easy to implement, and the code is available online.
Baiyu Pan, Jichao Jiao, Jianxin Pang, Jun Cheng 0002
IROS4
2025 An Efficient Hand Grasping Method Based on CVAE for Target Pose Estimation
abstract
With the advancement of humanoid robot industrialization, dexterous grasping has become a critical research area. Traditional two-finger grippers, while effective for regular geometries, struggle with complex shapes. Multi-finger dexterous hands offer significant advantages for adapting to diverse objects. This study proposes a grasp pose estimation method for multi-finger dexterous hands utilizing a Conditional Variational Autoencoder (CVAE) framework. Point cloud data and grasp poses from industrial components are used as inputs to the CVAE, with K-Nearest Neighbors (KNN) employed to enhance local feature extraction. Experimental results show that the proposed method achieves robust generalization and stability across various grasping scenarios.
Pengpeng Xu, Huaxi Zhang 0004, Wenlong Qin, Jianxin Pang, Jun Cheng 0002
IWCMC5
2023 DST: Deformable Speech Transformer for Emotion Recognition
abstract
Enabled by multi-head self-attention, Transformer has exhibited remarkable results in speech emotion recognition (SER). Compared to the original full attention mechanism, window-based attention is more effective in learning fine-grained features while greatly reducing model redundancy. However, emotional cues are present in a multi-granularity manner such that the pre-defined fixed window can severely degrade the model flexibility. In addition, it is difficult to obtain the optimal window settings manually. In this paper, we propose a Deformable Speech Transformer, named DST, for SER task. DST determines the usage of window sizes conditioned on in-put speech via a light-weight decision network. Meanwhile, data-dependent offsets derived from acoustic features are utilized to adjust the positions of the attention windows, allowing DST to adaptively discover and attend to the valuable in-formation embedded in the speech. Extensive experiments on IEMOCAP and MELD demonstrate the superiority of DST.
Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang
ICASSP4
2023 SpeechFormer++: A Hierarchical Efficient Framework for Paralinguistic Speech Processing
abstract
Paralinguistic speech processing is important in addressing many issues, such as sentiment and neurocognitive disorder analyses. Recently, Transformer has achieved remarkable success in the natural language processing field and has demonstrated its adaptation to speech. However, previous works on Transformer in the speech field have not incorporated the properties of speech, leaving the full potential of Transformer unexplored. In this paper, we consider the characteristics of speech and propose a general structure-based framework, called SpeechFormer++, for paralinguistic speech processing. More concretely, following the component relationship in the speech signal, we design a unit encoder to model the intra- and inter-unit information (i.e., frames, phones, and words) efficiently. According to the hierarchical relationship, we utilize merging blocks to generate features at different granularities, which is consistent with the structural pattern in the speech signal. Moreover, a word encoder is introduced to integrate word-grained features into each unit encoder, which effectively balances fine-grained and coarse-grained information. SpeechFormer++ is evaluated on the speech emotion recognition (IEMOCAP & MELD), depression classification (DAIC-WOZ) and Alzheimer's disease detection (Pitt) tasks. The results show that SpeechFormer++ outperforms the standard Transformer while greatly reducing the computational cost. Furthermore, it delivers superior results compared to the state-of-the-art approaches.
Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Context Sensing Attention Network for Video-based Person Re-identification
abstract
Video-based person re-identification (ReID) is challenging due to the presence of various interferences in video frames. Recent approaches handle this problem using temporal aggregation strategies. In this work, we propose a novel Context Sensing Attention Network (CSA-Net), which improves both the frame feature extraction and temporal aggregation steps. First, we introduce the Context Sensing Channel Attention (CSCA) module, which emphasizes responses from informative channels for each frame. These informative channels are identified with reference not only to each individual frame, but also to the content of the entire sequence. Therefore, CSCA explores both the individuality of each frame and the global context of the sequence. Second, we propose the Contrastive Feature Aggregation (CFA) module, which predicts frame weights for temporal aggregation. Here, the weight for each frame is determined in a contrastive manner: i.e., not only by the quality of each individual frame, but also by the average quality of the other frames in a sequence. Therefore, it effectively promotes the contribution of relatively good frames. Extensive experimental results on four datasets show that CSA-Net consistently achieves state-of-the-art performance.
Kan Wang 0004, Changxing Ding, Jianxin Pang, Xiangmin Xu 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 RA Loss: Relation-Aware Loss for Robust Person Re-identification
Kan Wang 0004, Shuping Hu, Jun Cheng 0002, Jianxin Pang, Huan Tan
ACCV (2)4
2022 Key-Sparse Transformer for Multimodal Speech Emotion Recognition
abstract
Speech emotion recognition is a challenging research topic that plays a critical role in human-computer interaction. Multimodal inputs further improve the performance as more emotional information is used. However, existing studies learn all the information in the sample while only a small portion of it is about emotion. The redundant information will become noises and limit the system performance. In this paper, a key-sparse Transformer is proposed for efficient emotion recognition by focusing more on emotion related information. The proposed method is evaluated on the IEMOCAP and LSSED. Experimental results show that the proposed method achieves better performance than the state-of-the-art approaches.
Xiaofeng Xing, Xiangmin Xu 0001, Jianxin Pang
ICASSP5
2022 SpeechFormer: A Hierarchical Efficient Framework Incorporating the Characteristics of Speech
abstract
Transformer has obtained promising results on cognitive speech signal processing field, which is of interest in various applications ranging from emotion to neurocognitive disorder analysis.However, most works treat speech signal as a whole, leading to the neglect of the pronunciation structure that is unique to speech and reflects the cognitive process.Meanwhile, Transformer has heavy computational burden due to its full attention operation.In this paper, a hierarchical efficient framework, called SpeechFormer, which considers the structural characteristics of speech, is proposed and can be served as a generalpurpose backbone for cognitive speech signal processing.The proposed SpeechFormer consists of frame, phoneme, word and utterance stages in succession, each performing a neighboring attention according to the structural pattern of speech with high computational efficiency.SpeechFormer is evaluated on speech emotion recognition (IEMOCAP & MELD) and neurocognitive disorder detection (Pitt & DAIC-WOZ) tasks, and the results show that SpeechFormer outperforms the standard Transformer-based framework while greatly reducing the computational cost.Furthermore, our SpeechFormer achieves comparable results to the state-of-the-art approaches.
Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang
INTERSPEECH4
2022 Triplet Ratio Loss for Robust Person Re-identification
Shuping Hu, Kan Wang 0004, Jun Cheng 0002, Huan Tan, Jianxin Pang
PRCV (1)5
2021 Reachability-based Push Recovery for Humanoid Robots with Variable-Height Inverted Pendulum
abstract
This paper studies push recovery for humanoid robots based on a variable-height inverted pendulum (VHIP) model. We first develop an approach for treating zero-step capturability of the VHIP with a novel methodology based on Hamilton-Jacobi (HJ) reachability analysis. Such an approach uses the sub-zero level set of a value function to encode capturability of the VHIP, where the value function is obtained by numerically solving a HJ variational inequality offline. Based on this analysis, a simple and effective method for adjusting foothold locations is then devised for cases where the VHIP state is not zero-step capturable. In addition, the HJ reachability analysis naturally induces an optimal control law that allows for rapid planning with the VHIP during push recovery online. To enable use of the strategy with a position-controlled humanoid robot, an associated differential inverse kinematics based tracking controller is employed. The effectiveness of the overall framework is demonstrated with the UBTECH Walker robot in the MuJoCo simulator. Simulation validations show a significant improvement in push robustness as compared to the methods based on the classical linear inverted pendulum model.
Shunpeng Yang, Hua Chen 0007, Zhefeng Cao, Patrick M. Wensing, Yizhang Liu, Jianxin Pang, Wei Zhang 0013
ICRA7
2021 A Capturability-based Control Framework for the Underactuated Bipedal Walking *
abstract
This work considers the control of underactuated bipedal walking, and a novel capturability-based control framework is presented. Compared with traditional approaches, the presented control method does not rely on the use of the Poincaré map, which may take significant computational cost. Firstly, a new definition of stable walking is presented, and a novel foot-placement based control method is proposed. Then, a controller design method is presented based on this control method. For the controller design, the foot placement adjustment is achieved by updating the virtual constraints using a heuristic method, and an improved virtual constraint control method is proposed to enforce the virtual constraints. Finally, the effectiveness of the presented control framework is illustrated on a five-link underactuated planar biped by numerical simulations.
Haihui Yuan, Sumian Song, Ruilong Du, Shiqiang Zhu, Jason Gu, Mingguo Zhao, Jianxin Pang
ICRA7
2018 Person re-identification by discriminant analytical least squares metric learning
Zhao Yang 0001, Fei Dai 0002, Jianxin Pang, Dapeng Tao
Mach. Vis. Appl.4
2018 Deep Multi-View Feature Learning for Person Re-Identification
abstract
Person re-identification aims to identify the same pedestrians across different camera views at different locations. This important yet difficult intelligent video analysis problem remains a vigorous area of research due to demands for performance improvements. Person re-identification involves two main steps: feature representation and metric learning. Handcrafted features, such as color and texture histograms, are frequently used for person re-identification, but most handcrafted features are limited by not being directly applicable to practical problems. Deep learning methods have obtained the state-of-the-art performance in a wide variety of applications, including image annotation, face recognition, and speech recognition. However, deep learning features are heavily dependent on large-scale labeling of samples. In this paper, by utilizing the Cross-view Quadratic Discriminant Analysis (XQDA) metric learning, we propose a novel scheme called deep multi-view feature learning (DMVFL), which exploits the collaboration between handcrafted and deep learning features in a simple but effective way. Furthermore, we prove that the XQDA is a robust algorithm. Extensive experiments on two challenging person re-identification data sets (VIPeR and GRID) demonstrate that DMVFL improves on current state-of-the-art methods.
Dapeng Tao, Yanan Guo 0003, Baosheng Yu, Jianxin Pang, Zhengtao Yu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2017 Cauchy Estimator Discriminant Learning for RGB-D Sensor-based Scene Classification
Dapeng Tao, Xipeng Yang, Weifeng Liu 0001, Shuifa Sun, Yanan Guo 0003, Jianxin Pang
Multim. Tools Appl.7
2014 Directional projection based image fusion quality metric
Richang Hong, Wenyi Cao, Jianxin Pang
Inf. Sci.3
2013 Real-time hand detection based on multi-stage HOG-SVM classifier
abstract
In this paper, we propose a real-time hand detection method with multi-stage HOG-SVM classifier. Unlike traditional methods based on learning which make decomposition of feature vector or combination of different types of features or classifiers, upon the division of background into several categories, we propose a multi-stage classifier which combines several SVM classifies each of which is trained to distinguish corresponding divisions of background and target. Furthermore, in order to improve speed performance, skin color information and integral histogram are also applied. Experiment results demonstrate that the proposed algorithm works well under multiple challenging backgrounds in real-time speed (16 frames per second).
Jun Cheng 0002, Jianxin Pang
ICIP3
2009 Image Fusion Quality Metrics by Directional Projection
abstract
Image fusion has been over-studied recently. Nevertheless, few works aim to how to evaluate the performance of image fusion algorithms. In this paper, we extend the work in image quality evaluation to a novel metric for objective evaluation of image fusion. Firstly the input images and the result image are converted into local sensitive intensity (LSI) by Radon transform. Then we use the sensitive intensity to measure how many information have been transferred from each source into the fused result by the difference of LSI. Finally all the LSI pairs are incorporated into the expression according to Weber-Fechner law. Experimental results demonstrate that our proposed metric is compliant with subjective evaluations and outperforms other recently developed objective metrics of image fusion.
Richang Hong, Yan Song 0001, Jinhui Tang 0001, Jianxin Pang
SMC4
2008 Image quality assessment metrics with Radon transform
abstract
Objective image quality assessment (QA) is a fundamental and challenging job in image processing which evaluates the image quality consistently with human perception automatically. Generally, an image can be segmented into two kinds of areas: structure and texture. And structural information plays a much more important role between the two. The pixels, edges and shape with directional characteristic contribute much more to the structural information. In this paper, the structural information is modeled as the energy of directional projection, which is shown in a directional projection-based map that is built by Radon transform. On the assumption that any image's distortion could be modeled as the difference between the directional projection-based maps of reference and distortion images, we propose a new objective quality assessment method with Radon transform for full reference model, and we also try to explore the feasibility to develop the image QA metrics with scalable efficiency and computational cost. Experimental results show that the proposed metrics are well consistent with the subjective quality score.
Jianxin Pang, Rong Zhang 0004, Zhengkai Liu
SMC1
2007 Quality Assessment for Image Coding Based on Matching Pursuit
abstract
Valuation of image coding relies on not only the efficiency of the coding, but also the quality of the coded image. We present a new objective quality assessment metric for image coding based on matching pursuit. First of all, we get the characteristics of the most important structure by projecting the reference image onto the base functions from a dictionary using matching pursuit. Secondly, we process the reference image and gain the structure information of the images in the order of importance, projecting the images onto the structural characteristics. Finally the objective quality score is given by comparing the differences of structure information between the reference and coded images. Experimental results show that the proposed approach is well consistent with the subjective quality score.
Jianxin Pang, Rong Zhang 0004, Lu Lu 0001, Zhengkai Liu
ICME1