Yaxin Peng

dblp:20/7643 · DBLP profile ↗
← Back
55ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0002-2983-555XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 3 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Interpretable feature modeling for robust color watermarking in the quaternion framework
Yong Chen 0019, Zhigang Jia, Hao Peng 0002, Yaxin Peng, Yan Peng 0001
Expert Syst. Appl.4
2026 STFNet: A Knowledge-Guided Spatial-Temporal Fusion Network for Low-SNR Modulation Recognition
abstract
Automatic modulation classification (AMC) in complex electromagnetic environments is essential for ensuring reliable spectrum surveillance. However, most existing AMC algorithms primarily rely on the data engineer and focus on improving accuracy under high signal-to-noise ratio (SNR) conditions, making it difficult to maintain robust and accurate performance in real-world unstable SNR scenarios. In order to address this challenge, a knowledge-guided spatial-temporal fusion network for low-SNR modulation recognition is proposed, named STFNet. Firstly, aiming at the instability of the model classification caused by the single input mode, the enhanced constellation diagram and the I/Q signal are simultaneously introduced as the inputs of the STFNet. Furthermore, the STFNet is designed as a dual-path adaptive multimodal fusion architecture to simultaneously exploit the spatial features of the enhanced constellation diagram and the temporal features of the I/Q signal. Secondly, an AMC-specific transfer learning (AMC-TL) strategy is introduced to enhance the global robustness of spatial representations through contrastive learning. More importantly, a domain knowledge-guided mixture of experts (DKG-MoE) is proposed to incorporate traditional features into the expert routing process, which improves temporal recognition accuracy. Then, an adaptive modality attention (AMA) module is developed to balance the accuracy under high SNR and the robustness under low SNR. Finally, extensive experiments on four benchmark datasets demonstrate that the proposed method achieves a 4.13% accuracy improvement under low SNR conditions (-20 dB ∼ 0 dB) and outperforms all state-of-the-art (SOTA) methods in overall average accuracy (65.57%). The code is publicly available at https://github.com/yoho78/STFNet.
Hongyu Wei, Hanqian Mo, Yunpeng Chen, Pingfan Wu, Yaxin Peng, Hao Kong 0004
IEEE Internet Things J.5
2026 Fast quaternion QR algorithm: Advancing watermarking with multifaceted capabilities
Yong Chen 0019, Zhigang Jia, Hao Peng 0002, Yaxin Peng, Yan Peng 0001
Signal Process.4
2026 Evidential Prior Guided Neural Collapse for Open World Object Detection
abstract
Open World Object Detection (OWOD) faces a fundamental dilemma: maintaining a stable representation for known classes while reserving flexible space for discovering unknown objects. Existing methods, while improving recall, often fail to assign discriminative confidence scores to unknown instances, resulting in critically low Average Precision (U-AP) and representation degradation during incremental learning. To remedy this, we propose the Evidential Prior Guided Neural Collapse (ENC) framework. ENC unifies representation learning and uncertainty quantification via a Geometric-Evidence Coupling mechanism. Unlike previous approaches, we map evidential support directly to the angular alignment with Simplex Equiangular Tight Frame (ETF) prototypes. Theoretically, the evidential prior functions as a geometric regularizer: it maximizes equiangular separation for confident known samples, while constraining ambiguous queries to approximate an isotropic uniform distribution via distributional regularization. Furthermore, to mitigate decision conflicts in self-supervised learning, we propose a dissonance-aware objectness optimization strategy that mines informative samples near the decision boundary. Extensive experiments on M-OWODB and S-OWODB benchmarks demonstrate that ENC sets a new state-of-the-art. Notably, it achieves a significant improvement in unknown class discovery, boosting U-AP from ≈ 1% to 9.2%, while exhibiting superior robustness against catastrophic forgetting in challenging incremental scenarios.
Kewen Xia, Xiaodong Yue 0002, Wei Liu 0303, Jianxiang Zhu, Yaxin Peng
IEEE Trans. Circuits Syst. Video Technol.7
2025 The Adaptive Q-Network for Recommendation Tasks with Dynamic Item Space
abstract
Reinforcement learning (RL) algorithms can improve recommendation performance by capturing long-term user-system interaction. However, current RL-based recommendation tasks seldom consider the dynamism of the environment, and standard RL algorithms are ineffective in recommending items dynamically. In addressing these issues, we design a novel task termed dynamic recommendation, which takes the emergence of real-world recommendable items into consideration. Meanwhile, we propose Adaptive Q-Network (AdaQN) to tackle the dynamic recommendation task. Firstly, AdaQN predicts the value of different action characteristics, particularly during the testing phase, which can capture emerging new action characteristics. The above procedure helps AdaQN in effectively adapting to the dynamic action space. Secondly, AdaQN establishes a stable mapping that projects the discrete action space onto a continuous characteristic space. Finally, AdaQN employs a lightweight Q-network design, which mitigates the complexity of the optimization process. Extensive experiments demonstrate that our approach has achieved state-of-the-art performance in the dynamic recommendation task.
Jianxiang Zhu, Dandan Lai, Zhongcui Ma, Yaxin Peng
AAAI4
2025 A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
abstract
Multimodal Large Language Models (MLLMs) have showcased impressive skills in tasks related to visual understanding and reasoning. Yet, their widespread application faces obstacles due to the high computational demands during both the training and inference phases, restricting their use to a limited audience within the research and user communities. In this paper, we investigate the design aspects of Multimodal Small Language Models (MSLMs) and propose an efficient multimodal assistant named Mipha, which is designed to create synergy among various aspects: visual representation, language models, and optimization strategies. We show that without increasing the volume of training data, our Mipha-3B outperforms the state-of-the-art large MLLMs, especially LLaVA-1.5-13B, on multiple benchmarks. Through detailed discussion, we provide insights and guidelines for developing strong MSLMs that rival the capabilities of MLLMs.
Minjie Zhu, Yichen Zhu 0001, Ning Liu 0007, Xin Liu 0086, Chaomin Shen 0001, Yaxin Peng
AAAI7
2025 ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
abstract
Zhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Yaxin Peng, Chaomin Shen, Feifei Feng, Yi Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhongyi Zhou, Yichen Zhu 0001, Minjie Zhu, Ning Liu 0007, Weibin Meng, Yaxin Peng, Chaomin Shen 0001, Feifei Feng
EMNLP8
2025 Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
abstract
Continual Learning enables models to learn and adapt to new tasks while retaining prior knowledge. Introducing new tasks, however, can naturally lead to feature entanglement across tasks, limiting the model’s capability to distinguish between new domain data. In this work, we propose a method called Feature Realignment through Experts on hyperSpHere in Continual Learning (Fresh-CL). By leveraging predefined and fixed simplex equiangular tight frame (ETF) classifiers on a hypersphere, our model improves feature separation both intra and inter tasks. However, the projection to a simplex ETF shifts with new tasks, disrupting structured feature representation of previous tasks and degrading performance. Therefore, we propose a dynamic extension of ETF through mixture of experts, enabling adaptive projections onto diverse subspaces to enhance feature representation. Experiments on 11 datasets demonstrate a 2% improvement in accuracy compared to the strongest baseline, particularly in fine-grained datasets, confirming the efficacy of combining ETF and MoE to improve feature distinction in continual learning scenarios.
Zhongyi Zhou, Yaxin Peng, Pin Yi, Minjie Zhu, Chaomin Shen 0001
ICASSP2
2025 CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
Yichen Zhu 0001, Zhibin Tang, Minjie Zhu, Chengmeng Li, Yaxin Peng, Yan Peng 0001, Feifei Feng
ICCV9
2025 DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression
abstract
In this paper, we present DiffusionVLA, a novel framework that integrates autoregressive reasoning with diffusion policies to address the limitations of existing methods: while autoregressive Vision-Language-Action (VLA) models lack precise and robust action generation, diffusion-based policies inherently lack reasoning capabilities. Central to our approach is autoregressive reasoning — a task decomposition and explanation process enabled by a pre-trained VLM — to guide diffusion-based action policies. To tightly couple reasoning with action generation, we introduce a reasoning injection module that directly embeds self-generated reasoning phrases into the policy learning process. The framework is simple, flexible, and efficient, enabling seamless deployment across diverse robotic platforms. We conduct extensive experiments using multiple real robots to validate the effectiveness of DiVLA. Our tests include a challenging factory sorting task, where DiVLA successfully categorizes objects, including those not seen during training. The reasoning injection module enhances interpretability, enabling explicit failure diagnosis by visualizing the model’s decision process. Additionally, we test DiVLA on a zero-shot bin-picking task, achieving \textbf{63.7\% accuracy on 102 previously unseen objects}. Our method demonstrates robustness to visual changes, such as distractors and new backgrounds, and easily adapts to new embodiments. Furthermore, DiVLA can follow novel instructions and retain conversational ability. Notably, DiVLA is data-efficient and fast at inference; our smallest DiVLA-2B runs 82Hz on a single A6000 GPU. Finally, we scale the model from 2B to 72B parameters, showcasing improved generalization capabilities with increased model size.
Yichen Zhu 0001, Minjie Zhu, Zhibin Tang, Zhongyi Zhou, Chaomin Shen 0001, Yaxin Peng, Feifei Feng
ICML9
2025 Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation
abstract
Diffusion Policy is a powerful technique tool for learning end-to-end visuomotor robot control. It is expected that Diffusion Policy possesses scalability, a key attribute for deep neural networks, typically suggesting that increasing model size would lead to enhanced performance. However, our observations indicate that Diffusion Policy in transformer architecture (DP-T) struggles to scale effectively; even minor additions of layers can deteriorate training outcomes. To address this issue, we introduce Scalable Diffusion Transformer Policy for visuomotor learning. Our proposed method, namely ScaleDP, introduces two modules that improve the training dynamic of Diffusion Policy and allow the network to better handle multimodal action distribution. First, we identify that DPT suffers from large gradient issues, making the optimization of Diffusion Policy unstable. To resolve this issue, we factorize the feature embedding of observation into multiple affine layers, and integrate it into the transformer blocks. Additionally, our utilize non-causal attention which allows the policy network to “see” future actions during prediction, helping to reduce compounding errors. We demonstrate that our proposed method successfully scales the Diffusion Policy from 10 million to 1 billion parameters. This new model, named ScaleDP, can effectively scale up the model size with improved performance and generalization. We benchmark ScaleDP across 50 different tasks from MetaWorld and find that our largest ScaleDP outperforms DP-T with an average improvement of 21.6%. Across 7 real-world robot tasks, our ScaleDP demonstrates an average improvement of 36. 25% over DP-T on four single-arm tasks and 75% on three bimanual tasks. We believe our work paves the way for scaling up models for visuomotor learning. The project page is available at https://scaling-diffusion-policy.github.io/.
Minjie Zhu, Yichen Zhu 0001, Ning Liu 0007, Chaomin Shen 0001, Yaxin Peng, Feifei Feng, Jian Tang 0008
ICRA9
2025 Efficient Feature Fusion for UAV Object Detection
abstract
Object detection in unmanned aerial vehicle (UAV) remote sensing images poses significant challenges due to unstable image quality, small object sizes, complex backgrounds, and environmental occlusions. Small objects, in particular, occupy small portions of images, making their accurate detection highly difficult. Existing multi-scale feature fusion methods address these challenges to some extent by aggregating features across different resolutions. However, they often fail to effectively balance the classification and localization performance for small objects, primarily due to insufficient feature representation and imbalanced network information flow. In this paper, we propose a novel feature fusion framework specifically designed for UAV object detection tasks to enhance both localization accuracy and classification performance. The proposed framework integrates hybrid upsampling and downsampling modules, enabling feature maps from different network depths to be flexibly adjusted to arbitrary resolutions. This design facilitates cross-layer connections and multi-scale feature fusion, ensuring improved representation of small objects. Our approach leverages hybrid downsampling to enhance fine-grained feature representation, improving spatial localization of small targets, even under complex conditions. Simultaneously, the upsampling module aggregates global contextual information, optimizing feature consistency across scales and enhancing classification robustness in cluttered scenes. Experimental results on two public UAV datasets demonstrate the effectiveness of the proposed framework. Integrated into the YOLO-v10 model, our method achieves a 2 percentage points improvement in average precision (AP) compared to the baseline YOLO-v10 model, while maintaining the same number of parameters. These results highlight the potential of our framework for accurate and efficient UAV object detection.
Yaxin Peng, Chaomin Shen 0001
IJCNN2
2025 ChatVLA-2: Vision-Language-Action Model with Open-World Reasoning
abstract
Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks. We argue that a generalizable VLA model should retain and expand upon the VLM's core competencies: 1) **Open-world reasoning** - the VLA should inherit the knowledge from VLM, i.e., recognize anything that the VLM can recognize, capable of solving math problems, possessing visual-spatial intelligence, 2) **Reasoning following** – effectively translating the open-world reasoning into actionable steps for the robot. In this work, we introduce **ChatVLA-2**, a novel mixture-of-expert VLA model coupled with a specialized three-stage training pipeline designed to preserve the VLM’s original strengths while enabling actionable reasoning. To validate our approach, we design a math-matching task wherein a robot interprets math problems written on a whiteboard and picks corresponding number cards from a table to solve equations. Remarkably, our method exhibits exceptional mathematical reasoning and OCR capabilities, despite these abilities not being explicitly trained within the VLA. Furthermore, we demonstrate that the VLA possesses strong spatial reasoning skills, enabling it to interpret novel directional instructions involving previously unseen objects. Overall, our method showcases reasoning and comprehension abilities that significantly surpass state-of-the-art imitation learning methods such as OpenVLA, DexVLA, and $\pi_0$. This work represents a substantial advancement toward developing truly generalizable robotic foundation models endowed with robust reasoning capacities.
Zhongyi Zhou, Yichen Zhu 0001, Zhibin Tang, Yaxin Peng, Chaomin Shen 0001
NeurIPS6
2025 SAB Net: A Semantic Attention Boosting Framework for Semantic Segmentation
abstract
Semantic segmentation has achieved great progress by effectively fusing features of contextual information. In this article, we propose an end-to-end semantic attention boosting (SAB) framework to adaptively fuse the contextual information iteratively across layers with semantic regularization. Specifically, we first propose a pixelwise semantic attention (SAP) block, with a semantic metric representing the pixelwise category relationship, to aggregate the nonlocal contextual information. In addition, we improve the computation complexity of SAP block from to for images with size . Second, we present a categorywise semantic attention (SAC) block to adaptively balance the nonlocal contextual dependencies and the local consistency with a categorywise weight, overcoming the contextual information confusion caused by the feature imbalance within intra-category. Furthermore, we propose the SAB module to refine the segmentation with SAC and SAP blocks. By applying the SAB module iteratively across layers, our model shrinks the semantic gap and enhances the structure reasoning by fully utilizing the coarse segmentation information. Extensive quantitative evaluations demonstrate that our method significantly improves the segmentation results and achieves superior performance on the PASCAL VOC 2012, Cityscapes, PASCAL Context, and ADE20K datasets.
Xiaofeng Ding 0003, Chaomin Shen 0001, Tieyong Zeng, Yaxin Peng
IEEE Trans. Neural Networks Learn. Syst.4
2024 Exploring Gradient Explosion in Generative Adversarial Imitation Learning: A Probabilistic Perspective
abstract
Generative Adversarial Imitation Learning (GAIL) stands as a cornerstone approach in imitation learning. This paper investigates the gradient explosion in two types of GAIL: GAIL with deterministic policy (DE-GAIL) and GAIL with stochastic policy (ST-GAIL). We begin with the observation that the training can be highly unstable for DE-GAIL at the beginning of the training phase and end up divergence. Conversely, the ST-GAIL training trajectory remains consistent, reliably converging. To shed light on these disparities, we provide an explanation from a theoretical perspective. By establishing a probabilistic lower bound for GAIL, we demonstrate that gradient explosion is an inevitable outcome for DE-GAIL due to occasionally large expert-imitator policy disparity, whereas ST-GAIL does not have the issue with it. To substantiate our assertion, we illustrate how modifications in the reward function can mitigate the gradient explosion challenge. Finally, we propose CREDO, a simple yet effective strategy that clips the reward function during the training phase, allowing the GAIL to enjoy high data efficiency and stable trainability.
Wanying Wang, Yichen Zhu 0001, Yirui Zhou, Chaomin Shen 0001, Jian Tang 0008, Yaxin Peng, Yangchun Zhang
AAAI7
2024 Object-Centric Instruction Augmentation for Robotic Manipulation
abstract
Humans interpret scenes by recognizing both the identities and positions of objects in their observations. For a robot to perform tasks such as "pick and place", understanding both what the objects are and where they are located is crucial. While the former has been extensively discussed in the literature that uses the large language model to enrich the text descriptions, the latter remains underexplored. In this work, we introduce the Object-Centric Instruction Augmentation (OCI) framework to augment highly semantic and information-dense language instruction with position cues. We utilize a Multi-modal Large Language Model (MLLM) to weave knowledge of object locations into natural language instruction, thus aiding the policy network in mastering actions for versatile manipulation. Additionally, we present a feature reuse mechanism to integrate the vision-language features from off-the-shelf pre-trained MLLM into policy networks. Through a series of simulated and real-world robotic tasks, we demonstrate that robotic manipulator imitation policies trained with our enhanced instructions outperform those relying solely on traditional language instructions.
Yichen Zhu 0001, Minjie Zhu, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Dong Liu 0058, Feifei Feng, Jian Tang 0008
ICRA8
2024 Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
abstract
The language-conditioned robotic manipulation aims to transfer natural language instructions into executable actions, from simple "pick-and-place" to tasks requiring intent recognition and visual reasoning. Inspired by the dual-process theory in cognitive science—which suggests two parallel systems of fast and slow thinking in human decision-making—we introduce Robotics with Fast and Slow Thinking (RFST), a framework that mimics human cognitive architecture to classify tasks and makes decisions on two systems based on instruction types. Our RFST consists of two key components: 1) an instruction discriminator to determine which system should be activated based on the current user’s instruction, and 2) a slow-thinking system that is comprised of a fine-tuned vision-language model aligned with the policy networks, which allow the robot to recognize user’s intention or perform reasoning tasks. To assess our methodology, we built a dataset featuring real-world trajectories, capturing actions ranging from spontaneous impulses to tasks requiring deliberate contemplation. Our results, both in simulation and real-world scenarios, confirm that our approach adeptly manages intricate tasks that demand intent recognition and reasoning.
Minjie Zhu, Yichen Zhu 0001, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Dong Liu 0058, Feifei Feng, Jian Tang 0008
ICRA8
2024 A memory pool variational autoencoder framework for cross-domain recommendation
abstract
Cross-domain recommendation (CDR) leverages knowledge from the source domain to make recommendations for the cold-start users in the target domain. On account of fully utilizing information, various relationships among users and items are taken into account, i.e., the interaction relationship between users and their corresponding items; the relationship among users or items; and the indirect relationship between the user and items related to other users. In order to process these relationships, we propose a novel framework named Memory Pool Variational AutoEncoder (MPVAE). The main advantages of the MPVAE model lie in three aspects: (1) it generates the embedding representations that incorporate more information by a memory pool mechanism in the source and target domains; (2) it involves the relationship among users or items efficiently by the similarity measurement, further, the indirect relationship can be explicitly described, which makes full use of information in the source domain; and (3) it leverages the superiority of the probability model from the perspective of the VAE structure, which ensures generation and robustness. Comprehensive experiments on three real datasets show that the proposed model achieves remarkable superiority over several competitive CDR models.
Jie Yang 0101, Jianxiang Zhu, Xiaofeng Ding 0003, Yaxin Peng, Yangchun Zhang
Expert Syst. Appl.4
2024 Generalization error for portable rewards in transfer imitation learning
Yirui Zhou, Mengxiao Lu, Jian Tang 0008, Yangchun Zhang, Yaxin Peng
Knowl. Based Syst.7
2023 CP3: Channel Pruning Plug-in for Point-Based Networks
abstract
Channel pruning can effectively reduce both computational cost and memory footprint of the original network while keeping a comparable accuracy performance. Though great success has been achieved in channel pruning for 2D image-based convolutional networks (CNNs), existing works seldom extend the channel pruning methods to 3D point-based neural networks (PNNs). Directly implementing the 2D CNN channel pruning methods to PNNs undermine the performance of PNNs because of the different representations of 2D images and 3D point clouds as well as the network architecture disparity. In this paper, we proposed CP3, which is a Channel Pruning Plugin for Point-based network. CP3is elaborately designed to leverage the characteristics of point clouds and PNNs in order to enable 2D channel pruning methods for PNNs. Specifically, it presents a coordinate-enhanced channel importance metric to reflect the correlation between dimensional information and individual channel features, and it recycles the discarded points in PNN's sampling process and reconsiders their potentially-exclusive information to enhance the robustness of channel pruning. Experiments on various PNN architectures show that CP3constantly improves state-of-the-art 2D CNN pruning approaches on different point cloud tasks. For instance, our compressed PointNeXt-S on ScanObjectNN achieves an accuracy of 88.52% with a pruning rate of 57.8%, outperforming the baseline pruning methods with an accuracy gain of 1.94%.
Yaomin Huang, Ning Liu 0007, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Guixu Zhang, Xinmei Liu, Feifei Feng, Jian Tang 0008
CVPR6
2023 Recognizable Information Bottleneck
abstract
Information Bottlenecks (IBs) learn representations that generalize to unseen data by information compression. However, existing IBs are practically unable to guarantee generalization in real-world scenarios due to the vacuous generalization bound. The recent PAC-Bayes IB uses information complexity instead of information compression to establish a connection with the mutual information generalization bound. However, it requires the computation of expensive second-order curvature, which hinders its practical application. In this paper, we establish the connection between the recognizability of representations and the recent functional conditional mutual information (f-CMI) generalization bound, which is significantly easier to estimate. On this basis we propose a Recognizable Information Bottleneck (RIB) which regularizes the recognizability of representations through a recognizability critic optimized by density ratio matching under the Bregman divergence. Extensive experiments on several commonly used datasets demonstrate the effectiveness of the proposed method in regularizing the model and estimating the generalization gap.
Yilin Lyu, Xin Liu 0086, Yaxin Peng, Tieyong Zeng, Liping Jing
IJCAI5
2023 Distributional generative adversarial imitation learning with reproducing kernel generalization
Yirui Zhou, Mengxiao Lu, Zhengping Che, Jian Tang 0008, Yangchun Zhang, Yan Peng 0001, Yaxin Peng
Neural Networks9
2023 Efficient Robust Watermarking Based on Structure-Preserving Quaternion Singular Value Decomposition
abstract
Quaternion singular value decomposition (QSVD) is a robust technique of digital watermarking that extracts high quality watermarks from watermarked images with low distortion. However, the existing QSVD-based watermarking schemes face the obstacle of "explosion of complexity" and have much room for improvement in terms of real-time, invisibility, and robustness. In this paper, we overcome such obstacle by introducing a new real structure-preserving QSVD algorithm and propose a novel QSVD-based watermarking scheme with high efficiency. Secret information is transmitted blindly by incorporating two new strategies: coefficient pair selection and adaptive embedding. The highly correlated coefficient pairs determined by the normalized cross-correlation method reduce the impact of embedding by reducing the maximum modification of the coefficient values, resulting in high fidelity of the watermarked image. Large-size 8-color binary watermark and QR code effectively verify that the proposed watermarking scheme can resist various image attacks in numerical experiments. Two keys designed by Logistic chaotic map ensure the security of the watermarking system. Under the premise of considering the correlation of color channels, the proposed watermarking scheme not only performs well in real-time and invisibility, but also has satisfactory advantages in robustness compared with the state-of-the-art methods.
Yong Chen 0019, Zhigang Jia, Yaxin Peng, Yan Peng 0001
IEEE Trans. Image Process.3
2023 SRRNet: A Semantic Representation Refinement Network for Image Segmentation
abstract
Semantic context has raised concerns in semantic segmentation. In most cases, it is applied to guide feature learning. Instead, this paper applies it to extract the semantic representation, which records the global feature information of each category with a memory tensor. Specifically, we propose a novel semantic representation (SR) module, which consists of semantic embedding (SE) and semantic attention (SA) blocks. The SE block adaptively embeds features into the semantic representation by calculating the memory similarity, and the SA block aggregates the embedded features with semantic attention. The main advantages of the SR module lie in three aspects: i) it enhances the representation ability of semantic context by employing global (cross-image) semantic information; ii) it improves the consistency of intraclass features by aggregating global features of the same categories; and iii) it can be extended to build a semantic representation refinement network (SRRNet) by iteratively applying the SR module across multiple scales, shrinking the semantic gap and enhancing the structural reasoning of the model. Extensive experiments demonstrate that our method significantly improves the segmentation results and achieves superior performance on the PASCAL VOC 2012, Cityscapes, and PASCAL Context datasets.
Xiaofeng Ding 0003, Tieyong Zeng, Jian Tang 0008, Zhengping Che, Yaxin Peng
IEEE Trans. Multim.5
2022 Label-Guided Auxiliary Training Improves 3D Object Detector
Yaomin Huang, Xinmei Liu, Yichen Zhu 0001, Chaomin Shen 0001, Zhengping Che, Guixu Zhang, Yaxin Peng, Feifei Feng, Jian Tang 0008
ECCV (9)8
2022 Generalization and Computation for Policy Classes of Generative Adversarial Imitation Learning
Yirui Zhou, Yangchun Zhang, Wanying Wang, Zhengping Che, Jian Tang 0008, Yaxin Peng
PPSN (1)8
2022 Scale robust point matching-Net: End-to-end scale point matching using Lie group
abstract
Abstract Point cloud matching is an important procedure in a variety of computer vision tasks. Traditional point cloud matching methods have made great progress, while neural network‐based approaches are becoming a trend, powered by their strong capabilities of feature extraction. Existing point matching neural networks, however, mainly focus on the rigid transformation. More complex transformations should also be considered in many scenarios. In this regard, the authors extend the rigid registration to non‐rigid cases and propose a network called the Scale Robust Point Matching (SRPM)‐Net for scale point matching. This robust structure‐preserving network is implemented by incorporating Lie group parametrisation. It is conducted by Lie group linearisation representation with the constraints of parameters under the corresponding basis of Lie algebra. SRPM‐Net preserves the structure of the solution and avoids degeneration. The contributions of this paper lie in two aspects: Most importantly, SRPM‐Net provides an extendable framework for handling complicated transformations. Secondly, it introduces a new feature learning module, which better preserves the shape structure by aggregating the high‐dimensional feature and calculating the normal vector of point cloud surface automatically. Experimental results show that SRPM‐Net is more robust and accurate than existing traditional and recent deep learning methods under various situations.
Xin Wang 0084, Guangwei Zhao, Yaxin Peng, Chaomin Shen 0001
IET Comput. Vis.4
2022 Semisupervised SAR image change detection based on a siamese variational autoencoder
Guangwei Zhao, Yaxin Peng
Inf. Process. Manag.2
2021 VAN: Voting and Attention Based Network for Unsupervised Medical Image Registration
Zhiang Zu, Guixu Zhang, Yaxin Peng, Chaomin Shen 0001
PRICAI (1)3
2021 A new structure-preserving quaternion QR decomposition method for color image blind watermarking
Yong Chen 0019, Zhigang Jia, Yan Peng 0001, Yaxin Peng, Dan Zhang 0001
Signal Process.4
2020 Label Smoothing Technique for Ordinal Classification in Cloud Assessment
abstract
Satellite image classification is a challenging task if the input labels are not sufficiently accurate. The automatic cloud cover assessment (ACCA), for example, aims to classify the cloud covers of satellite images as alphabetical categories from A to E showing the escalating levels of clouds; however, those labels for training are often obtained by a subjective qualitative assessment, i.e., they may be not accurate. Therefore, this paper studies how to conduct ACCA under this circumstance. We propose a label smoothing approach and improve the accuracy around 3 percentage points (e.g., from 75.9% to 78.4% for ResNet network) without changing other network structures and parameters.
Yuxuan Wei, Qixuan Liu, Guixu Zhang, Yaxin Peng, Chaomin Shen 0001
IGARSS4
2019 The Adversarial Attack and Detection under the Fisher Information Metric
abstract
Many deep learning models are vulnerable to the adversarial attack, i.e., imperceptible but intentionally-designed perturbations to the input can cause incorrect output of the networks. In this paper, using information geometry, we provide a reasonable explanation for the vulnerability of deep learning models. By considering the data space as a non-linear space with the Fisher information metric induced from a neural network, we first propose an adversarial attack algorithm termed one-step spectral attack (OSSA). The method is described by a constrained quadratic form of the Fisher information matrix, where the optimal adversarial perturbation is given by the first eigenvector, and the vulnerability is reflected by the eigenvalues. The larger an eigenvalue is, the more vulnerable the model is to be attacked by the corresponding eigenvector. Taking advantage of the property, we also propose an adversarial detection method with the eigenvalues serving as characteristics. Both our attack and detection algorithms are numerically optimized to work efficiently on large datasets. Our evaluations show superior performance compared with other methods, implying that the Fisher information is a promising approach to investigate the adversarial attacks and defenses.
Chenxiao Zhao, P. Thomas Fletcher, Mixue Yu, Yaxin Peng, Guixu Zhang, Chaomin Shen 0001
AAAI4
2019 Manifold Alignment and Distribution Adaptation for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation is a problem which exploits the knowledge learned from the resource-rich domain to obtain an accurate classifier for the resource-poor domain. Most of the existing methods lift performance by reducing the differences between distributions, such as the difference between marginal probability distributions, the difference between conditional probability distributions, or both. However, all these methods consider the two distributions to be equally important, which could lead to poor classification performance in practical applications. Therefore, a balanced factor is required to weigh the two distributions to compensate for the degraded performance. In this paper, we first introduce this balance factor to weigh the distribution importance. On this base, we utilize the marginal distribution, introduce the ideas of manifold regularization, and then preserve the neighboring structures of the data sets, with the dimension reduction as much as possible. By this way, we propose the manifold alignment and balanced distribution adaptation algorithm. A large number of experiments have also been conducted, showing that our algorithm behaves much better than the previous ones.
Ying Li 0028, Yaxin Peng, Zhijie Wen, Shihui Ying
ICME3
2019 Asymmetric Local Metric Learning with PSD Constraint for Person Re-identification
abstract
Person re-identification is one of the key issues in both machine learning and video monitor application. In particular, defining an appropriate distance metric between the person images is very important. Existing metric learning approaches used in person re-identification either learn a single measure, or ignore the positive semi-definite (PSD) of measurement matrix, at the same time, since the number of negative sample pairs largely exceeds the number of positive sample pairs, some metric learning methods are largely influenced by the sample imbalance. Considering the above issues, we propose a new adaptive local metric learning method with positive semi-definite (PSD) constraint. Unlike existing metric learning methods which learn a single distance metric, we use an approximation error bound of a smooth metric matrix function over the data manifold to learn local metrics as linear combinations of basis metrics defined on anchor points over different regions of the instance space. Besides, we develop an efficient two stage algorithm that first learns the anchor points and the linear combinations of each instance, then learns the metric matrices of the anchor points. We employ the fast iterative shrinkage-thresholding algorithm which is a fast first-order optimization algorithm in the learning process of the linear combinations as well as the basis metrics of the anchor points. Our metric learning method has excellent performance. We firstly apply the proposed method on 5 UCI databases, which are widely used in machine learning, to test and evaluate the effectiveness of the proposed method. Then the proposed approach is applied for person re-identification, achieving better performance on three challenging databases (GRID, VIPeR, CUHK01) than the existing methods. The experimental results show that the proposed method can prvide the theoretical and practical support for the person re-identification problem.
Zhijie Wen, Ying Li 0028, Shihui Ying, Yaxin Peng
ICRA5
2019 Deep Ordinal Classification for Automatic Cloud Assessment
abstract
We develop a deep neural network for ordinal classification and use this technique for the automatic cloud cover assessment (ACCA) of satellite images. We adopt a VGG+ResNet approach with a novel loss function to solve this problem. Results using Quicklook images from Centre for Remote Imaging, Sensing and Processing (CRISP) are promising, as approximately 97.3% of all sub-scenes are correctly labelled, potentially even higher when one ordinal error is tolerated.
Qixuan Liu, Jinsong Fan, Chaomin Shen 0001, Yaxin Peng
IGARSS4
2019 Geometric Understanding for Unsupervised Subspace Learning
abstract
In this paper, we address the unsupervised subspace learning from a geometric viewpoint. First, we formulate the subspace learning as an inverse problem on Grassmannian manifold by considering all subspaces as points on it. Then, to make the model computable, we parameterize the Grassmannian manifold by using an orbit of rotation group action on all standard subspaces, which are spanned by the orthonormal basis. Further, to improve the robustness, we introduce a low-rank regularizer which makes the dimension of subspace as low as possible. Thus, the subspace learning problem is transferred to a minimization problem with variables of rotation and dimension. Then, we adopt the alternately iterative strategy to optimize the variables, where a structure-preserving method, based on the geodesic structure of the rotation group, is designed to update the rotation. Finally, we compare the proposed approach with six state-of-the-art methods on three different kinds of real datasets. The experimental results validate that our proposed method outperforms all compared methods.
Shihui Ying, Lipeng Cai, Changzhou He, Yaxin Peng
IJCAI4
2019 Strict Subspace and Label-Space Structure for Domain Adaptation
Ying Li 0028, Yaxin Peng, Chaomin Shen 0001
KSEM (1)4
2019 Partial Alignment of Data Sets Based on Fast Intrinsic Feature Match
Yaxin Peng, Naiwu Wen, Xiaohuang Zhu, Chaomin Shen 0001
KSEM (1)1
2018 Parallel Hashing Using Representative Points in Hyperoctants
abstract
The goal of hashing is to learn a low-dimensional binary representation of high-dimensional information, leading to a tremendous reduction of computational cost. Previous studies usually achieved this goal by applying projection or quantization methods. However, the projection method fails to capture the intrinsic data structures, and the quantization method cannot make full use of complete information by its strategy of partitioning original space. To combine their advantages and avoid their drawbacks, we propose a novel algorithm, termed as representative points quantization (RPQ), by using the representative points defined as the barycenters of points in the hyperoctants. To settle the problem of exponential time complexity with the growth of the coding length, for long hashing codes, we further propose a parallel RPQ (PRPQ) algorithm, by separating a long code into several short codes, re-coding the short codes in different low dimensional subspaces, and then concatenating them to a long code. Experiments on image retrieval tasks demonstrate that RPQ and PRPQ can well capture the main topology structure of data, showing that our algorithm achieves better performance than state-of-the-art methods.
Chaomin Shen 0001, Mixue Yu, Chenxiao Zhao, Yaxin Peng, Guixu Zhang
CIKM4
2018 Enhanced Metric Learning via Dempster-Shafer Evidence Theory
Ying Li 0028, Yabo Zhang, Yaxin Peng
ICONIP (3)3
2018 Cloud Cover Assessment in Satellite Images Via Deep Ordinal Classification
abstract
The percentage of cloud cover is one of the key indices for satellite data products. To date, cloud cover assessment is performed manually in most groundstations. To facilitate the process, this paper proposes a deep learning approach for cloud cover assessment in quicklook satellite images. The quicklook images from Centre for Remote Imaging, Sensing and Processing (CRISP) are used for demonstration. Same as the manual operation, given a quicklook image, the algorithm returns 8 labels ranging from A to E and *, indicating the cloud percentages in different areas of the image. This is achieved by constructing 8 improved VGG-16 models, where parameters such as the loss function, learning rate and dropout are tailored for better performance. Results indicate that approach is promising, as around 85% of sub-scenes are correctly labelled, and the accuracy is even higher if one ordinal error is accepted. This paper demonstrates a new application in remote sensing using state-of-the-art deep learning techniques.
Chaomin Shen 0001, Chenxiao Zhao, Mixue Yu, Yaxin Peng
IGARSS4
2018 Global Nonlinear Metric Learning by Gluing Local Linear Metrics
abstract
We address the nonlinear metric learning by constructing a smooth nonlinear metric from the data. First, we locally define an initial linear metric on each cluster by principal component analysis. Second, we glue such local linear metrics to form a smooth nonlinear metric by a partition of unity on the sample space, and further learn the global nonlinear metric. Third, we conduct the intrinsic steepest descent algorithm on matrix manifolds for implementation. Finally, we compare our approach with several state-of-the-art methods on a variety of datasets. The results validate that the robustness and accuracy of classification are both improved under our nonlinear metric. The novelty of our global smooth nonlinear metric learning model lies in that it has completely overcome drawbacks of local metric learning methods: the partition coefficients obtained by the partition of unity is smooth, while the metric at any point on the manifold can be directly defined.
Yaxin Peng, Lingfang Hu, Shihui Ying, Chaomin Shen 0001
SDM1
2018 Nonlinear Semi-Supervised Metric Learning Via Multiple Kernels and Local Topology
abstract
Changing the metric on the data may change the data distribution, hence a good distance metric can promote the performance of learning algorithm. In this paper, we address the semi-supervised distance metric learning (ML) problem to obtain the best nonlinear metric for the data. First, we describe the nonlinear metric by the multiple kernel representation. By this approach, we project the data into a high dimensional space, where the data can be well represented by linear ML. Then, we reformulate the linear ML by a minimization problem on the positive definite matrix group. Finally, we develop a two-step algorithm for solving this model and design an intrinsic steepest descent algorithm to learn the positive definite metric matrix. Experimental results validate that our proposed method is effective and outperforms several state-of-the-art ML methods.
Yanqin Bai, Yaxin Peng, Shaoyi Du, Shihui Ying
Int. J. Neural Syst.3
2018 Performance Analysis for SVM Combining with Metric Learning
Lingfang Hu, Chaomin Shen 0001, Yaxin Peng
Neural Process. Lett.5
2018 Manifold Preserving: An Intrinsic Approach for Semisupervised Distance Metric Learning
abstract
In this paper, we address the semisupervised distance metric learning problem and its applications in classification and image retrieval. First, we formulate a semisupervised distance metric learning model by considering the metric information of inner classes and interclasses. In this model, an adaptive parameter is designed to balance the inner metrics and intermetrics by using data structure. Second, we convert the model to a minimization problem whose variable is symmetric positive-definite matrix. Third, in implementation, we deduce an intrinsic steepest descent method, which assures that the metric matrix is strictly symmetric positive-definite at each iteration, with the manifold structure of the symmetric positive-definite matrix manifold. Finally, we test the proposed algorithm on conventional data sets, and compare it with other four representative methods. The numerical results validate that the proposed method significantly improves the classification with the same computational efficiency.
Shihui Ying, Zhijie Wen, Jun Shi 0004, Yaxin Peng, Hong Qiao
IEEE Trans. Neural Networks Learn. Syst.4
2016 Joint distribution adaptation based TSK Fuzzy logic system for epileptic EEG signal identification
abstract
Transfer learning based method, which utilizes plenty labeled data in the source domain to build an accuracy classifier for the target domain, serves as an effective means in the epileptic detection by using electroencephalogram (EEG) signals. Among existing approaches, Fuzzy logic system (FLS) based on transductive transfer learning is an efficient method due to its superior interpretability and strong learning abilities. However, this kind of method cannot simultaneously reduce the differences in both marginal distributions and conditional distributions between the training and test datasets of EEG signals. To overcome this problem, in this paper, we construct a Takagi-Sugeno-Kang (TSK) FLS based on the joint distribution adaptation (JDA), which refers to TSK-JDA-FLS. It aims to match both marginal and conditional distributions, and we extend the algorithm to perform a multi-class classification for identifying epileptic EEG signals. Extensive experiments verify that TSK-JDA-FLS significantly outperforms competitive non-transfer learning and transfer learning methods in the epileptic EEG datasets.
Yaxin Peng, Guixu Zhang, Chaomin Shen 0001
BIBM2
2016 Object(s)-of-interest segmentation for images with inhomogeneous intensities based on curve evolution
Yaxin Peng, Lili Bao, Ling Pi
Neurocomputing1
2016 Compute Karcher means on SO(n) by the geometric conjugate gradient method
Shihui Ying, Han Qin, Yaxin Peng, Zhijie Wen
Neurocomputing3
2016 Virus image classification using multi-scale completed local binary pattern features extracted from filtered images by multi-scale principal component analysis
Zhijie Wen, Zhuojun Li, Yaxin Peng, Shihui Ying
Pattern Recognit. Lett.3
2014 Spectral unmixing using Lasso screening rules
abstract
Lasso (Least Absolute Shrinkage and Selection Operator) is a technique for selecting a sparse combination of given features. Spectral unmixing is a Lasso problem, regarding that the spectrum of every pixel is a linear combination of a small number of spectrums from a possibly very large spectral library. In this paper we apply a technique called screening to speedup the Lasso process for spectral unmixing. Our contribution is two-fold: we make use of the theoretical results to practical remote sensing problems; more importantly, we develop a tailored Lasso algorithm coupled with screening, as the unmixing requires that the fractions should be positive and sum to one. We also solve the high mutual coherence problem in the library by ticking out the spectrums with high mutual coherence. We use the AVIRIS data over Cuprite, Nevada (250 lines by 191 columns) to demonstrate the idea. Experiments demonstrate the effectiveness of the screening method.
Chaomin Shen 0001, Xiaoliang Shi, Yaxin Peng
IGARSS4
2014 LieTrICP: An improvement of trimmed iterative closest point algorithm
Yaxin Peng, Shihui Ying, Zhiyu Hu
Neurocomputing2
2013 Soft shape registration under Lie group frame
abstract
In this study, the authors address a two‐dimensional (2D) shape registration problem on data with anisotropic‐scale deformation and noise. First, the model is formulated under the iterative closest point (ICP) framework, which is one of the most popular methods for shape registration. To overcome the effect of noise, the expectation maximisation algorithm is used to improve the model. Then, the structure of Lie groups is adopted to parameterise the proposed model, which provides a unified framework to deal with the shape registration problems. Such representation makes it possible to introduce some suitable constraints to the model, which improves the robustness of the algorithm. Thereby, the 2D shape registration problem is turned to an optimisation problem on the matrix Lie group. Furthermore, a sequence of quadratic programming is designed to approximate the solution for the model. Finally, several comparative experiments are carried out to validate that the authors’ algorithm performs well in terms of robustness, especially in the presence of outliers.
Yaxin Peng, Shihui Ying
IET Comput. Vis.1
2011 Iwasawa decomposition: a new approach to 2D affine registration problem
Shihui Ying, Yaxin Peng, Zhijie Wen
Pattern Anal. Appl.2
2010 Variational Color Image Segmentation via Chromaticity-Brightness Decomposition
Yaxin Peng, Guixu Zhang
MMM3
2007 Variational-based speckle noise removal of SAR imagery
abstract
In this paper we present a variational method for synthetic aperture radar (SAR) speckle removal. Variational method is a newly developed technique for the removal of SAR's multiplicative noise. For an image, we could define an energy functional. The energy evolves as the original image changes, and the minimum energy corresponds to the speckle reduced result. Partial differential equation (PDE) technique is used to get the minimal solution. Our energy functional makes use of the statistical information of the multiplicative noise since it follows a Gamma law with mean mu = 1 and variance sigma2= 1/M for M-look SAR. Our energy is a regularization term with two constraints. The regularization term is the integral for the norm of image gradient; two constraints are the mean of noise should be 1 and the variance of noise should be 1/M. We use the method of Lagrange multipliers, Euler-Lagrange equation and heat flow method to obtain the minimizer of the energy. ERS Precision Image (PRI) data are to demonstrate our algorithm. Numerical result shows that the speckle reduced image preserves edges and point targets while smoothes homogenous regions in the original image. The algorithm is computationally efficient and easy to implement.
Chaomin Shen 0001, Yaxin Peng, Ling Pi
IGARSS2