Yuanxi Peng

dblp:44/10313 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0002-6647-8862ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Frequency domain-parametric spline learning driven dual-branch network for hyperspectral classification
Jianghe Zhai, Yuanxi Peng, Longlong Zhang, Tong Zhou 0008
Eng. Appl. Artif. Intell.4
2026 A Wavelet-Guided Robust Attention Network for small aircraft detection in complex scenes
Jianghe Zhai, Yuanxi Peng, Longlong Zhang, Tong Zhou 0008
Eng. Appl. Artif. Intell.5
2026 MADBD: Two-stage point cloud registration via multi-attention and dual branches decoupling
Longlong Zhang, Yuanxi Peng, Tong Zhou 0008
Neurocomputing4
2025 Enhancing Multimodal Entity Linking via Distillation and Multimodal Large Language Models
abstract
Multimodal entity linking (MEL) aims to link ambiguous multimodal mentions to their corresponding entities in a multimodal knowledge graph. Although many existing methods have been dedicated to exploring fine-grained intra- and cross-modal interactions between mentions and entities and have achieved good results, the discrepancies between the data distributions in training and real-world applications, as well as the noisy onehot labels, still impede the generalization of MEL models, which leads to poor performance when encountering unseen entities. Although general-purpose multimodal large language models (MLLMs) are powerful, it is costly and time-consuming to apply them directly to the MEL task. To address the above issues, we propose a Distillation-Enhanced framework for Multimodal Entity Linking (DEMEL). During training, DEMEL takes the best-trained MEL model so far as the teacher model, and distills the knowledge of the teacher model into the student model, i.e. the MEL model of the current iteration, when training it with onehot labels. This imposes regularization on the model, balances bias and variance in the training process, and improves the generalization ability of the MEL model. Moreover, DEMEL employs an MLLM to selectively rerank predictions for uncertain samples in the inference phase, improving accuracy while minimizing invocation costs. Extensive experiments on three public MEL datasets demonstrate that DEMEL outperforms state-of-the-art baselines, achieving 3.27% improvement with MLLM reranking for just 8.59% of test samples, and up to 4.8% H@1 enhancement in low-resource settings even without using MLLM reranking.
Yuanxi Peng, Ruochun Jin
CIKM4
2025 CLEAR: A Parser-Independent Disambiguation Framework for NL2SQL
abstract
Parsing Natural Language to SQL (NL2SQL) helps users who are not proficient in databases to efficiently query desired data through natural language. Although existing NL2SQL parsers demonstrate good capabilities in processing clear queries, ambiguity still remains an unresolved issue which makes parsers produce unstable outputs that deviate from the user's actual intent. To bridge the gap, this paper introduces the CLEAR framework, a systematic study of disambiguation for NL2SQL, including ambiguity detection, clarification, and reformulation, which benefits any NL2SQL parsers. Firstly, CLEAR employs a pipeline using Large Language Models (LLMs) and a series of rules to detect ambiguities, thus obtaining the “candidate mapping” for ambiguity representation. Secondly, an interactive selection module is employed to collect the clarification information from users through multiple-choice questions, thus obtaining the “selection mapping”. Finally, rewriting rules are employed to reformulate the question and schema, thus obtaining a clear input for parsers to generate clear SQLs. Furthermore, we construct CLAMBSQL, a novel benchmark for systematic evaluation for NL2SQL disambiguation, which contains fine-grained ambiguity and clarification annotations. Experiments on various datasets and baselines demonstrate that CLEAR can successfully address seven types of ambiguity. When parsers are integrated with CLEAR, the performance of ambiguous SQLs detection achieves a significant improvement of 30.5 % on AMBROSIA in the AllFound metric and 21.1 % on AmbiQT in the BothInTop-5 metric, the performance of ambiguity clarification achieves a remarkable improvement of 16.2 % on CLAMBSQL in the CEX metric, and the performance of the general prediction achieves an increase of 1.6 % in the EX metric and 7.7 % in the CSR metric on BIRD. The CLEAR code and CLAMBSQL dataset are available at https://github.com/mengzhang18/CLEAR.
Kexin Ma 0008, Kedi Zhang, Yuanxi Peng, Ruochun Jin
ICDE5
2025 Geometrically-Inspired Irregular Expansion Techniques for Graph-based Point Cloud Learning
abstract
Advanced deep learning methodologies have made notable progress in 3D point cloud tasks. Nevertheless, the absence of significant long-range correlation measurement in irregular point cloud entities limits the representation of 3D geometries. Several methodologies have been proposed to address this issue. Still, they fall short of altering the fundamental aggregation paradigm inherent in Graph Convolutional Networks (GCNs), which suffer the exclusive use of the summation operator to encapsulate the information from adjacent nodes, resulting in limited expressiveness and inefficient computation. This paper presents an efficient geometrically irregular expansion on Graph (GIEG) convolution network for point cloud analysis. More specifically, spatial features, formulated with a localized graph representation derived from multiple sequence expansions, are comprehensively exploited by capturing local point semantics while avoiding dense point trappings, which extend the traditional neighborhood with a path-based neighborhood better than native point cloud methods. Thanks to the novel convolution module, the GIEG can extract comprehensive and influential feature semantics for individual points. Extensive experimental results validate the effectiveness of our method against challenging benchmarks.
Qi Zhang 0082, Haoqian Wang, Yuanxi Peng, Teng Li 0011
ICME3
2025 Ultra-low power MoS2 optoelectronic synapse with wavelength sensitivity for color target recognition
Yabo Chen, Xiaotong Han, Bujia Liang, Xiaokuo Yang, Yuanxi Peng
Sci. China Inf. Sci.9
2025 An Efficient Point Network for Light Detection and Ranging point cloud perception in large-scale scene
Jialin Gui, Yuanxi Peng, Teng Li 0011
Eng. Appl. Artif. Intell.3
2025 RDHNet: addressing rotational and permutational symmetries in continuous multi-agent systems
Dongzi Wang 0002, Lilan Huang, Muning Wen, Yuanxi Peng, Minglong Li, Teng Li 0011
Frontiers Comput. Sci.4
2025 PCSViT: Efficient and hardware friendly Pyramid Vision Transformer with channel and spatial self-attentions
Xiaofeng Zou, Yuanxi Peng, Xinye Cao
Neurocomputing2
2025 Exploiting Complex-Valued Representations in Automatic Modulation Recognition: A Framework Integrating a Transformer With Relative Positional Encoding and Separable Convolution
abstract
Automatic modulation recognition plays a critical role in military applications, particularly in electronic warfare, spectrum surveillance, and secure communication systems. The precise identification of signal modulation modes is crucial for ensuring the efficacy, security, and efficiency of communications. Given the problems of limited feature extraction capability and performance degradation when dealing with low signal-to-noise ratio signals, this work proposes a novel model architecture that combines the Transformer with relative position encoding and separable convolution in the complex domain. The network can directly process the complex representation of signals to capture features in the time-frequency domain and enhance the ability to recognize complex signals. This method introduces relative position encoding in the complex domain into the Transformer framework, which uses complex attention mechanisms and adaptive position encoding to enhance the model’s long-range modeling capability. At the same time, it improves computational efficiency and local feature extraction capability by introducing separable convolution layers. Then, we construct an attention-driven feature fusion module, which can automatically adjust the weight ratio between features to achieve the optimal combination of features. The model shows excellent classification performance under various signal-to-noise ratio conditions in RadioML2016.10a and RadioML2016.10b datasets, especially in low signal-to-noise ratio environments, which is significantly improved compared to other methods. The research not only provides a new solution for AMR tasks but also expands new ideas for applying the Transformer model in signal processing.
Xuan Liao, Longlong Zhang, Yuanxi Peng, Tong Zhou 0008
IEEE Internet Things J.5
2025 TS-DANet: Truncated SVD Interaction and Dual-Attention Network for Hyperspectral and Multispectral Image Fusion
abstract
Fusing hyperspectral image (HSI) and multispectral image (MSI) is essential to merge HSI’s spectral richness with MSI’s spatial detail. This paper introduces TS-DANet, a novel two-part network for HSI-MSI fusion designed to address the limitations of current methods in model prior utilization and spectral-spatial feature extraction. In the initial part, we use truncated singular value decomposition (TSVD) interaction to extract spectral priors from low-resolution HSI (LR-HSI) and spatial sparsity priors from high-resolution MSI (HR-MSI), integrating these through a physics-based optimization for image fusion. We also incorporate a dual-attention mechanism, featuring a dynamic spectral attention module for detailed spectral features and a multi-scale spatial attention module for detailed spatial features. The second part employs a dynamic residual optimization module to further refine spatial and spectral information. Extensive experiments on three remote sensing datasets show that TS-DANet outperforms existing state-of-the-art algorithms in fusion performance. The code will be available at https://github.com.
Jun Li 0094, Mingze Peng, Yuanxi Peng, Yinuo Liu
IEEE Geosci. Remote. Sens. Lett.5
2025 Seesaw: A 4096-bit vector processor for accelerating Kyber based on RISC-V ISA extensions
Xiaofeng Zou, Yuanxi Peng, Lingjun Kong
Parallel Comput.2
2025 FLQ: Design and implementation of hybrid multi-base full logarithmic quantization neural network acceleration architecture based on FPGA
Longlong Zhang, Xuan Liao, Tong Zhou 0008, Yuanxi Peng
Signal Process. Image Commun.5
2025 Parameter-Free Spectral-Spatial Optimization Algorithm for Semiblind Hyperspectral and Multispectral Image Fusion
abstract
Semiblind fusion of hyperspectral images (HSIs) and multispectral images (MSIs) is a critical technique for generating high-resolution HSIs (HR-HSIs) without the need for complex point spread function (PSF) estimation. Despite the broad application potential of semiblind fusion algorithms, existing methods face three major challenges. First, as imaging technology advances, the spatial resolution gap between HSIs and MSIs continues to widen, making high-magnification fusion increasingly urgent. Second, especially for deep learning-based methods, existing algorithms require meticulous hyperparameter tuning to enhance fusion quality. Third, the complex fusion process impedes the speed of image fusion. To address these challenges, we propose a parameter-free spectral-spatial optimization algorithm specifically designed to handle high-magnification differences in the semiblind fusion of HSI and MSI. This method enables fast fusion using simple matrix operations and consists of three main steps: 1) rapidly computing an initial solution using the Moore-Penrose inverse of the spatial response function (SRF); 2) constructing spectral errors from the initial solution to effectively extract spectral information; and 3) reconstructing spatial details using MSI to form spatial errors, thereby accurately reconstructing HR-HSI for high-quality fusion. Comparative experiments on two simulated and three real datasets with state-of-the-art algorithms demonstrate that our proposed semiblind fusion method not only achieves$64\times $high-magnification fusion but also reduces the computation time to just 23.9%–49.5% of that required by the fastest competing methods. The code is available athttps://github.com/Long-ji/PFSSOA.
Jialin Gui, Yuanxi Peng, Yibing Zhan, Jun Li 0094, Yulei Tian
IEEE Trans. Geosci. Remote. Sens.3
2024 RoMAT: Role-based multi-agent transformer for generalizable heterogeneous cooperation
Dongzi Wang 0002, Fangwei Zhong, Minglong Li, Muning Wen, Yuanxi Peng, Teng Li 0011, Yaodong Yang 0001
Neural Networks5
2024 Channel-Layer-Oriented Lightweight Spectral-Spatial Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is commonly influenced by convolution neural networks (CNNs). However, the large number of parameters and computational complexity associated with CNNs can limit their practical application, particularly when computing and storage resources are limited. To address this challenge, we propose a channel-layer-oriented lightweight network for HSI classification. Motivated by existing structures that typically set large channels and stack multiple layers, we give more optimal solutions strategically to further compress the model. For intralayer feature extraction, we develop a channel-oriented spectral–spatial module (COS2M), which introduces a dual-single-channel (DSC) 3-D convolution that works in conjunction with depthwise convolution to fully extract spectral–spatial information. For interlayer information transmission, we propose a novel neighbor-pixel-aware activation function (NPAF), where the activation of a single pixel is determined by the learnable interaction with its neighbor range that enhances information transmission and improves the network’s fitting ability through the single activation layer. By implementing these strategies, we aim to overcome the limitations of traditional CNNs and enable efficient HSI classification within resource-constrained environments. The whole network is designed to be a compact end-to-end structure. It achieves better classification performance than other deep learning methods and lightweight models, even with limited training samples. The network parameters, model complexity, and inference time also demonstrate significant superiority, as confirmed by experiments on three benchmark datasets. The source codes are available publicly at:https://github.com/AchunLee/CLOLN_TGRS
Chunchao Li, Behnood Rasti, Xuebin Tang, Puhong Duan, Jun Li 0094, Yuanxi Peng
IEEE Trans. Geosci. Remote. Sens.6
2024 Joint Spatial-Spectral Optimization for the High-Magnification Fusion of Hyperspectral and Multispectral Images
abstract
The fusion of hyperspectral and multispectral images is an important strategy for enhancing the spatial resolution of hyperspectral images. With the rapid advancement of multispectral imaging technology, the disparity in spatial resolution between multispectral and hyperspectral images is increasing. In certain scenarios, termed high-magnification, this difference can exceed$32\times $. Previous methods do not perform well under high-magnification fusion, and naturally, a challenge arises in achieving effective high-magnification super-resolution fusion. In light of the above analysis, this article introduces a novel algorithm for high-magnification super-resolution fusion of hyperspectral and multispectral images based on the joint optimization of spatial and spectral information. Specifically, our algorithm consists of three stages: 1) a fast preliminary fusion stage based on the Moore-Penrose inverse and singular value correlation priors for the rapid acquisition of preliminary solutions; 2) a joint spatial-spectral optimization stage where a coupled optimization framework is constructed to achieve integrated optimization of spatial and spectral information; and 3) an error backpropagation optimization stage where an effective error optimization term is introduced to further refine the fusion performance. We conducted extensive experiments on widely employed publicly available simulated datasets and real datasets. The experimental results unequivocally indicate that our proposed methodology consistently exhibits superior fusion performance compared with state-of-the-art methods, even under the condition of${\geq }60\times $super-resolution.
Yibing Zhan, Zhengbin Pang, Tong Zhou 0008, Xueqiong Li, Long Lan, Yuanxi Peng
IEEE Trans. Geosci. Remote. Sens.7
2023 Semantic Segmentation of Spectral LiDAR Point Clouds Based on Neural Architecture Search
abstract
Multispectral light detection and ranging (LiDAR) point clouds, as a new kind of data, can be characterized by high consistency and integrity of spectral information geometric data. which makes it beneficial for land use classification. However, direct classification of multispectral LiDAR data remains challenging, since the data are irregular and unordered. By describing the point cloud in the form of graph data and using graph convolution to naturally model local geometric structures’ representation, higher processing accuracy and efficiency can be achieved. Notably, however, due to the complexity of graph data, using only a single graph convolution to learn the data features may limit the capacity of the architecture; in addition, selecting appropriate graph convolutions requires human expertise and massive numbers of trials. In this article, we propose a network that combines several outstanding graph convolutions by adapting neural architecture search (NAS) for point cloud segmentation, an approach referred to as GCNAS, which can adaptively achieve different levels of semantic expression. In the search module, we use the Monte Carlo tree search (MCTS) method to select the layerwise graph convolution kernels. Experiments on the Titan multispectral LiDAR dataset have verified the effectiveness of GCNAS.
Qi Zhang 0082, Yuanxi Peng, Teng Li 0011
IEEE Trans. Geosci. Remote. Sens.2
2022 Full-BNN: A Low Storage and Power Consumption Time-Domain Architecture based on FPGA
abstract
With the increasing demand for low power and storage consumption in mobile platforms, wearable devices, and Internet of Things devices, how to better apply lightweight neural networks in many edge computing scenarios and resource-limited settings is still facing challenges. This paper first proposes a novel binary convolution structure based on the time-domain to reduce resource and power consumption for the convolution process. Furthermore, through the joint design of convolution, batch normalization, and activation function in the time-domain, we propose a full-BNN model and hardware architecture, which keeps the values of all intermediate results as one bit to reduce storage requirements by 75%. Then, we optimize the above design with spatial and temporal parallelism to improve the overall computing efficiency. Finally, we built an accelerating system and take the MNIST data set as an example to test the optimized architecture on the DSP + FPGA platform. The results show that the model can be used as a neural network acceleration unit with low storage requirement and low power consumption for classification tasks with a small loss of accuracy. The joint design method in the time-domain may further inspire other computing architectures.
Longlong Zhang, Xuebin Tang, Yuanxi Peng, Tong Zhou 0008
ASAP4
2022 A Complementary Spectral-Spatial Method for Hyperspectral Image Classification
abstract
In hyperspectral image classification, using spatial information as a supplement to spectral information is an effective way to improve classification accuracy. In this paper, a novel robust complementary method using spectral–spatial information is proposed to reduce the information loss in feature extraction, thus improving the classification effect. In short, two complementary feature extraction stages are used to get the probability maps for decision fusion. In the stage of pre-processing feature extraction, we propose an adaptive cubic total variational smoothing method (ACTVSP), which is first proposed and applied in the remote sensing research field, to obtain the first-stage probability map. At the same time, we utilize edge-preserving filtering in the post-processing stage and obtain the second probability map by the means of pixel-level classifier. Finally, the probability-like maps obtained in the above two stages are integrated by decision fusion rules. Experiments on ten public data sets show that the effectiveness of our proposed method and demonstrate the superior of distinguishing different land covers on the basis of very few training samples. Therefore, it can be applied to practical applications in different scenes.
Lulu Shi, Chunchao Li, Teng Li 0011, Yuanxi Peng
IEEE Trans. Geosci. Remote. Sens.4
2022 Unsupervised Joint Adversarial Domain Adaptation for Cross-Scene Hyperspectral Image Classification
abstract
In practical hyperspectral image cross-scene classification (HSICC) tasks, the arduous work of obtaining labels and the distribution inconsistency caused by spectral shift leave deep learning methods to face great challenges. Unsupervised domain adaptation aims to exploit knowledge from the annotated source domain and transfer it to the unlabeled target domain, thereby boosting the performance of unsupervised classification. Nevertheless, existing HSICC approaches cannot effectively exploit class structure information from target data. Specially, this article proposes an unsupervised joint adversarial domain adaptation (UJADA) architecture for HSICC to further narrow the distribution gap between distinct domains. The proposed method contains two modules: domain adversarial module that learns domain-invariant features, bi-classifier adversarial module that explores task-specific decision boundaries between classes, and both share a feature generator consisting of the dense-based spectral-spatial convolution network. The UJADA simultaneously considers domain-level and class-level feature alignment between source and target hyperspectral images in a unified adversarial learning processing. Furthermore, the classifier determinacy disparity metric is introduced to fine-grained measure the output probabilistic discrepancy between two task-specific label predictors on target data, thus ensuring the discriminability of transferable features. Comprehensive experiments and ablation studies conducted on two public cross-scene data pairs and our newly acquired ultra-low-altitude hyperspectral images under different illumination conditions demonstrate the superior performance of the proposed algorithm, which will greatly promote the practical application of hyperspectral intelligent perception technology.
Xuebin Tang, Chunchao Li, Yuanxi Peng
IEEE Trans. Geosci. Remote. Sens.3
2021 FPGA Implementation of an Improved OMP for Compressive Sensing Reconstruction
abstract
This article proposes an improved orthogonal matching pursuit (OMP) algorithm and its implementation with Xilinx Vivado high-level synthesis (HLS). We use the Gram-Schmidt orthogonalization to improve the update process of signal residuals so that the signal recovery only needs to perform the least-squares solution once, which greatly reduces the number of matrix operations in a hardware implementation. Simulation results show that our OMP algorithm has the same signal reconstruction accuracy as the original OMP algorithm. Our approach provides a fast and reconfigurable implementation for different signal sizes, different measurement matrix sizes, and different sparsity levels. The proposed design can recover a 128-length signal with measurement number M = 32 and sparsity K = 5 and K = 8 in 13.2 and 21 μs, which is at least a 21.9% and 22.2% improvement compared with the existing HLS-based works; a 256-length signal with M = 64 and K = 8 in 20.6 μs, which is a 24% improvement compared with the existing work; and a 1024-length signal with measurement number M = 256 and sparsity K = 12 and K = 36 in 150.3 and 423 μs, respectively, which are close to the results of traditional hardware description language (HDL) implementations. Our results show that our improved OMP algorithm not only offers a superior reconstruction time compared with other recent HLS-based works but also can compete with existing works that are implemented using the traditional field-programmable gate array (FPGA) design route.
Jun Li 0094, Paul Chow, Yuanxi Peng
IEEE Trans. Very Large Scale Integr. Syst.3
2020 FPGA-Based Multi-precision Architecture for Accelerating Large-Scale Floating-Point Matrix Computing
Longlong Zhang, Yuanxi Peng, Xiao Hu 0004, Ahui Huang
NPC2
2014 Benefits of Adding Hardware Support for Broadcast and Reduce Operations in MPSoC Applications
abstract
MPI has been used as a parallel programming model for supercomputers and clusters and recently in MultiProcessor Systems-on-Chip (MPSoC). One component of MPI is collective communication and its performance is key for certain parallel applications to achieve good speedups. Previous work showed that, with synthetic communication-only benchmarks, communication improvements of up to 11.4-fold and 22-fold for broadcast and reduce operations, respectively, can be achieved by providing hardware support at the network level in a Network-on-Chip (NoC). However, these numbers do not provide a good estimation of the advantage for actual applications, as there are other factors that affect performance besides communications, such as computation. To this end, we extend our previous work by evaluating the impact of hardware support over a set of five parallel application kernels of varying computation-to-communication ratios. By introducing some useful computation to the performance evaluation, we obtain more representative results of the benefits of adding hardware support for broadcast and reduce operations. The experiments show that applications with lower computation-to-communication ratios benefit the most from hardware support as they highly depend on efficient collective communications to achieve better scalability. We also extend our work by doing more analysis on clock frequency, resource usage, power, and energy. The results show reasonable scalability for resource utilization and power in the network interfaces as the number of channels increases and that, even though more power is dissipated in the network interfaces due to the added hardware, the total energy used can still be less if the actual speedup is sufficient. The application kernels are executed in a 24-embedded-processor system distributed across four FPGAs.
Yuanxi Peng, Manuel Saldaña, Christopher A. Madill, Xiaofeng Zou, Paul Chow
ACM Trans. Reconfigurable Technol. Syst.1
2011 Hardware Support for Broadcast and Reduce in MPSoC
abstract
MPI has been used as a parallel programming model for supercomputers and clusters but also in Multiprocessor System-on-Chip. One component of MPI is collective communication and its performance is key for parallel applications to achieve good speedups. Considerable research has been done to optimize such communication by improving the MPI library algorithms. However, these optimizations are focused on the processing nodes (end-points in a network) rather than on the network itself. In this paper, we target a Network-on-Chip (NoC) and modify it to provide hardware support for broadcast and reduce operations for the ArchES-MPI library. This library is a subset implementation of the MPI standard targeting embedded processors and hardware accelerators implemented in FPGAs. The experimental results show that for a system with 24 embedded processors, the broadcast and reduce operations improved up to 11.4-fold and 22-fold, respectively. Higher benefits are expected for larger systems at the expense of a modest increase resource utilization.
Yuanxi Peng, Manuel Saldaña, Paul Chow
FPL1