Yaxing Li

dblp:121/0661 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distribution of statistics on separable permutations restricted by a flat POP
Alice L. L. Gao, Sergey Kitaev, Yaxing Li, Xuan Ruan
Discret. Appl. Math.3
2025 Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models
abstract
In recent years, there has been increasing attention on the capabilities of large-scale models, particularly in handling complex tasks that small-scale models are unable to perform. Notably, large language models (LLMs) have demonstrated ``intelligent'' abilities such as complex reasoning and abstract language comprehension, reflecting cognitive-like behaviors. However, current research on emergent abilities in large models predominantly focuses on the relationship between model performance and size, leaving a significant gap in the systematic quantitative analysis of the internal structures and mechanisms driving these emergent abilities. Drawing inspiration from neuroscience research on brain network structure and self-organization, we propose (i) a general network representation of large models, (ii) a new analytical framework — *Neuron-based Multifractal Analysis (NeuroMFA)* - for structural analysis, and (iii) a novel structure-based metric as a proxy for emergent abilities of large models. By linking structural features to the capabilities of large models, *NeuroMFA* provides a quantitative framework for analyzing emergent phenomena in large models. Our experiments show that the proposed method yields a comprehensive measure of the network's evolving heterogeneity and organization, offering theoretical foundations and a new perspective for investigating emergence in large models.
Xiongye Xiao, Heng Ping, Defu Cao, Yaxing Li, Yizhuo Zhou, Nikos Kanakaris, Paul Bogdan
ICLR5
2025 Deep Learning for Seismic Imaging in the Presence of Velocity Errors
abstract
Seismic migration is a tool to obtain images of underground structures; however, it requires accurate velocity models. Errors in estimated migration velocities lead to defocused and distorted migration images. We propose a deep learning method for accurate seismic imaging in the presence of velocity errors. Our idea is to correct the distorted common image gathers (CIGs) due to velocity errors by using a convolutional neural network (CNN). We design a CIG-to-CIG (CIG2CIG) CNN, in which both the inputs and outputs are CIGs. Furthermore, we apply velocity constraints to the CIG2CIG CNN, forming another velocity-constrained CIG2CIG (VC-CIG2CIG) CNN to perform the same task. To train the two CNNs, we create hundreds of true and wrong velocity models, which are applied to migration to produce true CIGs and distorted CIGs, respectively. Experiments demonstrate that the VC-CIG2CIG network is superior to the CIG2CIG network in correcting distorted CIGs and suppressing artifacts.
Sanfu Li, Yaxing Li, Yunzhi Shi, Xinming Wu
IEEE Geosci. Remote. Sens. Lett.2
2024 Neuro-Inspired Information-Theoretic Hierarchical Perception for Multimodal Learning
abstract
Integrating and processing information from various sources or modalities are critical for obtaining a comprehensive and accurate perception of the real world in autonomous systems and cyber-physical systems. Drawing inspiration from neuroscience, we develop the Information-Theoretic Hierarchical Perception (ITHP) model, which utilizes the concept of information bottleneck. Different from most traditional fusion models that incorporate all modalities identically in neural networks, our model designates a prime modality and regards the remaining modalities as detectors in the information pathway, serving to distill the flow of information. Our proposed perception model focuses on constructing an effective and compact information flow by achieving a balance between the minimization of mutual information between the latent state and the input modal state, and the maximization of mutual information between the latent states and the remaining modal states. This approach leads to compact latent state representations that retain relevant information while minimizing redundancy, thereby substantially enhancing the performance of multimodal representation learning. Experimental evaluations on the MUStARD, CMU-MOSI, and CMU-MOSEI datasets demonstrate that our model consistently distills crucial information in multimodal learning scenarios, outperforming state-of-the-art benchmarks. Remarkably, on the CMU-MOSI dataset, ITHP surpasses human-level performance in the multimodal sentiment binary classification task across all evaluation metrics (i.e., Binary Accuracy, F1 Score, Mean Absolute Error, and Pearson Correlation).
Xiongye Xiao, Gengshuo Liu, Defu Cao, Yaxing Li, Tianqing Fang, Mingxi Cheng, Paul Bogdan
ICLR6
2023 Performance Analysis of Broadband Countermeasure Cancellation in Multiple-access Datalink Networks
abstract
Multiple-access datalink communication networks (DCNs) are widely used in civil and military communication systems, such as global positioning system and joint tactical information distribution system. With the development of digital signal processing technologies, strong countermeasures, such as block jamming and aiming jamming bring significant challenges to the communication performance of multiple-access DCNs. Meanwhile, cancelling countermeasures whose direction and timing are unknown can be potentially enriched by antenna arrays, thanks to the strong spatial resolution of antenna arrays. This paper analyzes the cancelling performance of strong broadband countermeasures using antenna arrays in multiple-access DCNs. Signal and array models are established taking broadband countermeasures into consideration. Closeform solution expressions of discrete systems are derived, while the analytic expressions of the stability performance and convergence speed of broadband countermeasures cancellation are deduced. Extensive numerical results demonstrate the accuracy of the derivation. The elevational and horizontal dimension pattern of the antenna array are also presented through simulation.
Qiaran Lu, Fangmin He, Yaxing Li
VTC2023-Spring4
2022 Federated Reinforcement Learning for RIS-Aided Non-Orthogonal Multiple Access MEC
abstract
A novel reconfigurable intelligent surface (RIS) aided non-orthogonal multiple access mobile edge computing (NOMA-MEC) framework is proposed to release the heavy transmission delay of the edge devices (EDs) in next-generation wireless communication networks. We formulate a stochastic optimization problem that jointly optimizes the phase-shifter design of the RIS, the task offloading decision and the computing resource allocation of MEC, to minimize the overall transmission delay of all EDs in a long-term manner. The mathematical solution for the formulated optimization problem is a long-term offline policy which is non-trivial for conventional optimization approaches due to high computational complexity and stringent delay constraint. Therefore, we propose a federated reinforcement learning (FRL) approach for the formulated optimization problem to obtain the optimal solution taking advantages of the computing resource of all EDs. Moreover, a reputation-enabled ED selection scheme is proposed in the FRL approach that takes the task offloading history into consideration. The proposed RIS-aided NOMA-MEC framework is capable of outperforming conventional orthogonal multiple access (OMA) enabled RIS-MEC networks. The proposed FRL scheme achieves a near-optimal performance when the computional task is image classification in the MNIST and the IRIS dataset.
Yaxing Li, Fangmin He
VTC Fall2
2022 Deep Learning for Enhancing Multisource Reverse Time Migration
abstract
Reverse time migration (RTM) is a technique used to obtain high-resolution images of underground reflectors; however, this method is computationally intensive when dealing with large amounts of seismic data. Multi-source RTM can significantly reduce the computational cost by processing multiple shots simultaneously. However, multi-source-based methods frequently result in crosstalk artifacts in the migrated images, causing serious interference in the imaging signals. Plane-wave migration, as a mainstream multi-source method, can yield migrated images with plane waves in different angles by implementing phase encoding of the source and receiver wavefields; however, this method frequently requires a trade-off between computational efficiency and imaging quality. We propose a method based on deep learning for removing crosstalk artifacts and enhancing the image quality of plane-wave migration images. We designed a convolutional neural network that accepts an input of seven plane-wave images at different angles and outputs a clear and enhanced image. We built over 500 1024×256 velocity models, and employed each of them using plane-wave migration to produce raw images at 0°, ±10°, ±20°, and ±30° as input of the network. Labels are high-resolution images computed from the corresponding reflectivity models by convolving with a Ricker wavelet. Random sub-images with a size of 512×128 were used for training the network. Numerical examples demonstrated the effectiveness of the trained network in crosstalk removal and imaging enhancement. The proposed method is superior to both the conventional RTM and plane-wave RTM (PWRTM) in imaging resolution. Moreover, the proposed method requires only seven migrations, significantly improving the computational efficiency. In the numerical examples, the processing time required by our method was approximately 1.6% and 10% of that required by RTM and PWRTM, respectively.
Yaxing Li, Xinming Wu, Zhicheng Geng
IEEE Trans. Geosci. Remote. Sens.1
2021 Neural Noise Embedding for End-To-End Speech Enhancement with Conditional Layer Normalization
abstract
Most of the deep learning based speech enhancement methods focus on the modeling of complicated relationship between the noisy speech and the clean speech without the consideration of noise information. In order to cope with various complex noise scenes, we introduce a novel enhancement architecture that integrates a deep autoencoder with neural noise embedding. In this study, a new normalization method, termed conditional layer normalization (CLN), is introduced to improve the generalization of deep learning based speech enhancement approaches for unseen environments. The noise embedding is passed through the CLN layers to regularize the network for speech enhancement task. The proposed network can be adaptively adjusted according to different noise information extracted from the noisy speech input. The network in overall is trained in an end-to-end manner and the experimental results show that the proposed scheme produces satisfactory enhancement performance comparing the other methods. The visualization shows that our proposed network captures noise information, which is helpful to improve robustness to unseen environments for speech enhancement.
Xiaoqi Li 0011, Yaxing Li, Yuanjie Dong, Shengwu Xiong 0001
ICASSP3
2021 Variational Information Bottleneck Based Regularization for Speaker Recognition
Yuanjie Dong, Yaxing Li, Yunfei Zi, Xiaoqi Li 0011, Shengwu Xiong 0001
Interspeech3
2020 A Time-Frequency Network with Channel Attention and Non-Local Modules for Artificial Bandwidth Extension
abstract
Convolution neural networks (CNNs) have been achieving increasing attention for the artificial bandwidth extension (ABE) task recently. However, these methods use the flipped low-frequency phase to reconstruct speech signals, which may lead to the well-known invalid short-time Fourier Transform (STFT) problem. The convolutional operations only enable networks to construct informative features by fusing both channel-wise and spatial information within local receptive fields at each layer. In this paper, we introduce a Time-Frequency Network (TFNet) with channel attention (CA) and non-local (NL) modules for ABE. The TFNet exploits the information from both time and frequency domain branches concurrently to avoid the invalid STFT problem. To capture the channels and space dependencies, we incorporate the CA and NL modules to construct a proposed fully convolutional neural network for the time and frequency branches of TFNet. Experimental results demonstrate that the proposed method outperforms the competing method.
Yuanjie Dong, Yaxing Li, Xiaoqi Li 0011, Shan Xu 0007, Shengwu Xiong 0001
ICASSP2
2020 Bidirectional LSTM Network with Ordered Neurons for Speech Enhancement
Xiaoqi Li 0011, Yaxing Li, Yuanjie Dong, Shan Xu 0007, Shengwu Xiong 0001
INTERSPEECH2
2020 Incremental small sphere and large margin for online recognition of communication jamming
Yaxing Li, Songhu Ge, Jinling Xing
Appl. Intell.3
2020 Single and multiple frame coding of LSF parameters using deep neural network and pyramid vector quantizer
Yaxing Li, Ying Kang
Speech Commun.1
2019 Densely Connected Network with Time-frequency Dilated Convolution for Speech Enhancement
abstract
The data driven speech enhancement approaches using regression-based deep neural network usually result in enormous number of model parameters, which increase the computational load and the difficulty of model training. In order to improve the model efficiency, we propose a densely connected network with time-frequency (T-F) dilated convolution for speech enhancement. The T-F dilated convolution block is designed to enlarge the receptive field and capture the contextual information in both temporal and frequency domains. Considering the computational efficiency, the 1-D convolution with the bottleneck structure is exploited in the T-F convolution block. Each T-F convolution block is then densely connected to ensure maximum information flow between layers and alleviate the vanishing gradient problem of the network. The experimental results reveal that the proposed scheme not only improves the computational efficiency significantly but also produces satisfactory enhancement performance comparing the competing methods.
Yaxing Li, Xiaoqi Li 0011, Yuanjie Dong, Shan Xu 0007, Shengwu Xiong 0001
ICASSP1
2019 A Convolutional Neural Network with Non-Local Module for Speech Enhancement
Xiaoqi Li 0011, Yaxing Li, Shan Xu 0007, Yuanjie Dong, Xinrong Sun, Shengwu Xiong 0001
INTERSPEECH2
2018 Multi-frame Quantization of LSF Parameters Using a Deep Autoencoder and Pyramid Vector Quantizer
Yaxing Li, Eshete Derb Emiru, Shengwu Xiong 0001, Anna Zhu, Pengfei Duan 0005, Yichang Li
INTERSPEECH1
2018 Multi-frame Coding of LSF Parameters Using Block-Constrained Trellis Coded Vector Quantization
Yaxing Li, Shan Xu 0007, Shengwu Xiong 0001, Anna Zhu, Pengfei Duan 0005, Yueming Ding
INTERSPEECH1
2017 Deep neural network-based linear predictive parameter estimations for speech enhancement
abstract
This study presents a speech enhancement technique to improve noise corrupted speech via deep neural network (DNN)‐based linear predictive (LP) parameter estimations of speech and noise. With regard to the LP coefficient estimation, an enhanced estimation method using a DNN with multiple layers was proposed. Excitation variances were then estimated via a maximum‐likelihood scheme using observed noisy speech and estimated LP coefficients. A time‐smoothed Wiener filter was further introduced to improve the enhanced speech quality. Performance was evaluated via log spectral distance, a composite multivariate adaptive regression splines modelling‐based measure, and a segmental signal‐to‐noise ratio. The experimental results revealed that the proposed scheme outperformed competing methods.
Yaxing Li, Sangwon Kang
IET Signal Process.1
2016 Artificial bandwidth extension using deep neural network-based spectral envelope estimation and enhanced excitation estimation
abstract
The authors propose a robust artificial bandwidth extension (ABE) technique to improve narrowband (NB) speech signal quality using an enhanced spectrum envelope and excitation estimation. For envelope estimation, they propose an enhanced envelope estimation method using a deep neural network with multiple layers. For excitation estimation, they use a whitened NB excitation signal that is generated by passing the excitation signal through a whitening filter. An adaptive spectral double shifting method is introduced to obtain an enhanced wideband (WB) excitation signal. The proposed ABE system is applied to the decoded output of an adaptive multi‐rate (AMR) codec at 12.2 kbps. They evaluate its performance using log spectral distortion, a WB perceptual evaluation of speech quality, and a formal listening test. The objective and subjective evaluations confirm that the proposed ABE system provides better speech quality than AMR at the same bit rate.
Yaxing Li, Sangwon Kang
IET Signal Process.1