VLDB 2026 Research / reviewers in the wild / expert
Yanxin Hu
dblp:62/8208 · also Yanxing Hu
· DBLP profile ↗
16ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physics-augmented federated continual learning for rotating machinery fault diagnosis
Yanxin Hu, Yan Huang 0032, Zhenzhen Xie 0002, Junjie Pang |
Neurocomputing | 1 |
| 2026 | FedDiDy: Federated Class-Incremental Fault Diagnosis for Industrial IoT Rotating Machinery Under Dynamic Edge ParticipationabstractFederated learning (FL) is promising for privacy-sensitive fault diagnosis in the Industrial Internet of Things (IIoT). However, real-world deployments must address the coexistence of class-incremental fault evolution and dynamic edge participation, which leads to fragmented class exposure, catastrophic forgetting, and aggregation bias. To address this issue, we propose FedDiDy, a unified framework for federated class-incremental fault diagnosis under dynamic participation. FedDiDy combines a multi-head classifier for task decoupling, a raw-data-free conditional generator for historical knowledge replay, and a perception-aware aggregation mechanism for bias mitigation. Experiments on four datasets under Bernoulli, Cyclic, and Markov participation patterns show that FedDiDy consistently outperforms the compared baselines, with up to 15.1% absolute improvement in average accuracy and 14.44% reduction in forgetting rate. Yanxin Hu, Yan Huang 0032, Zhenzhen Xie 0002, Junjie Pang, Zhipeng Cai 0001 |
IEEE Internet Things J. | 1 |
| 2026 | KTR: Structure-aware replay for continual learning on hypergraphsabstractClass-incremental continual learning on hypergraphs is challenging under limited replay memory. Buffered samples do not contribute equally to preserving historical higher-order structures. Existing replay methods mainly use random selection or loss-based selection. However, they often ignore structural cohesiveness. As a result, structurally important samples may be missed. We propose KTR ( K nowledge-preserving T russ-based R eplay), a structure-aware replay framework for continual learning on hypergraphs. KTR prioritizes buffered samples by combining hypertuss-based structural importance with current loss-based utility. It further performs constrained replay through structural filtering and probabilistic sampling. The framework supports both node classification and hyperedge classification. Experiments on four continual hypergraph benchmarks show that KTR improves replay performance under limited-memory settings, with the clearest gains on temporal hyperedge-classification benchmarks. Under the HGNN+ backbone with memory budget b = 0.1 , KTR improves ACC by up to 20.32 percentage points and reduces forgetting by up to 24.07 percentage points relative to PBR on MAG-Top20K. On node-classification benchmarks, KTR remains competitive with strong replay baselines, but it does not uniformly dominate all methods on every dataset. These results support the use of structure-aware replay in continual hypergraph learning, especially when higher-order structural cohesion provides informative replay signals. Yanxin Hu, Zhenzhen Xie 0002, Junjie Pang |
Knowl. Based Syst. | 1 |
| 2025 | CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-car Speech Separation with Distributed Heterogeneous Arrays
Runduo Han, Yanxin Hu, Yihui Fu, Yukai Jv, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2025 | Retrieval of Soil Moisture and Vegetation Water Content From Passive Microwave Remote Sensing: A Local-Scale Evaluation via Ground-Based Multichannel Radiometry
Chunfeng Ma, Xin Li 0029, Shuguo Wang, Yang Zhang 0142, Yanxin Hu, Liyun Dai, Zengyan Wang, Tao Che |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | A Multifrequency Radiometry Experiment Over an Agricultural Field Toward Microwave Emission Model CalibrationabstractPassive microwave remote sensing has witnessed unprecedented progress in soil moisture (SM) estimation over the decades. However, it is challenging to estimate SM accurately due to the insufficient understanding of microwave emission mechanisms. A ground-based radiometry experiment is undertaken over an agricultural field, toward the reexamination and improvement of the microwave emission models and retrieval algorithms of SM and vegetation water content (VWC). This article reports the preliminary analysis of the experimental data and the calibration of the$\tau $–$\omega $model against the collected measurements. First, the collected multifrequency dual-polarized brightness temperature (TB) reflects the temporal variation of surface SM and a significantly negative correlation between them is observed, with the coefficient of determination ($R^{2}$) and slope (S) of a linear fitting line ranging from 0.036 to 0.367 and from −16.7 to −81.5, respectively. Surface roughness and VWC impact the relationship between TB and SM, with variable$R^{2}$and S observed. Second, the calibrated parameters have improved the model performance, with$R^{2}$greater than 0.65 and root mean standard error (RMSE) less than 4.7 K at all frequencies and polarizations. The parameter values are frequency- and polarization-dependent, and the best performance of the model simulation is observed at V-polarization of L- and Ku-bands, with$R^{2} =0.80$and RMSE =4.69 K at the L-band and$R^{2} =0.74$and RMSE =2.9 K at the Ku-band. Overall, the experiment has provided valuable datasets for calibrating forward models and the calibrated model will facilitate the improvement of surface parameters (e.g., SM and VWC) retrieval. Chunfeng Ma, Zengyan Wang, Liyun Dai, Yanxin Hu, Yang Zhang 0142, Tao Che, Leilei Dong, Xin Li 0029 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Federated Learning for Edge Heterogeneous Object Detection Algorithm
Yanxin Hu |
WASA (2) | 1 |
| 2023 | Lightweight object detection algorithm for robots with improved YOLOv5
Yanxin Hu, Zhiyu Chen 0006, Jianwei Guo 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | Overbooking-enabled Virtual Machine Deployment Approach in Mobile Edge ComputingabstractMobile Edge Computing (MEC) integrates computing, storage and other resources on the edge of the network and constructs a unified user service platform. Then, according to the principle of nearest service, MEC responds to the task requests of the edge nodes in time and effectively processes them. In MEC, edge servers are virtualized into several slots so that resources can be shared among different mobile users. However, there are many unpredictable risks in MEC, these risks can cause edge servers to fail, the virtual machine deployed in the server slot fails and the task cannot be executed normally. The introduction of primary-backup virtual machines solves this problem well. However, when the primary virtual machine is working normally, its backup virtual machine is idle, this will result in a waste of resources. In order to improve the resource utilization of the system, this paper firstly overbooks the backup virtual machine reasonably, and then formulates the virtual machine deployment problem as a combinatorial optimization problem. Finally, Virtual Machine Deployment Algorithm (VMDA) is proposed based on genetic algorithm. With the increase of the number of algorithm iterations and the population size of the virtual machine deployment scheme, there may be more optimal virtual machine deployment scheme individuals in the population. Therefore, the algorithm can obtain the approximate optimal value of resource utilization within the risk range allowed by the system, and the algorithm is compared with two other typical bin packing algorithms. The results confirm that VMDA outperforms the other two algorithms. Bingyi Hu, Jixun Gao, Quanzhen Huang, Huaichen Wang, Yanxin Hu, Jialei Liu, Yanmin Ge |
ICSS | 5 |
| 2021 | Conferencingspeech Challenge: Towards Far-Field Multi-Channel Speech Enhancement for Video ConferencingabstractThe ConferencingSpeech 2021 challenge is proposed to stimulate research on far-field multi-channel speech enhancement for video conferencing. The challenge consists of two separate tasks: 1) Task 1 is multi-channel speech enhancement with single microphone array and focusing on practical application with real-time requirement and 2) Task 2 is multi-channel speech enhancement with multiple distributed micro-phone arrays, which is a non-real-time track and does not have any constraints so that participants could explore any algorithms to obtain high speech quality. Targeting the real video conferencing room application, the challenge database was recorded from real speakers and all recording facilities were located by following the real setup of conferencing room. In this challenge, we open-sourced the list of open source clean speech and noise datasets, simulation scripts, and a baseline system for participants to develop their own system. The final ranking of the challenge will be decided by the subjective evaluation which is performed using Absolute Category Ratings (ACR) to estimate Mean Opinion Score (MOS), speech MOS (S-MOS), and noise MOS (N-MOS). This paper describes the challenge, tasks, datasets, subjective evaluation, and challenge results. The baseline system which is a complex ratio mask based neural network and its experimental results are also presented. Wei Rao 0002, Yihui Fu, Yanxin Hu, Yvkai Jv, Jiangyu Han, Zhongjie Jiang, Lei Xie 0001, Yannan Wang, Shinji Watanabe 0001, Zheng-Hua Tan, Hui Bu, Shidong Shang |
ASRU | 3 |
| 2021 | AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference ScenarioabstractIn this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech dataset collected by 8-channel circular microphone array for speech processing in conference scenario. The dataset consists of 211 recorded meeting sessions, each containing 4 to 8 speakers, with a total length of 120 hours. This dataset aims to bridge the advanced research on multi-speaker processing and the practical application scenario in three aspects. With real recorded meetings, AISHELL-4 provides realistic acoustics and rich natural speech characteristics in conversation such as short pause, speech overlap, quick speaker turn, noise, etc. Meanwhile, accurate transcription and speaker voice activity are provided for each meeting in AISHELL-4. This allows the researchers to explore different aspects in meeting processing, ranging from individual tasks such as speech front-end processing, speech recognition and speaker diarization, to multi-modality modeling and joint optimization of relevant tasks. Given most open source dataset for multi-speaker tasks are in English, AISHELL-4 is the only Mandarin dataset for conversation speech, providing additional value for data diversity in speech community. We also release a PyTorch-based training and evaluation framework as baseline system to promote reproducible research in this field. Yihui Fu, Luyao Cheng, Shubo Lv, Yukai Jv, Yuxiang Kong, Zhuo Chen 0006, Yanxin Hu, Lei Xie 0001, Jian Wu 0027, Hui Bu, Jun Du 0002, Jingdong Chen |
Interspeech | 7 |
| 2021 | DCCRN+: Channel-Wise Subband DCCRN with SNR Estimation for Speech EnhancementabstractDeep complex convolution recurrent network (DCCRN), which extends CRN with complex structure, has achieved superior performance in MOS evaluation in Interspeech 2020 deep noise suppression challenge (DNS2020).This paper further extends DCCRN with the following significant revisions.We first extend the model to sub-band processing where the bands are split and merged by learnable neural network filters instead of engineered FIR filters, leading to a faster noise suppressor trained in an end-to-end manner.Then the LSTM is further substituted with a complex TF-LSTM to better model temporal dependencies along both time and frequency axes.Moreover, instead of simply concatenating the output of each encoder layer to the input of the corresponding decoder layer, we use convolution blocks to first aggregate essential information from the encoder output before feeding it to the decoder layers.We specifically formulate the decoder with an extra a priori SNR estimation module to maintain good speech quality while removing noise.Finally a post-processing module is adopted to further suppress the unnatural residual noise.The new model, named DCCRN+, has surpassed the original DCCRN as well as several competitive models in terms of PESQ and DNSMOS, and has achieved superior performance in the new Interspeech 2021 DNS challenge. Shubo Lv, Yanxin Hu, Lei Xie 0001 |
Interspeech | 2 |
| 2021 | F-T-LSTM Based Complex Network for Joint Acoustic Echo Cancellation and Speech EnhancementabstractWith the increasing demand for audio communication and online conference, ensuring the robustness of Acoustic Echo Cancellation (AEC) under the complicated acoustic scenario including noise, reverberation and nonlinear distortion has become a top issue. Although there have been some traditional methods that consider nonlinear distortion, they are still inefficient for echo suppression and the performance will be attenuated when noise is present. In this paper, we present a real-time AEC approach using complex neural network to better modeling the important phase information and frequency-time-LSTMs (F-T-LSTM), which scan both frequency and time axis, for better temporal modeling. Moreover, we utilize modified SI-SNR as cost function to make the model to have better echo cancellation and noise suppression (NS) performance. With only 1.4M parameters, the proposed approach outperforms the AEC-challenge baseline by 0.27 in terms of Mean Opinion Score (MOS). Yuxiang Kong, Shubo Lv, Yanxin Hu, Lei Xie 0001 |
Interspeech | 4 |
| 2021 | DESNet: A Multi-Channel Network for Simultaneous Speech Dereverberation, Enhancement and SeparationabstractIn this paper, we propose a multi-channel network for simultaneous speech dereverberation, enhancement and separation (DESNet). To enable gradient propagation and joint optimization, we adopt the attentional selection mechanism of the multi-channel features, which is originally proposed in end-to-end unmixing, fixed-beamforming and extraction (E2E-UFE) structure. Furthermore, the novel deep complex convolutional recurrent network (DCCRN) is used as the structure of the speech unmixing and the neural network based weighted prediction error (WPE) is cascaded before-hand for speech dereverberation. We also introduce the staged SNR strategy and symphonic loss for the training of the network to further improve the final performance. Experiments show that in non-dereverberated case, the proposed DESNet outperforms DCCRN and most state-of-the-art structures in speech enhancement and separation, while in dereverberated scenario, DESNet also shows improvements over the cascaded WPE-DCCRN networks. Yihui Fu, Jian Wu 0027, Yanxin Hu, Mengtao Xing, Lei Xie 0001 |
SLT | 3 |
| 2020 | DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech EnhancementabstractSpeech enhancement has benefited from the success of deep learning in terms of intelligibility and perceptual quality.Conventional time-frequency (TF) domain methods focus on predicting TF-masks or speech spectrum, via a naive convolution neural network (CNN) or recurrent neural network (RNN).Some recent studies use complex-valued spectrogram as a training target but train in a real-valued network, predicting the magnitude and phase component or real and imaginary part, respectively.Particularly, convolution recurrent network (CRN) integrates a convolutional encoder-decoder (CED) structure and long short-term memory (LSTM), which has been proven to be helpful for complex targets.In order to train the complex target more effectively, in this paper, we design a new network structure simulating the complex-valued operation, called Deep Complex Convolution Recurrent Network (DCCRN), where both CNN and RNN structures can handle complex-valued operation.The proposed DCCRN models are very competitive over other previous networks, either on objective or subjective metric.With only 3.7M parameters, our DCCRN models submitted to the Interspeech 2020 Deep Noise Suppression (DNS) challenge ranked first for the real-time-track and second for the non-real-time track in terms of Mean Opinion Score (MOS). Yanxin Hu, Shubo Lv, Mengtao Xing, Yihui Fu, Jian Wu 0027, Bihong Zhang, Lei Xie 0001 |
INTERSPEECH | 1 |
| 2009 | SNR Degradation due to Carrier Frequency Offset in Amplify-and-Forward Relay System for Fading ChannelsabstractIn this paper, we analyze signal-to-noise ratio (SNR) performance of amplify-and-forward (AF) relay system in the presence of carrier frequency offset (CFO) for fading channels. The expression for the average SNR is derived under one-relay-node scenario, and is further extended to multiple-relay-node scenario. It is shown that the SNR is very sensitive to CFO and the sensitivity of SNR to CFO is mainly determined by the power of the corresponding link channel and gain factor. The analytical results are validated through Monte Carlo simulations for both flat and frequency-selective fading channels. Yanxin Hu, Yanxiang Jiang, Xiaohu You 0001 |
VTC Fall | 1 |