Feng Deng

dblp:14/358 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Utilizing large language models for integrating document-level contextual semantic into pseudo-relevance feedback
abstract
Pseudo-Relevance Feedback (PRF) is a key technique in information retrieval (IR). Traditional implementations rely on statistical information, such as term frequency, for precise matching and relevance assessment. However, these methods struggle to fully capture the deep semantic integrity of query terms, especially in handling polysemy, high semantic relevance, and long-document comprehension. To address these challenges, this paper innovatively proposes a large language model-assisted PRF probabilistic model. The model first employs a precise matching algorithm to evaluate and determine the term-level weights, and then uses a large language model to encode the contextual relationships within the query and feedback documents, thereby accurately acquiring the global semantic weights of terms relevant to the query at the document level. By adjusting a balancing factor to allocate weights between these two components, the model comprehensively selects expanded terms for constructing a new query representation and executing query expansion (QE). This model not only facilitates approximate matching through the integration of global semantic features of documents but also effectively combines with the precise matching information of traditional PRF models, enabling a comprehensive and accurate optimization of queries from a broader perspective. To validate effectiveness, extensive empirical analyses on five TREC datasets assess performance across key metrics such as MAP, P@10, NDCG, and MRR. Experimental results show significant improvements over baseline models. Comparative analyses and case studies confirm that the expanded terms maintain high semantic relevance and consistency with the original query while preserving diversity and effectively capturing global document semantics, establishing an efficient QE mechanism.
Min Pan, Wenrui Xiong, Junmei Wang, Feng Deng, Ellen Anne Huang, Jinguang Chen, Jimmy Huang 0001
Knowl. Based Syst.5
2025 Optical Flow-Augmented Dual-Stream Network for Left Ventricular Ejection Fraction Prediction
Feng Deng, Qinghua Fu, Lin Guo 0014, Ying An
ISBRA (2)1
2025 Region-aware discriminative learning GAN for super-resolution reconstruction of infrared imagery
Feng Deng, Shuaichao Wang
Neurocomputing1
2024 LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
Wenhao Guan, Wangjin Zhou, Feng Deng, Qingyang Hong
INTERSPEECH5
2023 HCoop: A Cooperative and Hybrid Resource Scheduling for Heterogeneous Jobs in Clouds
abstract
Heterogeneous workloads composed of long batch jobs and short term jobs are increasingly common in modern datacenters, which poses a challenge to resource scheduling for improving throughput and resource utilization. Achieving high resource utilization and throughput is important to the cloud provider for achieving high revenue. To address this challenge, we propose HCoop, a cooperative and hybrid resource scheduling for high resource utilization and throughput in clouds. HCoop first uses the machine learning algorithm to classify jobs into two categories (long jobs and short jobs) based on the extracted features. Then, it allocates the regular virtual machines (VMs) to tasks of long jobs with higher SLO (Service Level Objective) availability guarantee, and selectively allocates the spot instances to tasks of short jobs for improving the overall resource utilization while ensuring the SLO availability for short jobs. Also, HCoop leverages the complementary of tasks’ requirements on different resource types and the heterogeneity of jobs/tasks in job/task size, and it packs complementary tasks whose demands on multiple resource types are complementary to each other and allocates them to a VM to further increase the resource utilization. Extensive experimental results based on a real cluster and Amazon EC2 cloud service show that HCoop achieves high resource utilization and throughput compared to existing strategies.
Ying Mao 0001, Feng Deng
CloudCom7
2023 Dynamic TF-TDNN: Dynamic Time Delay Neural Network Based on Temporal-Frequency Attention for Dialect Recognition
abstract
Dialect recognition aims to recognize dialect categories in utterances, which has been applied in many audio applications. Recently, various Time Delayed Neural Network (TDNN) based AI models are proposed to solve dialect recognition problems, such as D-TDNN, DMC-TDNN, and ECAPA-TDNN, however, most of them only perform temporal attention in the last statistical pooling layer of the TDNN network, which ignores the importance of simultaneously capturing both frequency and temporal key information in utterances under different receptive fields. In contrast, we introduce a hybrid attention mechanism in both the temporal and frequency domain, called the TF-attention module, which adaptively pays more attention to the indeed important frames and the frame-level important information under different receptive fields for dialect recognition. Moreover, we are the first to introduce a dynamic architecture mechanism in the field of dialect recognition to dynamically reduce the computational cost and the number of parameters of models. We evaluate the proposed dynamic TF-TDNN on the OLR challenge AP20-OLR-dialect task and achieve State-Of-The-Art (SOTA) performance with fewer model parameters.
Chao Liao, Jinwen Huang, Huan Yuan, Jianchao Tan, Feng Deng, Chengru Song
ICASSP7
2023 NAS-DYMC: NAS-Based Dynamic Multi-Scale Convolutional Neural Network for Sound Event Detection
abstract
CNN+RNN models have become the mainstream approach for semi-supervised sound event detection, and the CNN part is mainly a stack of several 2D convolutional layers to capture the representations of the time-frequency features. However, conventional 2D convolution is of limited ability in capturing detailed information about acoustic events. In this paper, to enhance the representation ability of CNN, we propose NAS-DYMC, a NAS-based dynamic multi-scale convolutional neural network to extract a more effective acoustic representation. Specifically, multi-scale convolution can capture the characteristics of sound events with different time-frequency distributions and dynamic convolution enhances the representation capability of conventional convolution by adapting attention weights onto basis kernels. Furthermore, a neural architecture search (NAS) method is adopted to find the optimal network architecture from the search space consisting of various dynamic multi-scale convolutions for the DCASE 2021 Task4 dataset. Experimental results demonstrate the superiority of our proposed method.
Feng Deng, Jianchao Tan, Chengru Song
ICASSP3
2023 Image-driven Audio-visual Universal Source Separation
Chenxing Li, Ye Bai 0001, Feng Deng
INTERSPEECH4
2023 FASONet: A Feature Alignment-Based SAR and Optical Image Fusion Network for Land Use Classification
Feng Deng, Meiyu Huang, Xueshuang Xiang
PRCV (10)1
2022 CrossCas: A Novel Cross-Platform Approach for Predicting Cascades in Online Social Networks with Hidden Markov Model
abstract
Information sharing through online social networks (OSNs) facilitates quick discovery and consumption of information online. Many OSNs such as Facebook, Twitter provide resharing or reposting features, which allows users to share others' content with their own friends or followers. As content is shared from person to person, cascades of information-sharing can occur. There are many existing works focusing on analyzing and characterizing the cascades in OSNs. However, previous works focus on the analysis and characterization of cascades without providing a solution to accurately predict cascades. Although some methods for cascade prediction have been proposed recently, their methods work in social networks such as Facebook (or Twitter), and do not work well simultaneously in multiple OSNs such as Software Social Network (SSN) GitHub, Twitter and Reddit because GitHub, Twitter and Reddit have different social activity patterns. In this paper, we first perform a thorough analysis of cascades in multiple OSNs: GitHub, Twitter and Reddit, and identify the cascades of information-sharing. We then propose CrossCas, a novel cross-platform approach for predicting cascades in multiple OSNs with Hidden Markov Model (HMM). The experimental results show that our proposed method achieves high performance.
Xiaonan Zhang 0001, Richard A. Aló, Xiuzhen Huang, Long Cheng 0003, Feng Deng
GLOBECOM6
2022 EAD-Conformer: a Conformer-Based Encoder-Attention-Decoder-Network for Multi-Task Audio Source Separation
abstract
In this paper, we propose a Conformer-based network to improve the performance of multi-task audio source separation. This network, named EAD-Conformer, employs Conformer blocks to capture both local and global information, and an encoder-attention-decoder manner encourages the network to perform attentive modeling based on different sources. Specifically, EAD-Conformer first parses out the feature representations from the mixture by a Conformer-based encoder. Then, an attention module extracts selective information for each track and bridges encoder and decoders. Finally, three decoders respectively process attentive features and generate output masks for different sources. In addition, the proposed discriminate loss further enlarges the distance between different sources. Experiments demonstrate the effectiveness of EAD-Conformer, which achieves 13.37 dB, 11.41 dB, 10.56 dB signal-to-distortion ratio improvement on speech, music, noise track, respectively, and shows advantages over several well-known models.
Chenxing Li, Feng Deng, Zhongyuan Wang 0006
ICASSP3
2022 iCNN-Transformer: An improved CNN-Transformer with Channel-spatial Attention and Keyword Prediction for Automated Audio Captioning
Jun Wang 0077, Feng Deng
INTERSPEECH3
2022 Conformer Space Neural Architecture Search for Multi-Task Audio Separation
Shun Lu 0001, Chenxing Li, Jianchao Tan, Feng Deng, Chengru Song
INTERSPEECH6
2022 WA-Transformer: Window Attention-based Transformer with Two-stage Strategy for Multi-task Audio Source Separation
Chenxing Li, Feng Deng, Shun Lu 0001, Jianchao Tan, Chengru Song
INTERSPEECH3
2021 Multi-Task Audio Source Separation
abstract
The audio source separation tasks, such as speech enhancement, speech separation, and music source separation, have achieved impressive performance in recent studies. The powerful modeling capabilities of deep neural networks give us hope for more challenging tasks. This paper launches a new multi-task audio source separation (MTASS) challenge to separate the speech, music, and noise signals from the monaural mixture. First, we introduce the details of this task and generate a dataset of mixtures containing speech, music, and background noises. Then, we propose an MTASS model in the complex domain to fully utilize the differences in spectral characteristics of the three audio signals. In detail, the proposed model follows a two-stage pipeline, which separates the three types of audio signals and then performs signal compensation separately. After comparing different training targets, the complex ratio mask is selected as a more suitable target for the MTASS. The experimental results also indicate that the residual signal compensation module helps to recover the signals further. The proposed model shows significant advantages in separation performance over several well-known separation models.
Chenxing Li, Feng Deng
ASRU3
2021 SpeechNAS: Towards Better Trade-Off Between Latency and Accuracy for Large-Scale Speaker Verification
abstract
Recently, x-vector [1] has been a successful and popular approach for speaker verification, which employs a time delay neural network (TDNN) and statistics pooling to extract speaker characterizing embedding from variable-length utterances. Improvement upon the x-vector has been an active research area, and enormous neural networks have been elaborately designed based on the x-vector, e.g., extended TDNN (E-TDNN) [2], factorized TDNN (F-TDNN) [3], and densely connected TDNN (D-TDNN) [4]. In this work, we try to identify the optimal architectures from a TDNN based search space employing neural architecture search (NAS), named SpeechNAS. Leveraging the recent advances in the speaker recognition, such as high-order statistics pooling, multi-branch mechanism, D-TDNN and angular additive margin softmax (AAM) loss with a minimum hyper-spherical energy (MHE), SpeechNAS automatically discovers five network architectures, from SpeechNAS-1 to SpeechNAS-5, of various numbers of parameters and GFLOPs on the large-scale text-independent speaker recognition dataset VoxCelebl. Our derived best neural network achieves an equal error rate (EER) of 1.02% on the standard test set of VoxCelebl, which surpasses previous TDNN based state-of-the-art approaches by a large margin.
Wentao Zhu 0001, Tianlong Kong, Shun Lu 0001, Feng Deng, Sen Yang 0004, Ji Liu 0002
ASRU6
2020 NAAGN: Noise-Aware Attention-Gated Network for Speech Enhancement
Feng Deng
INTERSPEECH1
2019 Automatic Singing Evaluation without Reference Melody Using Bi-dense Neural Network
abstract
Automatic singing evaluation without reference melody has long been a difficult problem. This paper aims to pilot a novel data driven approach to tackle this artistic problem. We constructed a large scale dataset and designed an innovative Bi-Dense neural network which can address this task efficiently. Though the singing evaluation is quite a subjective task and depends a lot on listeners' preferences, we showed that a specific group has consistency on the singing evaluations, and it is possible to train a model to learn the subjective preferences of this group. In this paper, a large amount of singing clips and corresponding human gradings were collected. And an elaborate designed Bi-DenseNet was trained to discriminate the good singings from the poor ones. The experiments demonstrated the proposed network performs better than the existing networks for singing evaluation task.
Feng Deng
ICASSP3
2016 Probabilistic Network-Aware Task Placement for MapReduce Scheduling
abstract
Maximizing data locality in task scheduling is critical for the performance of MapReduce job execution. Manyexisting works on MapReduce scheduling decide the placementof map and reduce tasks on a coarse granularity of locationsmeasured by located machines and racks. They do not explicitlyconsider the network topology and data transmission cost, whichmay cause task straggling and degrade the job performance. Inorder to improve MapReduce job performance, in this paper, we consider the task placement with the goal of minimizing theoverall data transmission cost for a job execution while balancingthe transmission cost reduction and resource utilization. Wepropose a probabilistic network-aware scheduling algorithm thatselects a task (map task or reduce task) to be scheduled on a givenavailable task slot that leads to the minimum transmission costamong the task candidates, and then schedule the selected taskon the slot with a probability determined by its transmission cost, a lower expected transmission cost leads to a higher probabilityand vice versa. We also propose a method to more accuratelyestimate the intermediate data size based on the progress ofmap tasks, which is needed to calculate the transmission cost ofreduce tasks but is unknown at the time of reduce task scheduling. We implement our probabilistic network-aware schedulingalgorithm on Apache Hadoop and conduct experiments on ahigh-performance computing platform. The experimental resultsshow that our scheduling algorithm outperforms the previousapproaches in terms of job completion time and cluster resource utilization.
Haiying Shen, Ankur Sarker, Lei Yu 0002, Feng Deng
CLUSTER4
2016 Speech enhancement based on AR model parameters estimation
Feng Deng, Changchun Bao
Speech Commun.1
2015 Sparse HMM-based speech enhancement method for stationary and non-stationary noise environments
abstract
We propose a sparse hidden Markov model (HMM)-based single-channel speech enhancement method that models the speech and noise gains accurately in both stationary and nonstationary environments. The objective function is augmented with an lp regularization term resulting in a sparse autoregressive HMM (SARHMM). The method encourages sparsity in the speech- and noise- modeling, which eliminates the ambiguity between noise and speech spectra and, as a consequence, provides improved tracking of the changes of both spectral shapes and power levels of non-stationary noise. Using the modeled speech and noise SARHMMs, we first construct an estimator to estimate the noise spectrum. Then a Bayesian speech estimator is used to obtain the enhanced speech. The test results indicate that the proposed speech enhancement scheme performs much better than the reference methods in non-stationary environments, while providing state-of-the-art performance for stationary conditions.
Feng Deng, Changchun Bao, W. Bastiaan Kleijn
ICASSP1
2015 Enhancing Primary Students' Online Talks and Epistemic Beliefs through a Knowledge Building Community Approach
Ching Sing Chai, Feng Deng, Angela Lay Hong Koh, Guan Hui Quek
ICCE2
2015 A data-driven speech enhancement method based on modeled long-range temporal dynamics
Changchun Bao, Feng Bao 0003, Feng Deng
INTERSPEECH4
2015 Sparse Hidden Markov Models for Speech Enhancement in Non-Stationary Noise Environments
abstract
We propose a sparse hidden Markov model (HMM)-based single-channel speech enhancement method that models the speech and noise gains accurately in non-stationary noise environments. Autoregressive models are employed to describe the speech and noise in a unified framework and the speech and noise gains are modeled as random processes with memory. The likelihood criterion for finding the model parameters is augmented with an lp regularization term resulting in a sparse autoregressive HMM (SARHMM) system that encourages sparsity in the speech- and noise- modeling. In the SARHMM only a small number of HMM states contribute significantly to the model of each particular observed speech segment. As it eliminates ambiguity between noise and speech spectra, the sparsity of speech and noise modeling helps to improve the tracking of the changes of both spectral shapes and power levels of non-stationary noise. Using the modeled speech and noise SARHMMs, we first construct a noise estimator to estimate the noise power spectrum. Then, a Bayesian speech estimator is derived to obtain the enhanced speech signal. The subjective and objective test results indicate that the proposed speech enhancement scheme can achieve a larger segmental SNR improvement, a lower log-spectral distortion and a better speech quality in stationary noise conditions than state-of-the-art reference methods. The advantage of the new method is largest for non-stationary noise conditions .
Feng Deng, Changchun Bao, W. Bastiaan Kleijn
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Speech enhancement using generalized weighted β-order spectral amplitude estimator
Feng Deng, Feng Bao 0003, Changchun Bao
Speech Commun.1
2013 A speech enhancement method by coupling speech detection and spectral amplitude estimation
Feng Deng, Changchun Bao, Feng Bao 0003
INTERSPEECH1
2006 Locally adjusted cubic-spline capping for reconstructing seasonal trajectories of a satellite-derived surface parameter
abstract
Satellite-derived vegetation indices and their resulting surface parameters, such as the leaf area index (LAI), are inevitably affected by the atmosphere. Errors in the atmospheric corrections can often be easily identified in a seasonal trajectory of a surface parameter because the atmospheric effect generally causes erratic reductions in vegetation indices. A locally adjusted cubic-spline capping (LACC) method is developed here to screen affected data points in a pixel and to replace them through temporal interpolation. In LACC, a variable local smoothing parameter, which controls the local smoothness of the fitted curve, is automatically determined according to the local curvature of the original seasonal variation pattern. An iteration procedure is designed to produce a seasonal capping curve by progressively replacing abnormally low values with fitted values. This method has two advantages over existing methods based on harmonics, namely: 1) cubic splines are flexible for simulating a wide range of seasonal variation patterns and 2) a variable local smoothing parameter allows the fitted capping curve to mimic either rapid or slow variation patterns in various seasons. The capping curve is also mathematically differentiable for further applications. The effectiveness of this method is demonstrated through case studies for several cover types in China and processing a series of Moderate Resolution Imaging Spectroradiometer LAI images of China in 2001
Jing M. Chen, Feng Deng, Mingzhen Chen
IEEE Trans. Geosci. Remote. Sens.2
2006 Algorithm for global leaf area index retrieval using satellite imagery
abstract
Leaf area index (LAI) is one of the most important Earth surface parameters in modeling ecosystems and their interaction with climate. Based on a geometrical optical model (Four-Scale) and LAI algorithms previously derived for Canada-wide applications, this paper presents a new algorithm for the global retrieval of LAI where the bidirectional reflectance distribution function (BRDF) is considered explicitly in the algorithm and hence removing the need of doing BRDF corrections and normalizations to the input images. The core problem of integrating BRDF into the LAI algorithm is that nonlinear BRDF kernels that are used to relate spectral reflectances to LAI are also LAI dependent, and no analytical solution is found to derive directly LAI from reflectance data. This problem is solved through developing a simple iteration procedure. The relationships between LAI and reflectances of various spectral bands (red, near infrared, and shortwave infrared) are simulated with Four-Scale with a multiple scattering scheme. Based on the model simulations, the key coefficients in the BRDF kernels are fitted with Chebyshev polynomials of the second kind. Spectral indices - the simple ratio and the reduced simple ratio - are used to effectively combine the spectral bands for LAI retrieval. Example regional and global LAI maps are produced. Accuracy assessment on a Canada-wide LAI map is made in comparison with a previously validated 1998 LAI map and ground measurements made in seven Landsat scenes
Feng Deng, Jing M. Chen, Stephen Plummer, Mingzhen Chen, Jan Pisek
IEEE Trans. Geosci. Remote. Sens.1