Yonggang Hu

dblp:49/7213 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-0414-8697ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Systems, architecture and hardware · 2Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 MulSE: Integrating Dual-Path Modeling and Global Attention for Multi-Channel Speech Enhancement
Mao-shen Jia, Yonggang Hu
IEEE Signal Process. Lett.3
2025 Power Spectral Density Estimation for Acoustic Source Separation Using A Spherical Microphone Array
Mao-shen Jia, Yonggang Hu
INTERSPEECH3
2025 Direct-path Relative Harmonic Coefficients Detection for Multi-source Direction-of-Arrival Estimation in Reverberant Environments
Mao-shen Jia, Yonggang Hu
INTERSPEECH3
2024 Spatial Acoustic Enhancement Using Unbiased Relative Harmonic Coefficients
Mao-shen Jia, Yonggang Hu, Changchun Bao
INTERSPEECH3
2023 Generalized Relative Harmonic Coefficients
abstract
In literature, sound source localization under the far- and near-field scenarios are mostly addressed as independent tasks using different approaches. This causes a tedious task to detect the type of sound-field, whereas in practice there may not be a clear boundary between the far- and near-field soundfield. In contrast, this paper proposes a multi-channel feature denoted generalized relative harmonic coefficients (generalized RHC) in the spherical harmonics domain, which can equally localize both far- and near-field sound source without requiring any adjustments. We derive the analytical expression of this feature and summarize its unique properties, which facilitate two single-source directional-of-arrival estimators: (i) using a full grid search over the directional space; and (ii) a closed-form solution without any grid search. Experimental study in realistic noisy and reverberant environments under both near-field and far-field conditions validates the efficacy of the proposed algorithm.
Yonggang Hu, Sharon Gannot, Thushara D. Abhayapala
ICASSP1
2022 Closed-Form Single Source Direction-of-Arrival Estimator Using First-Order Relative Harmonic Coefficients
abstract
The relative harmonic coefficients (RHC), recently introduced as a multi-microphone spatial feature, demonstrates promising performance when applied to direction-of-arrival (DOA) estimation. All existing RHC-based DOA estimators suffer from a resolution limitation due to the inherent grid-based search. In contrast, this paper utilizes the first-order RHC to propose a closed-form DOA estimator by deriving a direction vector, which points towards to the desired source direction. Two objective metrics, namely localization accuracy and algorithm complexity, are adopted for the evaluation and comparison with existing RHC-based and intensity based localization approaches, in both simulated and real-life environments.
Yonggang Hu, Sharon Gannot
ICASSP1
2022 Decoupled Multiple Speaker Direction-of-Arrival Estimator Under Reverberant Environments
abstract
Direction-of-arrival (DOA) estimation for multiple simultaneous speakers in reverberant environments is still one of the challenging tasks in the audio signal processing field. A recent approach addresses this problem using a spherical harmonics domain feature namedrelative harmonic coefficients(RHC). Based on a bin-wise operation across the STFT (short-time Fourier transform) domain, this method detects the direct-path RHC in the first stage, followed by single source localization in the second stage. However, the method is computationally expensive as each STFT bin requires an exhaustive grid search over the two-dimensional (2-D) directional space. In this paper, we propose a significantly more computationally efficient alternative that decouples the azimuth and elevation 2-D search to two separate one-dimensional (1-D) search. The proposed multi-speaker localization algorithm comprises of two main steps, responsible for: (i) achieving a joint direct-path RHC detection and decoupled DOA estimation using 1-D search; and (ii) counting the number of speakers and estimating their DOAs based on the estimates from direct-path dominated STFT bins. Experiments using both simulated and real-life reverberant recordings confirm the significant computational complexity reduction while achieving competitive localization accuracy, compared to the baseline approaches. Although our proposed method performs in an unsupervised manner, it proves to be applicable even under unfavorable acoustic environments with a high reverberation level (e.g.,$T_{60}=1$second).
Yonggang Hu, Prasanga N. Samarasinghe, Sharon Gannot, Thushara D. Abhayapala
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Rethinking Bi-Level Optimization in Neural Architecture Search: A Gibbs Sampling Perspective
abstract
One-Shot architecture search, which aims to explore all possible operations jointly based on a single model, has been an active direction of Neural Architecture Search (NAS). As a well-known one-shot solution, Differentiable Architecture Search (DARTS) performs continuous relaxation on the architecture's importance and results in a bi-level optimization problem. However, as many recent studies have shown, DARTS cannot always work robustly for new tasks, which is mainly due to the approximate solution of the bi-level optimization. In this paper, one-shot neural architecture search is addressed by adopting a directed probabilistic graphical model to represent the joint probability distribution over data and model. Then, neural architectures are searched for and optimized by Gibbs sampling. We rethink the bi-level optimization problem as the task of Gibbs sampling from the posterior distribution, which expresses the preferences for different models given the observed dataset. We evaluate our proposed NAS method -- GibbsNAS on the search space used in DARTS/ENAS and the search space of NAS-Bench-201. Experimental results on multiple search space show the efficacy and stability of our approach.
Chao Xue 0003, Xiaoxing Wang, Junchi Yan, Yonggang Hu, Xiaokang Yang 0001, Kewei Sun
AAAI4
2021 ZipLine: An Optimized Algorithm for the Elastic Bulk Synchronous Parallel Model
abstract
The bulk synchronous parallel (BSP) is a celebrated synchronization model for distributed training of deep learning models. A shortcoming of the BSP is that it requires workers to wait for the straggler at every iteration. Therefore, employing BSP increases the waiting time of the faster workers of a cluster and results in an overall prolonged training time. To ameliorate this shortcoming of BSP, we proposed ElasticBSP [1], a model that aims to relax its strict synchronization requirement with an elastic synchronization by allowing delayed synchronization to minimize the waiting time. ELASTICBSP is realized by the algorithm named ZipLine. In this work, we show the theoretical proof of ZipLine and further propose algorithmic and implementation optimizations of ZipLine, namely ZipLineOpt and Ziplineoptbs, which reduce the time complexity of ZipLine to linearithmic time. The experiments show that ZipLineOpt and ZipLineOptBs enable the scalability of ElasticBSP. Further experimental evaluation on large deep neural networks on large ImageNet dataset demonstrate that our proposed Elas-ticbspmodel, materialized by the proposed optimized ZipLine variants, converges faster and to a higher accuracy than the predominant BSP.
Xing Zhao 0004, Manos Papagelis, Aijun An, Bao Xin Chen, Junfeng Liu 0005, Yonggang Hu
DSAA6
2021 Evaluation and Comparison of Three Source Direction-of-Arrival Estimators Using Relative Harmonic Coefficients
abstract
A spherical harmonics domain source feature called relative harmonic coefficients (RHC) has recently been applied to address the source direction-of-arrival (DOA) estimation problem. This paper presents a compact evaluation and comparison between two existing RHC based DOA estimators: (i) a method using a full grid search over the two-dimensional (2-D) directional space, (ii) a decoupled estimator which uses one-dimensional (1-D) search to separately localize the source's elevation and azimuth. We also propose a new estimator using a gradient descent search over the 2-D directional grid space. Extensive experiments in both simulated and real-life environments are conducted to examine and analyze the performance of all the underlying DOA estimators. Two objective metrics, including localization accuracy and algorithm complexity, are adopted for an evaluation and comparison between all estimators.
Yonggang Hu, Prasanga N. Samarasinghe, Sharon Gannot, Thushara D. Abhayapala
ICASSP1
2021 ZipLine: an optimized algorithm for the elastic bulk synchronous parallel model
Xing Zhao 0004, Manos Papagelis, Aijun An, Bao Xin Chen, Junfeng Liu 0005, Yonggang Hu
Mach. Learn.6
2021 Multiple Source Direction of Arrival Estimations Using Relative Sound Pressure Based MUSIC
abstract
Subspace approach of MUSIC (multiple signal classication) has become one of the most popular multi-source direction of arrival (DOA) estimations due to its easy implementation in practice. However, its localization accuracy is vulnerable to noise. This paper develops a novel MUSIC algorithm, more suitable in noisy environments, using the relative sound pressure measurements of a higher order microphone array. This proposed MUSIC approach is also decomposed into the spherical harmonics domain where a frequency smoothing technique is allowed to de-correlate the coherent source signals for improved localization accuracy. The proposed algorithm is also capable of estimating the number of active sound sources, which is pre-requisite knowledge for the traditional MUSIC approach. Extensive experimental results in diverse environments using both simulated and real recordings show advantages of the proposed algorithm over the traditional MUSIC method as well as another recently proposed multi-source localization approach.
Yonggang Hu, Thushara D. Abhayapala, Prasanga N. Samarasinghe
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Unsupervised Multiple Source Localization Using Relative Harmonic Coefficients
abstract
This paper presents an unsupervised multi-source localization algorithm using a recently introduced feature called the relative harmonic coefficients. We derive a closed-form expression of the feature and briefly summarize its unique properties. We then exploit this feature to develop a single-source frame/bin detector which simplifies the challenging problem of multiple source localization into a single source localization problem. We show that the underlying method is suitable for localization using overlapped, disjoint as well as simultaneous multi-source recordings. Experimental results in both simulated and real-life reverberant environments confirm improved localization accuracy of the proposed method in comparison with the existing state-of-art approach.
Yonggang Hu, Prasanga N. Samarasinghe, Thushara D. Abhayapala, Sharon Gannot
ICASSP1
2020 MergeNAS: Merge Operations into One for Differentiable Architecture Search
abstract
Differentiable architecture search (DARTS) has been a promising one-shot architecture search approach for its mathematical formulation and competitive results. However, besides its caused high memory utilization and a large computation requirement, many research works have shown that DARTS also often suffers notable over-fitting and thus does not work robustly for some new tasks. In this paper, we propose a one-shot neural architecture search method referred to as MergeNAS by merging different types of operations e.g. convolutions into one operation. This merge-based approach not only reduces the search cost (about half a GPU day), but also alleviates over-fitting by reducing the redundant parameters. Extensive experiments on different search space and various datasets have been conducted to verify our approach, showing that MergeNAS can converge to a stable architecture and achieve better performance with fewer parameters and search cost. For test accuracy and its stability, MergeNAS outperforms all NAS baseline methods implemented on NAS-Bench-201, including DARTS, ENAS, RS, BOHB, GDAS and hand-crafted ResNet.
Xiaoxing Wang, Chao Xue 0003, Junchi Yan, Xiaokang Yang 0001, Yonggang Hu, Kewei Sun
IJCAI5
2020 Acoustic Signal Enhancement Using Relative Harmonic Coefficients: Spherical Harmonics Domain Approach
abstract
Over recent years, spatial acoustic signal processing using higher order microphone arrays in the spherical harmonics domain has been a popular research topic. This paper uses a recently introduced source feature called the relative harmonic coefficients to develop an acoustic signal enhancement approach in noisy environments. This proposed method enables to extract the clean spherical harmonic coefficients from noisy higher order microphone recordings. Hence, this technique can be used as a pre-processing tool for noise-free measurements required by many spatial audio applications. We finally present a simulation study analyzing the performance of this approach in far field noisy environments.
Yonggang Hu, Prasanga N. Samarasinghe, Thushara D. Abhayapala
INTERSPEECH1
2020 Semi-Supervised Multiple Source Localization Using Relative Harmonic Coefficients Under Noisy and Reverberant Environments
abstract
This article develops a semi-supervised algorithm to address the challenging multi-source localization problem in a noisy and reverberant environment, using a spherical harmonics domain source feature of the relative harmonic coefficients. We present a comprehensive research of this source feature, including (i) an illustration confirming its sole dependence on the source position, (ii) a feature estimator in the presence of noise, (iii) a feature selector exploiting its inherent directivity over space. Source features at varied spherical harmonic modes, representing unique characterization of the soundfield, are fused by the Multi-Mode Gaussian Process modeling. Based on the unifying model, we then formulate the mapping function revealing the underlying relationship between the source feature(s) and position(s) using a Bayesian inference approach. Another issue of the overlapped components is addressed by a pre-processing technique performing overlapped frame detection, which in turn reduces this challenging problem to a single source localization. It is highlighted that this data-driven method has a strong potential to be implemented in practice because only a limited number of labeled measurements is required. We evaluate this proposed algorithm using simulated recordings between multiple speakers in diverse environments, and extensive results confirm improved performance in comparison with the state-of-art methods. Additional assessments using real-life recordings further prove the effectiveness of the method, even at unfavorable circumstances with severe source overlapping.
Yonggang Hu, Prasanga N. Samarasinghe, Sharon Gannot, Thushara D. Abhayapala
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Transferable AutoML by Model Sharing Over Grouped Datasets
abstract
Automated Machine Learning (AutoML) is an active area on the design of deep neural networks for specific tasks and datasets. Given the complexity of discovering new network designs, methods for speeding up the search procedure are becoming important. This paper presents a so-called transferable AutoML approach that Automated Machine Learning (AutoML) is an active area on the design of deep neural networks for specific tasks and datasets. Given the complexity of discovering new network designs, methods for speeding up the search procedure are becoming important. This paper presents a so-called transferable AutoML approach that leverages previously trained models to speed up the search process for new tasks and datasets. Our approach involves a novel meta-feature extraction technique based on the performance of benchmark models, and a dynamic dataset clustering algorithm based on Markov process and statistical hypothesis test. As such multiple models can share a common structure while with different learned parameters. The transferable AutoML can either be applied to search from scratch, search from predesigned models, or transfer from basic cells according to the difficulties of the given datasets. The experimental results on image classification show notable speedup in overall search time for multiple datasets with negligible loss in accuracy.
Chao Xue 0003, Junchi Yan, Stephen M. Chu, Yonggang Hu, Yonghua Lin
CVPR5
2019 Dynamic Graph Embedding via LSTM History Tracking
abstract
Many real world networks are very large and constantly change over time. These dynamic networks exist in various domains such as social networks, traffic networks and biological interactions. To handle large dynamic networks in downstream applications such as link prediction and anomaly detection, it is essential for such networks to be transferred into a low dimensional space. Recently, network embedding, a technique that converts a large graph into a low-dimensional representation, has become increasingly popular due to its strength in preserving the structure of a network. Efficient dynamic network embedding, however, has not yet been fully explored. In this paper, we present a dynamic network embedding method that integrates the history of nodes over time into the current state of nodes. The key contribution of our work is 1) generating dynamic network embedding by combining both dynamic and static node information 2) tracking history of neighbors of nodes using LSTM 3) significantly decreasing the time and memory by training an autoencoder LSTM model using temporal walks rather than adjacency matrices of graphs which are the common practice. We evaluate our method in multiple applications such as anomaly detection, link prediction and node classification in datasets from various domains.
Shima Khoshraftar, Sedigheh Mahdavi, Aijun An, Yonggang Hu, Junfeng Liu 0005
DSAA4
2019 Modeling Characteristics of Real Loudspeakers Using Various Acoustic Models: Modal-domain Approaches
abstract
The accuracy and perception of soundfields produced by loudspeaker arrays are strongly influenced by the inherent characteristics of the commercial loudspeakers. This paper analyzes such characteristics of loudspeakers by deriving equivalent theoretical models, and by studying their impact on soundfield reproduction. A number of acoustic models are investigated, including plane waves decomposition, point source decomposition and mixed source decomposition. Each proposed model employs three effective sparse decomposition algorithms for optimized solutions, including iteratively reweighted least squares (IRLS), matching pursuit (MP) and least absolute shrinkage and selection operator (LASSO). A successful model shall enable the prediction of the soundfield outside the original recording region. Therefore, we validate the effectiveness of the models by comparing the simulated soundfield with secondary measurements obtained beyond the original area. Experimental results have confirmed that both the plane wave and mixed source model achieve promising performance with respect to the proposed metrics.
Yonggang Hu, Prasanga N. Samarasinghe, Thushara D. Abhayapala, Glenn Dickins
ICASSP1
2019 Elastic Bulk Synchronous Parallel Model for Distributed Deep Learning
abstract
The bulk synchronous parallel (BSP) is a celebrated synchronization model for general-purpose parallel computing that has successfully been employed for distributed training of machine learning models. A prevalent shortcoming of the BSP is that it requires workers to wait for the straggler at every iteration. To ameliorate this shortcoming of classic BSP, we propose ELASTICBSP a model that aims to relax its strict synchronization requirement. The proposed model offers more flexibility and adaptability during the training phase, without sacrificing on the accuracy of the trained model. We also propose an efficient method that materializes the model, named ZIPLINE. The algorithm is tunable and can effectively balance the trade-off between quality of convergence and iteration throughput, in order to accommodate different environments or applications. A thorough experimental evaluation demonstrates that our proposed ELASTICBSP model converges faster and to a higher accuracy than the classic BSP. It also achieves comparable (if not higher) accuracy than the other sensible synchronization models.
Xing Zhao 0004, Manos Papagelis, Aijun An, Bao Xin Chen, Junfeng Liu 0005, Yonggang Hu
ICDM6
2016 Deep parallelization of parallel FP-growth using parent-child MapReduce
abstract
MapReduce is an important programming model for processing in distributed environments. Compared to other distributed programming models, MapReduce reduces communication overheads between computers and improves fault tolerance. However, the MapReduce model does not allow for automatic synchronization between jobs. A large number of data analytics algorithms use a recursive divide-and-conquer approach, which inherently allows for parallelism at each level of recursion. However, it is often difficult to parallelize such algorithms using the traditional MapReduce model if the process requires synchronization. In this paper we introduce Parent-Child MapReduce, a version of the MapReduce programming model that allows for MapReduce tasks to be created dynamically and synchronized in a hierarchical parent-child fashion. Using the Parallel FP-Growth (PFP) algorithm for mining frequent patterns as a reference, we show that Parent-Child MapReduce can be used to parallelize recursive divide-and-conquer algorithms using the MapReduce model and that this can lead to significant speed ups in the computational speed of such algorithms. Our evaluation shows that we can achieve 68% (or 3 times) performance gain when used with PFP.
Adetokunbo Makanju, Zahra Farzanyar, Aijun An, Nick Cercone, Zane Zhenhua Hu, Yonggang Hu
IEEE BigData6
2016 Distributed and parallel high utility sequential pattern mining
abstract
The problem of mining high utility sequential patterns (HUSP) has been studied recently. Existing solutions are mostly memory-based, which assume that data can fit into the main memory of a computer. However, with advent of big data, such an assumption does not hold any longer. Hence, existing algorithms are not applicable to the big data environments, where data are often distributed and too large to be dealt with by a single machine. In this paper, we propose a new framework for mining HUSPs in big data. A distributed and parallel algorithm called BigHUSP is proposed to discover HUSPs efficiently. At its heart, BigHUSP uses multiple MapReduce-like steps to process data in parallel. We also propose a number of pruning strategies to minimize search space in a distributed environment, and thus decrease computational and communication costs, while still maintaining correctness. Our experiments with real life and large synthetic datasets validate the effectiveness of BigHUSP for mining HUSPs from large sequence datasets.
Morteza Zihayat, Zane Zhenhua Hu, Aijun An, Yonggang Hu
IEEE BigData4
2016 A Novel Development Infrastructure for Scalable Video Coding/Transcoding Applications
abstract
Due to recent demand for playback of high quality video on mobile devices, there is the need for a scalable, error resilient framework with the ability to adjust to the network and receiver's specification. To cope with the bandwidth fluctuations in the network, the framework's scalability allows for high bit-rate video to be transcoded to a low bit-rate format while preserving quality as much as possible. The new High Efficiency Video Coding (HEVC/H.265) standard allows for high compression rates, however, it is computationally intensive. We propose a novel Development Infrastructure for Video coding/transcoding Applications (DIVA). This framework is capable of providing different quality of services (e.g., Bronze, Silver, Gold), and has some error-resilient capability. Taking advantage of IBM Platform Symphony, the computationally intensive task of HEVC encoding can be distributed on available local or cloud resources. Our experiments illustrate the feasibility of this approach.
Vida Movahedi, Amir Asif, Alicia Chin, Ihab Amer, Zane Zhenhua Hu, Yonggang Hu
DCC6
2014 DynMR: dynamic MapReduce with ReduceTask interleaving and MapTask backfilling
abstract
In order to improve the performance of MapReduce, we design DynMR. It addresses the following problems that persist in the existing implementations: 1) difficulty in selecting optimal performance parameters for a single job in a fixed, dedicated environment, and lack of capability to configure parameters that can perform optimally in a dynamic, multi-job cluster; 2) long job execution resulting from a task long-tail effect, often caused by ReduceTask data skew or heterogeneous computing nodes; 3) inefficient use of hardware resources, since ReduceTasks bundle several functional phases together and may idle during certain phases.
Jian Tan 0001, Alicia Chin, Zane Zhenhua Hu, Yonggang Hu, Shicong Meng, Xiaoqiao Meng, Li Zhang 0002
EuroSys4
2012 A hierarchical geostatistical model of walking style variety
abstract
This paper presents a new method on generating realistic human animation of various styles with given step constraints and a specific skeleton. Given a set of normal walking data captured from different subjects, a hierarchical geostatistical model is automatically learned to encode variety of walking styles by representing human body as a hierarchy of joint groups. For each child hierarchy level, there is a learned geostatistical model, whose low-dimensional control parameters are automatically constructed from the synthesized motion of its parent hierarchy level. The top level is controlled by the given step constraints. Also a realistic transition is finished during each three sequential stances from two interpolated cycles satisfying input step constraints thanks to the local model. We show that the normal walking styles are flexibly controlled by a simple graphlike user interface representing the given skeleton and steps. Our results demonstrate that this method would be helpful to remove the motion clones in group or crowd animation.
Chongzhao Han, Yonggang Hu
INDIN4
2009 Orthogonal Centroid Locally Linear Embedding for Classification
Yonggang Hu, Yi Wu 0003
ADMA2