VLDB 2026 Research / reviewers in the wild / expert
Hemin Yang
dblp:137/0146
· DBLP profile ↗
16ranked-venue papers
7as first author
7since 2021 · last 2024
0000-0001-7717-8163ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
Xiaofei Wang 0009, Sefik Emre Eskimez, Manthan Thakker, Hemin Yang, Zirun Zhu, Yufei Xia, Jinzhu Li, Sheng Zhao 0002, Jinyu Li 0001, Naoyuki Kanda |
INTERSPEECH | 4 |
| 2024 | Total-Duration-Aware Duration Modeling for Text-to-Speech SystemsabstractAccurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications.However, the impact of adjusting the speech rate on speech quality, such as intelligibility and speaker characteristics, has been underexplored.In this work, we propose a novel total-duration-aware (TDA) duration model for TTS, where phoneme durations are predicted not only from the text input but also from an additional input of the total target duration.We also propose a MaskGIT-based duration model that enhances the diversity and quality of the predicted phoneme durations.Our results demonstrate that the proposed TDA duration models achieve better intelligibility and speaker similarity for various speech rate configurations compared to the baseline models.We also show that the proposed MaskGIT-based model can generate phoneme durations with higher quality and diversity compared to its regression or flow-matching counterparts. Sefik Emre Eskimez, Xiaofei Wang 0009, Manthan Thakker, Chung-Hsien Tsai, Canrun Li, Hemin Yang, Zirun Zhu, Jinyu Li 0001, Sheng Zhao 0002, Naoyuki Kanda |
INTERSPEECH | 7 |
| 2024 | E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTSabstractThis paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-art speaker similarity and intelligibility. In the E2 TTS framework, the text input is converted into a character sequence with filler tokens. The flow-matching-based mel spectrogram generator is then trained based on the audio infilling task. Unlike many previous works, it does not require additional components (e.g., duration model, grapheme-to-phoneme) or complex techniques (e.g., monotonic alignment search). Despite its simplicity, E2 TTS achieves state-of-the-art zero-shot TTS capabilities that are comparable to or surpass previous works, including Voicebox and NaturalSpeech 3. The simplicity of E2 TTS also allows for flexibility in the input representation. We propose several variants of E2 TTS to improve usability during inference. See https://aka.ms/e2tts/for demo samples. Sefik Emre Eskimez, Xiaofei Wang 0009, Manthan Thakker, Canrun Li, Chung-Hsien Tsai, Hemin Yang, Zirun Zhu, Xu Tan 0003, Sheng Zhao 0002, Naoyuki Kanda |
SLT | 7 |
| 2024 | DDTSE: Discriminative Diffusion Model for Target Speech ExtractionabstractDiffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction under multispeaker noisy conditions remains relatively unexplored. Moreover, the superior quality of diffusion methods typically comes at the cost of slower inference speed. In this paper, we introduce the Discriminative Diffusion model for Target Speech Extraction (DDTSE). We apply the same forward process as diffusion models and utilize the reconstruction loss similar to discriminative methods. Furthermore, we devise a two-stage training strategy to emulate the inference process during model training. DDTSE not only works as a standalone system, but also can further improve the performance of discriminative models without additional retraining. Experimental results demonstrate that DDTSE not only achieves higher perceptual quality but also accelerates the inference process by 3 times compared to the conventional diffusion model. Leying Zhang, Yao Qian, Linfeng Yu, Heming Wang, Hemin Yang, Shujie Liu 0001, Yanmin Qian |
SLT | 5 |
| 2023 | DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation TasksabstractSelf-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem, in this paper, we propose data2vec-SG (Speech Generation), which is a teacher-student learning framework that addresses speech generation tasks. Our data2vec-SG introduces a reconstruction module into data2vec [1] and enforces the representations to contain not only the semantic information but also the acoustic knowledge to generate clean speech waveforms. Experimental results demonstrate that the proposed framework boosts the performance of various speech generation tasks including speech enhancement, speech separation, and packet loss concealment. Meanwhile, the learned representation is also capable of helping other downstream tasks, which is demonstrated by the good performance in the speech recognition task in both clean and noisy conditions. Heming Wang, Yao Qian, Hemin Yang, Naoyuki Kanda, Takuya Yoshioka, Xiaofei Wang 0009, Shujie Liu 0001, Zhuo Chen 0006, DeLiang Wang, Michael Zeng 0001 |
ICASSP | 3 |
| 2023 | Real-Time Audio-Visual End-To-End Speech EnhancementabstractAudio-visual speech enhancement (AV-SE) methods utilize auxiliary visual cues to enhance speakers’ voices. Therefore, technically they should be able to outperform the audio-only speech enhancement (SE) methods. However, there are few works in the literature on an AV-SE system that can work in real time on a CPU. In this paper, we propose a low-latency real-time audio-visual end-to-end enhancement (AV-E3Net) model based on the recently proposed end-to-end enhancement network (E3Net). Our main contribution includes two aspects: 1) We employ a dense connection module to solve the performance degradation caused by the deep model structure. This module significantly improves the model’s performance on the AV-SE task. 2) We propose a multi-stage gating-and-summation (GS) fusion module to merge audio and visual cues. Our results show that the proposed model provides better perceptual quality and intelligibility than the baseline E3net model with a negligible computational cost increase. Zirun Zhu, Hemin Yang, Sefik Emre Eskimez, Huaming Wang |
ICASSP | 2 |
| 2021 | Human Listening and Live Captioning: Multi-Task Training for Speech EnhancementabstractWith the surge of online meetings, it has become more critical than ever to provide high-quality speech audio and live captioning under various noise conditions.However, most monaural speech enhancement (SE) models introduce processing artifacts and thus degrade the performance of downstream tasks, including automatic speech recognition (ASR).This paper proposes a multi-task training framework to make the SE models unharmful to ASR.Because most ASR training samples do not have corresponding clean signal references, we alternately perform two model update steps called SE-step and ASR-step.The SEstep uses clean and noisy signal pairs and a signal-based loss function.The ASR-step applies a pre-trained ASR model to training signals enhanced with the SE model.A cross-entropy loss between the ASR output and reference transcriptions is calculated to update the SE model parameters.Experimental results with realistic large-scale settings using ASR models trained on 75,000-hour data show that the proposed framework improves the word error rate for the SE output by 11.82% with little compromise in the SE quality.Performance analysis is also carried out by changing the ASR model, the data used for the ASR-step, and the schedule of the two update steps. Sefik Emre Eskimez, Xiaofei Wang 0009, Hemin Yang, Zirun Zhu, Zhuo Chen 0006, Huaming Wang, Takuya Yoshioka |
Interspeech | 4 |
| 2020 | STEREOS: Smart Table EntRy Eviction for OpenFlow SwitchesabstractSoftware-defined networking (SDN) is fundamentally changing the way networks operate, enabling programmable and flexible network management and configuration. As the de facto standard southbound interface of SDN, OpenFlow defines how the control plane interacts with the data forwarding plane. In OpenFlow, flow tables play a significant role in packet forwarding. However, the size of the flow table is limited due to power, cost, and silicon area constraints and capacity-limited tables cannot hold all of the active flows in medium-to-large-scale SDN networks. Thus, when a flow table reaches capacity, an intelligent eviction strategy, which efficiently manages the limited flow table resource, is critical. In this paper, we propose Smart Table EntRy Eviction for OpenFlow Switches (STEREOS), which uses machine learning to classify flow entries as active or inactive and forms the basis for intelligent eviction. Trace-driven simulations demonstrate that STEREOS increases flow table usage by more than 50% and reduces incorrect flow entry evictions by up to 78%, compared with the dominant Least Recently Used eviction policy. Moreover, packet-level simulations of a datacenter network demonstrate that STEREOS can greatly reduce the control overhead, increase overall network throughput by 19%, and reduce packet loss rate by 70%. Hemin Yang, George F. Riley, Douglas M. Blough |
IEEE J. Sel. Areas Commun. | 1 |
| 2018 | Machine Learning Based Proactive Flow Entry Deletion for OpenFlowabstractOpenFlow is the de facto southbound interface for Software Defined Networking (SDN), which defines the interactions between the control plane and data plane. In OpenFlow, flow table is significant in packet forwarding. However, the capacity of the flow table is limited due to power, cost, and silicon area constraints. In this case, it is extremely important to efficiently manage flow tables. In this paper, we focus on one of the flow table managing mechanisms defined in OpenFlow, proactive flow entry deletion. This mechanism enables the controller to proactively delete flow entries in flow tables by explicitly sending specific OpenFlow messages to switches. The key challenge for this mechanism is to determine which flow entries should be removed. To address this challenge, we propose a machine learning based proactive flow entry deletion which can learn from the historical data of flow entries and thus predict the time when a flow entry will be last referred to. Based on the predictions, the flow entry with smallest last refer time will be deleted. Our simulations show that our proposal can achieve up to 23% fewer capacity misses compared with the random deletion and First-In-First-Out (FIFO) deletion policies, as well as a slightly decreased (2% ~ 8%) overhead. Hemin Yang, George F. Riley |
ICC | 1 |
| 2018 | Machine Learning Based Flow Entry Eviction for OpenFlow SwitchesabstractSoftware Defined Networking (SDN) is fundamentally changing the way networks work, which enables programmable and flexible network management and configuration. As the de facto southbound interface of SDN, OpenFlow defines how the control plane can directly interact with the forwarding plane. In OpenFlow, flow tables play a significant role in packet forwarding. However, the capacity of flow table is limited due to power, cost, and silicon area constraints. The capacity-limited flow table cannot hold the explosive flows generated by the fine- grained granularity control mechanism used in SDN. Thus the flow table is frequently overflowed. In the case of overflow, eviction strategy which replaces existing flow entries with the new ones is critical to guarantee the efficient usage of the flow table. In this paper, we present a machine learning based eviction approach which can identify whether a flow entry is active or inactive and thus timely evict the inactive flow entries when flow table overflow occurs. Our simulations based on real network packet traces show that the proposed method can increase the usage of flow table by more than 55% and reduce the number of capacity misses by up to 80%, compared with the Least Recently Used eviction policy. Hemin Yang, George F. Riley |
ICCCN | 1 |
| 2017 | Scalability comparison of SDN control plane architectures based on simulationsabstractSoftware Defined Networking (SDN) is an emerging networking paradigm which separates the control plane from the data forwarding plane. To control large networks, designing a scalable control plane for SDN is one of the most significant challenges. In this paper, we focus on comparing the scalability performance of SDN control planes with different architectures. We first compare simulation and emulation approach in scalability evaluation for SDN, and conclude that simulation is the most appropriate approach for our study. To apply simulation approach, we identify two critical processes (flow setup and statistics collection) which restrict the scalability of the control plane. Based on these two processes, we abstract switches and controllers in the five existing SDN control plane architectures (i.e., centralized, P2P with local view, P2P with global view, hierarchical, and hybrid). These abstractions allow us to build simulations based on ns-3 network simulator for SDN networks with different control plane architectures. Our experiments show that the hierarchical control plane achieves the best scalability performance, while the centralized and P2P with global view control planes get the worst performance. Furthermore, our simulations demonstrate the significance of statistics collection speed in scaling the SDN network. We believe this study can give implications about control plane architecture selection for operators (developers) who want to deploy (develop) their own SDN networks (controllers). Hemin Yang, Jared S. Ivey, George F. Riley |
IPCCC | 1 |
| 2016 | On design of Interference Self-Coordination (ISC) solution to enable mHealth services in HetNetsabstractAs defined by WHO and ITU, mobile health (mHealth) is expected to change healthcare service modes and will be at the core of the future healthcare system. As for considering human activities, hotspot and indoor environments are the major scenarios to carry mHealth services (e.g., homecare). To accommodate mHealth services, dense/superdense small-cell layout in Heterogeneous Networks (HetNets) is a typical deployment scenario. On the other hand, mHealth services should be guaranteed with reliability and security in any wireless networks, since health-risk or medical-related information involved. In this study, we focus on how to guarantee control-information delivery in HetNets, so that mHealth users can be allowed to access wireless facilities reliably. To handle this challenge, we extend Interference Self-Coordination theory into HetNets, in order to enable service-sensitive healthcare services in mobile networks if needed. The system-level simulation experiments demonstrate that our proposal can achieve significant performance gain: the SINR of control channel is increased by around +104%; the packet error-loss rate can be reduced by -73%; the packet delay is reduced by -46%. These gains are valuable for quality-sensitive applications in HetNets, such as mHealth. Yaxing Qiu, Hemin Yang, Anpeng Huang, Bingli Jiao |
ICC | 2 |
| 2016 | Comparing a Scalable SDN Simulation Framework Built on ns-3 and DCE with Existing SDN Simulators and EmulatorsabstractAs software-defined networking (SDN) grows beyond its original aim to simply separate the control and data network planes, it becomes useful both financially and analytically to provide adequate mechanisms for simulating this new paradigm. A number of simulation/emulation tools for modeling SDN, such as Mininet, are already available. A new, novel framework for providing SDN simulation has been provided in this work using the network simulator ns-3. The ns-3 module Direct Code Execution (DCE) allows real-world network applications to be run within a simulated network topology. This work employs DCE for running the SDN controller library POX and its applications on nodes in a simulated network topology. In this way, real-world controller applications can be completely portable between simulation and actual deployment. This work also describes a user-defined ns-3 application mimicking an SDN switch supporting OpenFlow 1.0 that can interact with real-world controllers. To evaluate its performance, this ns-3 DCE SDN framework is compared against Mininet as well as some other readily available SDN simulation/emulation tools. Metrics such as realtime performance, memory usage, and reliability in terms of packet loss are analyzed across the multiple simulation/emulation tools to gauge how they compare. Jared S. Ivey, Hemin Yang, Chuanji Zhang, George F. Riley |
SIGSIM-PADS | 2 |
| 2015 | Best-fit cell attachment for decoupling DL/UL to promote traffic offloading in HetNetsabstractHeterogeneous Networks (HetNets), in which small cells are ultra-densely deployed within a macro cell, is widely regarded as a key solution to mobile traffic offloading. However, this kind of traffic offloading is confronted by two challenges. Firstly, the number of users in any small cell is really limited due to the larger power gap between macros and small cells in HetNets. In return, this phenomenon may degrade traffic offloading performance. Secondly, spectrum efficiency may be deteriorated since uplink channel states are not involved in the conventional cell attachment strategies. To address these challenges, we propose a best-fit cell attachment where downlink and uplink are decoupled and attached to cells independently. Specifically, in this paper, we conceive best-fit strategies for downlink and uplink cell attachment, for taking into account of both channel quality and cell load. And then, their implementation issues in different scenarios are discussed in detail. System-level simulation results demonstrated that our proposal can effectively offload traffic from macros to small cells, and improve data rates for cell-edge users significantly. Hemin Yang, Anpeng Huang, Linzhen Xie |
ICC | 1 |
| 2014 | Interference Self-Coordination: A Proposal to Enhance Reliability of System-Level Information in OFDM-Based Mobile Networks via PCI PlanningabstractFor system-level information in OFDM-based mobile networks, the full-cell coverage requirement raises the issue of Inter-Cell Interference (ICI) in physical control channels. Unfortunately, conventional approaches (e.g., scheduling, beamforming) cannot effectively deal with this kind of ICI because of configuration limitations in physical control channels. To solve this challenge, this study defines a new concept of `Interference Self-Coordination (ISC)' for enhancing this system-level information delivery. Specifically, we apply ISC to relieve the ICI effect over the Physical Downlink Control Channel (PDCCH), the most important physical control channel in Long-Term Evolution (LTE/LTE-Advanced) networks, in which Physical Cell Identifier (PCI) planning naturally serves as a self-coordination mechanism. We prove that PCI planning is the most effective solution to relieve ICI over PDCCHs, under the constraint of non-orthogonal control regions between neighboring cells. Since PCI planning is an NP-Complete problem, heuristic-strategy algorithms are devised for implementation, which are self-adaptive to dynamic configurations of the network and system, and compatible with any applicable ICI solutions. The numerical analysis and system-level simulation experiment results demonstrate that this proposal can increase overall system throughput by up to 17%, and also improve cell-edge throughput by up to 42%. In the worst 5% of sectors, the proposal can still obtain 45% system throughput and 200% cell-edge throughput gain. Hemin Yang, Anpeng Huang, Ruipeng Gao, Tammy Chang, Linzhen Xie |
IEEE Trans. Wirel. Commun. | 1 |
| 2013 | A solution to relieve ICI effects on system control information in OFDM-based mobile networks: Conflict coordination on PDCCH via PCI planningabstractIn OFDM-based mobile networks, the full-cell coverage of SCI (System-level Control Information) must be guaranteed for a user that can communicate with its eNB (evolved Node Base station) system. The requirement of SCI coverage causes severe ICI (Inter-Cell Interference) effect. Furthermore, the negative effect is intensified owing to single-frequency networking and physical control region configuration in OFDM-based networks. To enhance the reliability of SCI under the coverage constraint, we investigated how SCI is carried in the Physical Downlink Control Channel (PDCCH) in the common search space, and proposed a solution of Conflict Coordination on PDCCH via PCI (Physical Cell Identifier) Planning (CCP3). In our proposal, PCI planning is used to relieve the ICI effect of the SCI. Thus, a heuristic algorithm is developed for the PCI planning since it is an NP-Complete optimization problem. The link-level simulation experiments and numerical analysis results demonstrate that the proposed CCP3can improve effective SINR (Signal to Interference plus Noise Ratio) of PDCCH significantly, and reduce RE (Resource Element)-occupation conflict probability of PDCCH around 50%. To the best of our knowledge, this study is the first effort to strengthen the SCI reliability from the view of networking optimization. Hemin Yang, Ruipeng Gao, Anpeng Huang, Linzhen Xie |
ICC | 1 |