Yang Yang 0010

dblp:48/450-10 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
7since 2021 · last 2024
0000-0003-4521-1437ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 6 since 2021Computer networks · 9 · 4 first-authorArtificial intelligence and machine learning · 4 · 2 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2024 STREAMVC: Real-Time Low-Latency Voice Conversion
abstract
We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces the resulting waveform at low latency from the input signal even on a mobile platform, making it applicable to real-time communication scenarios like calls and video conferencing, and addressing use cases such as voice anonymization in these scenarios. Our design leverages the architecture and training strategy of the SoundStream neural audio codec for lightweight high-quality speech synthesis. We demonstrate the feasibility of learning soft speech units causally, as well as the effectiveness of supplying whitened fundamental frequency information to improve pitch stability without leaking the source timbre information.
Yang Yang 0010, Yury Kartynnik, Jiuqiang Tang, George Sung, Matthias Grundmann 0002
ICASSP1
2024 Binaural Angular Separation Network
abstract
We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omnidirectional microphones without needing to collect real RIRs. By relying on specific angular regions and multiple room simulations, the model utilizes consistent time difference of arrival (TDOA) cues, or what we call delay contrast, to separate target and interference sources while remaining robust in various reverberation environments. We demonstrate the model is not only generalizable to a commercially available device with a slightly different microphone geometry, but also outperforms our previous work which uses one additional microphone on the same device. The model runs in real-time on-device and is suitable for low-latency streaming applications such as telephony and video conferencing.
Yang Yang 0010, George Sung, Shao-Fu Shih, Hakan Erdogan, Chehung Lee, Matthias Grundmann 0002
ICASSP1
2023 Guided Speech Enhancement Network
abstract
High quality speech capture has been widely studied for both voice communication and human computer interface reasons. To improve the capture performance, we can often find multi-microphone speech enhancement techniques deployed on various devices. Multi-microphone speech enhancement problem is often decomposed into two decoupled steps: a beamformer that provides spatial filtering and a single-channel speech enhancement model that cleans up the beamformer output. In this work, we propose a speech enhancement solution that takes both the raw microphone and beamformer outputs as the input for an ML model. We devise a simple yet effective training scheme that allows the model to learn from the cues of the beamformer by contrasting the two inputs and greatly boost its capability in spatial rejection, while conducting the general tasks of denoising and dereverberation. The proposed solution takes advantage of classical spatial filtering algorithms instead of competing with them. By design, the beamformer module then could be selected separately and does not require a large amount of data to be optimized for a given form factor, and the network model can be considered as a standalone module which is highly transferable independently from the microphone array. We name the ML module in our solution as GSENet, short for Guided Speech Enhancement Network. We demonstrate its effectiveness on real world data collected on multi-microphone devices in terms of the suppression of noise and interfering speech.
Yang Yang 0010, Shao-Fu Shih, Hakan Erdogan, Jamie Menjay Lin, Chehung Lee, George Sung, Matthias Grundmann 0002
ICASSP1
2022 Region-of-Interest Based Neural Video Compression
Yura Perugachi-Diaz, Guillaume Sautière, Davide Abati, Yang Yang 0010, AmirHossein Habibian, Taco Cohen
BMVC4
2022 Transformer-based Transform Coding
Yinhao Zhu, Yang Yang 0010, Taco Cohen
ICLR2
2022 MobileCodec: neural inter-frame video compression on mobile devices
abstract
Realizing the potential of neural codecs on real-world mobile devices is a big technological challenge due to the inherent conflict between the computational complexity of deep networks and the power-constrained mobile hardware performance. We demonstrate practical feasibility by leveraging Qualcomm's innovation and technology, bridging the gap from neural network-based model simulations to operation on a mobile device powered by Snapdragon® technology. We show the first-ever inter-frame neural video decoder running on a commercial mobile phone, decompressing high-definition videos in real-time while maintaining a low bitrate and high visual quality, comparable to conventional codecs.
Hoang Le, Amir Said, Guillaume Sautière, Yang Yang 0010, Pranav Shrestha, Reza Pourreza 0002, Auke J. Wiggers
MMSys5
2021 Progressive Neural Image Compression With Nested Quantization And Latent Ordering
abstract
We present PLONQ, a progressive neural image compression scheme which pushes the boundary of variable bitrate compression by allowing quality scalable coding with a single bitstream. In contrast to existing learned variable bitrate solutions which produce separate bitstreams for each quality, it enables easier rate-control and requires less storage. Leveraging the latent scaling based variable bitrate solution, we introduce nested quantization, a method that defines multiple quantization levels with nested quantization grids, and progressively refines all latents from the coarsest to the finest quantization level. To achieve finer progressiveness in between any two quantization levels, latent elements are incrementally refined with an importance ordering defined in the rate-distortion sense. To the best of our knowledge, PLONQ is the first learning-based progressive image coding scheme and it outperforms SPIHT, a well-known wavelet-based progressive image codec.
Yadong Lu, Yinhao Zhu, Yang Yang 0010, Amir Said, Taco Cohen
ICIP3
2020 Feedback Recurrent Autoencoder for Video Compression
Adam Golinski, Reza Pourreza 0002, Yang Yang 0010, Guillaume Sautière, Taco Cohen
ACCV (4)3
2020 Guided Variational Autoencoder for Disentanglement Learning
abstract
We propose an algorithm, guided variational autoencoder (Guided-VAE), that is able to learn a controllable generative model by performing latent representation disentanglement learning. The learning objective is achieved by providing signal to the latent encoding/embedding in VAE without changing its main backbone architecture, hence retaining the desirable properties of the VAE. We design an unsupervised and a supervised strategy in Guided-VAE and observe enhanced modeling and controlling capability over the vanilla VAE. In the unsupervised strategy, we guide the VAE learning by introducing a lightweight decoder that learns latent geometric transformation and principal components; in the supervised strategy, we use an adversarial excitation and inhibition mechanism to encourage the disentanglement of the latent variables. Guided-VAE enjoys its transparency and simplicity for the general representation learning task, as well as disentanglement learning. On a number of experiments for representation learning, improved synthesis/sampling, better disentanglement for classification, and reduced classification errors in meta learning have been observed.
Zheng Ding, Yifan Xu 0009, Weijian Xu, Gaurav Parmar, Yang Yang 0010, Max Welling, Zhuowen Tu
CVPR5
2020 Feedback Recurrent Autoencoder
abstract
In this work, we propose a new recurrent autoencoder architecture, termed Feedback Recurrent AutoEncoder (FRAE), for online compression of sequential data with temporal dependency. The recurrent structure of FRAE is designed to efficiently extract the redundancy along the time dimension and allows a compact discrete representation of the data to be learned. We demonstrate its effectiveness in speech spectrogram compression. Specifically, we show that the FRAE, paired with a powerful neural vocoder, can produce high-quality speech waveforms at a low, fixed bitrate. We further show that by adding a learned prior for the latent space and using an entropy coder, we can achieve an even lower variable bitrate.
Yang Yang 0010, Guillaume Sautière, J. Jon Ryu, Taco Cohen
ICASSP1
2019 Automatic Grammar Augmentation for Robust Voice Command Recognition
abstract
This paper proposes a novel pipeline for automatic grammar augmentation that provides a significant improvement in the voice command recognition accuracy for systems with small footprint acoustic model (AM). The improvement is achieved by augmenting the user-defined voice command set, also called grammar set, with alternate grammar expressions. For a given grammar set, a set of potential grammar expressions (candidate set) for augmentation is constructed from an AM-specific statistical pronunciation dictionary that captures the consistent patterns and errors in the decoding of AM induced by variations in pronunciation, pitch, tempo, accent, ambiguous spellings, and noise conditions. Using this candidate set, greedy optimization based and cross-entropy-method (CEM) based algorithms are considered to search for an augmented grammar set with improved recognition accuracy utilizing a command-specific dataset. Our experiments show that the proposed pipeline along with algorithms considered in this paper significantly reduce the mis-detection and mis-classification rate without increasing the false-alarm rate. Experiments also demonstrate the consistent superior performance of CEM method over greedy-based algorithms.
Yang Yang 0010, Anusha Lalitha, Chris Lott
ICASSP1
2019 Joint Antenna Allocation and Link Scheduling in FlexRadio Networks
abstract
FlexRadio, a recent breakthrough in wireless Multi-RF technology, has introduced a new way to unify MIMO and full-duplex into a single framework with a fully flexible design. FlexRadio allows a wireless node to use an arbitrary number of RF chains to support transmission and reception, which makes MIMO and full-duplex subset configurations of FlexRadio. This new architecture has greatly changed the feasibility constraint in wireless networks, which makes the design of high performance MAC layer algorithms even more challenging. First, the RF chain becomes a new resource that needs to be allocated across the network, and the optimal configuration depends not only on the network topology and flow demand, but also on the number of available RF chains at each node. Second, it is not clear how to jointly allocate links and RF chain resources based on the arrival rates and queueing dynamics. In this paper, we introduce a new virtual link model to characterize the feasibility constraint from the perspective of contending RF chain usage. Based on this novel model, a distributed CSMA-like framework is developed to fully leverage the flexibility of RF chain resource allocation.
Zhenzhi Qian, Yang Yang 0010, Kannan Srinivasan 0001, Ness Shroff
INFOCOM2
2017 Load-Adaptive Base-Station Management for Energy Reduction Including Operation-Cost and Turn-On-Cost
abstract
The energy consumption of cellular networks has increased dramatically due to high demand for wireless communication. Base-stations (BSs) use about 60% to 80% of the energy consumed by these networks. An attractive way to reduce energy consumption is to turn the BSs off during periods of under-utilization. However, turning a BS back on typically consumes a lot of energy, which has not been considered in previous works, but critical to good energy management strategies. In this work, we dynamically determine the on-off schedule of these base-stations by taking both operation-cost and turn-on-cost into account. We develop the first online algorithm that only uses future information to decide the on-off status of each BS and characterize its performance using competitive ratio analysis. We extend it by utilizing history information which helps improve the competitive ratio. A heuristic adaptive online algorithm is then designed to balance the utilization of history and future information. We then show via simulation results that the adaptive algorithm works well under a wide range of traffic intensities.
Jiashang Liu, Yang Yang 0010, Prasun Sinha, Ness Shroff
WCNC2
2016 Anonymous-query based rate control for wireless multicast: approaching optimality with constant feedback
abstract
For a multicast group of n receivers, existing techniques either achieve high throughput at the cost of prohibitively large (e.g., O(n)) feedback overhead, or achieve low feedback overhead but without either optimal or near-optimal throughput guarantees. Simultaneously achieving good throughput guarantees and low feedback overhead has been an open problem and could be the key reason why wireless multicast has not been successfully deployed in practice. In this paper, we develop a novel anonymous-query based rate control, which approaches the optimal throughput with a constant feedback overhead independent of the number of receivers. In addition to our theoretical results, through implementation on a software-defined ratio platform, we show that the anonymous-query based algorithm achieves low-overhead and robustness in practice.
Fei Wu 0008, Yang Yang 0010, Ouyang Zhang, Kannan Srinivasan 0001, Ness Shroff
MobiHoc2
2015 Scheduling in wireless networks with full-duplex cut-through transmission
abstract
The recent breakthrough in wireless full-duplex communication makes possible a brand new way of multi-hop wireless communication, namely full-duplex cut-through transmission, where for a traffic flow that traverses through multiple links, every node along the route can receive a new packet and simultaneously forward the previously received packet. This wireless transmission scheme brings new challenges in the design of MAC layer algorithms that aim to reap its full benefit. First, the MAC layer rate region of the cut-through enabled network is directly a function of the routing decision, leading to a strong coupling between routing and scheduling. Second, it is unclear how to dynamically form/change cut-through routes based on the traffic rates and patterns. In this work, we introduce a novel method to characterize the interference relationship between links in the network with cut-through transmission, which decouples the routing decision with the scheduling decision and enables a seamless adaptation of traditional half-duplex routing/scheduling algorithm into wireless networks with full-duplex cut-through capabilities. Based on this interference model, a queue-length based CSMA-type scheduling algorithm is proposed, which both leverages the flexibility of full-duplex cut-through transmission and permits distributed implementation.
Yang Yang 0010, Ness Shroff
INFOCOM1
2015 Constant-Delay and Constant-Feedback Moving Window Network Coding for Wireless Multicast: Design and Asymptotic Analysis
abstract
A major challenge of wireless multicast is being able to support a large number of users while simultaneously maintaining low delay and low feedback overhead. In this paper, we develop a joint coding and feedback scheme named moving window network coding with anonymous feedback (MWNC-AF) that simultaneously achieves constant decoding delay and constant feedback overhead, irrespective of the number of receivers n, without sacrificing either throughput or reliability. We explicitly characterize the asymptotic decay rate of the tail probability of the decoding delay and prove that injecting a fixed amount of information bits into the MWNC-AF encoder buffer in each time slot (called “constant data injection process”) achieves the fastest decay rate, thus showing how to obtain delay optimality in a large deviation sense. We then investigate the average decoding delay of MWNC-AF and show that, when the traffic load approaches capacity, the average decoding delay under the constant injection process is at most one half of that under a Bernoulli injection process. We prove that the per-packet encoding and decoding complexities of MWNC-AF both scale as O(logn) and are thus insensitive to the increase of the number of receivers n. Our simulations further underscore the performance of our scheme through comparisons with existing schemes and show that the delay, encoding, and decoding complexities are low even for a large number of receivers, demonstrating the efficiency, scalability, and ease of implementability of MWNC-AF.
Fei Wu 0008, Yin Sun 0001, Yang Yang 0010, Kannan Srinivasan 0001, Ness Shroff
IEEE J. Sel. Areas Commun.3
2015 Throughput of Rateless Codes Over Broadcast Erasure Channels
abstract
In this paper, we characterize the throughput of a broadcast network with n receivers using rateless codes with block size K. We assume that the underlying channel is a Markov modulated erasure channel that is i.i.d. across users, but can be correlated in time. We characterize the system throughput asymptotically in n. Specifically, we explicitly show how the throughput behaves for different values of the coding block size K as a function of n, as n→ ∞. For finite values of K and n, under the more restrictive assumption of Gilbert-Elliott erasure channels, we are able to provide a lower bound on the maximum achievable throughput. Using simulations, we show the tightness of the bound with respect to system parameters n and K and find that its performance is significantly better than the previously known lower bounds.
Yang Yang 0010, Ness Shroff
IEEE/ACM Trans. Netw.1
2014 Characterizing the achievable throughput in wireless networks with two active RF chains
abstract
Recent breakthroughs in wireless communication show that by using new signal processing techniques, a wireless node is capable of transmitting and receiving simultaneously on the same frequency band by activating both of its RF chains, thus achieving full-duplex communication and potentially doubling the link throughput. However, with two sets of RF chains, one can build a half-duplex multi-input and multi-output (MIMO) system that achieves the same gain. While this gain is the same between a pair of nodes, the gains are unclear when multiple nodes are involved, as in a general network. The key reason is that MIMO and full-duplex have different interference patterns. A MIMO transmission blocks transmissions around its receiver and receptions around its transmitter. A full-duplex bidirectional transmission blocks any transmission around the two communicating nodes, but allows a reception on one RF chain. Thus, in a general network, the requirements for the two technologies could result in potentially different achievable throughput regions. This work investigates the achievable throughput performance of MIMO, full-duplex and their variants that allow simultaneous activation of two RF chains. It is the first work of its kind to precisely characterize the conditions under which these technologies outperform each other for a general network topology under a binary interference model. The analytical results in this paper are validated using software-defined radios.
Yang Yang 0010, Kannan Srinivasan 0001, Ness Shroff
INFOCOM1
2014 A near-optimal randomized algorithm for uplink resource allocation in OFDMA systems
abstract
OFDMA has been selected as the multiple access scheme for emerging broadband wireless communication systems. However, designing efficient resource allocation algorithms for OFDMA systems is a challenging task, especially in the uplink, due to the combinatorial nature of subcarrier assignment and the distributed power budget for different users. Inspired by Glauber dynamics, in this paper, we propose a randomized iteration-based uplink OFDMA resource allocation algorithm. We show that our algorithm is near-optimal in the sense that by increasing the number of iterations (which scales up the complexity), with arbitrarily large probability, the algorithm can converge to the subcarrier/power allocation pattern with the maximum sum-utility. We also show that this algorithm can be generalized to solve a joint uplink-downlink allocation problem in full-duplex OFDMA systems. Simulations are conducted to compare the performance of our algorithm with existing ones.
Yang Yang 0010, Changwon Nam, Ness Shroff
WiOpt1
2014 Delay Asymptotics With Retransmissions and Incremental Redundancy Codes Over Erasure Channels
abstract
Recent studies have shown that retransmissions can cause heavy-tailed transmission delays even when packet sizes are light tailed. In addition, the impact of heavy-tailed delays persists even when packets size are upper bounded. The key question we study in this paper is how the use of coding techniques to transmit information, together with different system configurations, would affect the distribution of delay. To investigate this problem, we model the underlying channel as a Markov modulated binary erasure channel, where transmitted bits are either received successfully or erased. Erasure codes are used to encode information prior to transmission, which ensures that a fixed fraction of the bits in the codeword can lead to successful decoding. We use incremental redundancy codes, where the codeword is divided into codeword trunks and these trunks are transmitted one at a time to provide incremental redundancies to the receiver until the information is recovered. We characterize the distribution of delay under two different scenarios: 1) decoder uses memory to cache all previously successfully received bits and 2) decoder does not use memory, where received bits are discarded if the corresponding information cannot be decoded. In both cases, we consider codeword length with infinite and finite support. From a theoretical perspective, our results provide a benchmark to quantify the tradeoff between system complexity and the distribution of delay.
Yang Yang 0010, Jian Tan 0001, Ness Shroff, Hesham El Gamal
IEEE Trans. Inf. Theory1
2012 Throughput of rateless codes over broadcast erasure channels
abstract
In this paper, we characterize the throughput of a broadcast network with n receivers using rateless codes with block size K. We assume that the underlying channel is a Markov modulated erasure channel that is i.i.d. across users, but can be correlated in time. We characterize the system throughput asymptotically in n. Specifically, we explicitly show how the throughput behaves for different values of the coding block size K as a function of n, as n approaches infinity. Under the more restrictive assumption of memoryless channels, we are able to provide a lower bound on the maximum achievable throughput for any finite values of K and n. Using simulations we show the tightness of the bound with respect to system parameters n and K, and find that its performance is significantly better than the previously known lower bound.
Yang Yang 0010, Ness Shroff
MobiHoc1
2011 Delay asymptotics with retransmissions and fixed rate codes over erasure channels
abstract
Recent work has shown that retransmissions can cause heavy-tailed transmission delays even when packet sizes are light-tailed. Moreover, the impact of heavy tailed delays persist even when packets are of finite size. The key question we study in this paper is how the use of coding techniques to transmit information could mitigate delays. To investigate this problem, we consider an important communication channel called the Binary Erasure Channel, where transmitted bits are either received successfully or lost (called an erasure). This model is a good abstraction of not only the wireless channel but also the higher layer link, where erasure errors can happen. Many coding schemes, known as erasure codes, have been designed for this channel. Specifically, we focus on the fixed rate coding scheme, where decoding is said to be successful if a certain fraction β of the codeword is received correctly. We study two different scenarios: (I) A codeword of length Lcis retransmitted as a unit until the receiver successfully receives more than βLcbits in the last transmission. (II) All successfully received bits from every (re)transmissions are buffered at the receiver according to their positions in the codeword, and the transmission completes once the received bits become decodable for the first time. Our studies reveal that complicated and surprising relationships exist between the coding complexity and the transmission delay/throughput. From a theoretical perspective, our results provide a benchmark to quantify the tradeoffs between coding complexity and transmission throughput for receivers that use memory to buffer (re)transmissions until success and those that do not buffer intermediate transmissions.
Jian Tan 0001, Yang Yang 0010, Ness Shroff, Hesham El Gamal
INFOCOM2