Yangyang Shi

dblp:31/10057 · DBLP profile ↗
← Back
74ranked-venue papers
27as first author
47since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 18 first-author · 28 since 2021Artificial intelligence and machine learning · 39 · 13 first-author · 25 since 2021Systems, architecture and hardware · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 OmniEvent: Unified Event Representation Learning
abstract
Event cameras have gained increasing popularity in computer vision due to their ultra-high dynamic range and temporal resolution. However, event networks heavily rely on task-specific designs due to the unstructured data distribution and spatial-temporal (S-T) inhomogeneity, making it hard to reuse existing architectures for new tasks. We propose OmniEvent, an innovative unified event representation learning framework that achieves SOTA performance across diverse tasks, fully removing the need for task-specific designs. Unlike previous methods that treat event data as 3D point clouds with manually tuned S-T scaling weights, OmniEvent proposes a decouple-enhance-fuse paradigm, where the local feature aggregation and enhancement are done independently on the spatial and temporal domains to avoid inhomogeneity issues. Space-filling curves are applied to enable large receptive fields while improving memory and compute efficiency. The features from individual domains are then fused by attention to learn S-T interactions. The output of OmniEvent is a grid-shaped tensor, which enables standard vision models to process event data without architectural changes. With a unified framework and similar hyperparameters, OmniEvent outperforms (task-specific) SOTA by up to 68.2% across 3 representative tasks and 10 datasets (Fig. 1).
Weiqi Yan 0005, Chenlu Lin, Youbiao Wang, Zhipeng Cai 0003, Xiuhong Lin, Yangyang Shi, Weiquan Liu
AAAI6
2026 SOPSmith: Forging Executable SOPs for LLM-Driven GPU Cluster Network Diagnosis
Guoyao Yu, Xiaoqing Sun, Yangyang Shi, Yang Song 0031, Xing Li 0007, Biao Lyu, Zhenguang Liu, Qinming He
IWQoS4
2026 TFKAN: Time-frequency KAN for long-term time series forecasting
Xiaoyan Kui, Canwei Liu, Qinsong Li, Zhipeng Hu, Yangyang Shi, Weixin Si, Beiji Zou 0001
Neurocomputing5
2026 Learning across modalities: Multi-scale contrastive forecasting with time-frequency representations and adversarial augmentations
Yangyang Shi, Qianqian Ren
Inf. Sci.1
2026 Integrating frequency-aware mamba with diffusion for 4D volumetric image synthesis
Yangyang Shi, Beiji Zou 0001, Xiaonian Deng, Yucong Zhang, Zehua Liu, Xiaoyan Kui, Weixin Si
Pattern Recognit.1
2025 AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
abstract
Ernie Chang, Yang Li, Patrick Huber, Vish Vogeti, David Kant, Yangyang Shi, Vikas Chandra. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ernie Chang, Yang Li 0183, Patrick Huber, Vish Vogeti, David Kant, Yangyang Shi, Vikas Chandra
ACL (1)6
2025 MMW: Side Talk Rejection Multi-Microphone Whisper On Smart Glasses
abstract
Smart glasses are increasingly positioned as the nextgeneration interface for ubiquitous access to large language models (LLMs). Nevertheless, achieving reliable interaction in real-world noisy environments remains a major challenge, particularly due to interference from side speech. In this work, we introduce a novel side-talk rejection multimicrophone Whisper (MMW) framework for smart glasses, incorporating three key innovations. First, we propose a Mix Block based on a Tri-Mamba architecture to effectively fuse multi-channel audio at the raw waveform level, while maintaining compatibility with streaming processing. Second, we design a Frame Diarization Mamba Layer to enhance frame-level side-talk suppression, facilitating more efficient fine-tuning of Whisper models. Third, we employ a Multi-Scale Group Relative Policy Optimization (GRPO) strategy to jointly optimize frame-level and utterance-level side speech suppression. Experimental evaluations demonstrate that the proposed MMW system can reduce the word error rate (WER) by 4.95% in noisy conditions.
Yiteng Huang, Yangyang Shi, Saurabh Adya, Ming Sun 0013, Florian Metze
ASRU5
2025 Distance-Aware and Knowledge-Driven Vision Mamba U-Net for Radiotherapy Dose Prediction
abstract
Dose planning is essential in radiotherapy for cancer patients, yet current practice relies on iterative manual optimization, underscoring the need for automated prediction. Existing deep learning approaches remain limited because they often ignore the 3D spatial relationships between tumors and surrounding organs at risk (OARs), and clinical priors on safe dose thresholds. To overcome these limitations, we propose DKVMU-Net, a distance-aware and knowledge-driven Vision Mamba U-Net for automated dose prediction. Our framework incorporates Vision Mamba blocks to capture global, long-range dependencies from CT scans and OAR signed distance field (SDF) maps, which naturally encode spatial information. Additionally, we introduce a deformable dynamic feature enhancement module (DDFEM) for texture refinement, followed by a linear crossattention fusion module to improve cross-modality integration. A customized loss function is also designed to incorporate prior knowledge of OAR dose constraints, ensuring optimal target coverage and OAR protection. To alleviate the scarcity of doseplanning datasets, we collect an in-house radiotherapy lung cancer dataset (RLCD), consisting of CT volumes, OAR masks, and corresponding SDF maps from 116 patients. We evaluate our DKVMU-Net on both the in-house dataset and public available OpenKBP dataset. Compared with the sate-of-the-art method, our approach achieves an 11.6 % improvement in dose score (1.641 vs. 1.857) and 26.3 % in DVH score (6.481 vs. 8.799) on RLCD, and a 7.8 % improvement in dose score (2.421 vs. 2.626) and 13.9 % in DVH score (1.057 vs. 1.227) on OpenKBP. These results demonstrate the robustness and effectiveness of our approach.
Yangyang Shi, Xiaoyan Kui, Yucong Zhang, Shihao Zou, Zuheng Ming, Weixin Si, Azeddine Beghdadi, Beiji Zou 0001
BIBM1
2025 From Global to Local: Mamba-Based Hierarchical Registration for Respiratory Lung Deformation
abstract
Deformable image registration is essential in medical applications, as accurately estimating organ displacements across respiratory phases enables precise radiation dose planning in dynamic environments, mitigates damage to organs at risk (OARs), and thus improves patients' health-related quality of life. Although current learning-based methods have achieved impressive performance in small deformation registration, challenges remain due to their limited ability to capture large deformations occurring during respiration. To address this issue, we propose a novel Mamba-based hierarchical registration framework that effectively extracts both global and local features for accurate deformation prediction. Specifically, given a pair of source and target 3DCT volumes, we incorporate a foundation model pretrained on medical image registration tasks to enhance alignment accuracy. We further propose a directional-deformable Mamba scheme to facilitate global context extraction and local motion awareness. The directional Mamba component scans input features from multiple orientations to achieve broad contextual perception, while the deformable Mamba module employs adaptive directional scanning strategies to capture dynamic local variations. To overcome the scarcity of annotated respiratory data, we also collect a new respiratory lung cancer dataset comprising 100 annotated phases from 20 patients. Experimental results on our in-house dataset demonstrate that our method outperforms state-of-the-art approaches, achieving a 1.3 % improvement in overall Dice accuracy and a 1.6 dB increase in PSNR, underscoring its strong potential for clinical deployment. Code and test data are available at: https://github.com/yangyangshi806/Mamba_based_Registration.
Yangyang Shi, Yucong Zhang, Beiji Zou 0001, Xiaoyan Kui, Zexin Ji, Zuheng Ming, Azeddine Beghdadi, Weixin Si
BIBM1
2025 FusionBC: Contrastive Graph and Information Bottleneck Fusion Learning for Time Series Forecasting
abstract
Graph contrastive learning has demonstrated its effectiveness in time series modeling. However, existing approaches often face significant challenges, including sensitivity to noise, incompleteness in data, and the difficulty of balancing performance across long- and short-term forecasting tasks. To overcome these limitations, we propose a novel fusion-based framework, Contrastive Graph and Information Bottleneck Fusion Learning (FusionBC), tailored for robust multivariate time series forecasting. By integrating Contrastive Graph Learning with the Information Bottleneck principle, our method selectively filters out irrelevant or redundant information during the learning process, leading to refined, noise-resilient representations that enhance forecasting accuracy. Specifically, the FusionBC framework employs adaptive graph augmentation, intelligently dropping edges or nodes to optimize graph structures and reinforce robustness. Additionally, it combines scale-wise temporal convolution with dual-flow spatial attentive graph convolution, effectively capturing both inter-variable dynamics and multi-scale intra-variable dependencies. Extensive evaluations on real-world datasets show that FusionBC consistently outperforms state-of-the-art methods, achieving a 4.0% improvement in RSE on multivariate forecasting benchmarks. This improvement underscores FusionBC's ability to deliver enhanced, noise-resilient forecasting results that are robust across a range of time series scenarios.
Yangyang Shi, Qianqian Ren, Zhijuan Li
HPCC2
2025 Agent-as-a-Judge: Evaluate Agents with Agents
abstract
Contemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes—ignoring the step-by-step nature of the thinking done by agentic systems—or require excessive manual labour. To address this, we introduce the Agent-as-a-Judge framework, wherein agentic systems are used to evaluate agentic systems. This is a natural extension of the LLM-as-a-Judge framework, incorporating agentic features that enable intermediate feedback for the entire task-solving processes for more precise evaluations. We apply the Agent-as-a-Judge framework to the task of code generation. To overcome issues with existing benchmarks and provide a proof-of-concept testbed for Agent-as-a-Judge, we present DevAI, a new benchmark of 55 realistic AI code generation tasks. DevAI includes rich manual annotations, like a total of 365 hierarchical solution requirements, which make it particularly suitable for an agentic evaluator. We benchmark three of the top code-generating agentic systems using Agent-as-a-Judge and find that our framework dramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline. Altogether, we believe that this work represents a concrete step towards enabling vastly more sophisticated agentic systems. To help that, our dataset and the full implementation of Agent-as-a-Judge will be publically available at https://github.com/metauto-ai/agent-as-a-judge
Mingchen Zhuge, Changsheng Zhao 0002, Dylan R. Ashley, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, Jürgen Schmidhuber
ICML11
2025 MASV: Speaker Verification with Global and Local Context Mamba
Yiteng Huang, Ming Sun 0013, Xinhao Mei, Yangyang Shi, Florian Metze
INTERSPEECH7
2025 ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
abstract
The optimal bit-width for achieving the best trade-off between quantized model size and accuracy has been a subject of ongoing debate. While some advocate for 4-bit quantization, others propose that 1.58-bit offers superior results. However, the lack of a cohesive framework for different bits has left such conclusions relatively tenuous. We present ParetoQ, the first unified framework that facilitates rigorous comparisons across 1-bit, 1.58-bit, 2-bit, 3-bit, and 4-bit quantization settings. Our findings reveal a notable learning transition between 2 and 3 bits: For 3-bits and above, the fine-tuned models stay close to their original pre-trained distributions, whereas for learning 2-bit networks or below, the representations change drastically. By optimizing training schemes and refining quantization functions, ParetoQ surpasses all previous methods tailored to specific bit widths. Remarkably, our ParetoQ ternary 600M-parameter model even outperforms the previous SoTA ternary 3B-parameter model in accuracy, using only one-fifth of the parameters. Extensive experimentation shows that ternary, 2-bit, and 3-bit quantization maintains comparable performance in the size-accuracy trade-off and generally exceeds 4-bit and binary quantization. Considering hardware constraints, 2-bit quantization offers promising potential for memory reduction and speedup.
Zechun Liu, Changsheng Zhao 0002, Hanxian Huang, Scott Roy, Lisa Jin, Yunyang Xiong, Yangyang Shi, Yuandong Tian, Bilge Soran, Raghuraman Krishnamoorthi, Tijmen Blankevoort, Vikas Chandra
NeurIPS10
2024 Tumor Micro-Environment Interactions Guided Graph Learning for Survival Analysis of Human Cancers from Whole-Slide Pathological Images
abstract
The recent advance of deep learning technology brings the possibility of assisting the pathologist to predict the patients' survival from whole-slide pathological images (WSIs). However, most of the prevalent methods only worked on the sampled patches in specifically or randomly selected tumor areas of WSIs, which has very limited capability to capture the complex interactions between tumor and its surrounding micro-environment components. As a matter of fact, tumor is supported and nurtured in the heterogeneous tumor micro-environment(TME), and the detailed analysis of TME and their correlation with tumors are important to in-depth analyze the mechanism of cancer development. In this paper, we considered the spatial interactions among tumor and its two major TME components (i.e., lymphocytes and stromal fibrosis) and presented a Tumor Micro-environment Interactions Guided Graph Learning (TMEGL) algorithm for the prognosis prediction of human cancers. Specifically, we firstly selected different types of patches as nodes to build graph for each WSI. Then, a novel TME neighborhood organization guided graph embedding algorithm was proposed to learn node representations that can preserve their topological structure information. Finally, a Gated Graph Attention Network is applied to capture the survival-associated intersections among tumor and different TME components for clinical outcome prediction. We tested TMEGL on three cancer cohorts derived from The Cancer Genome Atlas (TCGA), and the experimental results indicated that TMEGL not only outperforms the existing WSI-based survival analysis models, but also has good explainable ability for survival prediction.
Wei Shao 0005, Yangyang Shi, Daoqiang Zhang, Peng Wan 0004
CVPR2
2024 Target-Aware Language Modeling via Granular Data Sampling
abstract
Ernie Chang, Pin-Jie Lin, Yang Li, Changsheng Zhao, Daeil Kim, Rastislav Rabatin, Zechun Liu, Yangyang Shi, Vikas Chandra. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Ernie Chang, Pin-Jie Lin, Yang Li 0183, Changsheng Zhao 0002, Daeil Kim, Rastislav Rabatin, Zechun Liu, Yangyang Shi, Vikas Chandra
EMNLP8
2024 Scheduled Execution-Based Binary Indirect Call Targets Refinement
Yangyang Shi, Linan Tian, Yanqi Yang
ESORICS (3)1
2024 In-Context Prompt Editing for Conditional Audio Generation
abstract
Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-audio generation where the encoded representations are easily undermined by unseen prompts, which leads to the degradation of generated audio — the limited set of the text-audio pairs remains inadequate for conditional audio generation in the wild as user prompts are under-specified. In particular, we observe a consistent audio quality degradation in generated audio samples with user prompts, as opposed to training set prompts. To this end, we present a retrieval-based in-context prompt editing framework that leverages the training captions as demonstrative exemplars to revisit the user prompts. We show that the framework enhanced the audio quality across the set of collected user prompts, which were edited with reference to the training captions as exemplars.
Ernie Chang, Pin-Jie Lin, Yang Li 0183, Sidd Srinivasan, Gaël Le Lan, David Kant, Yangyang Shi, Forrest N. Iandola, Vikas Chandra
ICASSP7
2024 On the Open Prompt Challenge in Conditional Audio Generation
abstract
Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is challenging as user-input prompts are often under-specified when compared to text descriptions used to train TTA models. In this work, we treat TTA models as a "blackbox" and address the user prompt challenge with two key insights: (1) User prompts are generally under-specified, leading to a large alignment gap between user prompts and training prompts. (2) There is a distribution of audio descriptions for which TTA models are better at generating higher quality audio, which we refer to as "audionese". To this end, we rewrite prompts with instruction-tuned models and propose utilizing text-audio alignment as feedback signals via margin ranking learning for audio improvements. On both objective and subjective human evaluations, we observed marked improvements in both text-audio alignment and music audio quality.
Ernie Chang, Sidd Srinivasan, Mahi Luthra, Pin-Jie Lin, Varun Nagaraja, Forrest N. Iandola, Zechun Liu, Zhaoheng Ni, Changsheng Zhao 0002, Yangyang Shi, Vikas Chandra
ICASSP10
2024 Stack-and-Delay: A New Codebook Pattern for Music Generation
abstract
Language modeling based music generation relies on discrete representations of audio frames. An audio frame (e.g. 20ms) is typically represented by a set of discrete codes (e.g. 4) computed by a neural codec. Autoregressive decoding typically generates a few thousands of codes per song, which is prohibitively slow and implies introducing some parallel decoding. In this paper we compare different decoding strategies that aim to understand what codes can be decoded in parallel without penalizing the quality too much. We propose a novel stack-and-delay style of decoding to improve upon the vanilla (flattened codes) decoding, with a 4 fold inference speedup. This brings inference speed close to that of the previous state of the art (delay strategy). For the same inference efficiency budget the proposed approach outperforms in objective evaluations, almost closing the gap with vanilla quality-wise. The results are supported by spectral analysis and listening tests, which demonstrate that the samples produced by the new model exhibit improved high-frequency rendering and better maintenance of harmonics and rhythm patterns.
Gaël Le Lan, Varun Nagaraja, Ernie Chang, David Kant, Zhaoheng Ni, Yangyang Shi, Forrest N. Iandola, Vikas Chandra
ICASSP6
2024 Folding Attention: Memory and Power Optimization for On-Device Transformer-Based Streaming Speech Recognition
abstract
Transformer-based models excel in speech recognition. Existing efforts to optimize Transformer inference, typically for long-context applications, center on simplifying attention score calculations. However, streaming speech recognition models usually process a limited number of tokens each time, making attention score calculation less of a bottleneck. Instead, the bottleneck lies in the linear projection layers of multi-head attention and feedforward networks, constituting a substantial portion of the model size and contributing significantly to computation, memory, and power usage.To address this bottleneck, we propose folding attention, a technique targeting these linear layers, significantly reducing model size and improving memory and power efficiency. Experiments on on-device Transformer-based streaming speech recognition models show that folding attention reduces model size (and corresponding memory consumption) by up to 24% and power consumption by up to 23%, all without compromising model accuracy or computation overhead.
Yang Li 0183, Liangzhen Lai, Yuan Shangguan, Forrest N. Iandola, Zhaoheng Ni, Ernie Chang, Yangyang Shi, Vikas Chandra
ICASSP7
2024 MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
abstract
This paper addresses the growing need for efficient large language models (LLMs) on mobile devices, driven by increasing cloud costs and latency concerns. We focus on designing top-quality LLMs with fewer than a billion parameters, a practical choice for mobile deployment. Contrary to prevailing belief emphasizing the pivotal role of data and parameter quantity in determining model quality, our investigation underscores the significance of model architecture for sub-billion scale LLMs. Leveraging deep and thin architectures, coupled with embedding sharing and grouped-query attention mechanisms, we establish a strong baseline network denoted as MobileLLM, which attains a remarkable 2.7%/4.3% accuracy boost over preceding 125M/350M state-of-the-art models. Additionally, we propose an immediate block-wise weight-sharing approach with no increase in model size and only marginal latency overhead. The resultant models, denoted as MobileLLM-LS, demonstrate a further accuracy enhancement of 0.7%/0.8% than MobileLLM 125M/350M. Moreover, MobileLLM model family shows significant improvements compared to previous sub-billion models on chat benchmarks, and demonstrates close correctness to LLaMA-v2 7B in API calling tasks, highlighting the capability of small models for common on-device use cases.
Zechun Liu, Changsheng Zhao 0002, Forrest N. Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, Vikas Chandra
ICML9
2024 Speech ReaLLM - Real-time Speech Recognition with Multimodal Language Models by Teaching the Flow of Time
Frank Seide, Yangyang Shi, Morrie Doulaty, Yashesh Gaur, Junteng Jia, Chunyang Wu
INTERSPEECH2
2024 Data Efficient Reflow for Few Step Audio Generation
abstract
Flow matching has been successfully applied onto generative models, particularly in producing high-quality images and audio. However, the iterative sampling required for the ODE solver in flow matching-based approaches can be time-consuming. Reflow finetune, a technique derived from Rectified flow, offers a promising solution by transforming the ODE trajectory into a straight one, thereby reducing the number of sampling steps. In this paper, we focus on developing data-efficient flow-based approaches for text-to-audio generation. We found that directly applying reflow to the pre-trained flow matching-based audio generation models is typically computationally expensive. It requires over 50,000 training iterations and five times the amount of training data to achieve satisfactory results. To address this issue, we introduce a novel data-efficient reflow (DEreflow) method. This method modifies the reflow data pairs and trajectory to align with the flow matching distribution. As a result of this alignment, our approach requires significantly fewer steps (8,000 compared to 50,000) and data pairs $(0.5$ times the scale of training data compared to 5 times). Results show that the proposed DEreflow consistently outperforms the original reflow method on the text-to-audio generation task.
Lemeng Wu, Zhaoheng Ni, Bowen Shi 0002, Gaël Le Lan, Anurag Kumar 0003, Varun Nagaraja, Xinhao Mei, Yunyang Xiong, Bilge Soran, Raghuraman Krishnamoorthi, Wei-Ning Hsu, Yangyang Shi, Vikas Chandra
SLT12
2024 StegoType: Surface Typing from Egocentric Cameras
abstract
Text input is a critical component of any general purpose computing system, yet efficient and natural text input remains a challenge in AR and VR. Headset based hand-tracking has recently become pervasive among consumer VR devices and affords the opportunity to enable touch typing on virtual keyboards. We present an approach for decoding touch typing on uninstrumented flat surfaces using only egocentric camera-based hand-tracking as input. While egocentric hand-tracking accuracy is limited by issues like self occlusion and image fidelity, we show that a sufficiently diverse training set of hand motions paired with typed text can enable a deep learning model to extract signal from this noisy input. Furthermore, by carefully designing a closed-loop data collection process, we can train an end-to-end text decoder that accounts for natural sloppy typing on virtual keyboards. We evaluate our work with a user study (n=18) showing a mean online throughput of 42.4 WPM with an uncorrected error rate (UER) of 7% with our method compared to a physical keyboard baseline of 74.5 WPM at 0.8% UER, showing progress towards unlocking productivity and high throughput use cases in AR/VR.
Fadi Botros, Yangyang Shi, Pinhao Guo, Bradford J. Snow, Linguang Zhang, Jingming Dong, Keith Vertanen, Shugao Ma, Robert Wang 0002
UIST3
2023 Binary and Ternary Natural Language Generation
abstract
Ternary and binary neural networks enable multiplication-free computation and promise multiple orders of magnitude efficiency gains over full-precision networks if implemented on specialized hardware.However, since both the parameter and the output space are highly discretized, such networks have proven very difficult to optimize.The difficulties are compounded for the class of transformer text generation models due to the sensitivity of the attention operation to quantization and the noise-compounding effects of autoregressive decoding in the high-cardinality output space.We approach the problem with a mix of statistics-based quantization for the weights and elastic quantization of the activations and demonstrate the first ternary and binary transformer models on the downstream tasks of summarization and machine translation.Our ternary BART base achieves an R1 score of 41 on the CNN/DailyMail benchmark, which is merely 3.9 points behind the full model while being 16x more efficient.Our binary model, while less accurate, achieves a highly nontrivial score of 35.6.For machine translation, we achieved BLEU scores of 21.7 and 17.6 on the WMT16 En-Ro benchmark, compared with a full precision mBART model score of 26.8.We also compare our approach in the 8-bit activation setting, where our ternary and even binary weight models can match or outperform the best existing 8-bit weight models in the literature.Our code and models are available at: https://github.com/facebookresearch/ Ternary_Binary_Transformer.
Zechun Liu, Barlas Oguz, Aasish Pappu, Yangyang Shi, Raghuraman Krishnamoorthi
ACL (1)4
2023 Towards Zero-Shot Multilingual Transfer for Code-Switched Responses
abstract
Ting-Wei Wu, Changsheng Zhao, Ernie Chang, Yangyang Shi, Pierce Chuang, Vikas Chandra, Biing Juang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ting-Wei Wu, Changsheng Zhao 0002, Ernie Chang, Yangyang Shi, Pierce Chuang, Vikas Chandra, Biing-Hwang Juang
ACL (1)4
2023 TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for Pytorch
abstract
TorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing well-designed, easy-to-use, and performant PyTorch components. Its contributors routinely engage with users to understand their needs and fulfill them by developing impactful features. Here, we survey TorchAudio’s development principles and contents and highlight key features we include in its latest version (2.1): self-supervised learning pre-trained pipelines and training recipes, high-performance CTC decoders, speech recognition models and training recipes, advanced media I/O capabilities, and tools for performing forced alignment, multi-channel speech enhancement, and reference-less speech assessment. For a selection of these features, through empirical studies, we demonstrate their efficacy and show that they achieve competitive or state-of-the-art performance.
Jeff Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang 0007, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma 0001, Ruizhe Huang, Vineel Pratap, Yuekai Zhang, Anurag Kumar 0003, Chin-Yun Yu, Chuang Zhu, Chunxi Liu, Jacob Kahn, Mirco Ravanelli, Shinji Watanabe 0001, Yangyang Shi, Yumeng Tao
ASRU19
2023 Improving fast-slow Encoder based Transducer with Streaming Deliberation
abstract
This paper introduces a fast-slow encoder based transducer with streaming deliberation for end-to-end automatic speech recognition. We aim to improve the recognition accuracy of the fast-slow encoder based transducer while keeping its latency low by integrating a streaming deliberation model. Specifically, the deliberation model leverages partial hypotheses from the streaming fast encoder and implicitly learns to correct recognition errors. We modify the parallel beam search algorithm for fast-slow encoder based transducer to be efficient and compatible with the deliberation model. In addition, the deliberation model is designed to process streaming data. To further improve the deliberation performance, a simple text augmentation approach is explored. We also compare LSTM and Conformer models for encoding partial hypotheses. Experiments on Librispeech and in-house data show relative WER reductions (WERRs) from 3% to 5% with a slight increase in model size and negligible extra token emission latency compared with fast-slow encoder based transducer. The deliberation model also reduces rare WERs by 3-7% on large-scale in-house data. Compared with vanilla neural transducers, the proposed deliberation model together with fast-slow encoder based transducer obtains relative 10-11% WERRs on Librispeech and around relative 6% WERR on in-house data with smaller emission delays.
Ke Li 0023, Jay Mahadeokar, Jinxi Guo, Yangyang Shi, Gil Keren, Ozlem Kalinli, Michael L. Seltzer
ICASSP4
2023 SCA: Streaming Cross-Attention Alignment For Echo Cancellation
abstract
End-to-End deep learning has shown promising results for speech enhancement tasks, such as noise suppression, dereverberation, and speech separation. However, most state-of-the-art methods for echo cancellation are either classical DSP-based or hybrid DSP-ML algorithms. Components such as the delay estimator and adaptive linear filter are based on traditional signal processing concepts, and deep learning algorithms typically only serve to replace the non-linear residual echo suppressor. This paper introduces an end-to-end echo cancellation network with a streaming cross-attention alignment (SCA). Our proposed method can handle unaligned inputs without requiring external alignment and generate high-quality speech without echoes. At the same time, the end-to-end algorithm simplifies the current echo cancellation pipeline for time-variant echo path cases. We test our proposed method on the ICASSP2022 and Inter-speech2021 Microsoft deep echo cancellation challenge evaluation dataset, where our method outperforms some of the other hybrid and end-to-end methods.
Yang Liu 0175, Yangyang Shi, Kaustubh Kalgaonkar, Sriram Srinivasan 0003
ICASSP2
2023 Multi-Head State Space Model for Speech Recognition
Yassir Fathullah, Chunyang Wu, Yuan Shangguan, Junteng Jia, Wenhan Xiong, Jay Mahadeokar, Chunxi Liu, Yangyang Shi, Ozlem Kalinli, Mike Seltzer, Mark J. F. Gales
INTERSPEECH8
2023 Biased Self-supervised Learning for ASR
Florian Kreyssig, Yangyang Shi, Jinxi Guo, Leda Sari, Abdel-rahman Mohamed, Philip C. Woodland
INTERSPEECH2
2023 Characterizing the Survival-Associated Interactions Between Tumor-Infiltrating Lymphocytes and Tumors From Pathological Images and Multi-Omics Data
abstract
The tumor-infiltrating lymphocytes (TILs) and its correlation with tumors have shown significant values in the development of cancers. Many observations indicated that the combination of the whole-slide pathological images (WSIs) and genomic data can better characterize the immunological mechanisms of TILs. However, the existing image-genomic studies evaluated the TILs by the combination of pathological image and single-type of omics data (e.g., mRNA), which is difficulty in assessing the underlying molecular processes of TILs holistically. Additionally, it is still very challenging to characterize the intersections between TILs and tumor regions in WSIs and the high dimensional genomic data also brings difficulty for the integrative analysis with WSIs. Based on the above considerations, we proposed an end-to-end deep learning framework i.e., IMO-TILs that can integrate pathological image with multi-omics data (i.e., mRNA and miRNA) to analyze TILs and explore the survival-associated interactions between TILs and tumors. Specifically, we firstly apply the graph attention network to describe the spatial interactions between TILs and tumor regions in WSIs. As to genomic data, the Concrete AutoEncoder (i.e., CAE) is adopted to select survival-associated Eigengenes from the high-dimensional multi-omics data. Finally, the deep generalized canonical correlation analysis (DGCCA) accompanied with the attention layer is implemented to fuse the image and multi-omics data for prognosis prediction of human cancers. The experimental results on three cancer cohorts derived from the Cancer Genome Atlas (TCGA) indicated that our method can both achieve higher prognosis results and identify consistent imaging and multi-omics bio-markers correlated strongly with the prognosis of human cancers.
Wei Shao 0005, Yingli Zuo, Yangyang Shi, Yawen Wu, Jiao Tang, Junyong Zhao, Liang Sun 0009, Zixiao Lu, Jianpeng Sheng, Qi Zhu 0001, Daoqiang Zhang
IEEE Trans. Medical Imaging3
2022 Gadgets Splicing: Dynamic Binary Transformation for Precise Rewriting
abstract
Many systems and applications depend on binary rewriting technology to analyze and retrofit software binaries when source code is not available, including binary instrumentation, profiling and security policy reinforcement. However, the investigations have found that many static binary rewriters still fail to accurately transform all legal instructions in binaries. Dynamic binary rewriters allow for accuracy, but coverage and rewriting efficiency are limited. Therefore, the existing binary rewriting technology cannot meet all the needs of binary rewriting. In this paper, we present GRIN, a novel binary rewriting tool that allows for high-precision instruction identification. In GRIN, we propose a gadget-based entry address analysis technique. It identifies the entry addresses of the basic blocks in the binary by gathering and executing the basic blocks related to the computation of the entry addresses of the basic blocks. By traversing from these entries as the new entries of the program, we guarantee the correctness of the identified instructions. We have implemented the prototype of GRIN and evaluated on the SPEC2006 and the whole set of GNU Coreutils. We demonstrate that the precision of GRIN is improved to 99.92% compared to current state-of the-art techniques.
Linan Tian, Yangyang Shi, Yanqi Yang
CGO2
2022 Streaming Transformer Transducer based Speech Recognition Using Non-Causal Convolution
abstract
This paper improves the streaming transformer transducer for speech recognition using non-causal convolution. Many works apply the causal convolution to improve streaming transformer ignoring the lookahead context. We propose to use non-causal convolution to process the center block and lookahead context separately. This method leverages the lookahead context in convolution and maintains similar training and decoding efficiency. Given the similar latency, using the non-causal convolution with lookahead context gives better accuracy than causal convolution, especially for open-domain dictation. Besides, this paper applies talking-head attention and a novel history context compression scheme to further improve the performance. The talking-head attention improves the multi-head self-attention by transferring information among different heads. The history context compression method introduces more extended history context compactly. On our in-house data, the proposed methods improve a small Emformer baseline with lookahead context by relative WERR 5.1%, 14.5%, 8.4% on open-domain dictation, assistant general scenarios, and assistant calling scenarios respectively.
Yangyang Shi, Chunyang Wu, Dilin Wang, Alex Xiao, Jay Mahadeokar, Xiaohui Zhang 0007, Chunxi Liu, Ke Li 0023, Yuan Shangguan, Varun Nagaraja, Ozlem Kalinli, Mike Seltzer
ICASSP1
2022 Streaming parallel transducer beam search with fast slow cascaded encoders
Jay Mahadeokar, Yangyang Shi, Ke Li 0023, Jiedan Zhu, Vikas Chandra, Ozlem Kalinli, Michael L. Seltzer
INTERSPEECH2
2022 Learning a Dual-Mode Speech Recognition Model VIA Self-Pruning
abstract
There is growing interest in unifying the streaming and full-context automatic speech recognition (ASR) networks into a single end-to-end ASR model to simplify the model training and deployment for both use cases. While in real-world ASR applications, the streaming ASR models typically operate under more storage and computational constraints - e.g., on embedded devices - than any server-side full-context models. Motivated by the recent progress in Omni-sparsity supernet training, where multiple subnetworks are jointly optimized in one single model, this work aims to jointly learn a compact sparse on-device streaming ASR model, and a large dense server non-streaming model, in a single supernet. Next, we present that, performing supernet training on both wav2vec 2.0 self-supervised learning and supervised ASR fine-tuning can not only substantially improve the large non-streaming model as shown in prior works, and also be able to improve the compact sparse streaming model.
Chunxi Liu, Yuan Shangguan, Haichuan Yang, Yangyang Shi, Raghuraman Krishnamoorthi, Ozlem Kalinli
SLT4
2022 Synergistic Digital Twin and Holographic Augmented-Reality-Guided Percutaneous Puncture of Respiratory Liver Tumor
abstract
Thermal ablation is an exciting new minimally invasive treatment that destroys liver tumors without removing them. It uses image guidance to place a needle through the skin into a liver tumor, which is highly dependent on surgeons’ experience. With the development of digital medicine, augmented reality (AR) has become a more intuitive and safer way to achieve real-time navigation. However, the technology is still in its infancy due to its limited accuracy and real-time performance. To address these problems, we syncretized the holographic AR with the digital twin technique to track the dynamic surgical scene and provide the 3-D navigation of heterogeneous target regions via internal motion prediction. To tackle the dilemma of real-time performance and precise internal motion estimation, a dynamic adaptation scheme is proposed to compensate for the time cost induced by the external/internal correlation model and data transmission. We carried out a series of experiments to validate our methods. With the proposed external/internal correlation model, the average estimation errors of the tumor and vessels are 2.18 and 2.79 mm, respectively. Besides, we performedin vivoexperiments on two beagle dogs with an artificial lesion in their liver, respectively, and the puncture accuracy of our method are 2.5 and 2.17 mm. The results show that on one hand, our method can fulfill the real-time requirement of AR via using the intraoperative data, which is also more precise than that with preoperative data. On the other hand, our method can provide more 3-D information for surgeons, such as vessels, which can well ensure the safety of operation.
Yangyang Shi, Xuesong Deng, Yuqi Tong, Ruotong Li, Lijie Ren, Weixin Si
IEEE Trans. Hum. Mach. Syst.1
2021 On Lattice-Free Boosted MMI Training of HMM and CTC-Based Full-Context ASR Models
abstract
Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implemented in different frameworks. In this paper, by decoupling the concepts of modeling units and label topologies and building proper numerator/denominator graphs accordingly, we establish a generalized framework for hybrid acoustic modeling (AM). In this framework, we show that LF-MMI is a powerful training criterion applicable to both limited-context and full-context models, for wordpiece/mono-char/bi-char/chenone units, with both HMM/CTC topologies. From this framework, we propose three novel training schemes: chenone(ch)/wordpiece(wp)-CTC-bMMI, and wordpiece(wp)-HMM-bMMI with different advantages in training performance, decoding efficiency and decoding time-stamp accuracy. The advantages of different training schemes are evaluated comprehensively on Librispeech, and wp-CTC-bMMI and ch-CTC-bMMI are evaluated on two real world ASR tasks to show their effectiveness. Besides, we also show bi-char(bc) HMM-MMI models can serve as better alignment models than traditional non-neural GMM-HMMs.
Xiaohui Zhang 0007, Vimal Manohar, Frank Zhang 0001, Yangyang Shi, Nayan Singhal, Julian Chan, Fuchun Peng, Yatharth Saraf, Mike Seltzer
ASRU5
2021 Emformer: Efficient Memory Transformer Based Acoustic Model for Low Latency Streaming Speech Recognition
abstract
This paper proposes an efficient memory transformer Emformer for low latency streaming speech recognition. In Emformer, the long-range history context is distilled into an augmented memory bank to reduce self-attention’s computation complexity. A cache mechanism saves the computation for the key and value in self-attention for the left context. Emformer applies a parallelized block processing in training to support low latency models. We carry out experiments on benchmark LibriSpeech data. Under average latency of 960 ms, Emformer gets WER 2.50% on test-clean and 5.62% on test-other. Comparing with a strong baseline augmented memory transformer (AM-TRF), Emformer gets 4.6 folds training speedup and 18% relative real-time factor (RTF) reduction in decoding with relative WER reduction 17% on test-clean and 9% on test-other. For a low latency scenario with an average latency of 80 ms, Emformer achieves WER 3.01% on test-clean and 7.09% on test-other. Comparing with the LSTM baseline with the same latency and model size, Emformer gets relative WER reduction 9% and 16% on test-clean and test-other, respectively.
Yangyang Shi, Yongqiang Wang 0005, Chunyang Wu, Ching-Feng Yeh, Julian Chan, Frank Zhang 0001, Mike Seltzer
ICASSP1
2021 Transformer in Action: A Comparative Study of Transformer-Based Acoustic Models for Large Scale Speech Recognition Applications
abstract
Transformer-based acoustic models have shown promising results very recently. In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model [1] for large scale speech recognition applications. We compare the transformer based acoustic models with their LSTM counterparts on industrial scale tasks. Specifically, we compare Emformer with latency-controlled BLSTM (LCBLSTM) on medium latency tasks and LSTM on low latency tasks. On a low latency voice assistant task, Emformer gets 24% to 26% relative word error rate reductions (WERRs). For medium latency scenarios, comparing with LCBLSTM with similar model size and latency, Emformer gets significant WERR across four languages in video captioning datasets with 2-3 times inference real-time factors reduction.
Yongqiang Wang 0005, Yangyang Shi, Frank Zhang 0001, Chunyang Wu, Julian Chan, Ching-Feng Yeh, Alex Xiao
ICASSP2
2021 Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
abstract
How to leverage dynamic contextual information in end-toend speech recognition has remained an active research area.Previous solutions to this problem were either designed for specialized use cases that did not generalize well to open-domain scenarios, did not scale to large biasing lists, or underperformed on rare long-tail words.We address these limitations by proposing a novel solution that combines shallow fusion, trie-based deep biasing, and neural network language model contextualization.These techniques result in significant 19.5% relative Word Error Rate improvement over existing contextual biasing approaches and 5.4%-9.3%improvement compared to a strong hybrid baseline on both open-domain and constrained contextualization tasks, where the targets consist of mostly rare long-tail words.Our final system remains lightweight and modular, allowing for quick modification without model re-training.
Mahaveer Jain, Gil Keren, Suyoun Kim, Yangyang Shi, Jay Mahadeokar, Julian Chan, Yuan Shangguan, Christian Fügen, Ozlem Kalinli, Yatharth Saraf, Michael L. Seltzer
Interspeech5
2021 Flexi-Transducer: Optimizing Latency, Accuracy and Compute for Multi-Domain On-Device Scenarios
Jay Mahadeokar, Yangyang Shi, Yuan Shangguan, Chunyang Wu, Alex Xiao, Ozlem Kalinli, Christian Fügen, Michael L. Seltzer
Interspeech2
2021 Collaborative Training of Acoustic Encoders for Speech Recognition
abstract
On-device speech recognition requires training models of different sizes for deploying on devices with various computational budgets. When building such different models, we can benefit from training them jointly to take advantage of the knowledge shared between them. Joint training is also efficient since it reduces the redundancy in the training procedure's data handling operations. We propose a method for collaboratively training acoustic encoders of different sizes for speech recognition. We use a sequence transducer setup where different acoustic encoders share a common predictor and joiner modules. The acoustic encoders are also trained using co-distillation through an auxiliary task for frame level chenone prediction, along with the transducer loss. We perform experiments using the LibriSpeech corpus and demonstrate that the collaboratively trained acoustic encoders can provide up to a 11% relative improvement in the word error rate on both the test partitions.
Varun Nagaraja, Yangyang Shi, Ganesh Venkatesh, Ozlem Kalinli, Michael L. Seltzer, Vikas Chandra
Interspeech2
2021 Dissecting User-Perceived Latency of On-Device E2E Speech Recognition
abstract
As speech-enabled devices such as smartphones and smart speakers become increasingly ubiquitous, there is growing interest in building automatic speech recognition (ASR) systems that can run directly on-device; end-to-end (E2E) speech recognition models such as recurrent neural network transducers and their variants have recently emerged as prime candidates for this task.Apart from being accurate and compact, such systems need to decode speech with low user-perceived latency (UPL), producing words as soon as they are spoken.This work examines the impact of various techniques -model architectures, training criteria, decoding hyperparameters, and endpointer parameters -on UPL.Our analyses suggest that measures of model size (parameters, input chunk sizes), or measures of computation (e.g., FLOPS, RTF) that reflect the model's ability to process input frames are not always strongly correlated with observed UPL.Thus, conventional algorithmic latency measurements might be inadequate in accurately capturing latency observed when models are deployed on embedded devices.Instead, we find that factors affecting token emission latency, and endpointing behavior have a larger impact on UPL.We achieve the best trade-off between latency and word error rate when performing ASR jointly with endpointing, while utilizing the recently proposed alignment regularization mechanism.
Yuan Shangguan, Rohit Prabhavalkar, Jay Mahadeokar, Yangyang Shi, Jiatong Zhou, Chunyang Wu, Ozlem Kalinli, Christian Fügen, Michael L. Seltzer
Interspeech5
2021 Dynamic Encoder Transducer: A Flexible Solution for Trading Off Accuracy for Latency
abstract
We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or finetuning. To trading off accuracy and latency, DET assigns different encoders to decode different parts of an utterance. We apply and compare the layer dropout and the collaborative learning for DET training. The layer dropout method that randomly drops out encoder layers in the training phase, can do on-demand layer dropout in decoding. Collaborative learning jointly trains multiple encoders with different depths in one single model. Experiment results on Librispeech and in-house data show that DET provides a flexible accuracy and latency trade-off. Results on Librispeech show that the full-size encoder in DET relatively reduces the word error rate of the same size baseline by over 8%. The lightweight encoder in DET trained with collaborative learning reduces the model size by 25% but still gets similar WER as the full-size baseline. DET gets similar accuracy as a baseline model with better latency on a large in-house data set by assigning a lightweight encoder for the beginning part of one utterance and a full-size encoder for the rest.
Yangyang Shi, Varun Nagaraja, Chunyang Wu, Jay Mahadeokar, Rohit Prabhavalkar, Alex Xiao, Ching-Feng Yeh, Julian Chan, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer
Interspeech1
2021 Transformer-Based Acoustic Modeling for Streaming Speech Synthesis
Chunyang Wu, Zhiping Xiu, Yangyang Shi, Ozlem Kalinli, Christian Fügen, Thilo Köhler
Interspeech3
2021 Streaming Attention-Based Models with Augmented Memory for End-To-End Speech Recognition
abstract
Attention-based models have been gaining popularity recently for their strong performance demonstrated in fields such as machine translation [1] and automatic speech recognition [2]. One major challenge of attention-based models is the need of access to the full sequence and the quadratically growing computational cost concerning the sequence length. These characteristics pose challenges, especially for low-latency scenarios, where the system is often required to be streaming. In this paper, we build a compact and streaming speech recognition system on top of the end-to-end neural transducer architecture [3] with attention-based modules augmented with convolution [2]. The proposed system equips the end-to-end models with the streaming capability and reduces the large footprint from the streaming attention-based model using augmented memory [4], [5]. On the LibriSpeech [6] dataset, our proposed system achieves word error rates 2.7% on test-clean and 5.8% on test-other, to our best knowledge the lowest among streaming approaches reported so far.
Ching-Feng Yeh, Yongqiang Wang 0005, Yangyang Shi, Chunyang Wu, Frank Zhang 0001, Julian Chan, Michael L. Seltzer
SLT3
2020 Mining Effective Negative Training Samples for Keyword Spotting
abstract
Max-pooling neural network architectures have been proven to be useful for keyword spotting (KWS), but standard training methods suffer from a class-imbalance problem when using all frames from negative utterances. To address the problem, we propose an innovative algorithm, Regional Hard-Example (RHE) mining, to find effective negative training samples, in order to control the ratio of negative vs. positive data. To maintain the diversity of the negative samples, multiple non-contiguous difficult frames per negative training utterance are dynamically selected during training, based on the model statistics at each training epoch. Further, to improve model learning, we introduce a weakly constrained max-pooling method for positive training utterances, which constrains max-pooling over the keyword ending frames only at early stages of training. Finally, data augmentation is combined to bring further improvement. We assess the algorithms by conducting experiments on wake-up word detection tasks with two different neural network architectures. The experiments consistently show that the proposed methods provide significant improvements compared to a strong baseline. At a false alarm rate of once per hour, our methods achieve 45-58% relative reduction in false rejection rates over a strong baseline.
Jingyong Hou, Yangyang Shi, Mari Ostendorf, Mei-Yuh Hwang, Lei Xie 0001
ICASSP2
2020 Weak-Attention Suppression for Transformer Based Speech Recognition
abstract
Transformers, originally proposed for natural language processing (NLP) tasks, have recently achieved great success in automatic speech recognition (ASR). However, adjacent acoustic units (i.e., frames) are highly correlated, and long-distance dependencies between them are weak, unlike text units. It suggests that ASR will likely benefit from sparse and localized attention. In this paper, we propose Weak-Attention Suppression (WAS), a method that dynamically induces sparsity in attention probabilities. We demonstrate that WAS leads to consistent Word Error Rate (WER) improvement over strong transformer baselines. On the widely used LibriSpeech benchmark, our proposed method reduced WER by 10%$ on test-clean and 5% on test-other for streamable transformers, resulting in a new state-of-the-art among streaming models. Further analysis shows that WAS learns to suppress attention of non-critical and redundant continuous acoustic frames, and is more likely to suppress past frames rather than future ones. It indicates the importance of lookahead in attention-based ASR models.
Yangyang Shi, Yongqiang Wang 0005, Chunyang Wu, Christian Fügen, Frank Zhang 0001, Ching-Feng Yeh, Michael L. Seltzer
INTERSPEECH1
2020 Streaming Transformer-Based Acoustic Models Using Self-Attention with Augmented Memory
abstract
Transformer-based acoustic modeling has achieved great success for both hybrid and sequence-to-sequence speech recognition.However, it requires access to the full sequence, and the computational cost grows quadratically with respect to the input sequence length.These factors limit its adoption for streaming applications.In this work, we proposed a novel augmented memory self-attention, which attends on a short segment of the input sequence and a bank of memories.The memory bank stores the embedding information for all the processed segments.On the librispeech benchmark, our proposed method outperforms all the existing streamable transformer methods by a large margin and achieved over 15% relative error reduction, compared with the widely used LC-BLSTM baseline.Our findings are also confirmed on some large internal datasets.
Chunyang Wu, Yongqiang Wang 0005, Yangyang Shi, Ching-Feng Yeh, Frank Zhang 0001
INTERSPEECH3
2020 Functional code clone detection with syntax and semantics fusion learning
abstract
Clone detection of source code is among the most fundamental software engineering techniques. Despite intensive research in the past decade, existing techniques are still unsatisfactory in detecting "functional" code clones. In particular, existing techniques cannot efficiently extract syntax and semantics information from source code. In this paper, we propose a novel joint code representation that applies fusion embedding techniques to learn hidden syntactic and semantic features of source codes. Besides, we introduce a new granularity for functional code clone detection. Our approach regards the connected methods with caller-callee relationships as a functionality and the method without any caller-callee relationship with other methods represents a single functionality. Then we train a supervised deep learning model to detect functional code clones. We conduct evaluations on a large dataset of C++ programs and the experimental results show that fusion learning can significantly outperform the state-of-the-art techniques in detecting functional code clones.
Chunrong Fang, Yangyang Shi, Jeff Huang 0001, Qingkai Shi
ISSTA3
2020 Incorporating Android Code Smells into Java Static Code Metrics for Security Risk Prediction of Android Applications
abstract
With the wide-spread use of Android applications in people's daily life, it becomes more and more important to timely identify the security problems of these applications. To enrich existing studies in guarding the security and privacy of Android applications, we attempted to predict the security risk levels of Android applications. Specifically, we proposed an approach that incorporated Android code smells into traditional Java code metrics to predict how secure an Android application is. With an evaluation of our technique on 3,680 Android applications, we found that: (1) Android code smells could help improve the performance of security risk prediction of Android applications; (2) By building a Random Forest model based on Android code smells and Java code metrics, we could achieve an Area Under Curve (AUC) of 0.97; (3) Android code smells such as member ignoring method (MIM) and leaking inner class (LIC) have a relatively-large influence on Android security risk prediction, to which developers should pay more attention during their application development.
Ai Gong, Weiqin Zou, Yangyang Shi, Chunrong Fang
QRS4
2019 End-to-end Speech Recognition Using a High Rank LSTM-CTC Based Model
abstract
Long Short Term Memory Connectionist Temporal Classification (LSTM-CTC) based end-to-end models are widely used in speech recognition due to its simplicity in training and efficiency in decoding. In conventional LSTM-CTC based models, a bottleneck projection matrix maps the hidden feature vectors obtained from LSTM to softmax output layer. In this paper, we propose to use a high rank projection layer to replace the projection matrix. The output from the high rank projection layer is a weighted combination of vectors that are projected from the hidden feature vectors via different projection matrices and non-linear activation function. The high rank projection layer is able to improve the expressiveness of LSTM-CTC models. The experimental results show that on Wall Street Journal (WSJ) corpus and LibriSpeech data set, the proposed method achieves 4% - 6% relative word error rate (WER) reduction over the baseline CTC system. They outperform other published CTC based end-to-end (E2E) models under the condition that no external data or data augmentation is applied. Code has been made available at https://github.com/mobvoi/lstm_ctc.
Yangyang Shi, Mei-Yuh Hwang
ICASSP1
2019 Knowledge Distillation for Recurrent Neural Network Language Modeling with Trust Regularization
abstract
Recurrent Neural Networks (RNNs) have dominated language modeling because of their superior performance over traditional N-gram based models. In many applications, a large Recurrent Neural Network language model (RNNLM) or an ensemble of several RNNLMs is used. These models have large memory footprints and require heavy computation. In this paper, we examine the effect of applying knowledge distillation in reducing the model size for RNNLMs. In addition, we propose a trust regularization method to improve the knowledge distillation training for RNNLMs. Using knowledge distillation with trust regularization, we reduce the parameter size to a third of that of the previously published best model while maintaining the state-of-the-art perplexity result on Penn Treebank data. In a speech recognition N-best rescoring task, we reduce the RNNLM model size to 18.5% of the baseline system, with no degradation in word error rate (WER) performance on Wall Street Journal data set.
Yangyang Shi, Mei-Yuh Hwang, Haoyu Sheng
ICASSP1
2019 Region Proposal Network Based Small-Footprint Keyword Spotting
abstract
We apply an anchor-based region proposal network (RPN) for end-to-end keyword spotting (KWS). RPNs have been widely used for object detection in image and video processing; here, it is used to jointly model keyword classification and localization. The method proposes several anchors as rough locations of the keyword in an utterance and jointly learns classification and transformation to the ground truth region for each positive anchor. Additionally, we extend the keyword/non-keyword binary classification to detect multiple keywords. We verify our proposed method on a hotword detection data set with two hotwords. At a false alarm rate of one per hour, our method achieved more than 15% relative reduction in false rejection of the two keywords over multiple recent baselines. In addition, our method predicts the location of the keyword with over 90% overlap, which can be important for many applications.
Jingyong Hou, Yangyang Shi, Mari Ostendorf, Mei-Yuh Hwang, Lei Xie 0001
IEEE Signal Process. Lett.2
2018 Robust Control for a Magnetically Suspended Control Moment Gyro with Strong Gyroscopic Effects
abstract
For a magnetically suspended control moment gyroscope (MSCMG) with strong gyroscopic effects, the nutation and precession stability control is an important part of system stability. In order to improve nutation and precession stability effectively, a robust control method based on parameter perturbation robust control and regional pole assignment is proposed in this paper. First, the precise state space model of the magnetic bearing rotor system is established, and the gyroscope effect term is regarded as a varying perturbation parameter with the rotor speed. Then, the basic constraints for the stability of the closed loop system in the full speed range are derived from the parameter perturbation robust control method. Moreover, the additional constraints of high stability for the nutation and the precession are derived through the method of the region pole assignment. Finally, all the constraints that the robust controllers need to satisfy are derived into linear matrix inequalities (LMIs), and the robust controller is obtained by solving the LMIs. The validity of the method is verified on the MSCMG simulation platform.
Bangcheng Han, Yangyang Shi
IECON5
2016 Recurrent Support Vector Machines For Slot Tagging In Spoken Language Understanding
abstract
Yangyang Shi, Kaisheng Yao, Hu Chen, Dong Yu, Yi-Cheng Pan, Mei-Yuh Hwang. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Yangyang Shi, Kaisheng Yao, Dong Yu 0001, Yi-Cheng Pan, Mei-Yuh Hwang
HLT-NAACL1
2016 Deep LSTM based Feature Mapping for Query Classification
abstract
Traditional convolutional neural network (CNN) based query classification uses linear feature mapping in its convolution operation.The recurrent neural network (RNN), differs from a CNN in representing word sequence with their ordering information kept explicitly.We propose using a deep long-short-term-memory (DLSTM) based feature mapping to learn feature representation for CNN.The DLSTM, which is a stack of LSTM units, has different order of feature representations at different depth of LSTM unit.The bottom LSTM unit equipped with input and output gates, extracts the first order feature representation from current word.To extract higher order nonlinear feature representation, the LSTM unit at higher position gets input from two parts.First part is the lower LSTM unit's memory cell from previous word.Second part is the lower LSTM unit's hidden output from current word.In this way, the DLSTM captures the nonlinear nonconsecutive interaction within n-grams.Using an architecture that combines a stack of the DLSTM layers with a tradition CNN layer, we have observed new state-of-the-art query classification accuracy on benchmark data sets for query classification.
Yangyang Shi, Kaisheng Yao, Daxin Jiang
HLT-NAACL1
2015 Semi-supervised slot tagging in spoken language understanding using recurrent transductive support vector machines
abstract
In this paper, we propose a recurrent transductive support vector machine (rtsvm) for semi-supervised slot tagging. Taking advantage of the superior sequence representation capability of recurrent neural networks (rnns) and the semi-supervised learning capability of transductive support vector machines (tsvms), the rtsvm is stacking a tsvm on top of a rnn. The performance of the traditional tsvm is sensitive to the regularization weight for unlabeled data in semi-supervised learning. In practice, a suitable unlabeled data regularization weight is difficult to determine. To make the rtsvm semi-supervised learning robust, we propose a confident subset regularization method enforcing that the new decision boundaries learned from unlabeled data would not separate the confident clusters learned from labeled data. The experiments based on two datasets show that without using unlabeled data, the supervised version of rtsvm achieves significant F1 score improvement over previous best methods. By taking the unlabeled data into account, the semi-supervised rtsvm gets significant improvement over its supervised opponent.
Yangyang Shi, Kaisheng Yao, Yi-Cheng Pan, Mei-Yuh Hwang
ASRU1
2015 A factorization network based method for multi-lingual domain classification
abstract
In many spoken language understanding systems (SLUS), domain classification is the most crucial component, as system responses based on wrong domains often yield very unpleasant user experiences. In multi-lingual domain classification, the training data for some poor-resource languages often comes from machine translation. Some of the higher order n-gram features are distorted during machine translation. Feature co-occurrence becomes reliable feature in multi-lingual domain classification. In this paper, in order to effectively model feature co-occurrences, we propose Factorization Networks that are combinations of Factorization Machines (FMs) with Neural Networks (NNs). FNs extend the linear connections from the input feature layer to the hidden layer in NNs to factorization connections that represent the weights of feature co-occurrences using factorized method. In addition to FNs, we also propose a hybrid model that integrates FNs, NNs and Maximum Entropy (ME) models together. The component models in the hybrid model share the same input features. Based on two data sets (ATIS data set and Microsoft Cortana Chinese data ), the proposed models shows promising results. Especially for large Microsoft Cortana Chinese data which is translated from well annotated English data, FNs using unigram, class and query length features achieve more than 20% relative error reduction over linear (SVMs).
Yangyang Shi, Yi-Cheng Pan, Mei-Yuh Hwang, Kaisheng Yao, Yuanhang Zou, Baolin Peng
ICASSP1
2015 Contextual spoken language understanding using recurrent neural networks
abstract
We present a contextual spoken language understanding (contextual SLU) method using Recurrent Neural Networks (RNNs). Previous work has shown that context information, specifically the previously estimated domain assignment, is helpful for domain identification. We further show that other context information such as the previously estimated intent and slot labels are useful for both intent classification and slot filling tasks in SLU. We propose a step-n-gram model to extract sentence-level features from RNNs, which extract sequential features. The step-n-gram model is used together with a stack of Convolution Networks for training domain/intent classification. Our method therefore exploits possible correlations among domain/intent classification and slot filling and incorporates context information from the past predictions of domain/intent and slots. The proposed method obtains new state-of-the-art results on ATIS and improved performances over baseline techniques such as conditional random fields (CRFs) on a large context-sensitive SLU dataset.
Yangyang Shi, Kaisheng Yao, Yi-Cheng Pan, Mei-Yuh Hwang, Baolin Peng
ICASSP1
2015 RNN-based labeled data generation for spoken language understanding
Yik-Cheung Tam, Yangyang Shi, Hunk Chen, Mei-Yuh Hwang
INTERSPEECH2
2015 Recurrent neural network language model adaptation with curriculum learning
Yangyang Shi, Martha A. Larson, Catholijn M. Jonker
Comput. Speech Lang.1
2015 Integrating meta-information into recurrent neural network language models
Yangyang Shi, Martha A. Larson, Joris Pelemans, Catholijn M. Jonker, Patrick Wambacq, Pascal Wiggers, Kris Demuynck
Speech Commun.1
2014 Cluster based Chinese abbreviation modeling
Yangyang Shi, Yi-Cheng Pan, Mei-Yuh Hwang
INTERSPEECH1
2014 Spoken language understanding using long short-term memory neural networks
abstract
Neural network based approaches have recently produced record-setting performances in natural language understanding tasks such as word labeling. In the word labeling task, a tagger is used to assign a label to each word in an input sequence. Specifically, simple recurrent neural networks (RNNs) and convolutional neural networks (CNNs) have shown to significantly outperform the previous state-of-the-art - conditional random fields (CRFs). This paper investigates using long short-term memory (LSTM) neural networks, which contain input, output and forgetting gates and are more advanced than simple RNN, for the word labeling task. To explicitly model output-label dependence, we propose a regression model on top of the LSTM un-normalized scores. We also propose to apply deep LSTM to the task. We investigated the relative importance of each gate in the LSTM by setting other gates to a constant and only learning particular gates. Experiments on the ATIS dataset validated the effectiveness of the proposed models.
Kaisheng Yao, Baolin Peng, Yu Zhang 0033, Dong Yu 0001, Geoffrey Zweig, Yangyang Shi
SLT6
2013 K-component recurrent neural network language models using curriculum learning
abstract
Conventional n-gram language models are known for their limited ability to capture long-distance dependencies and their brittleness with respect to within-domain variations. In this paper, we propose a k-component recurrent neural network language model using curriculum learning (CL-KRNNLM) to address within-domain variations. Based on a Dutch-language corpus, we investigate three methods of curriculum learning that exploit dedicated component models for specific sub-domains. Under an oracle situation in which context information is known during testing, we experimentally test three hypotheses. The first is that domain-dedicated models perform better than general models on their specific domains. The second is that curriculum learning can be used to train recurrent neural network language models (RNNLMs) from general patterns to specific patterns. The third is that curriculum learning, used as an implicit weighting method to adjust the relative contributions of general and specific patterns, outperforms conventional linear interpolation. Under the condition that context information is unknown during testing, the CL-KRNNLM also achieves improvement over conventional RNNLM by 13% relative in terms of word prediction accuracy. Finally, the CL-KRNNLM is tested in an additional experiment involving N-best rescoring on a standard data set. Here, the context domains are created by clustering the training data using Latent Dirichlet Allocation and k-means clustering.
Yangyang Shi, Martha A. Larson, Catholijn M. Jonker
ASRU1
2013 Speed up of recurrent neural network language models with sentence independent subsampling stochastic gradient descent
Yangyang Shi, Mei-Yuh Hwang, Kaisheng Yao, Martha A. Larson
INTERSPEECH1
2013 Exploiting the succeeding words in recurrent neural network language models
Yangyang Shi, Martha A. Larson, Pascal Wiggers, Catholijn M. Jonker
INTERSPEECH1
2013 Recurrent neural networks for language understanding
abstract
Recurrent Neural Network Language Models (RNN-LMs) have recently shown exceptional performance across a variety of applications. In this paper, we modify the architecture to perform Language Understanding, and advance the state-of-the-art for the widely used ATIS dataset. The core of our approach is to take words as input as in a standard RNN-LM, and then to predict slot labels rather than words on the output side. We present several variations that differ in the amount of word context that is used on the input side, and in the use of non-lexical features. Remarkably, our simplest model produces state-of-the-art results, and we advance state-of-the-art through the use of bagof-words, word embedding, named-entity, syntactic, and wordclass features. Analysis indicates that the superior performance is attributable to the task-specific word representations learned by the RNN.
Kaisheng Yao, Geoffrey Zweig, Mei-Yuh Hwang, Yangyang Shi, Dong Yu 0001
INTERSPEECH4
2013 Classifying the socio-situational settings of transcripts of spoken discourses
Yangyang Shi, Pascal Wiggers, Catholijn M. Jonker
Speech Commun.1
2012 Dynamic Bayesian socio-situational setting classification
abstract
We propose a dynamic Bayesian classifier for the socio-situational setting of a conversation. Knowledge of the socio-situational setting can be used to search for content recorded in a particular setting or to select context-dependent models in speech recognition. The dynamic Bayesian classifier has the advantage - compared to static classifiers such a naive Bayes and support vector machines - that it can continuously update the classification during a conversation. We experimented with several models that use lexical and part-of-speech information. Our results show that the prediction accuracy of the dynamic Bayesian classifier using the first 25% of a conversation is almost 98% of the final prediction accuracy, which is calculated on the entire conversation. The best final prediction accuracy, 88.85%, is obtained by bigram dynamic Bayesian classification using words and part-of-speech tags.
Yangyang Shi, Pascal Wiggers, Catholijn M. Jonker
ICASSP1
2012 Towards Recurrent Neural Networks Language Models with Linguistic and Contextual Features
Yangyang Shi, Pascal Wiggers, Catholijn M. Jonker
INTERSPEECH1
2011 Socio-situational setting classification based on language use
abstract
We present a method for automatic classification of the socio-situational setting of a conversation based on the language used. The socio-situational setting depicts the social background of a conversation which involves the communicative goals, number of speakers, number of listeners and the relationship among the speakers and the listeners. Knowledge of the socio-situational setting can be used to search for content recorded in a particular setting or to select context-dependent models for example for speech recognition. We investigated the performance of different feature sets of conversation level features and word level features and their combinations on this task. Our final system, that classifies the conversations in the Spoken Dutch Corpus in one of 14 socio-situational settings, achieves an accuracy of 89.55%.
Yangyang Shi, Pascal Wiggers, Catholijn M. Jonker
ASRU1