EDBT 2026 Demo / reviewers in the wild / expert
Ruizhe Huang
dblp:131/4079
· DBLP profile ↗
20ranked-venue papers
9as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Memory Hotness Gap in Edge Systems with Hotness-Segregated Object AllocationabstractKernel operations in resource-constrained edge systems, such as memory swapping and deduplication, use the access frequency (hotness) of memory pages to guide page placement and reclamation. However, these operations suffer from page-hotness skew: a page may contain a mix of highly accessed and infrequently accessed objects, which causes inaccurate page-level classification, wasted DRAM capacity, and expensive I/O. We attribute this skewness to a cross-layer mismatch: the kernel manages memory at page granularity, whereas user-level allocators place objects without considering access hotness. Ruizhe Huang, Jiahua Wang, Qihang Xu, Peng Jiang 0007, Zhida An, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yuxin Ren 0001, Ning Jia 0004 |
LCTES | 1 |
| 2025 | Long-Form Fuzzy Speech-to-Text Alignment for 1000+ LanguagesabstractConventional speech-to-text forced alignment typically operates at the utterance level. In practice, however, we do not usually have short segments (e.g., 10 seconds) of audio with exact, verbatim transcriptions (e.g., the LibriSpeech corpus) as in lab conditions. Instead, audio often comes in long-form (e.g., an hour-long lecture recording), and the available transcription may be non-verbatim or include unspoken annotations, making it misaligned with the actual speech. This motivates the need for long-form fuzzy speech-to-text alignment, which has practical applications - for example, preparing segmented supervised audio data for training machine learning models. We demonstrate the Torchaudio long-form aligner, which supports such use cases. Moreover, it can be equipped with any CTC model that predicts frame-wise labels, turning the model into a robust and powerful aligner. Ruizhe Huang, Xiaohui Zhang 0007, Zhaoheng Ni, Moto Hira, Jeff Hwang, Vineel Pratap, Ju Lin, Ming Sun 0013, Florian Metze |
ASRU | 1 |
| 2025 | Cost-Efficient Cloud Infrastructure with Hugepage-aware Memory DeduplicationabstractOptimizing memory cost-efficiency is the top demand for many cloud computing scenarios. Memory deduplication and hugepage are both essential techniques for reducing memory cost and improving efficiency. However, the simultaneous use of memory deduplication and hugepages faces a dillema. Existing approaches either split hugepages into small pages to achieve efficient memory deduplication or ignore redundant portions within hugepages to maintain hugepage performance. Ruizhe Huang, Xinyu Wang 0043, Zhida An, Hanwen Lei, Peng Jiang 0007, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu |
SoCC | 1 |
| 2025 | DPUaudit: DPU-assisted Pull-based Architecture for Near-Zero Cost System AuditingabstractSystem auditing frameworks are crucial for modern data center security, as they record system events to detect intrusions. However, existing software-based auditing frameworks are limited by their high runtime overhead. To address the limitations of software-based frameworks, researchers had proposed a hardware-based auditing framework that offloads log processing to isolated hardware. However, despite using powerful specialized hardware, this approach still suffers from high runtime overhead, which contradicts their efficiency goal. We have identified that the high overhead is due to the pushbased architecture, which involves operating a log sender on the monitored host. Consequently, the existing approach requires heavy software protection mechanisms to secure the log sender, resulting in high runtime overhead.In this paper, we propose a new DPU-assisted pull-based architecture called DPUaudit for hardware-based auditing, which achieves near-zero runtime overhead. Instead of using a log sender, DPUaudit utilizes DPU to actively pull system events from the monitored host. This eliminates the need for heavy mechanisms to handle and safeguard the log sender, achieving highly efficient system auditing. Experimental results show that, on average, DPUaudit only slows down applications on the monitored host by 2.1% for six mainstream data center applications under different workloads, which is at least one order of magnitude smaller than existing approaches, while still ensuring the integrity of audit logs. Peng Jiang 0007, Hanlin Jiang, Ruizhe Huang, Hanwen Lei, Zhineng Zhong, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Yao Guo 0001, Xiangqun Chen, Ding Li 0001 |
HPCA | 3 |
| 2025 | Constructing Datasets From Public Police Body Camera FootageabstractThe enormous potential of body-worn cameras to improve accountability in policing remains largely unrealized due to large volumes of unreviewed footage. Transcription and diarization tools could aid in reviewing footage, but lack of public data hinders their development. We develop a pipeline to construct public datasets, making use of the small number of videos publicly released by police departments, with capacity to update the data as footage gets released or removed. Our pipeline produces two datasets, a large one with transcriptions automatically extracted from department-generated captions, and a smaller test set where we manually validated transcripts and alignment. We benchmark ASR models, including models fine-tuned on our data, on our test set, to show applications of our datasets and continued challenges of this domain. Our work presents a new vision for leveraging public body-worn camera footage—even when it can’t be rereleased—to help address this critical social issue. Jamie Rosas-Smith, Martijn Bartelds, Ruizhe Huang, L. Paola García-Perera, Karen Livescu, Daniel Jurafsky, Anjalie Field |
ICASSP | 3 |
| 2025 | Predictable and Secure System Auditing for Real-Time SystemsabstractSystem auditing frameworks are essential for operating system security as they record system events to support intrusion detection, compliance verification and attack reconstruction. However, existing auditing frameworks fail to meet the stringent requirements of real-time systems, which demand security, predictability, and efficiency. Though current solutions are optimized for security or performance, they do not focus on bounding the worst-case execution time (WCET) and incorporating into response-time analysis (RTA). This paper presents RT-NODROP, a secure and predictable auditing framework tailored for real-time systems. RT-NODROP employs a lightweight threadlet-based architecture to isolate audit events processing, periodically invoking threadlets to simultaneously bound WCET and event residence time. By integrating with real-time schedule, RT-NODROP ensures no event dropping, system efficiency, and predictability. We further develop an overhead-aware RTA and a period selection algorithm to balance security, performance, and schedulability. The evaluations demonstrate that RT-NODROP is superior over state-of-the-art frameworks (Sysdig, OMNILOG, Ellipsis), improving schedulability by$\mathbf{8 0. 1 1 \%, ~} \mathbf{1 1 7. 9 \%}$and$\mathbf{5 1. 0 5 \%}$, respectively. For the latency-intensive application Redis, RT-NODROP achieves up to$\mathbf{7 5. 1 \%}(\mathbf{1 3 8. 8 6 \%}, \mathbf{3 2 4. 6 \%})$higher throughput and$\mathbf{2. 1 9} \times$(3.07x, 5.02x) lower 99.9th percentile tail latency than Sysdig (OMNILOG, Ellipsis) while maintaining a minimum event residence time around 10 ms without event dropping. Peng Jiang 0007, Fanhang Hu, Ruizhe Huang, Shuomin Xue, Zhaomeng Deng, Yuxin Ren 0001, Ning Jia 0004, Yao Guo 0001, Xiangqun Chen, Ding Li 0001 |
RTSS | 3 |
| 2025 | Robust, Efficient, and Widely Available Greybox Fuzzing for COTS Binaries with System Call Pattern Feedback
Jifan Xiao, Peng Jiang 0007, Zixi Zhao, Ruizhe Huang, Ding Li 0001 |
USENIX Security Symposium | 4 |
| 2024 | ConEC: Earnings Call Dataset with Real-world Contexts for Benchmarking Contextual Speech RecognitionabstractKnowing the particular context associated with a conversation can help improving the performance of an automatic speech recognition (ASR) system. For example, if we are provided with a list of in-context words or phrases — such as the speaker’s contacts or recent song playlists — during inference, we can bias the recognition process towards this list. There are many works addressing contextual ASR; however, there is few publicly available real benchmark for evaluation, making it difficult to compare different solutions. To this end, we provide a corpus (“ConEC”) and baselines to evaluate contextual ASR approaches, grounded on real-world applications. The ConEC corpus is based on public-domain earnings calls (ECs) and associated supplementary materials, such as presentation slides, earnings news release as well as a list of meeting participants’ names and affiliations. We demonstrate that such real contexts are noisier than artificially synthesized contexts that contain the ground truth, yet they still make great room for future improvement of contextual ASR technology Ruizhe Huang, Mahsa Yarmohammadi, Jan Trmal, Desh Raj, L. Paola García-Perera, Alexei V. Ivanov, Patrick Ehlen, Mingzhi Yu, Daniel Povey, Sanjeev Khudanpur |
LREC/COLING | 1 |
| 2024 | Less Peaky and More Accurate CTC Forced Alignment by Label PriorsabstractConnectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can cause inaccurate forced alignments (FA), especially at finer granularity, e.g., phoneme level. This paper aims at alleviating the peaky behavior for CTC and improve its suitability for forced alignment generation, by leveraging label priors, so that the scores of alignment paths containing fewer blanks are boosted and maximized during training. As a result, our CTC model produces less peaky posteriors and is able to more accurately predict the offset of the tokens besides their onset. It outperforms the standard CTC model and a heuristics-based approach for obtaining CTC’s token offset timestamps by 12 − 40% in phoneme and word boundary errors (PBE and WBE) measured on the Buckeye and TIMIT data. Compared with the most widely used FA toolkit Montreal Forced Aligner (MFA), our method performs similarly on PBE/WBE on Buckeye, yet falls behind MFA on TIMIT. Nevertheless, our method has a much simpler training pipeline and better runtime efficiency. Our training recipe and pretrained model are released in TorchAudio. Ruizhe Huang, Xiaohui Zhang 0007, Zhaoheng Ni, Li Sun 0010, Moto Hira, Jeff Hwang, Vimal Manohar, Vineel Pratap, Matthew Wiesner, Shinji Watanabe 0001, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 1 |
| 2024 | Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
Ruizhe Huang, Mahsa Yarmohammadi, Sanjeev Khudanpur, Daniel Povey |
INTERSPEECH | 1 |
| 2024 | Evaluating the Santa Barbara Corpus: Challenges of the Breadth of Conversational Spoken Language
Matthew Maciejewski, Dominik Klement, Ruizhe Huang, Matthew Wiesner, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2023 | TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for PytorchabstractTorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing well-designed, easy-to-use, and performant PyTorch components. Its contributors routinely engage with users to understand their needs and fulfill them by developing impactful features. Here, we survey TorchAudio’s development principles and contents and highlight key features we include in its latest version (2.1): self-supervised learning pre-trained pipelines and training recipes, high-performance CTC decoders, speech recognition models and training recipes, advanced media I/O capabilities, and tools for performing forced alignment, multi-channel speech enhancement, and reference-less speech assessment. For a selection of these features, through empirical studies, we demonstrate their efficacy and show that they achieve competitive or state-of-the-art performance. Jeff Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang 0007, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma 0001, Ruizhe Huang, Vineel Pratap, Yuekai Zhang, Anurag Kumar 0003, Chin-Yun Yu, Chuang Zhu, Chunxi Liu, Jacob Kahn, Mirco Ravanelli, Shinji Watanabe 0001, Yangyang Shi, Yumeng Tao |
ASRU | 8 |
| 2023 | Building Keyword Search System from End-To-End Asr SystemsabstractKeyword search (KWS) systems are commonly built on top of existing automatic speech recognition (ASR) systems. However, end-to-end (E2E) ASR models are not naturally equipped with word-level timing information or confidence. Existing methods for re-purposing E2E ASR systems for KWS are largely heuristic or model-specific. In this paper, we describe a general KWS pipeline, applicable to any ASR model that generates N-best lists. We extract timing information using either external word-aligners, or time-preserving weighted finite-state transducer-based decoders. We show that our light-weight, ASR-agnostic approach for confidence estimation based on N-best lists outperforms other commonly used heuristics, such as using the decoder’s softmax probability, and even a more complicated dedicated confidence estimation model (CEM). Finally, we compare our performance to hybrid ASR models, extensively evaluating the impact of word-level timing, confidence, and recall on KWS performance. Our KWS pipeline is available online1, suitable for evaluating the aforementioned ASR components as downstream tasks. Ruizhe Huang, Matthew Wiesner, L. Paola García-Perera, Daniel Povey, Jan Trmal, Sanjeev Khudanpur |
ICASSP | 1 |
| 2023 | Auditing Frameworks Need Resource Isolation: A Systematic Study on the Super Producer Threat to System Auditing and Its Mitigation
Peng Jiang 0007, Ruizhe Huang, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Jianhai Luan, Yuxin Ren 0001, Xinwei Hu |
USENIX Security Symposium | 2 |
| 2021 | Training Hybrid Models on Noisy Transliterated Transcripts for Code-Switched Speech Recognition
Matthew Wiesner, Mousmita Sarma, Ashish Arora, Desh Raj, Dongji Gao, Ruizhe Huang, Supreet Preet, Moris Johnson, Zikra Iqbal, Nagendra K. Goel, Jan Trmal, L. Paola García-Perera, Sanjeev Khudanpur |
Interspeech | 6 |
| 2020 | Efficient MDI Adaptation for n-Gram Language ModelsabstractThis paper presents an efficient algorithm for n-gram language model adaptation under the minimum discrimination information (MDI) principle, where an out-of-domain language model is adapted to satisfy the constraints of marginal probabilities of the in-domain data.The challenge for MDI language model adaptation is its computational complexity.By taking advantage of the backoff structure of n-gram model and the idea of hierarchical training method, originally proposed for maximum entropy (ME) language models [1], we show that MDI adaptation can be computed in linear-time complexity to the inputs in each iteration.The complexity remains the same as ME models, although MDI is more general than ME.This makes MDI adaptation practical for large corpus and vocabulary.Experimental results confirm the scalability of our algorithm on very large datasets, while MDI adaptation gets slightly worse perplexity but better word error rate results compared to simple linear interpolation. Ruizhe Huang, Ke Li 0018, Ashish Arora, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 1 |
| 2015 | Making pattern queries bounded in big graphsabstractIt is cost-prohibitive to find matches Q(G) of a pattern query Q in a big graph G. We approach this by fetching a small subgraph GQof G such that Q(GQ) = Q(G). We show that many practical patterns are effectively bounded under access constraints A commonly found in real life, such that GQcan be identified in time determined by Q and A only, independent of the size |G| of G. This holds no matter whether pattern queries are localized (e.g., via subgraph isomorphism) or non-localized (graph simulation). We provide algorithms to decide whether a pattern Q is effectively bounded, and if so, to generate a query plan that computes Q(G) by accessing GQ, in time independent of |G|. When Q is not effectively bounded, we give an algorithm to extend access constraints and make Q bounded in G. Using real-life data, we experimentally verify the effectiveness of the approach, e.g., about 60% of queries are effectively bounded for subgraph isomorphism, and for such queries our approach outperforms the conventional methods by 4 orders of magnitude. Yang Cao 0012, Wenfei Fan, Jinpeng Huai, Ruizhe Huang |
ICDE | 4 |
| 2014 | Natural language question answering over RDF: a graph data driven approachabstractRDF question/answering (Q/A) allows users to ask questions in natural languages over a knowledge base represented by RDF. To answer a national language question, the existing work takes a two-stage approach: question understanding and query evaluation. Their focus is on question understanding to deal with the disambiguation of the natural language phrases. The most common technique is the joint disambiguation, which has the exponential search space. In this paper, we propose a systematic framework to answer natural language questions over RDF repository (RDF Q/A) from a graph data-driven perspective. We propose a semantic query graph to model the query intention in the natural language question in a structural way, based on which, RDF Q/A is reduced to subgraph matching problem. More importantly, we resolve the ambiguity of natural language questions at the time when matches of query are found. The cost of disambiguation is saved if there are no matching found. We compare our method with some state-of-the-art RDF Q/A systems in the benchmark dataset. Extensive experiments confirm that our method not only improves the precision but also speeds up query performance greatly. Lei Zou 0001, Ruizhe Huang, Haixun Wang, Jeffrey Xu Yu, Wenqiang He, Dongyan Zhao 0001 |
SIGMOD Conference | 2 |
| 2014 | gStore: a graph-based SPARQL query engine
Lei Zou 0001, M. Tamer Özsu, Lei Chen 0002, Xuchuan Shen, Ruizhe Huang, Dongyan Zhao 0001 |
VLDB J. | 5 |
| 2013 | Natural language question answering over RDF dataabstractAs more and more RDF data becomes available, such as DBpedia, Yago and Freebase, it is desired to provide users with simple interfaces to access the datasets. Although the SPARQL query language is a standard way to query RDF data, it remains tedious and difficult even for expert users because of the formality of the language and the complexity of the underlying schema of RDF data. An ideal system should allow users to express queries in their own languages. In this work, we propose a methodology to translate natural language questions into SPARQL queries, which can be answered by existing RDF engines and fulfill users? information need. Ruizhe Huang, Lei Zou 0001 |
SIGMOD Conference | 1 |