Wenyong Huang

dblp:248/7868 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 44% Language models and text generation · 22% Speech recognition and synthesis · 17%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 50% Memory systems · 50%
Network and information security
1 paper
Systems and software security · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Systems and software security › isolation
software fault isolation
0.912025
Segue & ColorGuard: Optimizing SFI Performance and Scalability on Modern Architectures · ASPLOS (1) 2025
Processor architecture and microarchitecture
instruction set architecture
0.912025
Segue & ColorGuard: Optimizing SFI Performance and Scalability on Modern Architectures · ASPLOS (1) 2025
Memory systems › memory architecture
tagged memory
0.912025
Segue & ColorGuard: Optimizing SFI Performance and Scalability on Modern Architectures · ASPLOS (1) 2025
Natural language and speech › Language models and text generation
alignment
0.812024
Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis · ICLR 2024
Machine learning › Trustworthy machine learning › AI safety › content safety
harmful content mitigation
0.812024
Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis · ICLR 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.612022
SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training · ICLR 2022
Natural language and speech › Speech recognition and synthesis
speech pre-training
0.612022
SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training · ICLR 2022
Compilers and program optimization › program instrumentation
compiler instrumentation
0.312025
Segue & ColorGuard: Optimizing SFI Performance and Scalability on Modern Architectures · ASPLOS (1) 2025

Methods — techniques the papers use, named apart from their topics

x86-64 segmentation · 2.6memory protection keys · 2.6supervised fine-tuning · 0.8mistake analysis · 0.8perturbation-invariant representation learning · 0.6
YearPublicationVenuePosition
2025 Segue & ColorGuard: Optimizing SFI Performance and Scalability on Modern Architectures
abstract
Software-based fault isolation (SFI) enables in-process isolation through compiler instrumentation of memory accesses, and is a critical part of WebAssembly (Wasm). We present two optimizations that improve SFI performance and scalability: Segue uses x86-64 segmentation to reduce the cost of instrumentation on memory accesses, e.g., it eliminates 44.7% of Wasm's overhead on a Wasm-compatible subset of SPEC CPU 2006, and reduces overhead of Wasm-sandboxed font rendering in Firefox by 75%; ColorGuard leverages memory tagging (e.g., MPK), to enable up to a 15× increase in the number of Wasm instances that can run concurrently in a single address space, improving efficiency for high scale server-side workloads. We also explore the challenges of deploying these optimizations in three production toolchains: Wasm2c, WAMR and Wasmtime.
Shravan Narayan, Tal Garfinkel, Evan Johnson 0001, Zachary Yedidia, Yingchen Wang, Anjo Vahldiek-Oberwagner, Michael LeMay, Wenyong Huang, Xin Wang 0240, Mingqiu Sun, Dean M. Tullsen, Deian Stefan
ASPLOS (1)9
2024 Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
abstract
The rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Existing alignment methods usually direct LLMs toward the favorable outcomes by utilizing human-annotated, flawless instruction-response pairs. Conversely, this study proposes a novel alignment technique based on mistake analysis, which deliberately exposes LLMs to erroneous content to learn the reasons for mistakes and how to avoid them. In this case, mistakes are repurposed into valuable data for alignment, effectively helping to avoid the production of erroneous responses. Without external models or human annotations, our method leverages a model's intrinsic ability to discern undesirable mistakes and improves the safety of its generated responses. Experimental results reveal that our method outperforms existing alignment approaches in enhancing model safety while maintaining the overall utility.
Kai Chen 0023, Chunwei Wang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu 0004, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
ICLR9
2022 SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training
Wenyong Huang, Zhenhe Zhang, Yu Ting Yeung, Xin Jiang 0002, Qun Liu 0001
ICLR1
2022 CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and Diagnosis
abstract
Mispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems.End-to-end (e2e) approaches are becoming dominant in MDD.However an e2e MDD model usually requires entire speech utterances as input context, which leads to significant time latency especially for long paragraphs.We propose a streaming e2e MDD model called CoCA-MDD.We utilize conv-transformer structure to encode input speech in a streaming manner.A coupled cross-attention (CoCA) mechanism is proposed to integrate frame-level acoustic features with encoded reference linguistic features.CoCA also enables our model to perform mispronunciation classification with whole utterances.The proposed model allows system fusion between the streaming output and mispronunciation classification output for further performance enhancement.We evaluate CoCA-MDD on publicly available corpora.CoCA-MDD achieves F1 scores of 57.03% and 60.78% for streaming and fusion modes respectively on L2-ARCTIC.For phone-level pronunciation scoring, CoCA-MDD achieves 0.58 Pearson correlation coefficient (PCC) value on SpeechOcean762.
Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung, Baohua Xu, Yasheng Wang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
INTERSPEECH3
2020 Conv-Transformer Transducer: Low Latency, Low Frame Rate, Streamable End-to-End Speech Recognition
abstract
Transformer has achieved competitive performance against state-of-the-art end-to-end models in automatic speech recognition (ASR), and requires significantly less training time than RNN-based models.The original Transformer, with encoderdecoder architecture, is only suitable for offline ASR.It relies on an attention mechanism to learn alignments, and encodes input audio bidirectionally.The high computation cost of Transformer decoding also limits its use in production streaming systems.To make Transformer suitable for streaming ASR, we explore Transducer framework as a streamable way to learn alignments.For audio encoding, we apply unidirectional Transformer with interleaved convolution layers.The interleaved convolution layers are used for modeling future context which is important to performance.To reduce computation cost, we gradually downsample acoustic input, also with the interleaved convolution layers.Moreover, we limit the length of history context in self-attention to maintain constant computation cost for each decoding step.We show that this architecture, named Conv-Transformer Transducer, achieves competitive performance on LibriSpeech dataset (3.6% WER on test-clean) without external language models.The performance is comparable to previously published streamable Transformer Transducer and strong hybrid streaming ASR systems, and is achieved with smaller look-ahead window (140 ms), fewer parameters and lower frame rate.
Wenyong Huang, Wenchao Hu, Yu Ting Yeung, Xiao Chen 0012
INTERSPEECH1