Zhengyu Zhou

dblp:69/7824 · DBLP profile ↗
← Back
14ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0002-8546-5282ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Learning theory · 36% Trustworthy machine learning · 32% Generative modeling · 10%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 24 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization
1.722026
Rademacher Complexity for Distributionally Robust Learning · AAAI 2026
Sample Complexity for Distributionally Robust Learning under chi-square divergence · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › sampling
adaptive sampling
1.012026
On the Robustness of Bandit Multiple Testing · AAAI 2026
Machine learning › Reinforcement learning
bandit
1.012026
On the Robustness of Bandit Multiple Testing · AAAI 2026
Machine learning › Trustworthy machine learning › risk control
false discovery rate control
1.012026
On the Robustness of Bandit Multiple Testing · AAAI 2026
Machine learning › Learning theory
generalization bounds
1.012026
Rademacher Complexity for Distributionally Robust Learning · AAAI 2026
Machine learning › Learning theory › hypothesis testing
multiple hypothesis testing
1.012026
On the Robustness of Bandit Multiple Testing · AAAI 2026
Machine learning › Learning theory › generalization bounds
rademacher complexity
1.012026
Rademacher Complexity for Distributionally Robust Learning · AAAI 2026
Machine learning › Generative modeling › normalizing flow
continuous normalizing flow
0.912025
An Error Analysis of Flow Matching for Deep Generative Modeling · ICML 2025
Machine learning › Generative modeling
flow matching
0.912025
An Error Analysis of Flow Matching for Deep Generative Modeling · ICML 2025
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.812024
A Boosting-Type Convergence Result for AdaBoost.MH with Factorized Multi-Class Classifiers · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
certified robustness
0.812024
DRF: Improving Certified Robustness via Distributional Robustness Framework · AAAI 2024
Machine learning › Optimization for machine learning
convergence analysis
0.812024
A Boosting-Type Convergence Result for AdaBoost.MH with Factorized Multi-Class Classifiers · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness › certified robustness
randomized smoothing
0.812024
DRF: Improving Certified Robustness via Distributional Robustness Framework · AAAI 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
DRF: Improving Certified Robustness via Distributional Robustness Framework · AAAI 2024
Information theory › hypothesis testing
goodness-of-fit testing
0.812024
Sequential Kernel Goodness-of-fit Testing · ICML 2024
Information theory
hypothesis testing
0.812024
Sequential Kernel Goodness-of-fit Testing · ICML 2024
Information theory › statistical inference › sequential analysis
sequential detection
0.812024
Sequential Kernel Goodness-of-fit Testing · ICML 2024
Information theory › hypothesis testing
testing by betting
0.812024
Sequential Kernel Goodness-of-fit Testing · ICML 2024
Machine learning › Learning theory
sample complexity
0.712023
Sample Complexity for Distributionally Robust Learning under chi-square divergence · J. Mach. Learn. Res. 2023
Machine learning › Learning theory
statistical learning theory
0.712023
Sample Complexity for Distributionally Robust Learning under chi-square divergence · J. Mach. Learn. Res. 2023
Machine learning › Learning theory › sample complexity
VC dimension bounds
0.712023
Sample Complexity for Distributionally Robust Learning under chi-square divergence · J. Mach. Learn. Res. 2023
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.212024
DRF: Improving Certified Robustness via Distributional Robustness Framework · AAAI 2024
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.212024
DRF: Improving Certified Robustness via Distributional Robustness Framework · AAAI 2024
Machine learning › Learning theory › classification
multiclass classification
0.212024
A Boosting-Type Convergence Result for AdaBoost.MH with Factorized Multi-Class Classifiers · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

sample complexity analysis · 1.0huber contamination model · 1.0cressie-read divergence · 1.0lipschitz condition analysis · 0.9least squares regression · 0.9type-i error control · 0.8sequential testing · 0.8randomized smoothing · 0.8kernel methods · 0.8factorized base classifiers · 0.8distributional robustness · 0.8convergence proof · 0.8adversarial training · 0.8
YearPublicationVenuePosition
2026 Rademacher Complexity for Distributionally Robust Learning
abstract
The goal of distributionally robust learning is to learn models capable of performing well against distributional shifts, such as latent heterogeneous subpopulations, unknown covariate shifts, or unmodeled temporal effects. Recently, Duchi and Namkoong (2021) have proven an upper bound for the excess risk of distributionally robust learning through the lens of covering number argument. However, there are situations where the covering argument fails. This motivates us to study the generalization bound through the lens of Rademacher complexity. More specifically, we consider the Cressie-Read divergence, f sub k of t is proportional to t to the k minus one. Our theoretical results indicate that the excess risk is of the order big O sub P of n to the negative one over two k star, where k star equals k over k minus one. The decay rate of the excess risk increases with increasing k. As illustrative examples, we consider three learning settings: 1) linear classifier; 2) Gaussian reproducing kernel Hilbert space; 3) one-hidden-layer networks. The empirical results validate our theoretical findings.
Zhengyu Zhou, Weiwei Liu 0003
AAAI1
2026 On the Robustness of Bandit Multiple Testing
abstract
Bandit multiple hypothesis testing has broad applications in biological sciences, clinical testing for drug discovery, and online A/B/n testing. The framework utilizes an adaptive sampling strategy for multiple testing which aims to maximize statistical power while ensuring anytime false discovery rate control. This paper proposes a robust approach for bandit multiple testing, allowing for at most an epsilon fraction of arbitrary distribution corruption, as in Huber’s contamination model. Specifically, we introduce two adaptive sampling strategies designed to minimize the number of samples required to exceed a target true positive rate, while providing anytime control over the false discovery rate. We analyze the sample complexity of our proposed methods and perform numerical simulations to demonstrate their efficiency and robustness. Furthermore, we extend our methods to address scenarios where distributions have infinite variance and situations involving multiple agents collaborating on the same bandit task.
Zhengyu Zhou, Weiwei Liu 0003
AAAI1
2025 An Error Analysis of Flow Matching for Deep Generative Modeling
abstract
Continuous Normalizing Flows (CNFs) have proven to be a highly efficient technique for generative modeling of complex data since the introduction of Flow Matching (FM). The core of FM is to learn the constructed velocity fields of CNFs through deep least squares regression. Despite its empirical effectiveness, theoretical investigations of FM remain limited. In this paper, we present the first end-to-end error analysis of CNFs built upon FM. Our analysis shows that for general target distributions with bounded support, the generated distribution of FM is guaranteed to converge to the target distribution in the sense of the Wasserstein-2 distance. Furthermore, the convergence rate is significantly improved under an additional mild Lipschitz condition of the target score function.
Zhengyu Zhou, Weiwei Liu 0003
ICML1
2024 DRF: Improving Certified Robustness via Distributional Robustness Framework
abstract
Randomized smoothing (RS) has provided state-of-the-art (SOTA) certified robustness against adversarial perturbations for large neural networks. Among studies in this field, methods based on adversarial training (AT) achieve remarkably robust performance by applying adversarial examples to construct the smoothed classifier. These AT-based RS methods typically seek a pointwise adversary that generates the worst-case adversarial examples by perturbing each input independently. However, there are unexplored benefits to considering such adversarial robustness across the entire data distribution. To this end, we provide a novel framework called DRF, which connects AT-based RS methods with distributional robustness (DR), and show that these methods are special cases of their counterparts in our framework. Due to the advantages conferred by DR, our framework can control the trade-off between the clean accuracy and certified robustness of smoothed classifiers to a significant extent. Our experiments demonstrate that DRF can substantially improve the certified robustness of AT-based RS.
Zhengyu Zhou, Weiwei Liu 0003
AAAI2
2024 Sequential Kernel Goodness-of-fit Testing
abstract
Goodness-of-fit testing, a classical statistical tool, has been extensively explored in the batch setting, where the sample size is predetermined. However, practitioners often prefer methods that adapt to the complexity of a problem rather than fixing the sample size beforehand. Classical batch tests are generally unsuitable for streaming data, as valid inference after data peeking requires multiple testing corrections, resulting in reduced statistical power. To address this issue, we delve into the design of consistent sequential goodness-of-fit tests. Following the principle of testing by betting, we reframe this task as selecting a sequence of payoff functions that maximize the wealth of a fictitious bettor, betting against the null in a repeated game. We conduct experiments to demonstrate the adaptability of our sequential test across varying difficulty levels of problems while maintaining control over type-I errors.
Zhengyu Zhou, Weiwei Liu 0003
ICML1
2024 A Boosting-Type Convergence Result for AdaBoost.MH with Factorized Multi-Class Classifiers
abstract
AdaBoost is a well-known algorithm in boosting. Schapire and Singer propose, an extension of AdaBoost, named AdaBoost.MH, for multi-class classification problems. Kégl shows empirically that AdaBoost.MH works better when the classical one-against-all base classifiers are replaced by factorized base classifiers containing a binary classifier and a vote (or code) vector. However, the factorization makes it much more difficult to provide a convergence result for the factorized version of AdaBoost.MH. Then, Kégl raises an open problem in COLT 2014 to look for a convergence result for the factorized AdaBoost.MH. In this work, we resolve this open problem by presenting a convergence result for AdaBoost.MH with factorized multi-class classifiers.
Xin Zou 0002, Zhengyu Zhou, Weiwei Liu 0003
NeurIPS2
2023 Sample Complexity for Distributionally Robust Learning under chi-square divergence
abstract
This paper investigates the sample complexity of learning a distributionally robust predictor under a particular distributional shift based on $\chi^2$-divergence, which is well known for its computational feasibility and statistical properties. We demonstrate that any hypothesis class $\mathcal{H}$ with finite VC dimension is distributionally robustly learnable. Moreover, we show that when the perturbation size is smaller than a constant, finite VC dimension is also necessary for distributionally robust learning by deriving a lower bound of sample complexity in terms of VC dimension.
Zhengyu Zhou, Wei Liu 0005
J. Mach. Learn. Res.1
2021 Using Paralinguistic Information to Disambiguate User Intentions for Distinguishing Phrase Structure and Sarcasm in Spoken Dialog Systems
abstract
This paper aims at utilizing paralinguistic information usually hidden in speech signals, such as pitch, short pause and sarcasm, to disambiguate user intention not easily distinguishable from speech recognition and natural language understanding results provided by a state-of-the-art spoken dialog system (SDS). We propose two methods to address the ambiguities in understanding name entities and sentence structures based on relevant speech cues and nuances. We also propose an approach to capturing sarcasm in speech and generating sarcasm-sensitive responses using an end-to-end neural network. An SDS prototype that directly feeds signal information into the understanding and response generation components has also been developed to support the three proposed applications. We have achieved encouraging experimental results in this initial study, demonstrating the potential of this new research direction.
Zhengyu Zhou, In Gyu Choi, Yongliang He, Vikas Yadav
SLT1
2019 A Neural Network Based Ranking Framework to Improve ASR with NLU Related Knowledge Deployed
abstract
This work proposes a new neural network framework to simultaneously rank multiple hypotheses generated by one or more automatic speech recognition (ASR) engines for a speech utterance. Features fed in the framework not only include those calculated from the ASR information, but also involve natural language understanding (NLU) related features, such as trigger features capturing long-distance constraints between word/slot pairs and BLSTM features representing intent-sensitive sentence embedding. The framework predicts the ranking result of the input hypotheses, outputting the top-ranked hypothesis as the new ASR result together with its slot filling and intention detection results. We conduct the experiments on an in-car infotainment corpus and the ATIS (Airline Travel Information Systems) corpus, for which hypotheses are generated by different types of engines and a single engine, respectively. The experimental results achieved are encouraging on both data corpora (e.g., 21.9% relative reduction in word error rate over state-of-the-art Google cloud ASR performance on the ATIS testing data), proving the effectiveness of the proposed ranking framework.
Zhengyu Zhou, Xuchen Song, Rami Botros
ICASSP1
2008 Recasting the discriminative n-gram model as a pseudo-conventional n-gram model for LVCSR
abstract
Discriminative n-gram language modeling has been used to re-rank candidate recognition hypotheses for performance improvements in large vocabulary continuous speech recognition (LVCSR). Discriminative n-gram modeling is defined in a linear framework. This work demonstrates that the linear discriminative n-gram model can be recast as a pseudo-conventional n-gram model if the order of the discriminative n-gram model is no higher than the order of the n-gram model in the baseline recognizer. Thus the power of discriminative n-gram model can be captured by mature n-gram related techniques such as single-pass n-gram decoding or lattice rescoring. This work utilizes the pseudo-conventional n-gram model to rescore the recognition lattices that are generated during decoding. Compared to the discriminative N-best re-ranking, this process of discriminative lattice rescoring (DLR) has two positive advantages: (1) Those discriminatively top-ranked utterance hypotheses within the lattice search spaces can be efficiently identified by the A* algorithm; (2) The rescored lattices can be further enhanced with other post-processing techniques to achieve cumulative improvement conveniently. Experiments with Mandarin LVCSR show that DLR improves efficiency - the computation time for 1000-best re-ranking is reduced by more than three-fold. The discriminatively rescored lattices are further processed by re-ranking with word-based mutual information (MI). While the DLR achieves around 15% relative character error rate (CER) reductions over the recognizer baseline, the MI based re-ranking further brings 5% relative CER reductions over the DLR performances.
Zhengyu Zhou, Helen M. Meng
ICASSP1
2007 Complementarity and redundancy in multimodal user inputs with speech and pen gestures
abstract
We present a comparative analysis of multi-modal user inputs with speech and pen gestures, together with their semantically equivalent uni-modal (speech only) counterparts. The multimodal interactions are derived from a corpus collected with a Pocket PC emulator in the context of navigation around Beijing. We devise a cross-modality integration methodology that interprets a multi-modal input and paraphrases it as a semantically equivalent, uni-modal input. Thus we generate parallel multi-modal (MM) and uni-modal (UM) corpora for comparative study. Empirical analysis based on class trigram perplexities shows two categories of data: (PPMM = PPUM) and (PPMM < PPUM). The former involves complementarity across modalities in expressing the user’s intent, including occurrences of ellipses. The latter involves redundancy, which will be useful for handling recognition errors by exploring mutual reinforcements. We present explanatory examples of data in these two categories. Index terms: multi-modal input, spoken input, pen gesture, joint interpretation, human-computer interaction, perplexity 1.
Pui-Yu Hui, Zhengyu Zhou, Helen M. Meng
INTERSPEECH2
2006 A Comparative Study of Discriminative Methods for Reranking LVCSR N-Best Hypotheses in Domain Adaptation and Generalization
abstract
This paper is an empirical study on the performance of different discriminative approaches to reranking the N-best hypotheses output from a large vocabulary continuous speech recognizer (LVCSR). Four algorithms, namely perceptron, boosting, ranking support vector machine (SVM) and minimum sample risk (MSR), are compared in terms of domain adaptation, generalization and time efficiency. In our experiments on Mandarin dictation speech, we found that for domain adaptation, perceptron performs the best; for generalization, boosting performs the best. The best result on a domain-specific test set is achieved by the perceptron algorithm. A relative character error rate (CER) reduction of 11% over the baseline was obtained. The best result on a general test set is 3.4% CER reduction over the baseline, achieved by the boosting algorithm.
Zhengyu Zhou, Jianfeng Gao 0001, Frank K. Soong, Helen M. Meng
ICASSP (1)1
2006 A multi-pass error detection and correction framework for Mandarin LVCSR
abstract
We previously proposed a multi-pass framework for Large Vocabulary Continuous Speech Recognition (LVCSR). The objective of this framework is to apply sophisticated linguistic models for recognition, while maintaining a balance between complexity and efficiency. The framework is composed of three passes: initial recognition, error detection and error correction. This paper presents and evaluates a prototype of the multi-pass framework based on Mandarin dictation. In this prototype, the first pass recognizes speech with a well-trained state-of-the-art recognizer incorporating an efficient language model; the second pass detects recognition errors by a new three-step error detection procedure; and the third pass corrects errors detected in those lightly erroneous utterances by a novel error correction approach. The error correction algorithm corrects recognition errors by first creating candidate lists for errors, and then re-ranking the candidates with a combined model of mutual information and trigram. Mandarin dictation experiments show a relative reduction of 4 % in character error rate (CER) over the initial recognition performance based on those light erroneous utterances detected. Index Terms: speech recognition, error detection & correction 1.
Zhengyu Zhou, Helen M. Meng, Wai Kit Lo
INTERSPEECH1
2004 A two-level schema for detecting recognition errors
abstract
This paper proposes a two-level schema for the automatic detection of possible errors in speech recognition hypotheses. Given the recognition hypothesis of an utterance, the first level in our schema applies an utterance classifier (UC) to decide if the hypothesis is error-free or erroneous. In the latter case, the utterance is passed on to the second level in our schema for further processing. A word classifier (WC) is applied to each of the word hypotheses in the utterance to decide whether or not it is a misrecognition. Hence the two-level schema can locate error-containing regions in the recognition hypotheses. These are the target regions to which we can apply more sophisticated and expensive language models for error correction as a next step. We have developed UC and WC based on Support Vector Machines (SVM). Experiments on Mandarin Chinese speech recognition using the Speech-Lab-In-A-Box corpora showed that the UC has a detection error rate of 16.5 % for misrecognized utterances; the WC has a detection error rate of 19.8 % for erroneous word hypotheses; and the overall two-level schema can catch 44.5 % of the erroneous word hypotheses. 1.
Zhengyu Zhou, Helen M. Meng
INTERSPEECH1