VLDB 2026 Research / reviewers in the wild / expert
Chengyu Dong
dblp:14/3155
· DBLP profile ↗
15ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discard-Based Garbage Collection for Distributed Log-Structured Storage Systems in ByteDance
Runhua Bian, Jianong Zhong, Jiahao Gu, Zhihong Guo, Fenghao Zhang, Jiangkun Zhao, Yangming Chen, Ruwen Fan, Haijia Shen, Chengyu Dong, Yao Wang 0022, Jiwu Shu, Youyou Lu |
FAST | 16 |
| 2024 | Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text ClassificationabstractFor extremely weak-supervised text classification, pioneer research generates pseudo labels by mining texts similar to the class names from the raw corpus, which may end up with very limited or even no samples for the minority classes.Recent works have started to generate the relevant texts by prompting LLMs using the class names or definitions; however, there is a high risk that LLMs cannot generate indistribution (i.e., similar to the corpus where the text classifier will be applied) data, leading to ungeneralizable classifiers.In this paper, we combine the advantages of these two approaches and propose to bridge the gap via a novel framework, text grafting, which aims to obtain clean and near-distribution weak supervision for minority classes.Specifically, we first use LLM-based logits to mine masked templates from the raw corpus, which have a high potential for data synthesis into the target minority class.Then, the templates are filled by state-of-the-art LLMs to synthesize neardistribution texts falling into minority classes.Text grafting shows significant improvement over direct mining or synthesis on minority classes.We also use analysis and case studies to comprehend the property of text grafting. Letian Peng, Yi Gu 0002, Chengyu Dong, Zihan Wang 0001, Jingbo Shang |
EMNLP | 3 |
| 2024 | Fast-ELECTRA for Efficient Pre-trainingabstractELECTRA pre-trains language models by detecting tokens in a sequence that have been replaced by an auxiliary model. Although ELECTRA offers a significant boost in efficiency, its potential is constrained by the training cost brought by the auxiliary model. Notably, this model, which is jointly trained with the main model, only serves to assist the training of the main model and is discarded post-training. This results in a substantial amount of training cost being expended in vain. To mitigate this issue, we propose Fast-ELECTRA, which leverages an existing language model as the auxiliary model. To construct a learning curriculum for the main model, we smooth its output distribution via temperature scaling following a descending schedule. Our approach rivals the performance of state-of-the-art ELECTRA-style pre-training methods, while significantly eliminating the computation and memory cost brought by the joint training of the auxiliary model. Our method also reduces the sensitivity to hyper-parameters and enhances the pre-training stability. Chengyu Dong, Hao Cheng 0002, Jingbo Shang, Jianfeng Gao 0001, Xiaodong Liu 0003 |
ICLR | 1 |
| 2024 | Toward Student-oriented Teacher Network Training for Knowledge DistillationabstractHow to conduct teacher training for knowledge distillation is still an open problem. It has been widely observed that a best-performing teacher does not necessarily yield the best-performing student, suggesting a fundamental discrepancy between the current teacher training practice and the ideal teacher training strategy. To fill this gap, we explore the feasibility of training a teacher that is oriented toward student performance with empirical risk minimization (ERM). Our analyses are inspired by the recent findings that the effectiveness of knowledge distillation hinges on the teacher’s capability to approximate the true label distribution of training inputs. We theoretically establish that ERM minimizer can approximate the true label distribution of training data as long as the feature extractor of the learner network is Lipschitz continuous and is robust to feature transformations. In light of our theory, we propose a teacher training method SoTeacher which incorporates Lipschitz regularization and consistency regularization into ERM. Experiments on benchmark datasets using various knowledge distillation algorithms and teacher-student pairs confirm that SoTeacher can improve student accuracy consistently. Chengyu Dong, Jingbo Shang |
ICLR | 1 |
| 2023 | Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text ClassificationabstractRecent advances in weakly supervised text classification mostly focus on designing sophisticated methods to turn high-level human heuristics into quality pseudo-labels.In this paper, we revisit the seed matching-based method, which is arguably the simplest way to generate pseudo-labels, and show that its power was greatly underestimated.We show that the limited performance of seed matching is largely due to the label bias injected by the simple seed-match rule, which prevents the classifier from learning reliable confidence for selecting high-quality pseudo-labels.Interestingly, simply deleting the seed words present in the matched input texts can mitigate the label bias and help learn better confidence.Subsequently, the performance achieved by seed matching can be improved significantly, making it on par with or even better than the state-of-theart.Furthermore, to handle the case when the seed words are not made known, we propose to simply delete the word tokens in the input text randomly with a high deletion ratio.Remarkably, seed matching equipped with this random deletion method can often achieve even better performance than that with seed deletion.We refer to our method as SimSeed, which is publicly available 1 . Chengyu Dong, Zihan Wang 0001, Jingbo Shang |
EMNLP | 1 |
| 2023 | Learning Concise and Descriptive Attributes for Visual RecognitionabstractRecent advances in foundation models present new opportunities for interpretable visual recognition – one can first query Large Language Models (LLMs) to obtain a set of attributes that describe each class, then apply vision-language models to classify images via these attributes. Pioneering work shows that querying thousands of attributes can achieve performance competitive with image features. However, our further investigation on 8 datasets reveals that LLM-generated attributes in a large quantity perform almost the same as random words. This surprising finding suggests that significant noise may be present in these attributes. We hypothesize that there exist subsets of attributes that can maintain the classification performance with much smaller sizes, and propose a novel learning-to-search method to discover those concise sets of attributes. As a result, on the CUB dataset, our method achieves performance close to that of massive LLM-generated attributes (e.g., 10k attributes for CUB), yet using only 32 attributes in total to distinguish 200 bird species. Furthermore, our new paradigm demonstrates several additional benefits: higher interpretability and interactivity for humans, and the ability to summarize knowledge for a recognition task. An Yan 0003, Yu Wang 0170, Yiwu Zhong, Chengyu Dong, Zexue He, William Yang Wang, Jingbo Shang, Julian J. McAuley |
ICCV | 4 |
| 2023 | Understand and Modularize Generator Optimization in ELECTRA-style PretrainingabstractDespite the effectiveness of ELECTRA-style pre-training, their performance is dependent on the careful selection of the model size for the auxiliary generator, leading to high trial-and-error costs. In this paper, we present the first systematic study of this problem. Our theoretical investigation highlights the importance of controlling the generator capacity in ELECTRA-style training. Meanwhile, we found it is not handled properly in the original ELECTRA design, leading to the sensitivity issue. Specifically, since adaptive optimizers like Adam will cripple the weighing of individual losses in the joint optimization, the original design fails to control the generator training effectively. To regain control over the generator, we modularize the generator optimization by decoupling the generator optimizer and discriminator optimizer completely, instead of simply relying on the weighted objective combination. Our simple technique reduced the sensitivity of ELECTRA training significantly and obtains considerable performance gain compared to the original design. Chengyu Dong, Hao Cheng 0002, Jingbo Shang, Jianfeng Gao 0001, Xiaodong Liu 0003 |
ICML | 1 |
| 2023 | Bridging Discrete and Backpropagation: Straight-Through and BeyondabstractBackpropagation, the cornerstone of deep learning, is limited to computing gradients for continuous variables. This limitation poses challenges for problems involving discrete latent variables. To address this issue, we propose a novel approach to approximate the gradient of parameters involved in generating discrete latent variables. First, we examine the widely used Straight-Through (ST) heuristic and demonstrate that it works as a first-order approximation of the gradient. Guided by our findings, we propose ReinMax, which achieves second-order accuracy by integrating Heun’s method, a second-order numerical method for solving ODEs. ReinMax does not require Hessian or other second-order derivatives, thus having negligible computation overheads. Extensive experimental results on various tasks demonstrate the superiority of ReinMax over the state of the art. Chengyu Dong, Xiaodong Liu 0003, Bin Yu 0001, Jianfeng Gao 0001 |
NeurIPS | 2 |
| 2023 | Physics-Informed Data Denoising for Real-Life Sensing SystemsabstractSensors measuring real-life physical processes are ubiquitous in today's interconnected world. These sensors inherently bear noise that often adversely affects the performance and reliability of the systems they support. Classic filtering approaches introduce strong assumption on the time or frequency characteristics of sensory measurements, while learning-based denoising approaches typically rely on using ground truth clean data to train a denoising model, which is often challenging or prohibitive to obtain for many real-world applications. We observe that in many scenarios, the relationships between different sensor measurements (e.g., location and acceleration) are analytically described by laws of physics (e.g., second-order differential equation). By incorporating such physics constraints, we can guide the denoising process to improve performance even in the absence of ground truth data. In light of this, we design a physics-informed denoising model that leverages the inherent algebraic relationships between different measurements governed by the underlying physics. By obviating the need for ground truth clean data, our method offers a practical denoising solution for real-world applications. We conducted experiments in various domains, including inertial navigation, CO2 monitoring, and HVAC control, and achieved state-of-the-art performance compared with existing denoising methods. Our method can denoise data in real time (4ms for a sequence of 1s) for low-cost noisy sensors and produces results that closely align with those from high-precision, high-cost alternatives, leading to an efficient, cost-effective approach for more accurate sensor-based systems. Xiyuan Zhang 0001, Xiaohan Fu, Diyan Teng, Chengyu Dong, Keerthivasan Vijayakumar, Jiayun Zhang, Ranak Roy Chowdhury, Junsheng Han, Dezhi Hong, Rashmi Kulkarni, Jingbo Shang, Rajesh K. Gupta 0001 |
SenSys | 4 |
| 2022 | Label Noise in Adversarial Training: A Novel Perspective to Study Robust OverfittingabstractWe show that label noise exists in adversarial training. Such label noise is due to the mismatch between the true label distribution of adversarial examples and the label inherited from clean examples – the true label distribution is distorted by the adversarial perturbation, but is neglected by the common practice that inherits labels from clean examples. Recognizing label noise sheds insights on the prevalence of robust overfitting in adversarial training, and explains its intriguing dependence on perturbation radius and data quality. Also, our label noise perspective aligns well with our observations of the epoch-wise double descent in adversarial training. Guided by our analyses, we proposed a method to automatically calibrate the label to address the label noise and robust overfitting. Our method achieves consistent performance improvements across various models and datasets without introducing new hyper-parameters or additional tuning. Chengyu Dong, Jingbo Shang |
NeurIPS | 1 |
| 2021 | "Average" Approximates "First Principal Component"? An Empirical Analysis on Representations from Neural Language ModelsabstractContextualized representations based on neural language models have furthered the state of the art in various NLP tasks.Despite its great success, the nature of such representations remains a mystery.In this paper, we present an empirical property of these representations-"average" ≈ "first principal component".Specifically, experiments show that the average of these representations shares almost the same direction as the first principal component of the matrix whose columns are these representations.We believe this explains why the average representation is always a simple yet strong baseline.Our further examinations show that this property also holds in more challenging scenarios, for example, when the representations are from a model right after its random initialization.Therefore, we conjecture that this property is intrinsic to the distribution of representations and not necessarily related to the input structure.We realize that these representations empirically follow a normal distribution for each dimension, and by assuming this is true, we demonstrate that the empirical property can be in fact derived mathematically. Zihan Wang 0001, Chengyu Dong, Jingbo Shang |
EMNLP (1) | 2 |
| 2020 | Towards Adaptive Residual Network Training: A Neural-ODE PerspectiveabstractIn pursuit of resource-economical machine learning, attempts have been made to dynamically adjust computation workloads in different training stages, i.e., starting with a shallow network and gradually increasing the model depth (and computation workloads) during training. However, there is neither guarantee nor guidance on designing such network grow, due to the lack of its theoretical underpinnings. In this work, to explore the theory behind, we conduct theoretical analyses from an ordinary differential equation perspective. Specifically, we illustrate the dynamics of network growth and propose a novel performance measure specific to the depth increase. Illuminated by our analyses, we move towards theoretically sound growing operations and schedulers, giving rise to an adaptive training algorithm for residual networks, LipGrow, which automatically increases network depth thus accelerates training. In our experiments, it achieves comparable performance while reducing ∼ 50% of training time. Chengyu Dong, Zichao Li 0002, Jingbo Shang |
ICML | 1 |
| 2008 | Analysis of subspace within-class covariance normalization for SVM-based speaker verificationabstractNuisance attribute projection (NAP) and within-class covariance normalization (WCCN) are two effective techniques for intersession variability compensation in SVM based speaker verification systems. However, by normalizing or removing the nuisance subspace containing the session variability can not guarantee to enlarge the distance between speakers. In this paper, we investigated the probability of using linear discriminant analysis (LDA) for discriminative training. To cope with the small sample size problem which prevents us from using LDA directly, we adapted the subspace LDA approach, which first projects the whole feature space into a relatively low dimensional subspace by PCA, and then performs LDA in the subspace. By some modification, the subspace LDA can be degenerated into a kind of WCCN approach, which we called subspace WCCN. Experiments on NIST SRE tasks showed that, the subspace WCCN outperformed the conventional direct WCCN, especially in low dimensional feature space. Liang Lu 0002, Xianyu Zhao, Jian Zhao 0001, Chengyu Dong, Haila Wang |
INTERSPEECH | 5 |
| 2007 | Score Normalization Technique for Text-Prompted Speaker Verification with Chinese Digits
Chengyu Dong, Haila Wang |
ICIC (2) | 3 |
| 2006 | A Boosting Approach for Utterance Verification
Chengyu Dong, Dezhi Huang, Jun Guo 0002, Haila Wang |
ICIC (2) | 1 |