VLDB 2026 Research / reviewers in the wild / expert
Shengyu Zhu 0001
dblp:131/6555
· DBLP profile ↗
22ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0001-9793-662XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speculative Safety-Aware DecodingabstractDespite extensive efforts to align Large Language Models (LLMs) with human values and safety rules, jailbreak attacks that exploit certain vulnerabilities continuously emerge, highlighting the need to strengthen existing LLMs with additional safety properties to defend against these attacks.However, tuning large models has become increasingly resourceintensive and may have difficulty ensuring consistent performance.We introduce Speculative Safety-Aware Decoding (SSD), a lightweight decoding-time approach that equips LLMs with the desired safety property while accelerating inference.We assume that there exists a small language model that possesses this desired property.SSD integrates speculative sampling during decoding and leverages the match ratio between the small and composite models to quantify jailbreak risks.This enables SSD to dynamically switch between decoding schemes to prioritize utility or safety, to handle the challenge of different model capacities.The output token is then sampled from a new distribution that combines the distributions of the original and the small models.Experimental results show that SSD successfully equips the large model with the desired safety property, and also allows the model to remain helpful to benign queries.Furthermore, SSD accelerates the inference time, thanks to the speculative sampling design. Xuekang Wang, Shengyu Zhu 0001, Xueqi Cheng 0001 |
EMNLP | 2 |
| 2024 | WToE: Learning When to Explore in Multiagent Reinforcement LearningabstractExisting multiagent exploration works focus on how to explore in the fully cooperative task, which is insufficient in the environment with nonstationarity induced by agent interactions. To tackle this issue, we propose When to Explore (WToE), a simple yet effective variational exploration method to learn WToE under nonstationary environments. WToE employs an interaction-oriented adaptive exploration mechanism to adapt to environmental changes. We first propose a novel graphical model that uses a latent random variable to model the step-level environmental change resulting from interaction effects. Leveraging this graphical model, we employ the supervised variational auto-encoder (VAE) framework to derive a short-term inferred policy from historical trajectories to deal with the nonstationarity. Finally, agents engage in exploration when the short-term inferred policy diverges from the current actor policy. The proposed approach theoretically guarantees the convergence of the Q -value function. In our experiments, we validate our exploration mechanism in grid examples, multiagent particle environments and the battle game of MAgent environments. The results demonstrate the superiority of WToE over multiple baselines and existing exploration methods, such as MAEXQ, NoisyNets, EITI, and PR2. Shaokang Dong, Hangyu Mao, Shangdong Yang, Shengyu Zhu 0001, Wenbin Li 0006, Jianye Hao, Yang Gao 0001 |
IEEE Trans. Cybern. | 4 |
| 2024 | On Low-Rank Directed Acyclic Graphs and Causal Structure LearningabstractDespite several advances in recent years, learning causal structures represented by directed acyclic graphs (DAGs) remains a challenging task in high-dimensional settings when the graphs to be learned are not sparse. In this article, we propose to exploit a low-rank assumption regarding the (weighted) adjacency matrix of a DAG causal model to help address this problem. We utilize existing low-rank techniques to adapt causal structure learning methods to take advantage of this assumption and establish several useful results relating interpretable graphical conditions to the low-rank assumption. Specifically, we show that the maximum rank is highly related to hubs, suggesting that scale-free (SF) networks, which are frequently encountered in practice, tend to be low rank. Our experiments demonstrate the utility of the low-rank adaptations for a variety of data models, especially with relatively large and dense graphs. Moreover, with a validation procedure, the adaptations maintain a superior or comparable performance even when graphs are not restricted to be low rank. Zhuangyan Fang, Shengyu Zhu 0001, Jiji Zhang, Zhitang Chen, Yangbo He |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Provably Invariant Learning without Domain InformationabstractTypical machine learning applications always assume the data follows independent and identically distributed (IID) assumptions. In contrast, this assumption is frequently violated in real-world circumstances, leading to the Out-of-Distribution (OOD) generalization problem and a major drop in model robustness. To mitigate this issue, the invariant learning technique is leveraged to distinguish between spurious features and invariant features among all input features and to train the model purely on the basis of the invariant features. Numerous invariant learning strategies imply that the training data should contain domain information. Such information includes the environment index or auxiliary information acquired from prior knowledge. However, acquiring these information is typically impossible in practice. In this study, we present TIVA for environment-independent invariance learning, which requires no environment-specific information in training data. We discover and prove that, given certain mild data conditions, it is possible to train an environment partitioning policy based on attributes that are independent of the targets and then conduct invariant risk minimization. We examine our method in comparison to other baseline methods, which demonstrate superior performance and excellent robustness under OOD, using multiple benchmarks. Xiaoyu Tan, Lin Yong, Shengyu Zhu 0001, Chao Qu, Xihe Qiu, Yinghui Xu 0001, Peng Cui 0001, Yuan Qi 0001 |
ICML | 3 |
| 2023 | Conditional counterfactual causal effect for individual attributionabstractIdentifying the causes of an event, also termed as causal attribution, is a commonly encountered task in many application problems. Available methods, mostly in Bayesian or causal inference literature, suffer from two main drawbacks: 1) cannot attribute for individuals, and 2) attributing one single cause at a time and cannot deal with the interaction effect among multiple causes. In this paper, based on our proposed new measurement, called conditional counterfactual causal effect (CCCE), we introduce an individual causal attribution method, which is able to utilize the individual observation as the evidence and consider common influence and interaction effect of multiple causes simultaneously. We discuss the identifiability of CCCE and also give the identification formulas under proper assumptions. Finally, we conduct experiments on simulated and real data to illustrate the effectiveness of CCCE and the results show that our proposed method outperforms significantly state-of-the-art methods. Lei Zhang 0006, Shengyu Zhu 0001, Zitong Lu, Zhenhua Dong, Chaoliang Zhang, Jun Xu 0001, Zhi Geng, Yangbo He |
UAI | 3 |
| 2023 | A Unified Framework for Layout Pattern Analysis With Deep Causal EstimationabstractThe decrease of feature size and the growing complexity of the fabrication process lead to more failures in manufacturing semiconductor devices. Therefore, identifying the root cause layout patterns of failures becomes increasingly crucial for yield improvement. In this article, a novel layout-aware diagnosis-based layout pattern analysis framework is proposed to identify the root cause efficiently. At the first stage of the framework, an encoder network trained using contrastive learning is used to extract representations of layout snippets that are invariant to trivial transformations, including shift, rotation, and mirroring, which are then clustered to form layout patterns. At the second stage, we model the causal relationship between any potential root cause layout patterns and the systematic defects by a structural causal model, which is then used to estimate the average causal effect (ACE) of candidate layout patterns on the systematic defect to identify the true root cause. Experimental results on real industrial cases demonstrate that our framework outperforms a commercial tool with higher accuracies and around$\times 8.4$speedup on average. Ran Chen 0001, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Bei Yu 0001, Pengyun Li, Yu Huang 0005, Jianye Hao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Out-of-distribution Generalization with Causal Invariant TransformationsabstractIn real-world applications, it is important and desirable to learn a model that performs well on out-of-distribution (OOD) data. Recently, causality has become a powerful tool to tackle the OOD generalization problem, with the idea resting on the causal mechanism that is invariant across domains of interest. To leverage the generally unknown causal mechanism, existing works assume a linear form of causal feature or require sufficiently many and diverse training domains, which are usually restrictive in practice. In this work, we obviate these assumptions and tackle the OOD problem without explicitly recovering the causal feature. Our approach is based on transformations that modify the non-causal feature but leave the causal part unchanged, which can be either obtained from prior knowledge or learned from the training data in the multi-domain scenario. Under the setting of invariant causal mechanism, we theoretically show that if all such transformations are available, then we can learn a minimax optimal model across the domains using only single domain data. Noticing that knowing a complete set of these causal invariant transformations may be impractical, we further show that it suffices to know only a subset of these transformations. Based on the theoretical findings, a regularized training procedure is proposed to improve the OOD generalization capability. Extensive experimental results on both synthetic and real datasets verify the effectiveness of the proposed algorithm, even with only a few causal invariant transformations. Ruoyu Wang 0016, Mingyang Yi, Zhitang Chen, Shengyu Zhu 0001 |
CVPR | 4 |
| 2022 | RCANet: Root Cause Analysis via Latent Variable Interaction Modeling for Yield ImprovementabstractIdentifying root causes of systematic defects is a crucial step in yield enhancement process of integrated circuit (IC) manufacturing. With increasing complexity of fabrication processes and decreasing sizes of pattern features, more systematic defects occur at advanced technology nodes, and traditional methods are unfeasible to directly identify failure causes, due to expensive time and labor costs. Root cause analysis (RCA) technology is thus studied to automatically identify common root causes in a short time. In this paper, we develop RCANet, an end-to-end unsupervised learning-based RCA framework, which analyses diagnosis reports of failing dies within a wafer and identifies both layout-aware and cell-internal root causes efficiently. Experimental results on designs with different technologies demonstrate that RCANet outperforms both a commercial tool and the state-of-the-art method. Xiaopeng Zhang 0009, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Evangeline F. Y. Young, Pengyun Li, Yu Huang 0005, Jianye Hao |
ITC | 4 |
| 2022 | ZIN: When and How to Learn Invariance Without Environment Partition?abstractIt is commonplace to encounter heterogeneous data, of which some aspects of the data distribution may vary but the underlying causal mechanisms remain constant. When data are divided into distinct environments according to the heterogeneity, recent invariant learning methods have proposed to learn robust and invariant models using this environment partition. It is hence tempting to utilize the inherent heterogeneity even when environment partition is not provided. Unfortunately, in this work, we show that learning invariant features under this circumstance is fundamentally impossible without further inductive biases or additional information. Then, we propose a framework to jointly learn environment partition and invariant representation, assisted by additional auxiliary information. We derive sufficient and necessary conditions for our framework to provably identify invariant features under a fairly general setting. Experimental results on both synthetic and real world datasets validate our analysis and demonstrate an improved performance of the proposed framework. Our findings also raise the need of making the role of inductive biases more explicit when learning invariant models without environment partition in future works. Codes are available at https://github.com/linyongver/ZIN_official . Shengyu Zhu 0001, Peng Cui 0001 |
NeurIPS | 2 |
| 2022 | Para-CFlows: $C^k$-universal diffeomorphism approximators as superior neural surrogatesabstractInvertible neural networks based on Coupling Flows (CFlows) have various applications such as image synthesis and data compression. The approximation universality for CFlows is of paramount importance to ensure the model expressiveness. In this paper, we prove that CFlows}can approximate any diffeomorphism in $C^k$-norm if its layers can approximate certain single-coordinate transforms. Specifically, we derive that a composition of affine coupling layers and invertible linear transforms achieves this universality. Furthermore, in parametric cases where the diffeomorphism depends on some extra parameters, we prove the corresponding approximation theorems for parametric coupling flows named Para-CFlows. In practice, we apply Para-CFlows as a neural surrogate model in contextual Bayesian optimization tasks, to demonstrate its superiority over other neural surrogate models in terms of optimization performance and gradient approximations. Junlong Lyu, Zhitang Chen, Chang Feng, Wenjing Cun, Shengyu Zhu 0001, Yanhui Geng, Zhijie Xu, Chen Yongwei |
NeurIPS | 5 |
| 2022 | Masked Gradient-Based Causal Structure LearningabstractThis paper studies the problem of learning causal structures from observational data. We reformulate the Structural Equation Model (SEM) with additive noises in a form parameterized by binary graph adjacency matrix and show that, if the original SEM is identifiable, then the binary adjacency matrix can be identified up to super-graphs of the true causal graph under mild conditions. We then utilize the reformulated SEM to develop a causal structure learning method that can be efficiently trained using gradient-based optimization, by leveraging a smooth characterization on acyclicity and the Gumbel-Softmax approach to approximate the binary adjacency matrix. It is found that the obtained entries are typically near zero or one and can be easily thresholded to identify the edges. We conduct experiments on synthetic and real datasets to validate the effectiveness of the proposed method, and show that it readily includes different smooth model functions and achieves a much improved performance on most datasets considered. Ignavier Ng, Shengyu Zhu 0001, Zhuangyan Fang, Haoyang Li 0002, Zhitang Chen, Jun Wang 0012 |
SDM | 2 |
| 2022 | Reframed GES with a neural conditional dependence measureabstractIn a nonparametric setting, the causal structure is often identifiable only up to Markov equivalence, and for the purpose of causal inference, it is useful to learn a graphical representation of the Markov equivalence class (MEC). In this paper, we revisit the Greedy Equivalence Search (GES) algorithm, which is widely cited as a score-based algorithm for learning the MEC of the underlying causal structure. We observe that in order to make the GES algorithm consistent in a nonparametric setting, it is not necessary to design a scoring metric that evaluates graphs. Instead, it suffices to plug in a consistent estimator of a measure of conditional dependence to guide the search. We therefore present a reframing of the GES algorithm, which is more flexible than the standard score-based version and readily lends itself to the nonparametric setting with a general measure of conditional dependence. In addition, we propose a neural conditional dependence (NCD) measure, which utilizes the expressive power of deep neural networks to characterize conditional independence in a nonparametric manner. We establish the optimality of the reframed GES algorithm under standard assumptions and the consistency of using our NCD estimator to decide conditional independence. Together these results justify the proposed approach. Experimental results demonstrate the effectiveness of our method in causal discovery, as well as the advantages of using our NCD measure over kernel-based measures. Xinwei Shen 0002, Shengyu Zhu 0001, Jiji Zhang, Shoubo Hu, Zhitang Chen |
UAI | 2 |
| 2022 | A local method for identifying causal relations under Markov equivalence
Zhuangyan Fang, Zhi Geng, Shengyu Zhu 0001, Yangbo He |
Artif. Intell. | 4 |
| 2021 | A Unified Framework for Layout Pattern Analysis with Deep Causal EstimationabstractThe decrease of feature size and the growing complexity of the fabrication process lead to more failures in manufacturing semiconductor devices. Therefore, identifying the root cause layout patterns of failures becomes increasingly crucial for yield improvement. In this paper, a novel layout-aware diagnosis-based layout pattern analysis framework is proposed to identify the root cause efficiently. At the first stage of the framework, an encoder network trained using contrastive learning is used to extract representations of layout snippets that are invariant to trivial transformations including shift, rotation, and mirroring, which are then clustered to form layout patterns. At the second stage, we model the causal relationship between any potential root cause layout patterns and the systematic defects by a structural causal model, which is then used to estimate the Average Causal Effect (ACE) of candidate layout patterns on the systematic defect to identify the true root cause. Experimental results on real industrial cases demonstrate that our framework outperforms a commercial tool with higher accuracies and around x8.4 speedup on average. Ran Chen 0001, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Bei Yu 0001, Pengyun Li, Yu Huang 0005, Jianye Hao |
ICCAD | 4 |
| 2021 | Ordering-Based Causal Discovery with Reinforcement LearningabstractIt is a long-standing question to discover causal relations among a set of variables in many empirical sciences. Recently, Reinforcement Learning (RL) has achieved promising results in causal discovery from observational data. However, searching the space of directed graphs and enforcing acyclicity by implicit penalties tend to be inefficient and restrict the existing RL-based method to small scale problems. In this work, we propose a novel RL-based approach for causal discovery, by incorporating RL into the ordering-based paradigm. Specifically, we formulate the ordering search problem as a multi-step Markov decision process, implement the ordering generating process with an encoder-decoder architecture, and finally use RL to optimize the proposed model based on the reward mechanisms designed for each ordering. A generated ordering would then be processed using variable selection to obtain the final causal graph. We analyze the consistency and computational complexity of the proposed method, and empirically show that a pretrained model can be exploited to accelerate training. Experimental results on both synthetic and real data sets shows that the proposed method achieves a much improved performance over existing RL-based method. Xiaoqiang Wang 0003, Yali Du 0001, Shengyu Zhu 0001, Liangjun Ke, Zhitang Chen, Jianye Hao, Jun Wang 0012 |
IJCAI | 3 |
| 2021 | Asymptotically Optimal One- and Two-Sample Testing With KernelsabstractWe characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test achieves the maximum rate subject to a constant level constraint on the type-I error probability. With Sanov's theorem, we derive a sufficient condition for one-sample tests to achieve the optimal error exponent in the universal setting, i.e., for any distribution defining the alternative hypothesis. We then show that two classes of Maximum Mean Discrepancy (MMD) based tests attain the optimal type-II error exponent on \mathbb Rd, while the quadratic-time Kernel Stein Discrepancy (KSD) based tests achieve this optimality with an asymptotic level constraint. For general two-sample testing, however, Sanov's theorem is insufficient to obtain a similar sufficient condition. We proceed to establish an extended version of Sanov's theorem and derive an exact error exponent for the quadratic-time MMD based two-sample tests. The obtained error exponent is further shown to be optimal among all two-sample tests satisfying a given level constraint. Our work hence provides an achievability result for optimal nonparametric one- and two-sample testing in the universal setting. Application to off-line change detection and related issues are also discussed. Shengyu Zhu 0001, Biao Chen 0001, Zhitang Chen, Pengfei Yang 0003 |
IEEE Trans. Inf. Theory | 1 |
| 2020 | Causal Discovery with Reinforcement Learning
Shengyu Zhu 0001, Ignavier Ng, Zhitang Chen |
ICLR | 1 |
| 2019 | Universal Hypothesis Testing with Kernels: Asymptotically Optimal Tests for Goodness of FitabstractWe characterize the asymptotic performance of nonparametric goodness of fit testing. The exponential decay rate of the type-II error probability is used as the asymptotic performance metric, and a test is optimal if it achieves the maximum rate subject to a constant level constraint on the type-I error probability. We show that two classes of Maximum Mean Discrepancy (MMD) based tests attain this optimality on $\mathbb R^d$, while the quadratic-time Kernel Stein Discrepancy (KSD) based tests achieve the maximum exponential decay rate under a relaxed level constraint. Under the same performance metric, we proceed to show that the quadratic-time MMD based two-sample tests are also optimal for general two-sample problems, provided that kernels are bounded continuous and characteristic. Key to our approach are Sanov’s theorem from large deviation theory and the weak metrizable properties of the MMD and KSD. Shengyu Zhu 0001, Biao Chen 0001, Pengfei Yang 0003, Zhitang Chen |
AISTATS | 1 |
| 2018 | Distributed Detection in Ad Hoc Networks Through Quantized ConsensusabstractWe study the asymptotic performance of distributed detection in large scale connected sensor networks. Contrasting to the canonical parallel network where a single node has access to local decisions from all other nodes, each node can only exchange information with its direct neighbors in the present setting. We establish that, with each node employing an identical one-bit quantizer for local information exchange, a novel consensus reaching approach can achieve the optimal asymptotic performance of centralized detection as the network size scales. The statement is true under three different detection frameworks: 1) the Bayesian criterion where the maximum a posteriori detector is optimal; 2) the Neyman-Pearson criterion with a constant type-I error probability constraint; and 3) the Neyman-Pearson criterion with an exponential type-I error probability constraint. Leveraging recent development in distributed consensus reaching using bounded quantizers with possibly unbounded data (which are log-likelihood ratios of local observations in the context of distributed detection), we design a one-bit deterministic quantizer with a controllable threshold that leads to desirable consensus error bounds. The obtained bounds are key to establishing the optimal asymptotic detection performance. In addition, we examine the non-asymptotic performance of the proposed approach and show that the type-I and type-II error probabilities at each node can be made arbitrarily close to the centralized ones simultaneously when a continuity condition is satisfied. Shengyu Zhu 0001, Biao Chen 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2016 | Quantized consensus ADMM for multi-agent distributed optimizationabstractThis paper considers multi-agent distributed optimization with quantized communication which is needed when inter-agent communications are subject to finite capacity and other practical constraints. To minimize the global objective formed by a sum of local convex functions, we develop a quantized distributed algorithm based on the alternating direction method of multipliers (ADMM). Under certain convexity assumptions, it is shown that the proposed algorithm converges to a consensus within log1+ηΩ iterations, where η > 0 depends on the network topology and the local objectives, and O is a polynomial fraction depending on the quantization resolution, the distance between initial and optimal variable values, the local objectives, and the network topology. We also obtain a tight upper bound on the consensus error which does not depend on the size of the network. Shengyu Zhu 0001, Mingyi Hong 0001, Biao Chen 0001 |
ICASSP | 1 |
| 2016 | Distributed detection over connected networks via one-bit quantizerabstractThis paper considers distributed detection over large scale connected networks with arbitrary topology. Contrasting to the canonical parallel fusion network where a single node has access to the outputs from all other sensors, each node can only exchange one-bit information with its direct neighbors in the present setting. Our approach adopts a novel consensus reaching algorithm using asymmetric bounded quantizers that allow controllable consensus error. Under the Neyman-Pearson criterion, we show that, with each sensor employing an identical one-bit quantizer for local information exchange, this approach achieves the optimal error exponent of centralized detection provided that the algorithm converges. Simulations show that the algorithm converges when the network is large enough. Shengyu Zhu 0001, Biao Chen 0001 |
ISIT | 1 |
| 2013 | Interactive distributed detection with conditionally independent observationsabstractThis paper deals with interactive distributed detection with conditionally independent observations where the fusion center may exchange information with a local sensor. Using a two sensor system, we demonstrate that this two-way interaction provides improvement in detection performance compared with the classical tandem detection system where only one-way communication is allowed. An important observation is that, contrary to that of the tandem network, the fusion rule is no longer a simple likelihood ratio test due to the correlation introduced in the initial feedback from the fusion center to the sensor. Shengyu Zhu 0001, Earnest Akofor, Biao Chen 0001 |
WCNC | 1 |