VLDB 2026 Research / reviewers in the wild / expert
Shengyang Sun
dblp:173/5093
· DBLP profile ↗
27ranked-venue papers
14as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Computer networks · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An incentive mechanism based on Multi-Agent Attentive Dueling Double Deep Q-Networks for Mobile Crowdsensing with incomplete information
Xiangling Wu, Wenming Ma, Xiaoyan Dai, Shengyang Sun, Xiaoang Zhu |
Comput. Networks | 5 |
| 2026 | Data quality guided multi task allocation based on bi-level deep reinforcement learning in crowdsensing
Wenming Ma, Xiangling Wu, Xiaoang Zhu, Shengyang Sun |
Comput. Commun. | 5 |
| 2026 | Enhancing Weakly Supervised Multimodal Video Anomaly Detection Through Text GuidanceabstractIn recent years, weakly supervised multimodal video anomaly detection, which leverages RGB, optical flow, and audio modalities, has garnered significant attention from researchers, emerging as a vital subfield within video anomaly detection. However, previous studies have inadequately explored the role of text modality in this domain. With the proliferation of large-scale text-annotated video datasets and the advent of video captioning models, obtaining text descriptions from videos has become increasingly feasible. Text modality, carrying explicit semantic information, can more accurately characterize events within videos and identify anomalies, thereby enhancing the model's detection capabilities and reducing false alarms. However, text feature extraction challenges anomaly detection. Pre-trained large language models often struggle to effectively capture the nuances associated with anomalies, as their training is based on generalized datasets. Directly fine-tuning the text feature extractor is also challenging, as anomaly-related text descriptions are sparse. Furthermore, due to the varying amounts of information carried by different modalities, issues such as modality redundancy and modality imbalance arise during feature fusion. To address the challenges of text feature extraction and the issues of modality redundancy and imbalance, we propose a novel text-guided weakly supervised multimodal video anomaly detection framework. Specifically, we introduce an in-context learning based multi-stage text augmentation mechanism to generate high-quality anomaly text samples. These high-quality samples are then used to fine-tune the text feature extractor, aiming to obtain a more effective text feature extractor for anomaly detection. Additionally, we present a multi-scale bottleneck Transformer fusion module to enhance multimodal integration, utilizing a set of reduced bottleneck tokens to progressively transmit compressed information across modalities, aiming to address the issues of modality redundancy and imbalance. Experimental results on large-scale datasets UCF-Crime and XD-Violence demonstrate that our proposed approach achieves state-of-the-art performance. This project is publicly available at https://shengyangsun.github.io/TGMVAD. Shengyang Sun, Jiashen Hua, Junyi Feng, Xiaojin Gong |
IEEE Trans. Multim. | 1 |
| 2025 | Activity-based capability updating method for task assignment in mobile crowdsensing
Wenming Ma, Xiangling Wu, Shengyang Sun, Xiaoang Zhu |
Comput. Networks | 4 |
| 2025 | A Pareto-based genetic algorithm for online task allocation in mobile crowdsensing
Xiangling Wu, Wenming Ma, Shengyang Sun, Xiaoang Zhu |
Comput. Commun. | 4 |
| 2025 | Delving Into Instance Modeling for Weakly Supervised Video Anomaly DetectionabstractWeakly-supervised video anomaly detection (WS-VAD) aims to identify fine-grained anomalies from sparse video-level labels, which has gained increasing attention in recent years due to its various applications such as disaster warning and public security. Recent studies typically formulate WS-VAD as a multi-instance learning (MIL) problem. However, they neglect the instance creation process and simply apply a uniform temporal pooling (UTP) operation to obtain the training instances, leading to severe anomaly contamination and dilution. In this paper, we emphasize the importance of the instance modeling procedure and propose two simple yet effective modules, i.e., the dynamic segment merging (DSM) module and the retrieval-augmented anomaly restoration (RA2R) module, to tackle the problem from segment-level and feature-level, respectively. We equip various state-of-the-art WS-VAD models with the proposed methods and conduct thorough experiments on the challenging datasets, e.g., UCF-Crime, and XD-Violence. Results demonstrate the proposed method brings consistent performance improvement and establishes new state-of-the-art. Shengyang Sun, Jiashen Hua, Junyi Feng, Dongxu Wei, Baisheng Lai, Xiaojin Gong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence DetectionabstractWeakly supervised multimodal violence detection aims to learn a violence detection model by leveraging multiple modalities such as RGB, optical flow, and audio, while only video-level annotations are available. In the pursuit of effective multimodal violence detection (MVD), information redundancy, modality imbalance, and modality asynchrony are identified as three key challenges. In this work, we propose a new weakly supervised MVD method that explicitly addresses these challenges. Specifically, we introduce a multi-scale bottleneck transformer (MSBT) based fusion module that employs a reduced number of bottleneck tokens to gradually condense information and fuse each pair of modalities and utilizes a bottleneck token-based weighting scheme to highlight more important fused features. Furthermore, we propose a temporal consistency contrast loss to semantically align pairwise fused features. Experiments on the largest-scale XD-Violence dataset demonstrate that the proposed method achieves state-of-the-art performance. Code is available at https://github.com/shengyangsun/MSBT. Shengyang Sun, Xiaojin Gong |
ICME | 1 |
| 2024 | TDSD: Text-Driven Scene-Decoupled Weakly Supervised Video Anomaly Detection
Shengyang Sun, Jiashen Hua, Junyi Feng, Dongxu Wei, Baisheng Lai, Xiaojin Gong |
ACM Multimedia | 1 |
| 2024 | Event-driven weakly supervised video anomaly detection
Shengyang Sun, Xiaojin Gong |
Image Vis. Comput. | 1 |
| 2023 | Hierarchical Semantic Contrast for Scene-aware Video Anomaly DetectionabstractIncreasing scene-awareness is a key challenge in video anomaly detection (VAD). In this work, we propose a hierarchical semantic contrast (HSC) method to learn a scene-aware VAD model from normal videos. We first incorporate foreground object and background scene features with high-level semantics by taking advantage of pre-trained video parsing models. Then, building upon the autoencoder-based reconstruction framework, we introduce both scene-level and object-level contrastive learning to enforce the encoded latent features to be compact within the same semantic classes while being separable across different classes. This hierarchical semantic contrast strategy helps to deal with the diversity of normal patterns and also increases their discrimination ability. Moreover, for the sake of tackling rare normal activities, we design a skeleton-based motion augmentation to increase samples and refine the model further. Extensive experiments on three public datasets and scene-dependent mixture datasets validate the effectiveness of our proposed method. Shengyang Sun, Xiaojin Gong |
CVPR | 1 |
| 2023 | Long-Short Temporal Co-Teaching for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection (WS-VAD) is a challenging problem that aims to learn VAD models only with video-level annotations. In this work, we propose a Long-Short Temporal Co-teaching (LSTC) method to address the WS-VAD problem. It constructs two tubelet-based spatio-temporal transformer networks to learn from short- and long-term video clips respectively. Each network is trained with respect to a multiple instance learning (MIL)-based ranking loss, together with a cross-entropy loss when clip-level pseudo labels are available. A co-teaching strategy is adopted to train the two networks. That is, clip-level pseudo labels generated from each network are used to supervise the other one at the next training round, and the two networks are learned alternatively and iteratively. Our proposed method is able to better deal with the anomalies with varying durations as well as subtle anomalies. Extensive experiments on three public datasets demonstrate that our method outperforms state-of-the-art WS-VAD methods. Code is available at https://github.com/shengyangsun/LSTC_VAD. Shengyang Sun, Xiaojin Gong |
ICME | 1 |
| 2022 | Understanding the Variance Collapse of SVGD in High Dimensions
Jimmy Ba, Murat A. Erdogdu, Marzyeh Ghassemi, Shengyang Sun, Taiji Suzuki, Denny Wu, Tianzong Zhang |
ICLR | 4 |
| 2022 | Information-theoretic Online Memory Selection for Continual Learning
Shengyang Sun, Daniele Calandriello, Huiyi Hu, Michalis K. Titsias |
ICLR | 1 |
| 2021 | Beyond Marginal Uncertainty: How Accurately can Bayesian Regression Models Estimate Posterior Predictive Correlations?abstractWhile uncertainty estimation is a well-studied topic in deep learning, most such work focuses on marginal uncertainty estimates, i.e. the predictive mean and variance at individual input locations. But it is often more useful to estimate predictive correlations between the function values at different input locations. In this paper, we consider the problem of benchmarking how accurately Bayesian models can estimate predictive correlations. We first consider a downstream task which depends on posterior predictive correlations: transductive active learning (TAL). We find that TAL makes better use of models’ uncertainty estimates than ordinary active learning, and recommend this as a benchmark for evaluating Bayesian models. Since TAL is too expensive and indirect to guide development of algorithms, we introduce two metrics which more directly evaluate the predictive correlations and which can be computed efficiently: meta-correlations (i.e. the correlations between the models correlation estimates and the true values), and cross-normalized likelihoods (XLL). We validate these metrics by demonstrating their consistency with TAL performance and obtain insights about the relative performance of current Bayesian neural net and Gaussian process models. Chaoqi Wang, Shengyang Sun, Roger B. Grosse |
AISTATS | 2 |
| 2021 | Scalable Variational Gaussian Processes via Harmonic Kernel DecompositionabstractWe introduce a new scalable variational Gaussian process approximation which provides a high fidelity approximation while retaining general applicability. We propose the harmonic kernel decomposition (HKD), which uses Fourier series to decompose a kernel as a sum of orthogonal kernels. Our variational approximation exploits this orthogonality to enable a large number of inducing points at a low computational cost. We demonstrate that, on a range of regression and classification problems, our approach can exploit input space symmetries such as translations and reflections, and it significantly outperforms standard variational methods in scalability and accuracy. Notably, our approach achieves state-of-the-art results on CIFAR-10 among pure GP models. Shengyang Sun, Jiaxin Shi, Andrew Gordon Wilson, Roger B. Grosse |
ICML | 1 |
| 2021 | In-situ learning in multilayer locally-connected memristive spiking neural network
Hui Xu 0010, Shengyang Sun, Zhiwei Li 0008, Qingjiang Li, Haijun Liu 0003, Nan Li 0020 |
Neurocomputing | 3 |
| 2020 | Enhanced Spiking Neural Network with forgetting phenomenon based on electronic synaptic devices
Hui Xu 0010, Shengyang Sun, Sen Liu 0006, Nan Li 0020, Qingjiang Li, Haijun Liu 0003, Zhiwei Li 0008 |
Neurocomputing | 3 |
| 2019 | Aggregated Momentum: Stability Through Passive Damping
James Lucas, Shengyang Sun, Richard S. Zemel, Roger B. Grosse |
ICLR (Poster) | 2 |
| 2019 | Functional variational Bayesian Neural Networks
Shengyang Sun, Guodong Zhang 0006, Jiaxin Shi, Roger B. Grosse |
ICLR (Poster) | 1 |
| 2019 | Cascaded Neural Network for Memristor based Neuromorphic ComputingabstractRecent years, several memristor-based neuromorphic processing chips have been proposed. However, there is few architectures to consider the cascading problems, and the scalability is not strong while dealing some tasks. To address this issue, we present a memristor-based cascaded method with some basic computation unit, several neural network processing chips can be cascaded by this means to improve the processing capability of the dataset. Compared with VGGNet and GoogLeNet, the proposed cascaded framework can achieve 93.54% Fashion-MNIST accuracy under the 4.15M parameters. Extensive experiments are conducted show that the circuit simulation results can still provide a high recognition accuracy, the recognition accuracy loss after circuit simulation can be controlled at around 0.26%. Shengyang Sun, Hui Xu 0010, Haijun Liu 0003, Qingjiang Li |
IJCNN | 1 |
| 2019 | Fast-rate PAC-Bayes Generalization Bounds via Shifted Rademacher ProcessesabstractThe developments of Rademacher complexity and PAC-Bayesian theory have been largely independent. One exception is the PAC-Bayes theorem of Kakade, Sridharan, and Tewari (2008), which is established via Rademacher complexity theory by viewing Gibbs classifiers as linear operators. The goal of this paper is to extend this bridge between Rademacher complexity and state-of-the-art PAC-Bayesian theory. We first demonstrate that one can match the fast rate of Catoni's PAC-Bayes bounds (Catoni, 2007) using shifted Rademacher processes (Wegkamp, 2003; Lecué and Mitchell, 2012; Zhivotovskiy and Hanneke, 2018). We then derive a new fast-rate PAC-Bayes bound in terms of the "flatness" of the empirical risk surface on which the posterior concentrates. Our analysis establishes a new framework for deriving fast-rate PAC-Bayes bounds and yields new insights on PAC-Bayesian theory. Shengyang Sun, Daniel M. Roy 0001 |
NeurIPS | 2 |
| 2018 | Kernel Implicit Variational Inference
Jiaxin Shi, Shengyang Sun, Jun Zhu 0001 |
ICLR (Poster) | 2 |
| 2018 | A Spectral Approach to Gradient Estimation for Implicit DistributionsabstractRecently there have been increasing interests in learning and inference with implicit distributions (i.e., distributions without tractable densities). To this end, we develop a gradient estimator for implicit distributions based on Stein’s identity and a spectral decomposition of kernel operators, where the eigenfunctions are approximated by the Nystr{ö}m method. Unlike the previous works that only provide estimates at the sample points, our approach directly estimates the gradient function, thus allows for a simple and principled out-of-sample extension. We provide theoretical results on the error bound of the estimator and discuss the bias-variance tradeoff in practice. The effectiveness of our method is demonstrated by applications to gradient-free Hamiltonian Monte Carlo and variational inference with implicit distributions. Finally, we discuss the intuition behind the estimator by drawing connections between the Nystr{ö}m method and kernel PCA, which indicates that the estimator can automatically adapt to the geometry of the underlying distribution. Jiaxin Shi, Shengyang Sun, Jun Zhu 0001 |
ICML | 2 |
| 2018 | Differentiable Compositional Kernel Learning for Gaussian ProcessesabstractThe generalization properties of Gaussian processes depend heavily on the choice of kernel, and this choice remains a dark art. We present the Neural Kernel Network (NKN), a flexible family of kernels represented by a neural network. The NKN’s architecture is based on the composition rules for kernels, so that each unit of the network corresponds to a valid kernel. It can compactly approximate compositional kernel structures such as those used by the Automatic Statistician (Lloyd et al., 2014), but because the architecture is differentiable, it is end-to-end trainable with gradient- based optimization. We show that the NKN is universal for the class of stationary kernels. Empirically we demonstrate NKN’s pattern discovery and extrapolation abilities on several tasks that depend crucially on identifying the underlying structure, including time series and texture extrapolation, as well as Bayesian optimization. Shengyang Sun, Guodong Zhang 0006, Chaoqi Wang, Wenyuan Zeng, Jiaman Li, Roger B. Grosse |
ICML | 1 |
| 2018 | Noisy Natural Gradient as Variational InferenceabstractVariational Bayesian neural nets combine the flexibility of deep learning with Bayesian uncertainty estimation. Unfortunately, there is a tradeoff between cheap but simple variational families (e.g. fully factorized) or expensive and complicated inference procedures. We show that natural gradient ascent with adaptive weight noise implicitly fits a variational posterior to maximize the evidence lower bound (ELBO). This insight allows us to train full-covariance, fully factorized, or matrix-variate Gaussian variational posteriors using noisy versions of natural gradient, Adam, and K-FAC, respectively, making it possible to scale up to modern-size ConvNets. On standard regression benchmarks, our noisy K-FAC algorithm makes better predictions and matches Hamiltonian Monte Carlo’s predictive variances better than existing methods. Its improved uncertainty estimates lead to more efficient exploration in active learning, and intrinsic motivation for reinforcement learning. Guodong Zhang 0006, Shengyang Sun, David Duvenaud, Roger B. Grosse |
ICML | 2 |
| 2018 | Low-Consumption Neuromorphic Memristor Architecture Based on Convolutional Neural NetworksabstractWith the rapid development of VLSI industry, the research of intelligent applications moves towards IoT edge computing. While the power consumption and area cost of deep neural networks usually exceed the hardware limitation of edge devices. In this paper, we propose a low-power neural network architecture to address such problem. We simplify the current popular convolutional neural networks structure, and utilize the memristor crossbar to store weights to execute convolution operation in parallel, and we present the spiking convolutional neural networks. At the same time, we proposed a performance metrics V to help provide design guidelines for choosing the parameters of the network. Shengyang Sun, Zhiwei Li 0008, Haijun Liu 0003, Qingjiang Li, Hui Xu 0010 |
IJCNN | 1 |
| 2017 | Learning Structured Weight Uncertainty in Bayesian Neural NetworksabstractDeep neural networks (DNNs) are increasingly popular in modern machine learning. Bayesian learning affords the opportunity to quantify posterior uncertainty on DNN model parameters. Most existing work adopts independent Gaussian priors on the model weights, ignoring possible structural information. In this paper, we consider the matrix variate Gaussian (MVG) distribution to model structured correlations within the weights of a DNN. To make posterior inference feasible, a reparametrization is proposed for the MVG prior, simplifying the complex MVG-based model to an equivalent yet simpler model with independent Gaussian priors on the transformed weights. Consequently, we develop a scalable Bayesian online inference algorithm by adopting the recently proposed probabilistic backpropagation framework. Experiments on several synthetic and real datasets indicate the superiority of our model, achieving competitive performance in terms of model likelihood and predictive root mean square error. Importantly, it also yields faster convergence speed compared to related Bayesian DNN models. Shengyang Sun, Changyou Chen, Lawrence Carin |
AISTATS | 1 |