Ruiying Lu

dblp:255/5995 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-8825-6064ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MaskAD: Parallel Masked Autoencoder for Multi-class Unsupervised Anomaly Detection
abstract
Multi-class unsupervised anomaly detection endeavors to establish a unified model capable of identifying anomalies across multiple classes when only normal data is accessible. However, widely employed reconstruction-based networks often struggle with the 'identical shortcut' issue of both normal and anomalous samples being reconstructed equally well, consequently failing to identify outliers. Although current methodologies attempt to tackle this problem, they remain susceptible to infiltration of anomalous information. In contrast, we introduce a novel scheme to make use of the `identical shortcut' phenomenon rather than pursue to eliminate it. Firstly, inspired by our interesting observation that normal and abnormal regions manifest distinct behaviors when encountering diverse masks, we devise a multi-branch masked autoencoder tailored for multi-class image reconstruction. Subsequently, we introduce a parallel masking scheme to magnify the reconstruction disparity between normal and abnormal regions when confronted with various masks. Ultimately, we propose a reconstruction association discrepancy learning method as a new anomaly localization criterion. The effectiveness of our approach is validated both quantitatively and qualitatively, achieving state-of-the-art results.
Ruiying Lu
AAAI1
2026 PDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection
abstract
Jinhan Liu, Yibo Yang, Ruiying Lu, Piotr Piękos, Yimeng Chen, Peng Wang, Dandan Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jinhan Liu, Ruiying Lu, Piotr Piekos, Dandan Guo
ACL (1)3
2026 IDCFace: Identity Consistent Face anonymization for secure recognition
Ruiying Lu, Shuang Wan, Zimin Miao, Nannan Wang 0001, Chunlei Peng
Pattern Recognit.1
2025 SP-Mamba: Spatial-Perception State Space Model for Unsupervised Medical Anomaly Detection
Ruiying Lu
ACM Multimedia2
2025 Posture-Aware Robust Person Re-Identification via Optimal Transport Calibration
Ruiying Lu, Yalin Sun, Chunlei Peng, Yu Zheng 0006
IEEE Trans. Inf. Forensics Secur.1
2025 Advancing Hyperspectral and Multispectral Image Fusion: An Information-Aware Transformer-Based Unfolding Network
abstract
In hyperspectral image (HSI) processing, the fusion of the high-resolution multispectral image (HR-MSI) and the low-resolution HSI (LR-HSI) on the same scene, known as MSI-HSI fusion, is a crucial step in obtaining the desired high-resolution HSI (HR-HSI). With the powerful representation ability, convolutional neural network (CNN)-based deep unfolding methods have demonstrated promising performances. However, limited receptive fields of CNN often lead to inaccurate long-range spatial features, and inherent input and output images for each stage in unfolding networks restrict the feature transmission, thus limiting the overall performance. To this end, we propose a novel and efficient information-aware transformer-based unfolding network (ITU-Net) to model the long-range dependencies and transfer more information across the stages. Specifically, we employ a customized transformer block to learn representations from both the spatial and frequency domains as well as avoid the quadratic complexity with respect to the input length. For spatial feature extractions, we develop an information transfer guided linearized attention (ITLA), which transmits high-throughput information between adjacent stages and extracts contextual features along the spatial dimension in linear complexity. Moreover, we introduce frequency domain learning in the feedforward network (FFN) to capture token variations of the image and narrow the frequency gap. Via integrating our proposed transformer blocks with the unfolding framework, our ITU-Net achieves state-of-the-art (SOTA) performance on both synthetic and real hyperspectral datasets.
Bo Chen 0001, Ruiying Lu, Ziheng Cheng 0001, Chunhui Qu, Xin Yuan 0002
IEEE Trans. Neural Networks Learn. Syst.3
2024 Latent Diffusion Prior Enhanced Deep Unfolding for Snapshot Spectral Compressive Imaging
Zongliang Wu, Ruiying Lu, Ying Fu 0001, Xin Yuan 0002
ECCV (33)2
2024 FOCT: Few-shot Industrial Anomaly Detection with Foreground-aware Online Conditional Transport
abstract
Few-Shot Industrial Anomaly Detection (FS-IAD) has drawn great attention most recently since data efficiency and the ability to design algorithms for fast migration across products have become the main concerns. The difficulty of memory-based IAD in low-data regime primarily lies in inefficient measurement between the memory bank and query images. We address such a pivotal issue from a new perspective of optimal matching between features of image regions. Taking the unbalanced nature of query features into consideration, we adopt Conditional Transport (CT) as a metric to compute the structural distance between representations of the two sets to determine feature relevance. CT distance generates the optimal matching flows between unbalanced structural elements that achieve the minimum matching cost, which can be directly used for IAD since it well reflects the differences of query images compared with the normal memory. Realizing the fact that query images usually come one-by-one or batch-by-batch, we further propose an Online Conditional Transport (OCT) by making full use of the current and historical query images for IAD via simultaneously calibrating the memory bank using the online query images and matching features between the calibrated memory and the current query image. Go one step further, for sparse foreground products, we employ a predominant segment model to implement Foreground-aware OCT (FOCT) for improving the effectiveness and efficiency of OCT by forcing the model to pay more attention to diverse targets rather than redundant backgrounds when calibrating the memory bank. FOCT can improve the diversity of calibrated memory, which is critical for robust FS-IAD in practice. Besides, FOCT is flexible since it can be friendly plugged and played with any pre-trained backbones, such as WRN, and any pre-trained segment models, such as SAM. The effectiveness and efficiency of our model is demonstrated across diverse datasets, including benchmarks of MVTec and MPDD, achieving SOTA performance.
Hongyi Zhao, Ruiying Lu, Yujie Wu 0008, Xiongpeng He
ACM Multimedia3
2024 Topic-Aware Sensitive Information Detection in Chinese Large Language Model
abstract
With the rapid advancement of deep learning, generative AI models have emerged as a prominent area of focus. However, these developments bring potential security concerns. China’s guiding document of government on generative AI security identifies 31 specific security risks across five categories, including content that violates socialist core values and discriminatory content. Detecting the sensitivity of both user input and model-generated content has therefore become a critical challenge for the security of generative AI models. This paper proposes a robust scheme for detecting sensitive information in the Chinese languages generated from the large language models. Initially, we detect and classify the security risks outlined in the document "Basic Security Requirements for Generative Artificial Intelligence Service" and develop a comprehensive dataset, named Chinese Sensitive Language Detection (CSLD). Specifically, in order to leverage the semantics of languages, we introduce the topic model to pre-analyze text data and construct topic-augmented classification vectors (TACV) that supply effective contextual information for sensitive content detection. Additionally, we propose a topic-infused attention mechanism (TIAM) to provide richer contextual information and relevant topics to guide sensitive information detection. At the same time, the proposed framework is designed to integrate with various classes of Chinese pre-trained models, enabling accurate classification of sensitive content while maintaining low latency and memory usage. Furthermore, the proposed dataset surpasses existing ones in terms of coverage, data volume, and its focus on the security challenges specific to Chinese large language models. Without bells and whistles, our experiments demonstrate that our model outperforms existing models in terms of accuracy and efficiency on the CSLD dataset.
Yalin Sun, Ruiying Lu
TrustCom2
2024 Motion-Aware Dynamic Graph Neural Network for Video Compressive Sensing
abstract
Video snapshot compressive imaging (SCI) utilizes a 2D detector to capture sequential video frames and compress them into a single measurement. Various reconstruction methods have been developed to recover the high-speed video frames from the snapshot measurement. However, most existing reconstruction methods are incapable of efficiently capturing long-range spatial and temporal dependencies, which are critical for video processing. In this paper, we propose a flexible and robust approach based on the graph neural network (GNN) to efficiently model non-local interactions between pixels in space and time regardless of the distance. Specifically, we develop a motion-aware dynamic GNN for better video representation, i.e., represent each node as the aggregation of relative neighbors under the guidance of frame-by-frame motions, which consists of motion-aware dynamic sampling, cross-scale node sampling, global knowledge integration, and graph aggregation. Extensive results on both simulation and real data demonstrate both the effectiveness and efficiency of the proposed approach, and the visualization illustrates the intrinsic dynamic sampling operations of our proposed model for boosting the video SCI reconstruction results. The code and model will be released.
Ruiying Lu, Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Hierarchical Topic-Aware Contextualized Transformers
abstract
Training on disjoint fixed-length segments, Transformers convert static word embeddings into contextualized word representations. However, they often restrict the context of a token to the segment it resides in and hence neglect the contextual information across segments, failing to capture longer-term dependencies beyond the predefined segment length. This article uses a probabilistic deep topic model to provide hierarchical contextualized embeddings at both the token and segment levels, and integrate topic information through a constrained attention mechanism. The proposed method not only injects contextualized topic information into Transformers, but also controls languages generation guided by specific topics, styles, and sentiments. Three plug-and-play modules are proposed, including the contextual topical token embedding, the segment embedding, and the multi-head topic attention mechanism. We aim to capture the semantic coherence and word concurrence patterns at the global level, and also enrich the representation of each token by adapting to its local context, with negligible increased memory footprint and computational time. Experiments on various corpora show that by adding marginal extra parameters, the proposed hierarchical topic-aware contextualized Transformers consistently outperform their conventional counterparts, and generate sentences and paragraphs according to human preferences.
Ruiying Lu, Bo Chen 0001, Dandan Guo, Dongsheng Wang 0003, Mingyuan Zhou
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 ConZIC: Controllable Zero-shot Image Captioning by Sampling-Based Polishing
abstract
Zero-shot capability has been considered as a new revolution of deep learning, letting machines work on tasks without curated training data. As a good start and the only existing outcome of zero-shot image captioning (IC), ZeroCap abandons supervised training and sequentially searches every word in the caption using the knowledge of large-scale pre-trained models. Though effective, its autoregressive generation and gradient-directed searching mechanism limit the diversity of captions and inference speed, respectively. Moreover, ZeroCap does not consider the controllability issue of zero-shot IC. To move forward, we propose a framework for Controllable Zero-shot IC, named ConZIC. The core of ConZIC is a novel sampling-based non-autoregressive language model named Gibbs-BERT, which can generate and continuously polish every word. Extensive quantitative and qualitative results demonstrate the superior performance of our proposed ConZIC for both zero-shot IC and controllable zero-shot IC. Especially, ConZIC achieves about$5\times$generation speed than ZeroCap, and about$1.5\times$diversity scores, with accurate generation given different control signals. Our code is available at https://github.com/joeyz0z/ConZIC.
Zequn Zeng, Hao Zhang 0050, Ruiying Lu, Dongsheng Wang 0003, Bo Chen 0001, Zhengjue Wang
CVPR3
2023 PatchCT: Aligning Patch Set and Label Set with Conditional Transport for Multi-Label Image Classification
abstract
Multi-label image classification is a prediction task that aims to identify more than one label from a given image. This paper considers the semantic consistency of the latent space between the visual patch and linguistic label domains and introduces the conditional transport (CT) theory to bridge the acknowledged gap. While recent cross-modal attention-based studies have attempted to align such two representations and achieved impressive performance, they required carefully-designed alignment modules and extra complex operations in the attention computation. We find that by formulating the multi-label classification as a CT problem, we can exploit the interactions between the image and label efficiently by minimizing the bidirectional CT cost. Specifically, after feeding the images and textual labels into the modality-specific encoders, we view each image as a mixture of patch embeddings and a mixture of label embeddings, which capture the local region features and the class prototypes, respectively. CT is then employed to learn and align those two semantic sets by defining the forward and backward navigators. Importantly, the defined navigators in CT distance model the similarities between patches and labels, which provides an interpretable tool to visualize the learned prototypes. Extensive experiments on three public image benchmarks show that the proposed model consistently outperforms the previous methods.
Miaoge Li, Dongsheng Wang 0003, Zequn Zeng, Ruiying Lu, Bo Chen 0001, Mingyuan Zhou
ICCV5
2023 Hierarchical Vector Quantized Transformer for Multi-class Unsupervised Anomaly Detection
abstract
Unsupervised image Anomaly Detection (UAD) aims to learn robust and discriminative representations of normal samples. While separate solutions per class endow expensive computation and limited generalizability, this paper focuses on building a unified framework for multiple classes. Under such a challenging setting, popular reconstruction-based networks with continuous latent representation assumption always suffer from the "identical shortcut" issue, where both normal and abnormal samples can be well recovered and difficult to distinguish. To address this pivotal issue, we propose a hierarchical vector quantized prototype-oriented Transformer under a probabilistic framework. First, instead of learning the continuous representations, we preserve the typical normal patterns as discrete iconic prototypes, and confirm the importance of Vector Quantization in preventing the model from falling into the shortcut. The vector quantized iconic prototypes are integrated into the Transformer for reconstruction, such that the abnormal data point is flipped to a normal data point. Second, we investigate an exquisite hierarchical framework to relieve the codebook collapse issue and replenish frail normal patterns. Third, a prototype-oriented optimal transport method is proposed to better regulate the prototypes and hierarchically evaluate the abnormal score. By evaluating on MVTec-AD and VisA datasets, our model surpasses the state-of-the-art alternatives and possesses good interpretability. The code is available at https://github.com/RuiyingLu/HVQ-Trans.
Ruiying Lu, Dongsheng Wang 0003, Bo Chen 0001, Ruimin Hu
NeurIPS1
2023 Recurrent Neural Networks for Snapshot Compressive Imaging
abstract
Conventional high-speed and spectral imaging systems are expensive and they usually consume a significant amount of memory and bandwidth to save and transmit the high-dimensional data. By contrast, snapshot compressive imaging (SCI), where multiple sequential frames are coded by different masks and then summed to a single measurement, is a promising idea to use a 2-dimensional camera to capture 3-dimensional scenes. In this paper, we consider the reconstruction problem in SCI, i.e., recovering a series of scenes from a compressed measurement. Specifically, the measurement and modulation masks are fed into our proposed network, dubbed BIdirectional Recurrent Neural networks with Adversarial Training (BIRNAT) to reconstruct the desired frames. BIRNAT employs a deep convolutional neural network with residual blocks and self-attention to reconstruct the first frame, based on which a bidirectional recurrent neural network is utilized to sequentially reconstruct the following frames. Moreover, we build an extended BIRNAT-color algorithm for color videos aiming at joint reconstruction and demosaicing. Extensive results on both video and spectral, simulation and real data from three SCI cameras demonstrate the superior performance of BIRNAT.
Ziheng Cheng 0001, Bo Chen 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Ziyi Meng 0001, Xin Yuan 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 HyperMiner: Topic Taxonomy Mining with Hyperbolic Embedding
abstract
Embedded topic models are able to learn interpretable topics even with large and heavy-tailed vocabularies. However, they generally hold the Euclidean embedding space assumption, leading to a basic limitation in capturing hierarchical relations. To this end, we present a novel framework that introduces hyperbolic embeddings to represent words and topics. With the tree-likeness property of hyperbolic space, the underlying semantic hierarchy among words and topics can be better exploited to mine more interpretable topics. Furthermore, due to the superiority of hyperbolic geometry in representing hierarchical data, tree-structure knowledge can also be naturally injected to guide the learning of a topic hierarchy. Therefore, we further develop a regularization term based on the idea of contrastive learning to inject prior structural knowledge efficiently. Experiments on both topic taxonomy discovery and document representation demonstrate that the proposed framework achieves improved performance against existing embedded topic models.
Yishi Xu, Dongsheng Wang 0003, Bo Chen 0001, Ruiying Lu, Zhibin Duan, Mingyuan Zhou
NeurIPS4
2022 Matching Visual Features to Hierarchical Semantic Topics for Image Paragraph Captioning
Dandan Guo, Ruiying Lu, Bo Chen 0001, Zequn Zeng, Mingyuan Zhou
Int. J. Comput. Vis.2
2021 Memory-Efficient Network for Large-Scale Video Compressive Sensing
abstract
Video snapshot compressive imaging (SCI) captures a sequence of video frames in a single shot using a 2D detector. The underlying principle is that during one exposure time, different masks are imposed on the high-speed scene to form a compressed measurement. With the knowledge of masks, optimization algorithms or deep learning methods are employed to reconstruct the desired high-speed video frames from this snapshot measurement. Unfortunately, though these methods can achieve decent results, the long running time of optimization algorithms or huge training memory occupation of deep networks still preclude them in practical applications. In this paper, we develop a memory-efficient network for large-scale video SCI based on multi-group reversible 3D convolutional neural networks. In addition to the basic model for the grayscale SCI system, we take one step further to combine demosaicing and SCI reconstruction to directly recover color video from Bayer measurements. Extensive results on both simulation and real data captured by SCI cameras demonstrate that our proposed model outperforms previous state-of-the-art with less memory and thus can be used in large-scale problems. The code is at https: //github.com/BoChenGroup/RevSCI-net.
Ziheng Cheng 0001, Bo Chen 0001, Guanliang Liu, Hao Zhang 0050, Ruiying Lu, Zhengjue Wang, Xin Yuan 0002
CVPR5
2021 Dual-view Snapshot Compressive Imaging via Optical Flow Aided Recurrent Neural Network
Ruiying Lu, Bo Chen 0001, Guanliang Liu, Ziheng Cheng 0001, Xin Yuan 0002
Int. J. Comput. Vis.1
2020 BIRNAT: Bidirectional Recurrent Neural Networks with Adversarial Training for Video Snapshot Compressive Imaging
Ziheng Cheng 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Bo Chen 0001, Ziyi Meng 0001, Xin Yuan 0002
ECCV (24)2
2020 Recurrent Hierarchical Topic-Guided RNN for Language Generation
abstract
To simultaneously capture syntax and global semantics from a text corpus, we propose a new larger-context recurrent neural network (RNN) based language model, which extracts recurrent hierarchical semantic structure via a dynamic deep topic model to guide natural language generation. Moving beyond a conventional RNN-based language model that ignores long-range word dependencies and sentence order, the proposed model captures not only intra-sentence word dependencies, but also temporal transitions between sentences and inter-sentence topic dependencies. For inference, we develop a hybrid of stochastic-gradient Markov chain Monte Carlo and recurrent autoencoding variational Bayes. Experimental results on a variety of real-world text corpora demonstrate that the proposed model not only outperforms larger-context RNN-based language models, but also learns interpretable recurrent multilayer topics and generates diverse sentences and paragraphs that are syntactically correct and semantically coherent.
Dandan Guo, Bo Chen 0001, Ruiying Lu, Mingyuan Zhou
ICML3
2020 RAFnet: Recurrent attention fusion network of hyperspectral and multispectral images
Ruiying Lu, Bo Chen 0001, Ziheng Cheng 0001
Signal Process.1
2020 FusionNet: An Unsupervised Convolutional Variational Network for Hyperspectral and Multispectral Image Fusion
abstract
Due to hardware limitations of the imaging sensors, it is challenging to acquire images of high resolution in both spatial and spectral domains. Fusing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to obtain an HR-HSI in an unsupervised manner has drawn considerable attention. Though effective, most existing fusion methods are limited due to the use of linear parametric modeling for the spectral mixture process, and even the deep learning-based methods only focus on deterministic fully-connected networks without exploiting the spatial correlation and local spectral structures of the images. In this paper, we propose a novel variational probabilistic autoencoder framework implemented by convolutional neural networks, in order to fuse the spatial and spectral information contained in the LR-HSI and HR-MSI, called FusionNet. The FusionNet consists of a spectral generative network, a spatial-dependent prior network, and a spatial-spectral variational inference network, which are jointly optimized in an unsupervised manner, leading to an end-to-end fusion system. Further, for fast adaptation to different observation scenes, we give a meta-learning explanation to the fusion problem, and combine the FusionNet with meta-learning in a synergistic manner. Effectiveness and efficiency of the proposed method are evaluated based on several publicly available datasets, demonstrating that the proposed FusionNet outperforms the state-of-the-art fusion methods.
Zhengjue Wang, Bo Chen 0001, Ruiying Lu, Hao Zhang 0050, Hongwei Liu 0001, Pramod K. Varshney
IEEE Trans. Image Process.3