EDBT 2026 Demo / reviewers in the wild / expert
Ruoqi Li
dblp:81/10112
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Asynchronous Intermittent Control Methodology for Cyber-Physical Systems Under Dynamic Actuator Faults
Ruoqi Li, Bingbing Zhang 0001, Yang Yang 0052, Qi-He Shan, Lei Liu 0006 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | SwinFVO: Self-Supervised Visual Odometry With Enhanced Global Spatiotemporal PerceptionabstractPose estimation using visual sensors has become a fundamental component in robotic navigation and autonomous driving systems. Learning-based monocular visual odometry (VO) has attracted substantial attention due to its resilience to camera parameter variations and dynamic environments. Given that camera movement manifests as pixel-level motion across the entire image in optical flow data, capturing both global contextual information and local feature details is crucial for accurate pose estimation. To address this challenge, we propose SwinFVO, a novel self-supervised visual odometry framework that incorporates enhanced motion perception to achieve global spatial dependency modeling with temporal continuity. Leveraging quadrant-based motion characteristics, we perform cross-regional feature interaction through a refined Swin Transformer architecture. Two robust spatiotemporal feature extractors are designed to extend the single-frame-based Swin Transformer to a temporally-aware framework for sequential understanding. Through the exploration of long-range spatial correlations and preservation of temporal consistency, SwinFVO delivers accurate and consistent pose estimation. Extensive experiments across multiple datasets demonstrate the superior performance and generalization capability of SwinFVO in both pose and depth estimation tasks. It achieves competitive results against classical algorithms and outperforms related state-of-the-art (SOTA) methods by up to 20.6% and 72.4% on average translational and rotational evaluations, respectively. Rujun Song, Ruoqi Li, Zhuoling Xiao, Bo Yan 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Hierarchical Gaussian Mixture Normalizing Flow Modeling for Unified Anomaly Detection
Xincheng Yao, Ruoqi Li, Zefeng Qian |
ECCV (32) | 2 |
| 2024 | Generative Representation and Discriminative Classification for Few-shot Open-set Object DetectionabstractOpen-Set Object Detection (OSOD) aims to train detectors on closed-set datasets to detect known objects and identify unknown objects in open-set conditions. Traditional discriminative classifier-based OSOD methods struggle to accurately learn the decision boundary between known and unknown classes, often resulting in the misclassification of unknown samples. In this work, we aim to combine generative representation with discriminative classification to alleviate the issue of misclassification by transforming known-unknown recognition into a binary classification problem. The proposed two-stage OSOD approach proceeds as follows: during the generative representation stage, we employ Class-Conditioned Normalizing Flow (CCNF) to establish distribution mapping for each known category; In the discriminative classification stage, by utilizing a small number of unknown class samples, semi-push-pull supervised learning and entropy contrast learning are used to separate known and unknown classes. Extensive experiments demonstrate that our method significantly enhances OSOD performance, evidenced by a 25.8%-28.6% reduction in the Wilderness Index and a decrease of 4391-8870 units in Absolute Open-Set Errors on the test set VOC-COCO-T1. Peixue Shen, Ruoqi Li, Yan Luo 0003, Yiru Zhao |
VCIP | 2 |
| 2023 | One-for-All: Proposal Masked Cross-Class Anomaly DetectionabstractOne of the most challenges for anomaly detection (AD) is how to learn one unified and generalizable model to adapt to multi-class especially cross-class settings: the model is trained with normal samples from seen classes with the objective to detect anomalies from both seen and unseen classes. In this work, we propose a novel Proposal Masked Anomaly Detection (PMAD) approach for such challenging multi- and cross-class anomaly detection. The proposed PMAD can be adapted to seen and unseen classes by two key designs: MAE-based patch-level reconstruction and prototype-guided proposal masking. First, motivated by MAE (Masked AutoEncoder), we develop a patch-level reconstruction model rather than the image-level reconstruction adopted in most AD methods for this reason: the masked patches in unseen classes can be reconstructed well by using the visible patches and the adaptive reconstruction capability of MAE. Moreover, we improve MAE by ViT encoder-decoder architecture, combinational masking, and visual tokens as reconstruction objectives to make it more suitable for anomaly detection. Second, we develop a two-stage anomaly detection manner during inference. In the proposal masking stage, the prototype-guided proposal masking module is utilized to generate proposals for suspicious anomalies as much as possible, then masked patches can be generated from the proposal regions. By masking most likely anomalous patches, the “shortcut reconstruction” issue (i.e., anomalous regions can be well reconstructed) can be mostly avoided. In the reconstruction stage, these masked patches are then reconstructed by the trained patch-level reconstruction model to determine if they are anomalies. Extensive experiments show that the proposed PMAD can outperform current state-of-the-art models significantly under the multi- and especially cross-class settings. Code will be publicly available at https://github.com/xcyao00/PMAD. Xincheng Yao, Ruoqi Li, Jun Sun 0005 |
AAAI | 3 |
| 2023 | Explicit Boundary Guided Semi-Push-Pull Contrastive Learning for Supervised Anomaly DetectionabstractMost anomaly detection (AD) models are learned using only normal samples in an unsupervised way, which may result in ambiguous decision boundary and insufficient discriminability. In fact, a few anomaly samples are often available in real-world applications, the valuable knowledge of known anomalies should also be effectively exploited. However, utilizing a few known anomalies during training may cause another issue that the model may be biased by those known anomalies and fail to generalize to unseen anomalies. In this paper, we tackle supervised anomaly detection, i.e., we learn AD models using a few available anomalies with the objective to detect both the seen and unseen anomalies. We propose a novel explicit boundary guided semi-push-pull contrastive learning mechanism, which can enhance model's discriminability while mitigating the bias issue. Our approach is based on two core designs: First, we find an explicit and compact separating boundary as the guidance for further feature learning. As the boundary only relies on the normal feature distribution, the bias problem caused by a few known anomalies can be alleviated. Second, a boundary guided semi-push-pull loss is developed to only pull the normal features together while pushing the abnormal features apart from the separating boundary beyond a certain margin region. In this way, our model can form a more explicit and discriminative decision boundary to distinguish known and also unseen anomalies from normal samples more effectively. Code will be available at https://github.com/xcyao00/BGAD. Xincheng Yao, Ruoqi Li, Jun Sun 0005 |
CVPR | 2 |
| 2023 | Adaptive Semantic Fusion Framework for Unsupervised Monocular Depth EstimationabstractUnsupervised monocular depth estimation plays an important role in autonomous driving, and has been received considerable research attention in recent years. Nevertheless, numerous existing methods relying on photometric consistency are excessively susceptible to variations in illumination and suffer in the regions with strong reflection. To overcome this limitation, we propose a novel unsupervised depth estimation framework named ColorDepth, which forces the model to explore object semantic to infer depth. Specifically, we extract pixel-level semantic prior clues of objects using the semantic segmentation network. These priors and the original image are then adaptively fused into color data by a learnable parameter for depth estimation. The incorporation of semantics endows our model with the ability to perceive scene structure information. The fused data effectively alleviates the depth ambiguity within the same semantic block, leading to improved consistency and robustness in challenging scenarios. Extensive experiments on the KITTI and Make3D datasets show that our method surpasses the previous state-of-the-art methods even those supervised by additional constraints, and brings significant performance improvement particularly in the regions of high reflection. Ruoqi Li, Kaiyang Du, Zhuoling Xiao, Bo Yan 0007, Zhengxi Yuan |
ICASSP | 1 |
| 2023 | Focus the Discrepancy: Intra- and Inter-Correlation Learning for Image Anomaly DetectionabstractHumans recognize anomalies through two aspects: larger patch-wise representation discrepancies and weaker patch-to-normal-patch correlations. However, the previous AD methods didn’t sufficiently combine the two complementary aspects to design AD models. To this end, we find that Transformer can ideally satisfy the two aspects as its great power in the unified modeling of patch-wise representations and patch-to-patch correlations. In this paper, we propose a novel AD framework: FOcus-the-Discrepancy (FOD), which can simultaneously spot the patch-wise, intra- and inter-discrepancies of anomalies. The major characteristic of our method is that we renovate the self-attention maps in transformers to Intra-Inter-Correlation (I2Correlation). The I2Correlation contains a two-branch structure to first explicitly establish intra-and inter-image correlations, and then fuses the features of two-branch to spotlight the abnormal patterns. To learn the intra- and inter-correlations adaptively, we propose the RBF-kernel-based target-correlations as learning targets for self-supervised learning. Besides, we introduce an entropy constraint strategy to solve the mode collapse issue in optimization and further amplify the normal-abnormal distinguishability. Extensive experiments on three unsupervised real-world AD benchmarks show the superior performance of our approach. Code will be available at https://github.com/xcyao00/FOD. Xincheng Yao, Ruoqi Li, Zefeng Qian, Yan Luo 0003 |
ICCV | 2 |
| 2023 | GlobalDepth: Global-Aware Attention Model for Unsupervised Monocular Depth EstimationabstractMonocular depth estimation is a significant task in computer vision, which can be widely used in Simultaneous Localization and Mapping (SLAM) and navigation. However, the current unsupervised approaches have limitations in global information perception, especially at distant objects and the boundaries of the objects. To overcome this weakness, we propose a global-aware attention model called GlobalDepth for depth estimation, which includes two essential modules: Global Feature Extraction (GFE) and Selective Feature Fusion (SFF). GFE considers the correlation among multiple channels and refines the encoder feature by extending the receptive field of the network. Furthermore, we restructure the skip connection by employing SFF between the low-level and the high-level features in element wise, rather than simply concatenation or addition at the feature level. Our model excavates the key information and enhances the ability of global perception to predict details of the scene. Extensive experimental results demonstrate that our method reduces the absolute relative error by 10.32% compared with other state-of-the-art models on KITTI datasets. Ruoqi Li, Zhuoling Xiao, Bo Yan 0007 |
ISCAS | 2 |
| 2022 | You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object SegmentationabstractWe present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image and language can increase feature complexity and thus may be sub-optimal for RVOS. To this end, we propose a meta-transfer module, which is trained in a learning-to-learn fashion and aims to transfer the target-specific information from the language domain to the image domain, while discarding the uncorrelated complex variations of language description. To bridge the gap between the image and language domains, we develop a multi-scale cross-modal feature mining block that aggregates all the essential features required by RVOS from both domains and generates regression labels for the meta-transfer module. The whole system can be trained in an end-to-end manner and shows competitive performance against state-of-the-art two-stage approaches. Dezhuang Li, Ruoqi Li, Lijun Wang 0001, Yifan Wang 0004, Jinqing Qi, Lu Zhang 0053, Ting Liu 0018, Qingquan Xu, Huchuan Lu |
AAAI | 2 |
| 2022 | Out-of-Distribution Identification: Let Detector Tell Which I Am Not Sure
Ruoqi Li, Hao Zhou 0014, Yan Luo 0003 |
ECCV (10) | 1 |
| 2022 | From Pixels to Semantics: Self-Supervised Video Object Segmentation With Multiperspective Feature MiningabstractExisting self-supervised methods pose one-shot video object segmentation (O-VOS) as pixel-level matching to enable segmentation mask propagation across frames. However, the two tasks are not fully equivalent since O-VOS is more reliant on semantic correspondence rather than accurate pixel matching. To remedy this issue, we explore a new self-supervised framework that integrates pixel-level correspondence learning with semantic-level adaptation. The pixel-level correspondence learning is performed through photometric reconstruction of adjacent RGB frames during offline training, while semantic-level adaption operates at test-time by enforcing a bi-directional agreement of the predicted segmentation masks. In addition, we further propose a new network architecture with multi-perspective feature mining mechanism which can not only enhance reliable features but also suppress noisy ones to facilitate more robust image matching. By training the network using the proposed self-supervised framework, we achieve state-of-the-art performance on widely adopted datasets, further closing up the gap between self-supervised learning methods and their fully supervised counterparts. Ruoqi Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Xiaopeng Wei, Qiang Zhang 0008 |
IEEE Trans. Image Process. | 1 |
| 2019 | A Knowledge Graph Framework for Software-Defined Industrial Cyber-Physical SystemsabstractAutomatic code generation is a critical step towards flexible manufacturing processes. In the last decade, Model-Driven Engineering is commonly adopted for code generation while requiring tight coupling between the model and its corresponding code, which increases the difficulty of dynamic reconfiguration in the system. To improve the flexibility and efficiency of industrial software design and development processes, Knowledge Graph is proposed to be applied in knowledge-driven code generation process in this paper. The knowledge-driven query system can conduct parameter searching, variable calculation, ontology reasoning, and code invocation to assist code generation. In addition, a structure of domain-specific knowledge graph combined with SQL database and reasoning rules are proposed to improve performance. The feasibility of our method is demonstrated through the dynamic AGV route planning example. Ruoqi Li, Wenbin William Dai, Xiaosheng Chen, Genke Yang |
IECON | 1 |