EDBT 2026 Demo / reviewers in the wild / expert
Yan Luo 0003
dblp:25/2208-3
· DBLP profile ↗
18ranked-venue papers
4as first author
16since 2021 · last 2027
0000-0002-1394-4452ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Visual behavior understanding for smart livestock management: A survey from spatial perception to temporal reasoning
Yan Luo 0003, Xingjian Gu, Bo Li 0070, Longshen Liu, Dong Liu 0059, Tomas Norton, Mingzhou Lu |
Expert Syst. Appl. | 1 |
| 2026 | Hierarchical kernel decoupling for graph convolution: Enhancing skeleton-based action recognition through structured representation
Ying Li 0016, Hao Zhou 0014, Chuanping Hu, Mingzhou Lu, Yan Luo 0003 |
Pattern Recognit. | 6 |
| 2025 | Make Unseen Clear: Occluder Removal for Complete 3D Pedestrian DetectionabstractIn autonomous driving, the ability to detect pedestrians accurately is crucial for safety. Some detectors, however, often struggle with occlusions, where pedestrians partially hidden behind objects appear incomplete and are harder to be identified accurately. To alleviate this issue, we introduce Clear3D, which make previous partially unseen pedestrians visible by effectively removing occluders to enhance 3D pedestrian detection. The first essential step in Clear3D is to accurately identify occluded regions. This is achieved by calculating the intensity variation along the LiDAR ray direction, where regions with small changes are classified as occluded. After identifying the occluded areas, they are transformed as masks and fed into the inpainting module, reconstructing the originally unseen parts. Above process, encompassing both the perception and reconstruction of occluded regions, is termed occluder removal. Finally, by fusing the recovered occlusion-free image features with LiDAR features, Clear3D generates more complete representation, enhancing the performance of downstream 3D classification and localization tasks. Extensive evaluations show that Clear3D achieves state-of-the-art accuracy across various challenging settings on KITTI, particularly in heavy occlusion environments. Yan Luo 0003, Zefeng Qian, Yi Wang 0070, Xiaolei Hu |
ICASSP | 1 |
| 2025 | ADPretrain: Advancing Industrial Anomaly Detection via Anomaly Representation PretrainingabstractThe current mainstream and state-of-the-art anomaly detection (AD) methods are
substantially established on pretrained feature networks yielded by ImageNet pre-
training. However, regardless of supervised or self-supervised pretraining, the
pretraining process on ImageNet does not match the goal of anomaly detection
(i.e., pretraining in natural images doesn’t aim to distinguish between normal and
abnormal). Moreover, natural images and industrial image data in AD scenarios
typically have the distribution shift. The two issues can cause ImageNet-pretrained
features to be suboptimal for AD tasks. To further promote the development of
the AD field, pretrained representations specially for AD tasks are eager and very
valuable. To this end, we propose a novel AD representation learning framework
specially designed for learning robust and discriminative pretrained representa-
tions for industrial anomaly detection. Specifically, closely surrounding the goal
of anomaly detection (i.e., focus on discrepancies between normals and anoma-
lies), we propose angle- and norm-oriented contrastive losses to maximize the
angle size and norm difference between normal and abnormal features simulta-
neously. To avoid the distribution shift from natural images to AD images, our
pretraining is performed on a large-scale AD dataset, RealIAD. To further alle-
viate the potential shift between pretraining data and downstream AD datasets,
we learn the pretrained AD representations based on the class-generalizable repre-
sentation, residual features. For evaluation, based on five embedding-based AD
methods, we simply replace their original features with our pretrained represen-
tations. Extensive experiments on five AD datasets and five backbones consis-
tently show the superiority of our pretrained features. The code is available at
https://github.com/xcyao00/ADPretrain. Xincheng Yao, Yan Luo 0003, Zefeng Qian |
NeurIPS | 2 |
| 2024 | IRAD: Input-Reference Joint Driven Reconstruction for Unified Anomaly DetectionabstractUnified (multi-class and cross-class) anomaly detection (AD) is a growing area of interest in real-world applications. However, the popular reconstruction-based AD approach usually faces two significant challenges: the "identical shortcut" issue (copying the input as output) and the lack of class adaptability (the AD model cannot be directly applied to new classes). To address these challenges, we propose a novel unified AD method, named IRAD (Input-Reference Joint Driven). Our core insight is to effectively incorporate both input and references into the reconstruction process. Our IRAD consists of three components: 1) A Suspicious Anomaly Substituting module that replaces the potential abnormal regions of input with anomaly-free reference patches to prevent abnormal information leakage, effectively addressing the "identical shortcut". 2) An Input-Reference Fusing module that merges reference embeddings with input, which urges the subsequent Decoder to effectively utilize the normal reference patterns to reconstruct anomaly-free samples, making our model more class-adaptive. 3) A Rich Feature Preserving Decoder that efficiently preserves low-level details, mitigating low-level information degradation during reverse construction from high to low level. In multi-class AD, IRAD achieves better results on Mvtec-AD, BTAD, and VisA. In cross-class AD, IRAD also outperforms the baesline methods on Mvtec-AD and VisA. Zixin Chen, Xincheng Yao, Yan Luo 0003, Baozhu Zhang |
VCIP | 3 |
| 2024 | Generative Representation and Discriminative Classification for Few-shot Open-set Object DetectionabstractOpen-Set Object Detection (OSOD) aims to train detectors on closed-set datasets to detect known objects and identify unknown objects in open-set conditions. Traditional discriminative classifier-based OSOD methods struggle to accurately learn the decision boundary between known and unknown classes, often resulting in the misclassification of unknown samples. In this work, we aim to combine generative representation with discriminative classification to alleviate the issue of misclassification by transforming known-unknown recognition into a binary classification problem. The proposed two-stage OSOD approach proceeds as follows: during the generative representation stage, we employ Class-Conditioned Normalizing Flow (CCNF) to establish distribution mapping for each known category; In the discriminative classification stage, by utilizing a small number of unknown class samples, semi-push-pull supervised learning and entropy contrast learning are used to separate known and unknown classes. Extensive experiments demonstrate that our method significantly enhances OSOD performance, evidenced by a 25.8%-28.6% reduction in the Wilderness Index and a decrease of 4391-8870 units in Absolute Open-Set Errors on the test set VOC-COCO-T1. Peixue Shen, Ruoqi Li, Yan Luo 0003, Yiru Zhao |
VCIP | 3 |
| 2024 | S²CNet: Semantic and Structure Completion Network for 3D Object DetectionabstractLiDAR has become one of the primary 3D object detection sensors in autonomous driving. However, due to the inherent sparsity of point clouds, certain objects exhibit structure incompleteness in occluded and distant areas, which hampers the accurate perception of objects in 3D space. To tackle this challenge, we propose Semantic and Structure Completion Network (S2CNet) for 3D object detection. Concretely, we design the Semantic Completion (SeC) module to generate semantic features in Bird’s-Eye-View (BEV) space, utilizing a teacher-student paradigm. Notably, we adopt a coarse-to-fine guidance strategy to encourage student network to generate semantic features specifically within foreground regions. This ensures that the student network focuses on the generation of foreground object features. Besides, we introduce an attention-based module to adaptively fuse the generated features and raw features. SeC module faces particular limitation when dealing with objects containing only a few points, in such case, the network is prone to generating low quality proposals with inaccurate localization. Complementary to SeC module, we introduce the Structure Completion (StC) module, in which a group of structural proposals are obtained by traversing most structures in a structure-guided manner, and thus at least one proposal with ground truth similar structure can be guaranteed. Extensive experiments on the KITTI and nuScenes benchmarks demonstrate the effectiveness of our method, especially for the hard setting objects with fewer points. Yan Luo 0003, Zefeng Qian, Muming Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Consistent GT-Proposal Assignment for Challenging Pedestrian DetectionabstractAccurate pedestrian classification and localization has garnered significant attention due to their extensive applications in various multimedia applications such as security monitoring, autonomous driving, and more. We have observed that the commonly employed Intersection over Union (IoU) metric in many pedestrian detectors is susceptible to an inconsistent GT-Proposal assignment issue. This issue arises when spatially adjacent proposals, which have highly similar features, are assigned to distinct ground-truth boxes, leading to confusion during the training process and an increased number of false positives during inference. To address this challenge, our work presents a novel algorithm namedDirectionalAssignmentStrategy (DAS). Firstly, in conjunction with depth distribution, our approach transforms the assignment metric from a two-dimensional (2D) view into a three-dimensional (3D) space, enabling the optimization of the regression head under the constraint of depth direction. Secondly, in contrast to the conventional IoU-basedone-to-oneassignment of one proposal to one ground-truth box, our method aims to establish a more reasoned matching between sets of proposals and ground-truth boxes. By doing so, the detector is less reliant on the setting of a specific threshold. Leveraging this strategy as a plug-in module within state-of-the-art pedestrian detectors, we demonstrate a notable improvement in performance. Yan Luo 0003, Muming Zhao, Jun Sun 0005, Guangtao Zhai |
IEEE Trans. Multim. | 1 |
| 2023 | Focus the Discrepancy: Intra- and Inter-Correlation Learning for Image Anomaly DetectionabstractHumans recognize anomalies through two aspects: larger patch-wise representation discrepancies and weaker patch-to-normal-patch correlations. However, the previous AD methods didn’t sufficiently combine the two complementary aspects to design AD models. To this end, we find that Transformer can ideally satisfy the two aspects as its great power in the unified modeling of patch-wise representations and patch-to-patch correlations. In this paper, we propose a novel AD framework: FOcus-the-Discrepancy (FOD), which can simultaneously spot the patch-wise, intra- and inter-discrepancies of anomalies. The major characteristic of our method is that we renovate the self-attention maps in transformers to Intra-Inter-Correlation (I2Correlation). The I2Correlation contains a two-branch structure to first explicitly establish intra-and inter-image correlations, and then fuses the features of two-branch to spotlight the abnormal patterns. To learn the intra- and inter-correlations adaptively, we propose the RBF-kernel-based target-correlations as learning targets for self-supervised learning. Besides, we introduce an entropy constraint strategy to solve the mode collapse issue in optimization and further amplify the normal-abnormal distinguishability. Extensive experiments on three unsupervised real-world AD benchmarks show the superior performance of our approach. Code will be available at https://github.com/xcyao00/FOD. Xincheng Yao, Ruoqi Li, Zefeng Qian, Yan Luo 0003 |
ICCV | 4 |
| 2022 | Structure Guided Proposal Completion for 3D Object Detection
Yan Luo 0003 |
ACCV (1) | 3 |
| 2022 | Out-of-Distribution Identification: Let Detector Tell Which I Am Not Sure
Ruoqi Li, Hao Zhou 0014, Yan Luo 0003 |
ECCV (10) | 5 |
| 2022 | Informed Patch Enhanced HyperGCN for skeleton-based action recognition
Ying Li 0016, Hao Zhou 0014, Yan Luo 0003, Chuanping Hu |
Inf. Process. Manag. | 5 |
| 2022 | Thinking Inside Uncertainty: Interest Moment Perception for Diverse Temporal GroundingabstractGiven a language query, temporal grounding task is to localize temporal boundaries of the described event in an untrimmed video. There is a long-standing challenge that multiple moments may be associated with one same video-query pair, termed label uncertainty. However, existing methods struggle to localize diverse moments due to the lack of multi-label annotations. In this paper, we propose a novel Diverse Temporal Grounding framework (DTG) to achieve diverse moment localization with only single-label annotations. By delving into the label uncertainty, we find the diverse moments retrieved tend to involve similar actions/objects, driving us to perceive these interest moments. Specifically, we construct soft multi-label through semantic similarity of multiple video-query pairs. These soft labels reveal whether multiple moments in the intra-videos contain similar verbs/nouns, thereby guiding interest moment generation. Meanwhile, we put forward a diverse moment regression network (DMRNet) to achieve multiple predictions in a single pass, where plausible moments are dynamically picked out from the interest moments for joint optimization. Moreover, we introduce new metrics that better reveal multi-output performance. Extensive experiments conducted on Charades-STA and ActivityNet Captions show that our method achieves state-of-the-art performance in terms of both standard and new metrics. Hao Zhou 0014, Yan Luo 0003, Chuanping Hu, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Sequential Attention-Based Distinct Part Modeling for Balanced Pedestrian DetectionabstractDespite pedestrian detectors having made significant progress by introducing convolutional neural networks, their performance still suffers degradation, especially in occlusion scenes with more false positives (FPs) and false negatives (FNs). To alleviate the problem, we propose a novel Sequential Attention-based Distinct Part Modeling (SA-DPM) for balanced pedestrian detection. It takes one step further in constructing more robust representation that supports detection with fewer FNs and FPs. Specifically, the Sequential Attention serves as one internal perception process that captures several distinct part areas step by step from each pedestrian proposal (full-body). Different from the previous either-or feature selection, the following Joint Learning attempts to seek a reasonable trade-off between part and full-body features, and combines both features for more accurate classification and regression. Evaluation on the widely used pedestrian datasets including Caltech and Citypersons shows that the proposed SA-DPM achieves promising performance for both non-occluded and occluded pedestrian detection tasks, especially on Caltech Heavy Occlusion set, which yields a new state-of-the-art miss rate by 30.18% and outperforms the second best detector by 6.32%. Yan Luo 0003, Weiyao Lin, Xiaokang Yang 0001, Jun Sun 0005 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Embracing Uncertainty: Decoupling and De-Bias for Robust Temporal GroundingabstractTemporal grounding aims to localize temporal boundaries within untrimmed videos by language queries, but it faces the challenge of two types of inevitable human uncertainties: query uncertainty and label uncertainty. The two uncertainties stem from human subjectivity, leading to limited generalization ability of temporal grounding. In this work, we propose a novel DeNet (Decoupling and Debias) to embrace human uncertainty: Decoupling — We explicitly disentangle each query into a relation feature and a modified feature. The relation feature, which is mainly based on skeleton-like words (including nouns and verbs), aims to extract basic and consistent information in the presence of query uncertainty. Meanwhile, modified feature assigned with style-like words (including adjectives, adverbs, etc) represents the subjective information, and thus brings personalized predictions; De-bias — We propose a de-bias mechanism to generate diverse predictions, aim to alleviate the bias caused by single-style annotations in the presence of label uncertainty. Moreover, we put forward new multi-label metrics to diversify the performance evaluation. Extensive experiments show that our approach is more effective and robust than state-of-the-arts on Charades-STA and ActivityNet Captions datasets. Hao Zhou 0014, Yan Luo 0003, Chuanping Hu |
CVPR | 3 |
| 2021 | Improving Visual Relationship Detection With Two-Stage Correlation ExploitationabstractVisual relationship detection, as a challenging task used to find and distinguish interactions between object-pairs in one image, has received much attention recently. In this work, we devise a unified visual relationship detection framework with two types of correlation exploitation to address the combination explosion problem in the object-pairs proposing stage and the non-exclusive label problem in the predicate recognition stage. In the object-pairs proposing stage, with the exploitation of relative location correlation between two objects in one pair, one location-embedded rating module (LRM) is developed to effectively select plausible proposals. In the predicate recognition stage, one label-correlation graph module (LGM) is introduced to measure the implicit semantic correlation among predicates; and then assign discrete distributed labels to predicates to improve the precision of top-n recall. Experiments on the two widely used VRD and VG datasets show that our proposed method outperforms current state-of-the-art methods. Hao Zhou 0014, Muming Zhao, Yan Luo 0003, Chuanping Hu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Comer-Line-Prediction based Water-tank Detection and LocalizationabstractWater tanks on the roof of buildings require regular labor-costing inspection, and object detection can be used to automate the task. Current detection frameworks have several drawbacks when they are applied: (1) The output horizontal rectangular boxes cannot provide arbitrary quadrilateral detection representations; (2) False positive results may easily appear when key-point based models are used. In this paper, we propose a novel detection framework: Corner-Line-Prediction, which generates tight quadrilateral detection results of the tank blocks. Our model is built on key point detection network to detect corner points precisely. And an original line predictor is integrated to recognize unique tank edges, such that numerous false positive detections can be suppressed. Experimental results show that our Corner-Line-Prediction (CLP) framework outperforms state- of-the-art detection algorithms in average-precision (AP) and produces better localization results, compared with mainstream general detection models. Haoning Chen, Yan Luo 0003, Bingkun Zhao, Jiahao Bao |
VCIP | 3 |
| 2019 | Improving Small-Scale Pedestrian Detection Using Informed ContextabstractFinding small objects is fundamentally challenging because there is little signal on the object to exploit. For the small-scale pedestrian detection, one must use image evidence beyond the pedestrian extent, which is often formulated as context. Unlike existing object detection methods that use adjacent regions or whole image as the context simply, we focus on more informed contexts exploiting and utilizing to improve small-scale pedestrian detection: firstly, one relationship network is developed to utilize the correlation among pedestrian instances in one image; secondly, two spatial regions, overhead area and feet bottom area, are taken as spatial context to exploit the relevance between pedestrian and scenes; at last, GRU [7] (Gated Recurrent Units) modules are introduced to take encoded contexts as input to guide the feature selection and fusion of each proposal. Instead of getting all of the outputs at once, we also iterate twice to refine the detection incrementally. Comprehensive experiments on Caltech Pedestrian [8] and SJTU-SPID [9] datasets, indicate that, with more informed context, the detection performance can be improved significantly, especially for the small-scale pedestrians. Zexiang Liu, Yan Luo 0003, Qiping Zhou, Yunyu Lai |
VCIP | 3 |