VLDB 2026 Research / reviewers in the wild / expert
Lu Zhang 0054
dblp:82/10609-54
· DBLP profile ↗
19ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-6240-5300ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CARD: Control-Driven Autoregressive Reconstruction with Decoupled Learning for Multi-Class Anomaly DetectionabstractMulti-class unsupervised anomaly detection (UAD) is challenging due to the difficulty of harmonizing distributional differences across categories within a unified framework. While recent diffusion-based methods have demonstrated promising performance by leveraging denoising processes for anomaly reconstruction, the lack of explicit causal constraints limits their ability to handle complex logical inconsistencies. Inspired by the success of autoregressive models in enforcing local-to-global logical consistency, we propose a Control-driven Autoregressive Reconstruction with Decoupled learning (CARD) framework for multi-class UAD. It first tokenizes images using vector quantization and employs a vision autoregressive model to capture causal dependencies within normal patterns as explicit prior knowledge. Then, we introduce a Control-Driven Reconstruction (CDR) network, which aligns input features as control signals into the frozen autoregressive model to adjust the predicted distribution, enabling a generative reconstruction process. Additionally, we apply perturbations to the CDR input to simulate anomalous conditions, facilitating the model to correct out-of-distribution anomaly features under prior knowledge constraints. By decoupling normal pattern learning from reconstruction, CARD prevents identity mapping caused by forgetting implicit priors in the conventional reconstruction-based method. Comprehensive experiments on several benchmark datasets validate the effectiveness of our approach. On the MVTecAD dataset, CARD achieved AUROC scores of 98.7% at the image level and 98.5% at the pixel level. Mingqing Wang, Boyi Sun, Qianfan Zhao, Lu Zhang 0054, Zhiyong Liu 0001, Xu Yang 0004, Suiwu Zheng |
MMAsia | 5 |
| 2025 | Weakly Aligned Feature Fusion for Multimodal Object DetectionabstractTo achieve accurate and robust object detection in the real-world scenario, various forms of images are incorporated, such as color, thermal, and depth. However, multimodal data often suffer from the position shift problem, i.e., the image pair is not strictly aligned, making one object has different positions in different modalities. For the deep learning method, this problem makes it difficult to fuse multimodal features and puzzles the convolutional neural network (CNN) training. In this article, we propose a general multimodal detector named aligned region CNN (AR-CNN) to tackle the position shift problem. First, a region feature (RF) alignment module with adjacent similarity constraint is designed to consistently predict the position shift between two modalities and adaptively align the cross-modal RFs. Second, we propose a novel region of interest (RoI) jitter strategy to improve the robustness to unexpected shift patterns. Third, we present a new multimodal feature fusion method that selects the more reliable feature and suppresses the less useful one via feature reweighting. In addition, by locating bounding boxes in both modalities and building their relationships, we provide novel multimodal labeling named KAIST-Paired. Extensive experiments on 2-D and 3-D object detection, RGB-T, and RGB-D datasets demonstrate the effectiveness and robustness of our method. Lu Zhang 0054, Zhiyong Liu 0001, Xiangyu Zhu 0001, Zhan Song, Xu Yang 0004, Zhen Lei 0001, Hong Qiao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Enhancing class-incremental object detection in remote sensing through instance-aware distillation
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001 |
Neurocomputing | 2 |
| 2024 | Active domain adaptation for semantic segmentation via dynamically balancing domainness and uncertainty
Lu Zhang 0054, Zhiyong Liu 0001 |
Image Vis. Comput. | 2 |
| 2024 | Automatically Discovering Novel Visual Categories With Adaptive Prototype LearningabstractThis article targets the task of novel category discovery (NCD), which aims to discover unknown categories when a certain number of classes are already known. The NCD task is challenging due to its closeness to real-world scenarios, where we have only encountered some partial classes and corresponding images. Unlike previous approaches to NCD, we propose a novel adaptive prototype learning method that leverages prototypes to emphasize category discrimination and alleviate the issue of missing annotations for novel classes. Concretely, the proposed method consists of two main stages: prototypical representation learning and prototypical self-training. In the first stage, we develop a robust feature extractor that could effectively handle images from both base and novel categories. This ability of instance and category discrimination of the feature extractor is boosted by self-supervised learning and adaptive prototypes. In the second stage, we utilize the prototypes again to rectify offline pseudo labels and train a final parametric classifier for category clustering. We conduct extensive experiments on four benchmark datasets, demonstrating our method's effectiveness and robustness with state-of-the-art performance. Lu Zhang 0054, Lu Qi 0001, Xu Yang 0004, Hong Qiao, Ming-Hsuan Yang 0001, Zhiyong Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Refined Pseudo Labeling for Source-Free Domain Adaptive Object DetectionabstractDomain adaptive object detection (DAOD) assumes that both labeled source data and unlabeled target data are available for training, but this assumption does not always hold in real-world scenarios. Thus, source-free DAOD is proposed to adapt the source-trained detectors to target domains with only unlabeled target data. Existing source-free DAOD methods typically utilize pseudo labeling, where the performance heavily relies on the selection of confidence threshold. However, most prior works adopt a single fixed threshold for all classes to generate pseudo labels, which ignore the imbalanced class distribution, resulting in biased pseudo labels. In this work, we propose a refined pseudo labeling framework for source-free DAOD. First, to generate unbiased pseudo labels, we present a category-aware adaptive threshold estimation module, which adaptively provides the appropriate threshold for each category. Second, to alleviate incorrect box regression, a localization-aware pseudo label assignment strategy is introduced to divide labels into certain and uncertain ones and optimize them separately. Finally, extensive experiments on four adaptation tasks demonstrate the effectiveness of our method. Lu Zhang 0054, Zhiyong Liu 0001 |
ICASSP | 2 |
| 2023 | Unseen Object Instance Segmentation with Fully Test-time RGB-D Embeddings AdaptationabstractSegmenting unseen objects is a crucial ability for the robot since it may encounter new environments during the operation. Recently, a popular solution is leveraging RGB-D features of large-scale synthetic data and directly applying the model to unseen real-world scenarios. However, the domain shift caused by the sim2real gap is inevitable, posing a crucial challenge to the segmentation model. In this paper, we em-phasize the adaptation process across sim2real domains and model it as a learning problem on the BatchNorm param-eters of a simulation-trained model. Specifically, we propose a novel non-parametric entropy objective, which formulates the learning objective for the test-time adaptation in an open-world manner. Then, a cross-modality knowledge distillation objective is further designed to encourage the test-time knowledge transfer for feature enhancement. Our approach can be efficiently implemented with only test images, without requiring annotations or revisiting the large-scale synthetic training data. Besides significant time savings, the proposed method consistently improves segmentation results on the overlap and boundary metrics, achieving state-of-the-art performance on unseen object instance segmentation. Lu Zhang 0054, Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001 |
ICRA | 1 |
| 2023 | Zero-Shot Object Goal Visual NavigationabstractObject goal visual navigation is a challenging task that aims to guide a robot to find the target object based on its visual observation, and the target is limited to the classes pre-defined in the training stage. However, in real households, there may exist numerous target classes that the robot needs to deal with, and it is hard for all of these classes to be contained in the training stage. To address this challenge, we study the zero-shot object goal visual navigation task, which aims at guiding robots to find targets belonging to novel classes without any training samples. To this end, we also propose a novel zero-shot object navigation framework called semantic similarity network (SSNet). Our framework use the detection results and the cosine similarity between semantic word embeddings as input. Such type of input data has a weak correlation with classes and thus our framework has the ability to generalize the policy to novel classes. Extensive experiments on the AI2-THOR platform show that our model outperforms the baseline models in the zero-shot object navigation task, which proves the generalization ability of our model. Our code is available at: https://github.com/pioneer-innovation/Zero-Shot-Object-Navigation. Qianfan Zhao, Lu Zhang 0054, Bin He 0003, Hong Qiao, Zhiyong Liu 0001 |
ICRA | 2 |
| 2023 | Incremental Few-Shot Object Detection with scale- and centerness-aware weight generation
Lu Zhang 0054, Xu Yang 0004, Lu Qi 0001, Shaofeng Zeng, Zhiyong Liu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2023 | Frequency-based pseudo-domain generation for domain generalizable object detection
Lu Zhang 0054, Zhiyong Liu 0001 |
Neurocomputing | 2 |
| 2023 | RTDOD: A large-scale RGB-thermal domain-incremental object detection dataset for UAVs
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001 |
Image Vis. Comput. | 2 |
| 2022 | FIT: Frequency-Based Image Translation for Domain Adaptive Object Detection
Lu Zhang 0054, Zhiyong Liu 0001, Hangtao Feng |
ICONIP (3) | 2 |
| 2022 | Incremental few-shot object detection via knowledge transfer
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001 |
Pattern Recognit. Lett. | 2 |
| 2020 | GSDCN: A Customized Two-Stage Neural Network for Benthonic Organism Detection
Zhaoliang Wan, Lu Zhang 0054, Hai Huang 0004, Xu Yang 0004 |
ICONIP (2) | 2 |
| 2020 | Knowledge-Experience Graph with Denoising Autoencoder for Zero-Shot Learning in Visual Cognitive Development
Xu Yang 0004, Zhiyong Liu 0001, Lu Zhang 0054, Dongchun Ren, Mingyu Fan |
ICONIP (5) | 4 |
| 2020 | MixedFusion: 6D Object Pose Estimation from Decoupled RGB-Depth FeaturesabstractEstimating the 6D pose of objects is an important process for intelligent systems to achieve interaction with the real-world. As the RGB-D sensors become more accessible, the fusion-based methods have prevailed, since the point clouds provide complementary geometric information with RGB values. However, due to the difference in feature space between color image and depth image, the network structures that directly perform point-to-point matching fusion do not effectively fuse the features of the two. In this paper, we propose a simple but effective approach, named MixedFusion. Different from the prior works, we argue that the spatial correspondence of color and point clouds could be decoupled and reconnected, thus enabling a more flexible fusion scheme. By performing the proposed method, more informative points can be mixed and fused with rich color features. Extensive experiments are conducted on the challenging LineMod and YCB-Video datasets, which shows that our method significantly boosts the performance without introducing extra overheads. Furthermore, when the minimum tolerance of metric narrows, the proposed approach performs better for the high-precision demands. Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001 |
ICPR | 2 |
| 2019 | Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not strictly aligned, making one object has different positions in different modalities. In deep learning based methods, this problem makes it difficult to fuse the feature maps from both modalities and puzzles the CNN training. In this paper, we propose a novel Aligned Region CNN (AR-CNN) to handle the weakly aligned multispectral data in an end-to-end way. Firstly, we design a Region Feature Alignment (RFA) module to capture the position shift and adaptively align the region features of the two modalities. Secondly, we present a new multimodal fusion method, which performs feature re-weighting to select more reliable features and suppress the useless ones. Besides, we propose a novel RoI jitter strategy to improve the robustness to unexpected shift patterns of different devices and system settings. Finally, since our method depends on a new kind of labelling: bounding boxes that match each modality, we manually relabel the KAIST dataset by locating bounding boxes in both modalities and building their relationships, providing a new KAIST-Paired Annotation. Extensive experimental validations on existing datasets are performed, demonstrating the effectiveness and robustness of the proposed method. Code and data are available at https://github.com/luzhang16/AR-CNN. Lu Zhang 0054, Xiangyu Zhu 0001, Xu Yang 0004, Zhen Lei 0001, Zhiyong Liu 0001 |
ICCV | 1 |
| 2019 | Faster R-CNN for marine organisms detection and recognition using data augmentation
Hai Huang 0004, Hao Zhou 0014, Xu Yang 0004, Lu Zhang 0054, Lu Qi 0001, Ai-Yun Zang |
Neurocomputing | 4 |
| 2018 | Single Shot Feature Aggregation Network for Underwater Object DetectionabstractThe rapidly developing ocean exploration and observation make the demand for underwater object detection become increasingly urgent. Recently, deep convolutional neural networks (CNN) have shown strong ability in feature representation and CNN-based detectors also achieve remarkable performance, but still facing the big challenge when detecting multi-scale objects in a complex underwater environment. To address this challenge, we propose a novel underwater object detector, introducing multiscale features and complementary context information for better classification and location ability. In the auto-grabbing contest of 2017 Underwater Robot Picking Contest sponsored by National Natural Science Foundation of China (NSFC), we won the 1-st place by using proposed method for real coastal underwater object detection. Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001, Lu Qi 0001, Hao Zhou 0014, Charles Chiu |
ICPR | 1 |