VLDB 2026 Research / reviewers in the wild / expert
Feng Xue 0001
dblp:03/517-1
· DBLP profile ↗
19ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-4101-3401ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Direction-aware deep policy learning for efficient capacitated arc routing
Feng Xue 0001, Runze Guo, Anlong Ming, Nicu Sebe |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation With Mutual Scoring of Unlabeled SamplesabstractZero-shot anomaly classification (AC) and segmentation (AS) methods aim to identify and outline defects without using any labeled samples. In this paper, we reveal a key property that is overlooked by existing methods: normal image patches across industrial products typically find many other similar patches, not only in 2D appearance but also in 3D shapes, while anomalies remain diverse and isolated. To explicitly leverage this discriminative property, we propose a Mutual Scoring framework (MuSc-V2) for zero-shot AC/AS, which flexibly supports single 2D/3D or multimodality. Specifically, our method begins by improving 3D representation through Iterative Point Grouping (IPG), which reduces false positives from discontinuous surfaces. Then we use Similarity Neighborhood Aggregation with Multi-Degrees (SNAMD) to fuse 2D/3D neighborhood cues into more discriminative multi-scale patch features for mutual scoring. The core comprises a Mutual Scoring Mechanism (MSM) that lets samples within each modality to assign score to each other, and Cross-modal Anomaly Enhancement (CAE) that fuses 2D and 3D scores to recover modality-specific missing anomalies. Finally, Re-scoring with Constrained Neighborhood (RsCon) suppresses false classification based on similarity to more representative samples. Our framework flexibly works on both the full dataset and smaller subsets with consistently robust performance, ensuring seamless adaptability across diverse product lines. In aid of the novel framework, MuSc-V2 achieves significant performance improvements: a $\mathbf{+23.7\%}$+23.7% AP gain on the MVTec 3D-AD dataset and a $\mathbf{+19.3\%}$+19.3% boost on the Eyecandies dataset, surpassing previous zero-shot benchmarks and even outperforming most few-shot methods. Feng Xue 0001, Yu Zhou 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | AnomalyNCD: Towards Novel Anomaly Class Discovery in Industrial ScenariosabstractRecently, multi-class anomaly classification has garnered increasing attention. Previous methods directly cluster anomalies but often struggle due to the lack of anomaly-prior knowledge. Acquiring this knowledge faces two issues: the non-prominent and weak-semantics anomalies. In this paper, we propose AnomalyNCD, a multi-class anomaly classification network compatible with different anomaly detection methods. To address the non-prominence of anomalies, we design main element binarization (MEBin) to obtain anomaly-centered images, ensuring anomalies are learned while avoiding the impact of incorrect detections. Next, to learn anomalies with weak semantics, we design mask-guided representation learning, which focuses on isolated anomalies guided by masks and reduces confusion from erroneous inputs through corrected pseudo labels. Finally, to enable flexible classification at both region and image levels, we develop a region merging strategy that determines the overall image category based on the classified anomaly regions. Our method outperforms the state-of-the-art works on the MVTec AD and MTD datasets. Compared with the current methods, AnomalyNCD combined with zero-shot anomaly detection method achieves a 10.8% F1gain, 8.8% NMI gain, and 9.5% ARI gain on MVTec AD, and 12.8% F1gain, 5.7% NMI gain, and 10.8% ARI gain on MTD. Code is available at https://github.com/HUST-SLOW/AnomalyNCD. Ziming Huang, Feng Xue 0001, Yu Zhou 0016 |
CVPR | 4 |
| 2025 | SeaS: Few-Shot Industrial Anomaly Image Generation with Separation and Sharing Fine-TuningabstractWe introduce SeaS, a unified industrial generative model for automatically creating diverse anomalies, authentic normal products, and precise anomaly masks. While extensive research exists, most efforts either focus on specific tasks, i.e., anomalies or normal products only, or require separate models for each anomaly type. Consequently, prior methods either offer limited generative capability or depend on a vast array of anomaly-specific models. We demonstrate that U-Net's differentiated learning ability captures the distinct visual traits of slightly-varied normal products and diverse anomalies, enabling us to construct a unified model for all tasks. Specifically, we first introduce an Unbalanced Abnormal (UA) Text Prompt, comprising one normal token and multiple anomaly tokens. More importantly, our Decoupled Anomaly Alignment (DA) loss decouples anomaly attributes and binds them to distinct anomaly tokens of UA, enabling SeaS to create unseen anomalies by recombining these attributes. Furthermore, our Normal-image Alignment (NA) loss aligns the normal token to normal patterns, making generated normal products globally consistent and locally varied. Finally, SeaS produces accurate anomaly masks by fusing discriminative U-Net features with high-resolution VAE features. SeaS sets a new benchmark for industrial generation, significantly enhancing downstream applications, with average improvements of $+8.66\%$ pixel-level AP for synthesis-based AD approaches, $+1.10\%$ image-level AP for unsupervised AD methods, and $+12.79\%$ IoU for supervised segmentation models. Code is available at \href{https://github.com/HUST-SLOW/SeaS}{https://github.com/HUST-SLOW/SeaS}. Zhewei Dai, Shilei Zeng, Feng Xue 0001, Yu Zhou 0016 |
ICCV | 5 |
| 2025 | Superpowering Open-Vocabulary Object Detectors for X-ray Vision
Pablo Garcia-Fernandez, Lorenzo Vaquero, Feng Xue 0001, Daniel Cores, Nicu Sebe, Manuel Mucientes, Elisa Ricci 0001 |
ICCV | 4 |
| 2025 | Evidence-Based Real-Time Road Segmentation With RGB-D Data AugmentationabstractDespite significant progress in RGB-D based road segmentation in recent years, the latest methods cannot achieve both state-of-the-art accuracy and real time due to the high-performance reliance on heavy structures. We argue that this reliance is due to unsuitable multimodal fusion. To be specific, RGB and depth data in road scenes are each sensitive to different regions, but current RGB-D based road segmentation methods generally combine features within sensitive regions which preserves false road representation from one of the data. Based on such findings, we design an Evidence-based Road Segmentation Method (Evi-RoadSeg), which incorporates prior knowledge of the modal-specific characteristics. Firstly, we abandon the cross-modal fusion operation commonly used in existing multimodal based methods. Instead, we collect the road evidence from RGB and depth inputs separately via two low-latency subnetworks, and fuse the road representation of the two subnetworks by taking both modalities’ evidence as a measure of confidence. Secondly, we propose an RGB-D data augmentation scheme tailored to road scenes to enhance the unique properties of RGB and depth data. It facilitates learning by adding more sensitive regions to the samples. Finally, the proposed method is evaluated on the widely used KITTI-road, ORFD, and R2D datasets. Our method achieves state-of-the-art accuracy at over 70 FPS, 5$\times$faster than comparable RGB-D methods. Furthermore, extensive experiments illustrate that our method can be deployed on a Jetson Nano 2GB with a speed of 8$+$FPS. The code will be released in https://github.com/xuefeng-cvr/Evi-RoadSeg. Feng Xue 0001, Yicong Chang, Wenzhuang Xu, Wenteng Liang, Fei Sheng, Anlong Ming |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | MuSc: Zero-Shot Industrial Anomaly Classification and Segmentation with Mutual Scoring of the Unlabeled ImagesabstractThis paper studies zero-shot anomaly classification (AC) and segmentation (AS) in industrial vision.
We reveal that the abundant normal and abnormal cues implicit in unlabeled test images can be exploited for anomaly determination, which is ignored by prior methods.
Our key observation is that for the industrial product images, the normal image patches could find a relatively large number of similar patches in other unlabeled images,
while the abnormal ones only have a few similar patches.
We leverage such a discriminative characteristic to design a novel zero-shot AC/AS method by Mutual Scoring (MuSc) of the unlabeled images,
which does not need any training or prompts.
Specifically, we perform Local Neighborhood Aggregation with Multiple Degrees (LNAMD) to obtain the patch features that are capable of representing anomalies in varying sizes.
Then we propose the Mutual Scoring Mechanism (MSM) to leverage the unlabeled test images to assign the anomaly score to each other.
Furthermore, we present an optimization approach named Re-scoring with Constrained Image-level Neighborhood (RsCIN) for image-level anomaly classification to suppress the false positives caused by noises in normal images.
The superior performance on the challenging MVTec AD and VisA datasets demonstrates the effectiveness of our approach.
Compared with the state-of-the-art zero-shot approaches,
MuSc achieves a $\textbf{21.1}$% PRO absolute gain (from 72.7\% to 93.8\%) on MVTec AD, a $\textbf{19.4}$% pixel-AP gain and a $\textbf{14.7}$% pixel-AUROC gain on VisA.
In addition, our zero-shot approach outperforms most of the few-shot approaches and is comparable to some one-class methods.
Code is available at https://github.com/xrli-U/MuSc. Ziming Huang, Feng Xue 0001, Yu Zhou 0016 |
ICLR | 3 |
| 2024 | SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One ModelabstractIn the last year, universal monocular metric depth estimation (universal MMDE) has gained considerable attention, serving as the foundation model for various multimedia tasks, such as video and image editing. Nonetheless, current approaches face challenges in maintaining consistent accuracy across diverse scenes without scene-specific parameters and pre-training, hindering the practicality of MMDE. Furthermore, these methods rely on extensive datasets comprising millions, if not tens of millions, of data for training, leading to significant time and hardware expenses. This paper presents SM4Depth, a model that seamlessly works for both indoor and outdoor scenes, without needing extensive training data and GPU clusters. Firstly, to obtain consistent depth across diverse scenes, we propose a novel metric scale modeling, i.e., variation- based unnormalized depth bins. It reduces the ambiguity of the conventional metric bins and enables better adaptation to large depth gaps of scenes during training. Secondly, we propose a ''divide and conquer'' solution to reduce reliance on massive training data. Instead of estimating directly from the vast solution space, the metric bins are estimated from multiple solution sub-spaces to reduce complexity. Additionally, we introduce an uncut depth dataset, BUPT Depth, to evaluate the depth accuracy and consistency across various indoor and outdoor scenes. Trained on a consumer-grade GPU using just 150K RGB-D pairs, SM4Depth achieves outstanding performance on the most never-before-seen datasets, especially maintaining consistent accuracy across indoors and outdoors. The code can be found here. Feng Xue 0001, Anlong Ming, Mingshuai Zhao, Huadong Ma, Nicu Sebe |
ACM Multimedia | 2 |
| 2024 | Indoor Obstacle Discovery on Reflective Ground via Monocular Camera
Feng Xue 0001, Yicong Chang, Tianxi Wang, Yu Zhou 0016, Anlong Ming |
Int. J. Comput. Vis. | 1 |
| 2024 | Exploiting Low-Level Representations for Ultra-Fast Road SegmentationabstractAchieving real-time and accuracy on embedded platforms has always been the pursuit of road segmentation methods. To this end, they have proposed many lightweight networks. However, they ignore the fact that roads are “stuff” (background or environmental elements) rather than “things” (specific identifiable objects), which inspires us to explore the feasibility of representing roads with low-level instead of high-level features. Surprisingly, we find that the primary stage of mainstream network models is sufficient to represent most pixels of the road for segmentation. Motivated by this, we propose a Low-level Feature Dominated Road Segmentation network (LFD-RoadSeg). Specifically, LFD-RoadSeg employs a bilateral structure. The spatial detail branch is firstly designed to extract low-level feature representation for the road by the first stage of ResNet-18. To suppress texture-less regions mistaken as the road in the low-level feature, the context semantic branch is then designed to extract the context feature in a fast manner. To this end, in the second branch, we asymmetrically downsample the input image and design an aggregation module to achieve comparable receptive fields to the third stage of ResNet-18 but with less time consumption. Finally, to segment the road from the low-level feature, a selective fusion module is proposed to calculate pixel-wise attention between the low-level representation and context feature, and suppress the non-road low-level response by this attention. On KITTI-Road, LFD-RoadSeg achieves a maximum F1-measure (MaxF) of 95.21% and an average precision of 93.71%, while reaching 238 FPS on a single TITAN Xp and 54 FPS on a Jetson TX2, all with a compact model size of just 936k parameters. The source code is available at https://github.com/zhouhuan-hust/LFD-RoadSeg. Feng Xue 0001, Yucong Li, Shi Gong, Yu Zhou 0016 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Unknown Sniffer for Object Detection: Don't Turn a Blind Eye to Unknown ObjectsabstractThe recently proposed open-world object and open-set detection have achieved a breakthrough in finding never-seen-before objects and distinguishing them from known ones. However, their studies on knowledge transfer from known classes to unknown ones are not deep enough, resulting in the scanty capability for detecting unknowns hidden in the background. In this paper, we propose the unknown sniffer (UnSniffer) to find both unknown and known objects. Firstly, the generalized object confidence (GOC) score is introduced, which only uses known samples for supervision and avoids improper suppression of unknowns in the back-ground. Significantly, such confidence score learned from known objects can be generalized to unknown ones. Additionally, we propose a negative energy suppression loss to further suppress the non-object samples in the background. Next, the best box of each unknown is hard to obtain during inference due to lacking their semantic information in training. To solve this issue, we introduce a graph-based determination scheme to replace hand-designed non-maximum suppression (NMS) post-processing. Finally, we present the Unknown Object Detection Benchmark, the first publicly benchmark that encompasses precision evaluation for unknown detection to our knowledge. Experiments show that our method is far better than the existing state-of-the-art methods. Code is available at: https://github.com/Went-Liang/UnSniffer. Wenteng Liang, Feng Xue 0001, Guofeng Zhong, Anlong Ming |
CVPR | 2 |
| 2023 | MARF: Multiscale Adaptive-Switch Random Forest for Leg Detection With 2-D Laser ScannersabstractFor the 2-D laser-based tasks, e.g., people detection and people tracking, leg detection is usually the first step. Thus, it carries great weight in determining the performance of people detection and people tracking. However, many leg detectors ignore the inevitable noise and the multiscale characteristics of the laser scan, which makes them sensitive to the unreliable features of point cloud and further degrades the performance of the leg detector. In this article, we propose a multiscale adaptive-switch random forest (MARF) to overcome these two challenges. First, the adaptive-switch decision tree is designed to use noise-sensitive features to conduct weighted classification and noise-invariant features to conduct binary classification, which makes our detector perform more robust to noise. Second, considering the multiscale property that the sparsity of the 2-D point cloud is proportional to the length of laser beams, we design a multiscale random forest structure to detect legs at different distances. Moreover, the proposed approach allows us to discover a sparser human leg from point clouds than others. Consequently, our method shows an improved performance compared to other state-of-the-art leg detectors on the challenging Moving Legs dataset and retains the entire pipeline at a speed of 60+ FPS on low-computational laptops. Moreover, we further apply the proposed MARF to the people detection and tracking system, achieving a considerable gain in all metrics. Tianxi Wang, Feng Xue 0001, Yu Zhou 0016, Anlong Ming |
IEEE Trans. Cybern. | 2 |
| 2022 | Fast Road Segmentation via Uncertainty-aware Symmetric NetworkabstractThe high performance of RGB-D based road segmentation methods contrasts with their rare application in commercial autonomous driving, which is owing to two reasons: 1) the prior methods cannot achieve high inference speed and high accuracy in both ways; 2) the different properties of RGB and depth data are not well-exploited, limiting the reliability of predicted road. In this paper, based on the evidence theory, an uncertainty-aware symmetric network (USNet) is proposed to achieve a trade-off between speed and accuracy by fully fusing RGB and depth data. Firstly, cross-modal feature fusion operations, which are indispensable in the prior RGB-D based methods, are abandoned. We instead separately adopt two light-weight subnetworks to learn road representations from RGB and depth inputs. The light-weight structure guarantees the real-time inference of our method. Moreover, a multi-scale evidence collection (MEC) module is designed to collect evidence in multiple scales for each modality, which provides sufficient evidence for pixel class determination. Finally, in uncertainty-aware fusion (UAF) module, the uncertainty of each modality is perceived to guide the fusion of the two sub-networks. Experimental results demonstrate that our method achieves a state-of-the-art accuracy with real-time inference speed of$43+$FPS. The source code is available at https://github.com/morancyc/USNet. Yicong Chang, Feng Xue 0001, Fei Sheng, Wenteng Liang, Anlong Ming |
ICRA | 2 |
| 2022 | Monocular Depth Distribution Alignment with Low ComputationabstractThe performance of monocular depth estimation generally depends on the amount of parameters and computational cost. It leads to a large accuracy contrast between light-weight networks and heavy-weight networks, which limits their application in the real world. In this paper, we model the majority of accuracy contrast between them as the difference of depth distribution, which we call 'Distribution drift'. To this end, a distribution alignment network (DANet) is proposed. We firstly design a pyramid scene transformer (PST) module to capture inter-region interaction in multiple scales. By perceiving the difference of depth features between every two regions, DANet tends to predict a reasonable scene structure, which fits the shape of distribution to ground truth. Then, we propose a local-global optimization (LGO) scheme to realize the supervision of global range of scene depth. Thanks to the alignment of depth distribution shape and scene depth range, DANet sharply alleviates the distribution drift, and achieves a comparable performance with prior heavy-weight methods, but uses only 1% floating-point operations per second (FLOPs) of them. The experiments on two datasets, namely the widely used NYUDv2 dataset and the more challenging iBims-1 dataset, demonstrate the effectiveness of our method. The source code is available at https://github.com/YiLiM1/DANet. Fei Sheng, Feng Xue 0001, Yicong Chang, Wenteng Liang, Anlong Ming |
ICRA | 2 |
| 2022 | FastRoadSeg: Fast Monocular Road Segmentation NetworkabstractMonocular road segmentation is a fundamental component of autonomous driving. High accuracy has always been a focus of the community for this task. However, the exploration of high-speed methods is still lacking. This limits the prevalence of the monocular road segmentation method. In this paper, an efficient encoder-decoder network is proposed for segmenting the road from a single image. Specifically, to achieve high-speed inference, we leverage a shallow encoder for feature extraction and design a lightweight decoder for feature resolution recovery. To avoid the accuracy drop caused by network simplification, we introduce an efficient asymmetric dilated block to improve the ability of features to distinguish drivable/nondrivable roads to maintain the performance without excess computation. Experiments on three datasets verify the effectiveness of our method. On KITTI-Road, our approach ranks second among all monocular methods and even achieves a competitive performance compared to multimodal methods, with an inference speed over 135 FPS on a single TITAN Xp, which is at least 20 times faster than the state-of-the-art monocular method. Furthermore, our approach can be deployed on the embedded platform Jetson TX2 and achieves a speed of over 80 FPS. Codes will be available athttps://github.com/gongshichina/FastRoadSeg. Shi Gong, Feng Xue 0001, Cong Fang 0005, Yu Zhou 0016 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Boundary-induced and scene-aggregated network for monocular depth prediction
Feng Xue 0001, Junfeng Cao, Yu Zhou 0016, Fei Sheng, Anlong Ming |
Pattern Recognit. | 1 |
| 2020 | Tiny Obstacle Discovery by Occlusion-Aware Multilayer RegressionabstractEdges are the fundamental visual element for discovering tiny obstacles using a monocular camera. Nevertheless, tiny obstacles often have weak and inconsistent edge cues due to various properties such as small size and similar appearance to the free space, making it hard to capture them. To this end, we propose an occlusion-based multilayer approach, which specifies the scene prior as multilayer regions and utilizes these regions in each obstacle discovery module, i.e., edge detection and proposal extraction. Firstly, an obstacle-aware occlusion edge is generated to accurately capture the obstacle contour by fusing the edge cues inside all the multilayer regions, which intensifies the object characteristics of these obstacles. Then, a multistride sliding window strategy is proposed for capturing proposals that enclose the tiny obstacles as completely as possible. Moreover, a novel obstacle-aware regression model is proposed for effectively discovering obstacles. It is formed by a primary-secondary regressor, which can learn two dissimilarities between obstacles and other categories separately, and eventually generate an obstacle-occupied probability map. The experiments are conducted on two datasets to demonstrate the effectiveness of our approach under different scenarios. And the results show that the proposed method can approximately improve accuracy by 19% over FPHT and PHT, and achieves comparable performance to MergeNet. Furthermore, multiple experiments with different variants validate the contribution of our method. The source code is available at https://github.com/XuefengBUPT/TOD_OMR. Feng Xue 0001, Anlong Ming, Yu Zhou 0016 |
IEEE Trans. Image Process. | 1 |
| 2019 | Occlusion-Shared and Feature-Separated Network for Occlusion Relationship ReasoningabstractOcclusion relationship reasoning demands closed contour to express the object, and orientation of each contour pixel to describe the order relationship between objects. Current CNN-based methods neglect two critical issues of the task: (1) simultaneous existence of the relevance and distinction for the two elements, i.e, occlusion edge and occlusion orientation; and (2) inadequate exploration to the orientation features. For the reasons above, we propose the Occlusion-shared and Feature-separated Network (OFNet). On one hand, considering the relevance between edge and orientation, two sub-networks are designed to share the occlusion cue. On the other hand, the whole network is split into two paths to learn the high semantic features separately. Moreover, a contextual feature for orientation prediction is extracted, which represents the bilateral cue of the foreground and background areas. The bilateral cue is then fused with the occlusion cue to precisely locate the object regions. Finally, a stripe convolution is designed to further aggregate features from surrounding scenes of the occlusion edge. The proposed OFNet remarkably advances the state-of-the-art approaches on PIOD and BSDS ownership dataset. Feng Xue 0001, Menghan Zhou, Anlong Ming, Yu Zhou 0016 |
ICCV | 2 |
| 2019 | A Novel Multi-layer Framework for Tiny Obstacle DiscoveryabstractFor tiny obstacle discovery in a monocular image, edge is a fundamental visual element. Nevertheless, because of various reasons, e.g., noise and similar color distribution with background, it is still difficult to detect the edges of tiny (b) obstacles at long distance. In this paper, we propose an obstacle-aware discovery method to recover the missing contours of these obstacles, which helps to obtain obstacle proposals as much as possible. First, by using visual cues in monocular images, several multi-layer regions are elaborately inferred to reveal the distances from the camera. Second, several novel obstacle-aware occlusion edge maps are constructed to well capture the contours of tiny obstacles, which combines cues from each layer. Third, to ensure the existence of the tiny obstacle proposals, the maps from all layers are used for proposals extraction. Finally, based on these proposals containing tiny obstacles, a novel obstacle-aware regressor is proposed to generate an obstacle occupied probability map with high confidence. The convincing experimental results with comparisons on the Lost and Found dataset demonstrate the effectiveness of our approach, achieving around 9.5% improvement on the accuracy than FPHT and PHT, it even gets comparable performance to MergeNet. Moreover, our method outperforms the state-of-the-art algorithms and significantly improves the discovery ability for tiny obstacles at long distance. Feng Xue 0001, Anlong Ming, Menghan Zhou, Yu Zhou 0016 |
ICRA | 1 |