EDBT 2026 Demo / reviewers in the wild / expert
Yu Zhou 0016
dblp:36/2728-16
· DBLP profile ↗
44ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0002-6674-6484ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 19 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An All-in-One Quality Assessment Agent for 4D digital human: Bridging talking heads and animated human
Yingjie Zhou 0003, Farong Wen, Li Xu 0008, Yu Zhou 0016, Jiezhang Cao, Xiaohong Liu 0001, Xiongkuo Min, Yu Wang 0002, Guangtao Zhai |
Inf. Process. Manag. | 6 |
| 2026 | MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation With Mutual Scoring of Unlabeled SamplesabstractZero-shot anomaly classification (AC) and segmentation (AS) methods aim to identify and outline defects without using any labeled samples. In this paper, we reveal a key property that is overlooked by existing methods: normal image patches across industrial products typically find many other similar patches, not only in 2D appearance but also in 3D shapes, while anomalies remain diverse and isolated. To explicitly leverage this discriminative property, we propose a Mutual Scoring framework (MuSc-V2) for zero-shot AC/AS, which flexibly supports single 2D/3D or multimodality. Specifically, our method begins by improving 3D representation through Iterative Point Grouping (IPG), which reduces false positives from discontinuous surfaces. Then we use Similarity Neighborhood Aggregation with Multi-Degrees (SNAMD) to fuse 2D/3D neighborhood cues into more discriminative multi-scale patch features for mutual scoring. The core comprises a Mutual Scoring Mechanism (MSM) that lets samples within each modality to assign score to each other, and Cross-modal Anomaly Enhancement (CAE) that fuses 2D and 3D scores to recover modality-specific missing anomalies. Finally, Re-scoring with Constrained Neighborhood (RsCon) suppresses false classification based on similarity to more representative samples. Our framework flexibly works on both the full dataset and smaller subsets with consistently robust performance, ensuring seamless adaptability across diverse product lines. In aid of the novel framework, MuSc-V2 achieves significant performance improvements: a $\mathbf{+23.7\%}$+23.7% AP gain on the MVTec 3D-AD dataset and a $\mathbf{+19.3\%}$+19.3% boost on the Eyecandies dataset, surpassing previous zero-shot benchmarks and even outperforming most few-shot methods. Feng Xue 0001, Yu Zhou 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | AnomalyNCD: Towards Novel Anomaly Class Discovery in Industrial ScenariosabstractRecently, multi-class anomaly classification has garnered increasing attention. Previous methods directly cluster anomalies but often struggle due to the lack of anomaly-prior knowledge. Acquiring this knowledge faces two issues: the non-prominent and weak-semantics anomalies. In this paper, we propose AnomalyNCD, a multi-class anomaly classification network compatible with different anomaly detection methods. To address the non-prominence of anomalies, we design main element binarization (MEBin) to obtain anomaly-centered images, ensuring anomalies are learned while avoiding the impact of incorrect detections. Next, to learn anomalies with weak semantics, we design mask-guided representation learning, which focuses on isolated anomalies guided by masks and reduces confusion from erroneous inputs through corrected pseudo labels. Finally, to enable flexible classification at both region and image levels, we develop a region merging strategy that determines the overall image category based on the classified anomaly regions. Our method outperforms the state-of-the-art works on the MVTec AD and MTD datasets. Compared with the current methods, AnomalyNCD combined with zero-shot anomaly detection method achieves a 10.8% F1gain, 8.8% NMI gain, and 9.5% ARI gain on MVTec AD, and 12.8% F1gain, 5.7% NMI gain, and 10.8% ARI gain on MTD. Code is available at https://github.com/HUST-SLOW/AnomalyNCD. Ziming Huang, Feng Xue 0001, Yu Zhou 0016 |
CVPR | 6 |
| 2025 | SeaS: Few-Shot Industrial Anomaly Image Generation with Separation and Sharing Fine-TuningabstractWe introduce SeaS, a unified industrial generative model for automatically creating diverse anomalies, authentic normal products, and precise anomaly masks. While extensive research exists, most efforts either focus on specific tasks, i.e., anomalies or normal products only, or require separate models for each anomaly type. Consequently, prior methods either offer limited generative capability or depend on a vast array of anomaly-specific models. We demonstrate that U-Net's differentiated learning ability captures the distinct visual traits of slightly-varied normal products and diverse anomalies, enabling us to construct a unified model for all tasks. Specifically, we first introduce an Unbalanced Abnormal (UA) Text Prompt, comprising one normal token and multiple anomaly tokens. More importantly, our Decoupled Anomaly Alignment (DA) loss decouples anomaly attributes and binds them to distinct anomaly tokens of UA, enabling SeaS to create unseen anomalies by recombining these attributes. Furthermore, our Normal-image Alignment (NA) loss aligns the normal token to normal patterns, making generated normal products globally consistent and locally varied. Finally, SeaS produces accurate anomaly masks by fusing discriminative U-Net features with high-resolution VAE features. SeaS sets a new benchmark for industrial generation, significantly enhancing downstream applications, with average improvements of $+8.66\%$ pixel-level AP for synthesis-based AD approaches, $+1.10\%$ image-level AP for unsupervised AD methods, and $+12.79\%$ IoU for supervised segmentation models. Code is available at \href{https://github.com/HUST-SLOW/SeaS}{https://github.com/HUST-SLOW/SeaS}. Zhewei Dai, Shilei Zeng, Feng Xue 0001, Yu Zhou 0016 |
ICCV | 6 |
| 2025 | CDHQA: A Quality Assessment Database for Conversational Digital Human
Yingjie Zhou 0003, Yinghan Xia, Zhixiang Lu, Farong Wen, Yu Wang 0002, Yu Zhou 0016, Xiaohong Liu 0001, Xiongkuo Min, Jiezhang Cao, Guangtao Zhai |
ICIG (3) | 9 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 34 |
| 2024 | MuSc: Zero-Shot Industrial Anomaly Classification and Segmentation with Mutual Scoring of the Unlabeled ImagesabstractThis paper studies zero-shot anomaly classification (AC) and segmentation (AS) in industrial vision.
We reveal that the abundant normal and abnormal cues implicit in unlabeled test images can be exploited for anomaly determination, which is ignored by prior methods.
Our key observation is that for the industrial product images, the normal image patches could find a relatively large number of similar patches in other unlabeled images,
while the abnormal ones only have a few similar patches.
We leverage such a discriminative characteristic to design a novel zero-shot AC/AS method by Mutual Scoring (MuSc) of the unlabeled images,
which does not need any training or prompts.
Specifically, we perform Local Neighborhood Aggregation with Multiple Degrees (LNAMD) to obtain the patch features that are capable of representing anomalies in varying sizes.
Then we propose the Mutual Scoring Mechanism (MSM) to leverage the unlabeled test images to assign the anomaly score to each other.
Furthermore, we present an optimization approach named Re-scoring with Constrained Image-level Neighborhood (RsCIN) for image-level anomaly classification to suppress the false positives caused by noises in normal images.
The superior performance on the challenging MVTec AD and VisA datasets demonstrates the effectiveness of our approach.
Compared with the state-of-the-art zero-shot approaches,
MuSc achieves a $\textbf{21.1}$% PRO absolute gain (from 72.7\% to 93.8\%) on MVTec AD, a $\textbf{19.4}$% pixel-AP gain and a $\textbf{14.7}$% pixel-AUROC gain on VisA.
In addition, our zero-shot approach outperforms most of the few-shot approaches and is comparable to some one-class methods.
Code is available at https://github.com/xrli-U/MuSc. Ziming Huang, Feng Xue 0001, Yu Zhou 0016 |
ICLR | 4 |
| 2024 | Indoor Obstacle Discovery on Reflective Ground via Monocular Camera
Feng Xue 0001, Yicong Chang, Tianxi Wang, Yu Zhou 0016, Anlong Ming |
Int. J. Comput. Vis. | 4 |
| 2024 | Exploiting Low-Level Representations for Ultra-Fast Road SegmentationabstractAchieving real-time and accuracy on embedded platforms has always been the pursuit of road segmentation methods. To this end, they have proposed many lightweight networks. However, they ignore the fact that roads are “stuff” (background or environmental elements) rather than “things” (specific identifiable objects), which inspires us to explore the feasibility of representing roads with low-level instead of high-level features. Surprisingly, we find that the primary stage of mainstream network models is sufficient to represent most pixels of the road for segmentation. Motivated by this, we propose a Low-level Feature Dominated Road Segmentation network (LFD-RoadSeg). Specifically, LFD-RoadSeg employs a bilateral structure. The spatial detail branch is firstly designed to extract low-level feature representation for the road by the first stage of ResNet-18. To suppress texture-less regions mistaken as the road in the low-level feature, the context semantic branch is then designed to extract the context feature in a fast manner. To this end, in the second branch, we asymmetrically downsample the input image and design an aggregation module to achieve comparable receptive fields to the third stage of ResNet-18 but with less time consumption. Finally, to segment the road from the low-level feature, a selective fusion module is proposed to calculate pixel-wise attention between the low-level representation and context feature, and suppress the non-road low-level response by this attention. On KITTI-Road, LFD-RoadSeg achieves a maximum F1-measure (MaxF) of 95.21% and an average precision of 93.71%, while reaching 238 FPS on a single TITAN Xp and 54 FPS on a Jetson TX2, all with a compact model size of just 936k parameters. The source code is available at https://github.com/zhouhuan-hust/LFD-RoadSeg. Feng Xue 0001, Yucong Li, Shi Gong, Yu Zhou 0016 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Dual-distribution discrepancy with self-supervised refinement for anomaly detection in medical images
Yu Cai 0005, Hao Chen 0011, Xin Yang 0008, Yu Zhou 0016, Kwang-Ting Cheng |
Medical Image Anal. | 4 |
| 2023 | MARF: Multiscale Adaptive-Switch Random Forest for Leg Detection With 2-D Laser ScannersabstractFor the 2-D laser-based tasks, e.g., people detection and people tracking, leg detection is usually the first step. Thus, it carries great weight in determining the performance of people detection and people tracking. However, many leg detectors ignore the inevitable noise and the multiscale characteristics of the laser scan, which makes them sensitive to the unreliable features of point cloud and further degrades the performance of the leg detector. In this article, we propose a multiscale adaptive-switch random forest (MARF) to overcome these two challenges. First, the adaptive-switch decision tree is designed to use noise-sensitive features to conduct weighted classification and noise-invariant features to conduct binary classification, which makes our detector perform more robust to noise. Second, considering the multiscale property that the sparsity of the 2-D point cloud is proportional to the length of laser beams, we design a multiscale random forest structure to detect legs at different distances. Moreover, the proposed approach allows us to discover a sparser human leg from point clouds than others. Consequently, our method shows an improved performance compared to other state-of-the-art leg detectors on the challenging Moving Legs dataset and retains the entire pipeline at a speed of 60+ FPS on low-computational laptops. Moreover, we further apply the proposed MARF to the people detection and tracking system, achieving a considerable gain in all metrics. Tianxi Wang, Feng Xue 0001, Yu Zhou 0016, Anlong Ming |
IEEE Trans. Cybern. | 3 |
| 2023 | Focal Inverse Distance Transform Maps for Crowd LocalizationabstractIn this paper, we focus on the crowd localization task, a crucial topic of crowd analysis. Most regression-based methods utilize convolution neural networks (CNN) to regress a density map, which can not accurately locate the instance in the extremely dense scene, attributed to two crucial reasons: 1) the density map consists of a series of blurry Gaussian blobs, 2) severe overlaps exist in the dense region of the density map. To tackle this issue, we propose a novel Focal Inverse Distance Transform (FIDT) map for the crowd localization task. Compared with the density maps, the FIDT maps accurately describe the persons' locations without overlapping in dense regions. Based on the FIDT maps, a Local-Maxima-Detection-Strategy (LMDS) is derived to effectively extract the center point for each individual. Furthermore, we introduce an Independent SSIM (I-SSIM) loss to make the model tend to learn the local structural information, better recognizing local maxima. Extensive experiments demonstrate that the proposed method reports state-of-the-art localization performance on six crowd datasets and one vehicle dataset. Additionally, we find that the proposed method shows superior robustness on the negative and extremely dense scenes, which further verifies the effectiveness of the FIDT maps. Dingkang Liang, Wei Xu 0037, Yingying Zhu 0005, Yu Zhou 0016 |
IEEE Trans. Multim. | 4 |
| 2022 | Dual-Distribution Discrepancy for Anomaly Detection in Chest X-Rays
Yu Cai 0005, Hao Chen 0011, Xin Yang 0008, Yu Zhou 0016, Kwang-Ting Cheng |
MICCAI (3) | 4 |
| 2022 | TransCrowd: weakly-supervised crowd counting with transformers
Dingkang Liang, Xiwu Chen, Wei Xu 0037, Yu Zhou 0016, Xiang Bai |
Sci. China Inf. Sci. | 4 |
| 2022 | Semi-Supervised Text Detection With Accurate Pseudo-LabelsabstractRecent scene text detection methods have made great progress. However, existing methods rely heavily on extensive labeled data, which is very time-consuming and expensive. In this letter, we propose a novel semi-supervised text detection method to alleviate the dependence of text detectors on labeled data by generating accurate pseudo-labels and performing effective data augmentations. Specifically, a dual-threshold pseudo-label generation algorithm is designed to divide the prediction results into background regions, text regions, and uncertain regions. The definition of uncertain regions obviously improves the accuracy of pseudo-labels. To obtain accurate pseudo-labels for text at various scales, we first design a scale-aware loss function to adaptively adjust the loss weight of different scale texts. Then, a multi-scale feature extraction module is proposed to extract multi-scale text features and adaptively weight these features according to the scale of the text. Moreover, effective data augmentations are explored to use unlabeled data to improve the robustness of the model to various texts. Experiments show that our method achieves state-of-the-art performance on several datasets(e.g., outperforms existing methods by 2.0% on TD500). Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Yongdong Zhang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2022 | FastRoadSeg: Fast Monocular Road Segmentation NetworkabstractMonocular road segmentation is a fundamental component of autonomous driving. High accuracy has always been a focus of the community for this task. However, the exploration of high-speed methods is still lacking. This limits the prevalence of the monocular road segmentation method. In this paper, an efficient encoder-decoder network is proposed for segmenting the road from a single image. Specifically, to achieve high-speed inference, we leverage a shallow encoder for feature extraction and design a lightweight decoder for feature resolution recovery. To avoid the accuracy drop caused by network simplification, we introduce an efficient asymmetric dilated block to improve the ability of features to distinguish drivable/nondrivable roads to maintain the performance without excess computation. Experiments on three datasets verify the effectiveness of our method. On KITTI-Road, our approach ranks second among all monocular methods and even achieves a competitive performance compared to multimodal methods, with an inference speed over 135 FPS on a single TITAN Xp, which is at least 20 times faster than the state-of-the-art monocular method. Furthermore, our approach can be deployed on the embedded platform Jetson TX2 and achieves a speed of over 80 FPS. Codes will be available athttps://github.com/gongshichina/FastRoadSeg. Shi Gong, Feng Xue 0001, Cong Fang 0005, Yu Zhou 0016 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | TDI TextSpotter: Taking Data Imbalance into Account in Scene Text SpottingabstractRecent scene text spotters that integrate text detection module and recognition module have made significant progress. However, existing methods encounter two problems. 1). The data imbalance issue between text detection module and text recognition module limits the performance of text spotters. 2). The default left-to-right reading direction leads to errors in unconventional text spotting. In this paper, we propose a novel scene text spotter TDI to solve these problems. Firstly, in order to solve the data imbalance problem, a sample generation algorithm is proposed to generate plenty of samples online for training the text recognition module by using character features and character labels. Secondly, a weakly supervised character generation algorithm is designed to generate character-level labels from word-level labels for the sample generation algorithm and the training of the text detection module. Finally, in order to spot arbitrarily arranged text correctly, a direction perception module is proposed to perceive the reading direction of text instance. Experiments on several benchmarks show that these designs can significantly improve the performance of text spotter. Specifically, our method outperforms state-of-the-art methods on three public datasets in both text detection and end-to-end text recognition, which fully proves the effectiveness and robustness of our method. Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Jing Wang 0221, Zhengjun Zha, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2021 | Deep learning for predicting COVID-19 malignant progression
Cong Fang 0005, Song Bai 0001, Qianlan Chen, Yu Zhou 0016, Liming Xia, Lixin Qin, Shi Gong, Xudong Xie, Chunhua Zhou, Dandan Tu, Changzheng Zhang, Xiaowu Liu, Xiang Bai, Philip Torr 0001 |
Medical Image Anal. | 4 |
| 2021 | Boundary-induced and scene-aggregated network for monocular depth prediction
Feng Xue 0001, Junfeng Cao, Yu Zhou 0016, Fei Sheng, Anlong Ming |
Pattern Recognit. | 3 |
| 2021 | Video Text Tracking With a Spatio-Temporal Complementary ModelabstractText tracking is to track multiple texts in a video, and construct a trajectory for each text. Existing methods tackle this task by utilizing the tracking-by-detection framework, i.e., detecting the text instances in each frame and associating the corresponding text instances in consecutive frames. We argue that the tracking accuracy of this paradigm is severely limited in more complex scenarios, e.g., owing to motion blur, etc., the missed detection of text instances causes the break of the text trajectory. In addition, different text instances with similar appearance are easily confused, leading to the incorrect association of the text instances. To this end, a novel spatio-temporal complementary text tracking model is proposed in this paper. We leverage a Siamese Complementary Module to fully exploit the continuity characteristic of the text instances in the temporal dimension, which effectively alleviates the missed detection of the text instances, and hence ensures the completeness of each text trajectory. We further integrate the semantic cues and the visual cues of the text instance into a unified representation via a text similarity learning network, which supplies a high discriminative power in the presence of text instances with similar appearance, and thus avoids the mis-association between them. Our method achieves state-of-the-art performance on several public benchmarks. The source code is available at https://github.com/lsabrinax/VideoTextSCM. Yuzhe Gao, Jiajian Zhang, Yu Zhou 0016, Jing Wang 0221, Shenggao Zhu, Xiang Bai |
IEEE Trans. Image Process. | 4 |
| 2020 | TANet: Robust 3D Object Detection from Point Clouds with Triple AttentionabstractIn this paper, we focus on exploring the robustness of the 3D object detection in point clouds, which has been rarely discussed in existing approaches. We observe two crucial phenomena: 1) the detection accuracy of the hard objects, e.g., Pedestrians, is unsatisfactory, 2) when adding additional noise points, the performance of existing approaches decreases rapidly. To alleviate these problems, a novel TANet is introduced in this paper, which mainly contains a Triple Attention (TA) module, and a Coarse-to-Fine Regression (CFR) module. By considering the channel-wise, point-wise and voxel-wise attention jointly, the TA module enhances the crucial information of the target while suppresses the unstable cloud points. Besides, the novel stacked TA further exploits the multi-level feature attention. In addition, the CFR module boosts the accuracy of localization without excessive computation cost. Experimental results on the validation set of KITTI dataset demonstrate that, in the challenging noisy cases, i.e., adding additional random noisy points around each object, the presented approach goes far beyond state-of-the-art approaches. Furthermore, for the 3D object detection task of the KITTI benchmark, our approach ranks the first place on Pedestrian class, by using the point clouds as the only input. The running speed is around 29 frames per second. Zhe Liu 0033, Xin Zhao 0012, Tengteng Huang, Ruolan Hu, Yu Zhou 0016, Xiang Bai |
AAAI | 5 |
| 2020 | CRNet: A Center-aware Representation for Detecting Text of Arbitrary ShapesabstractExisting scene text detection methods achieve state-of-the-art performance by designing elaborate anchors or complex post-processing. Nonetheless, most methods still face the dilemma of detecting adjacent texts as one instance and long text with large character spacing as multiple fragments. To tackle these problems, we propose an anchor-free scene text detector leveraging Center-aware Representation to achieve accurate arbitrary-shaped scene text detection namely CRNet. Firstly, we propose a center-aware location algorithm to explicitly learn center regions and center points of text instances, which is able to separate adjacent text instances effectively. Then, a multi-scale context extraction module capable of extracting local context, long-range dependencies and global context adaptively is designed to effectively perceive long text with large character spacing. Finally, a low-level features enhancement block is introduced to enhance the geometric information of text. Extensive experiments conducted on several benchmarks including SCUT-CTW1500, Total-Text, ICDAR2015, ICDAR2017 MLT, and MSRA-TD500 demonstrate the effectiveness of our method. Specifically, without any anchor and complicated post-processing, our CRNet achieves 84.2% and 85.1% on CTW1500 and MSRA-TD500 in F-measure, outperforming all state-of-the-art anchor-based and anchor-free methods. Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Yan Li 0068, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2020 | Self-Similarity Action ProposalabstractTemporal action proposal generation, which aims to locate temporal segments that may contain actions, is a key prepositive step of various video analysis tasks, like temporal action detection. In this letter, we present Self-Similarity Action Proposal (SSAP), a simple method that generates action proposals using the self-similarity of videos. Specifically, a basic low-level index, structural similarity, is adopted to measure the similarity between adjacent frames. Potential action boundaries are located by thresholding the similarity values and candidate action segments are successively generated by grouping the boundaries. A segment evaluation module (SEM) is further employed to score and refine the segments. The framework achieves state-of-the-art performance on THUMOS14 and competitive results on ActivityNet v1.3. Notably, on THUMOS14, it achieves over 4% improvement on the average recall at 50 proposals and 3.3% gain in [email protected] when combined with an existing action classifier for temporal action detection. Yuchao Sun, Jianghu Lu, Cong Yao, Yu Zhou 0016 |
IEEE Signal Process. Lett. | 5 |
| 2020 | Tiny Obstacle Discovery by Occlusion-Aware Multilayer RegressionabstractEdges are the fundamental visual element for discovering tiny obstacles using a monocular camera. Nevertheless, tiny obstacles often have weak and inconsistent edge cues due to various properties such as small size and similar appearance to the free space, making it hard to capture them. To this end, we propose an occlusion-based multilayer approach, which specifies the scene prior as multilayer regions and utilizes these regions in each obstacle discovery module, i.e., edge detection and proposal extraction. Firstly, an obstacle-aware occlusion edge is generated to accurately capture the obstacle contour by fusing the edge cues inside all the multilayer regions, which intensifies the object characteristics of these obstacles. Then, a multistride sliding window strategy is proposed for capturing proposals that enclose the tiny obstacles as completely as possible. Moreover, a novel obstacle-aware regression model is proposed for effectively discovering obstacles. It is formed by a primary-secondary regressor, which can learn two dissimilarities between obstacles and other categories separately, and eventually generate an obstacle-occupied probability map. The experiments are conducted on two datasets to demonstrate the effectiveness of our approach under different scenarios. And the results show that the proposed method can approximately improve accuracy by 19% over FPHT and PHT, and achieves comparable performance to MergeNet. Furthermore, multiple experiments with different variants validate the contribution of our method. The source code is available at https://github.com/XuefengBUPT/TOD_OMR. Feng Xue 0001, Anlong Ming, Yu Zhou 0016 |
IEEE Trans. Image Process. | 3 |
| 2019 | Occlusion-Shared and Feature-Separated Network for Occlusion Relationship ReasoningabstractOcclusion relationship reasoning demands closed contour to express the object, and orientation of each contour pixel to describe the order relationship between objects. Current CNN-based methods neglect two critical issues of the task: (1) simultaneous existence of the relevance and distinction for the two elements, i.e, occlusion edge and occlusion orientation; and (2) inadequate exploration to the orientation features. For the reasons above, we propose the Occlusion-shared and Feature-separated Network (OFNet). On one hand, considering the relevance between edge and orientation, two sub-networks are designed to share the occlusion cue. On the other hand, the whole network is split into two paths to learn the high semantic features separately. Moreover, a contextual feature for orientation prediction is extracted, which represents the bilateral cue of the foreground and background areas. The bilateral cue is then fused with the occlusion cue to precisely locate the object regions. Finally, a stripe convolution is designed to further aggregate features from surrounding scenes of the occlusion edge. The proposed OFNet remarkably advances the state-of-the-art approaches on PIOD and BSDS ownership dataset. Feng Xue 0001, Menghan Zhou, Anlong Ming, Yu Zhou 0016 |
ICCV | 5 |
| 2019 | Context-Constrained Accurate Contour Extraction for Occlusion Edge DetectionabstractOcclusion edge detection requires both accurate locations and context constraints of the contour. Existing CNN-based pipeline does not utilize adaptive methods to filter the noise introduced by low-level features. To address this dilemma, we propose a novel Context-constrained accurate Contour Extraction Network (CCENet). Spatial details are retained and contour-sensitive context is augmented through two extraction blocks, respectively. Then, an elaborately designed fusion module is available to integrate features, which plays a complementary role to restore details and remove clutter. Weight response of attention mechanism is eventually utilized to enhance occluded contours and suppress noise. The proposed CCENet significantly surpasses state-of-the-art methods on PIOD and BSDS ownership dataset of object edge detection and occlusion orientation detection. Menghan Zhou, Anlong Ming, Yu Zhou 0016 |
ICME | 4 |
| 2019 | MLTS: A Multi-Language Scene Text SpotterabstractScene text detection and recognition are popular research topics in computer vision due to its various applications such as autonomous driving, blind assistance and text translation. However, many methods currently can only detect or recog-nize the text of one language. In scene text images, we can often see text in multi-language appearing on the same image. However, there is no valid model for multi-language text spotting. In this paper, an end-to-end method for multi-language scene text detection, recognition and script identification is proposed. The method, called MLTS, is an abbreviation of a Multi-Language Scene Text Spotter. By designing a special backbone for text and combining two different kinds of attention. MLTS achieves state-of-the-art performance for both joint localization and script identification in natural images and in cropped word script identification, the precision, recall and F-measure are 0.7145, 0.6583 and 0.6852 respectively, while the corresponding values of the best existing methods are 0.5759, 0.6207, 0.5974 respectively. Additionally, our MLTS achieves comparable performance on ICDAR2013 and ICDAR2015, which proves the effectiveness of the model. Yu Zhou 0016, Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
ICME | 1 |
| 2019 | A Novel Multi-layer Framework for Tiny Obstacle DiscoveryabstractFor tiny obstacle discovery in a monocular image, edge is a fundamental visual element. Nevertheless, because of various reasons, e.g., noise and similar color distribution with background, it is still difficult to detect the edges of tiny (b) obstacles at long distance. In this paper, we propose an obstacle-aware discovery method to recover the missing contours of these obstacles, which helps to obtain obstacle proposals as much as possible. First, by using visual cues in monocular images, several multi-layer regions are elaborately inferred to reveal the distances from the camera. Second, several novel obstacle-aware occlusion edge maps are constructed to well capture the contours of tiny obstacles, which combines cues from each layer. Third, to ensure the existence of the tiny obstacle proposals, the maps from all layers are used for proposals extraction. Finally, based on these proposals containing tiny obstacles, a novel obstacle-aware regressor is proposed to generate an obstacle occupied probability map with high confidence. The convincing experimental results with comparisons on the Lost and Found dataset demonstrate the effectiveness of our approach, achieving around 9.5% improvement on the accuracy than FPHT and PHT, it even gets comparable performance to MergeNet. Moreover, our method outperforms the state-of-the-art algorithms and significantly improves the discovery ability for tiny obstacles at long distance. Feng Xue 0001, Anlong Ming, Menghan Zhou, Yu Zhou 0016 |
ICRA | 4 |
| 2018 | Neural Abstract Style Transfer for Chinese Traditional Painting
Bo Li 0031, Caiming Xiong, Tianfu Wu 0001, Yu Zhou 0016, Rufeng Chu |
ACCV (2) | 4 |
| 2018 | Objectness-Aware Tracking via Double-Layer ModelabstractThe prediction drifts to the non-object backgrounds is a critical issue in conversional correlation filter (CF) based trackers. The key insight of this paper is to propose a double-layer model to address this problem. Specifically, the first layer is a CF tracker, which is employed to predict a rough position of the target, and the objectness layer, which is regarded as the second layer, is utilized to reveal the object characteristics of the predicted target. The novel objectness layer firstly constructs a set of target-related object proposals, which satisfy both the spatial and temporal constraints. And then an objectness classifier is learned upon the proposal set to best separate the target from the noise background proposals. The convincing experimental results on the challenging OTB100 and TC128 dataset demonstrate the effectiveness of the presented approach. Menghan Zhou, Jianxiang Ma, Anlong Ming, Yu Zhou 0016 |
ICIP | 4 |
| 2018 | Learning Training Samples for Occlusion Edge Detection and Its Application in Depth Ordering InferenceabstractThis paper studies the problem of occlusion edge detection, which is applied to infer the depth order of objects in a monocular image. The key observation is that, given the fixed regression objective, the accuracy of occlusion edge detection is effectively boosted by selecting appropriate training samples in a discriminative feature subspace. Specifically, the ℓ1-regularized logistic regression is employed to learn a more sparse yet discriminative feature subspace, while the training sample selection is formulated as a quadratic optimization with the robust Huber loss. The presented formulation avoids the noises efficiently, and hence the desirable occlusion edges can be detected. We validate the effectiveness of our approach on depth order inference problem. Experiments are conducted on two famous datasets, i.e., the Cornell depth-order dataset and the NYU2 dataset. Promising results demonstrate the superiority of our approach over the state-of-the-art approaches. Yu Zhou 0016, Jianxiang Ma, Anlong Ming, Xiang Bai |
ICPR | 1 |
| 2018 | Visual Homing via Guided Locality Preserving MatchingabstractThis study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoramic images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from hundreds of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost true matches without sacrifice in accuracy. To apply our GLPM to the visual homing problem, we develop a method for dense motion flow estimation from sparse feature matches based on Tikhonov regularization. Moreover, the focus-of-contraction/focus-of-expansion is derived to determine homing directions. The effectiveness of our method is demonstrated on a panoramic database in both feature matching and visual homing. Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Yu Zhou 0016, Zheng Wang 0007, Xiaojie Guo 0001 |
ICRA | 5 |
| 2017 | Object-Level ProposalsabstractEdge and surface are two fundamental visual elements of an object. The majority of existing object proposal approaches utilize edge or edge-like cues to rank candidates, while we consider that the surface cue containing the 3D characteristic of objects should be captured effectively for proposals, which has been rarely discussed before. In this paper, an object-level proposal model is presented, which constructs an occlusion-based objectness taking the surface cue into account. Specifically, the better detection of occlusion edges is focused on to enrich the surface cue into proposals, namely, the occlusion-dominated fusion and normalization criterion are designed to obtain the approximately overall contour information, to enhance the occlusion edge map at utmost and thus boost proposals. Experimental results on the PASCAL VOC 2007 and MS COCO 2014 dataset demonstrate the effectiveness of our approach, which achieves around 6% improvement on the average recall than Edge Boxes at 1000 proposals and also leads to a modest gain on the performance of object detection. Jianxiang Ma, Anlong Ming, Xinggang Wang, Yu Zhou 0016 |
ICCV | 5 |
| 2016 | Similarity Fusion for Visual Tracking
Yu Zhou 0016, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki |
Int. J. Comput. Vis. | 1 |
| 2016 | Human action recognition with skeleton induced discriminative approximate rigid part model
Yu Zhou 0016, Anlong Ming |
Pattern Recognit. Lett. | 1 |
| 2015 | Learning discriminative occlusion feature for depth ordering inference on monocular imageabstractIn this paper, a novel depth ordering inference approach is presented. Our main insight is to integrate the discriminative feature selection, occlusion feature learning and same-layer (S-L) relationship judgement into a uniform sparsity based classification objective, which cannot only supply the precise segmentation for the occlusion edge, but also reduce the solution space for the depth ordering inference efficiently. In addition, a novel triple descriptor is adopted to judge the foreground relationship, which is more discriminative than conversional local cues and can further reduce the solution space. The inference is executed by finding a valid path on a directed graph model. We validate our approach on the Cornell depth-order dataset and the NYU 2 dataset, and the convincing experimental results demonstrate the effectiveness of our approach. Anlong Ming, Baofeng Xun, Jia Ni, Mingfei Gao, Yu Zhou 0016 |
ICIP | 5 |
| 2014 | Real-time object tracking via optimal feature subspaceabstractIn this paper, we present a real-time tracking approach based on the Optimal Feature Subspace (OFS). OFS is an optimal subspace of a random feature space, which can best represent the target and making it most distinguished in the whole scene. Initially, we randomly crop patches inside the bounding box to generate an efficient feature template set. Then a greedy algorithm fusing the cues of both target and background is proposed to seek the OFS at every frame. In the forthcoming frame, considering the correlation of different dimensions, we compute the Mahalanobis distance of candidate patches to the appearance model in the obtained subspace to locate the target. The experimental results on several challenging video clips demonstrate that our approach outperforms the state-of-the-art methods, in terms of both speed and robustness. Xu Min, Yu Zhou 0016, Xiang Bai |
ICIP | 2 |
| 2014 | Online Multiple targets Detection and Tracking from Mobile robot in Cluttered indoor Environments with Depth CameraabstractIndoor environment is a common scene in our everyday life, and detecting and tracking multiple targets in this environment is a key component for many applications. However, this task still remains challenging due to limited space, intrinsic target appearance variation, e.g. full or partial occlusion, large pose deformation, and scale change. In the proposed approach, we give a novel framework for detection and tracking in indoor environments, and extend it to robot navigation. One of the key components of our approach is a virtual top view created from an RGB-D camera, which is named ground plane projection (GPP). The key advantage of using GPP is the fact that the intrinsic target appearance variation and extrinsic noise is far less likely to appear in GPP than in a regular side-view image. Moreover, it is a very simple task to determine free space in GPP without any appearance learning even from a moving camera. Hence GPP is very different from the top-view image obtained from a ceiling mounted camera. We perform both object detection and tracking in GPP. Two kinds of GPP images are utilized: gray GPP, which represents the maximal height of 3D points projecting to each pixel, and binary GPP, which is obtained by thresholding the gray GPP. For detection, a simple connected component labeling is used to detect footprints of targets in binary GPP. For tracking, a novel Pixel Level Association (PLA) strategy is proposed to link the same target in consecutive frames in gray GPP. It utilizes optical flow in gray GPP, which to our best knowledge has never been done before. Then we "back project" the detected and tracked objects in GPP to original, side-view (RGB) images. Hence we are able to detect and track objects in the side-view (RGB) images. Our system is able to robustly detect and track multiple moving targets in real time. The detection process does not rely on any target model, which means we do not need any training process. Moreover, tracking does not require any manual initialization, since all entering objects are robustly detected. We also extend the novel framework to robot navigation by tracking. As our experimental results demonstrate, our approach can achieve near prefect detection and tracking results. The performance gain in comparison to state-of-the-art trackers is most significant in the presence of occlusion and background clutter. Yu Zhou 0016, Yinfei Yang, Meng Yi, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2013 | On contrast combinations for visual saliency detectionabstractSaliency detection is an important task in computer vision and image processing. The most influential factor in bottom-up visual saliency is contrast operation. In this paper, we propose a unified model to combine widely used contrast measurements, namely, center-surround, corner-surround and global contrast to detect visual saliency. The proposed model benefits from the advantages of each individual contrast operation, and thus produces more robust and accurate saliency maps. Extensive experimental results on natural images show the effectiveness of the proposed model for visual saliency detection task, and demonstrate the combination is superior than individual subcomponent. Quan Zhou 0004, Shiwei Ren, Yu Zhou 0016, Jun Chen 0019, Wenyu Liu 0001 |
ICIP | 4 |
| 2012 | Navigation toward Non-static Target Object Using Footprint Detection Based Tracking
Meng Yi, Yinfei Yang, Wenjing Qi, Yu Zhou 0016, Zygmunt Pizlo, Longin Jan Latecki |
ACCV (3) | 4 |
| 2012 | Mismatch removal via coherent spatial mappingabstractWe propose a method for removing mismatches from given putative point correspondences in image pairs. Our algorithm aims to recover the underlying coherent spatial mapping which related to inliers. The thin-plate spline (TPS) is chosen to parameterize the coherent spatial mapping, and we formulate the solution of it as a maximum likelihood problem. The mismatches could be successfully removed after the EM algorithm, which we used for solving the problem, converges. The quantitative results on various experimental data demonstrate that our method outperforms many state-of-the-art methods. Moreover, the proposed method is also able to handle the case that image pairs contain non-rigid motions. Jiayi Ma 0001, Ji Zhao 0001, Yu Zhou 0016, Jinwen Tian |
ICIP | 3 |
| 2012 | Fusion with Diffusion for Robust Visual TrackingabstractA weighted graph is used as an underlying structure of many algorithms like semi-supervised learning and spectral clustering. The edge weights are usually deter-mined by a single similarity measure, but it often hard if not impossible to capture all relevant aspects of similarity when using a single similarity measure. In par-ticular, in the case of visual object matching it is beneficial to integrate different similarity measures that focus on different visual representations. In this paper, a novel approach to integrate multiple similarity measures is pro-posed. First pairs of similarity measures are combined with a diffusion process on their tensor product graph (TPG). Hence the diffused similarity of each pair of ob-jects becomes a function of joint diffusion of the two original similarities, which in turn depends on the neighborhood structure of the TPG. We call this process Fusion with Diffusion (FD). However, a higher order graph like the TPG usually means significant increase in time complexity. This is not the case in the proposed approach. A key feature of our approach is that the time complexity of the dif-fusion on the TPG is the same as the diffusion process on each of the original graphs, Moreover, it is not necessary to explicitly construct the TPG in our frame-work. Finally all diffused pairs of similarity measures are combined as a weighted sum. We demonstrate the advantages of the proposed approach on the task of visual tracking, where different aspects of the appearance similarity between the target object in frame t and target object candidates in frame t+1 are integrated. The obtained method is tested on several challenge video sequences and the experimental results show that it outperforms state-of-the-art tracking methods. Yu Zhou 0016, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki |
NIPS | 1 |
| 2011 | Shape Matching Using Points Co-occurrence PatternabstractShape matching is a very critical problem in computer vision, and many smart features have been designed in recent literature for improving the similarity measure between pairs of shapes, and most of them consider either distribution of the sample contour points, or convexity/concavity property of the contour. In this paper, we design a novel shape feature to capture the Co-Occurrence Pattern (COP) of the points sampled from any given shape contour, and each pattern is described by \textbf{Self-Similarity} which investigates the spatial co-occurrence relation among all the sample points. We test our feature on three famous shape databases: MPEG-7 CE-Shape-1 part B, Tari1000, and Kimia99 data set for shape matching and retrieval. The experimental results show that the proposed descriptor achieves higher computational efficiency with no significant performance loss. Yu Zhou 0016, Quan Zhou 0004, Xiang Bai, Wenyu Liu 0001 |
ICIG | 1 |
| 2011 | Shape Matching and Recognition Using Group-Wised Points
Yu Zhou 0016, Xiang Bai, Wenyu Liu 0001 |
PSIVT (2) | 2 |