Yue Zhou 0005

dblp:78/6191-5 · DBLP profile ↗
← Back
69ranked-venue papers
10as first author
31since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 7 first-author · 15 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Text4Seg: Reimagining Image Segmentation as Text Generation
abstract
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce Text4Seg, a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. This unified representation allows seamless integration into the auto-regressive training pipeline of MLLMs for easier optimization. We demonstrate that representing an image with $16\times16$ semantic descriptors yields competitive segmentation performance. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptors by 74\% and accelerating inference by $3\times$, without compromising performance. Extensive experiments across various vision tasks, such as referring expression segmentation and comprehension, show that Text4Seg achieves state-of-the-art performance on multiple datasets by fine-tuning different MLLM backbones. Our approach provides an efficient, scalable solution for vision-centric tasks within the MLLM framework.
Mengcheng Lan, Chaofeng Chen, Yue Zhou 0005, Jiaxing Xu, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001
ICLR3
2025 Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai
Sci. China Inf. Sci.21
2025 PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection
Peiyuan Zhang, Xue Yang 0005, Yi Yu 0010, Qingyun Li, Yue Zhou 0005, Xiaosong Jia, Jingdong Chen, Xiang Li 0041, Junchi Yan, Yansheng Li 0001
Int. J. Comput. Vis.6
2025 AirSpatialBot: A Spatially Aware Aerial Agent for Fine-Grained Vehicle Attribute Recognition and Retrieval
abstract
Despite notable advancements in remote sensing vision-language models (VLMs), existing models often struggle with spatial understanding, limiting their effectiveness in real-world applications. To push the boundaries of VLMs in remote sensing, we specifically address vehicle imagery captured by drones and introduce a spatially-aware dataset AirSpatial, which comprises over 206K instructions and introduces two novel tasks: Spatial Grounding and Spatial Question Answering. It is also the first remote sensing grounding dataset to provide 3DBB. To effectively leverage existing image understanding of VLMs to spatial domains, we adopt a two-stage training strategy comprising Image Understanding Pre-training and Spatial Understanding Fine-tuning. Utilizing this trained spatially-aware VLM, we develop an aerial agent, AirSpatialBot, which is capable of fine-grained vehicle attribute recognition and retrieval. By dynamically integrating task planning, image understanding, spatial understanding, and task execution capabilities, AirSpatialBot adapts to diverse query requirements. Experimental results validate the effectiveness of our approach, revealing the spatial limitations of existing VLMs while providing valuable insights. The model, code, and datasets will be released at https://github.com/VisionXLab/AirSpatialBot.
Yue Zhou 0005, Xue Yang 0005, Xue Jiang 0001, Xingzhao Liu
IEEE Trans. Geosci. Remote. Sens.1
2024 Enhancing Photo Animation: Augmented Stylistic Modules and Prior Knowledge Integration
Zhanyi Lu, Yue Zhou 0005, Ao Chen 0003
ACCV (6)2
2024 Alpharotate: A Rotation Detection Benchmark Using Tensorflow
abstract
AlphaRotate is an open-source TensorFlow benchmark for performing scalable rotation detection on various datasets. It currently provides more than 16 popular rotation detection models under a single, well-documented API designed for use by both practitioners and researchers. AlphaRotate regards high performance, robustness, sustainability and scalability as the core concept of design, and all models are covered by continuous integration, code coverage, maintainability checks, and visual monitoring and analysis. AlphaRotate can be installed from PyPI and is released under the Apache-2.0 License. Source code is available at https://github.com/yangxue0827/RotationDetection.
Xue Yang 0005, Yue Zhou 0005, Wenlong Liao, Junchi Yan
ICASSP2
2024 Efficient Hierarchical Stripe Attention for Lightweight Image Super-Resolution
abstract
Following the successful integration of Transformers in computer vision, several Transformer-based approaches have emerged, surpassing the previously dominant CNN-based techniques. However, it is observed that conventional self-attention involves redundant operations that can be further optimized. In this context, we propose a novel framework for lightweight super-resolution, termed as Hierarchical Stripe-attention Super-Resolution Network (HSSRNet). Specifically, our approach involves the fusion of correlated patches in each layer to access global pixel information. Subsequently, we introduce a stripe intra-patch self-attention (stripe IPSA) mechanism to model long-range dependencies efficiently. To further reduce computational complexity, we approximate the original attention map using two low-rank matrices. Finally, we employ PixelShuffle in conjunction with convolutions for upsampling. Extensive experiments demonstrate the effectiveness of our proposed modules, and the results indicate that our method achieves commendable performance on four benchmark datasets.
Xiaying Chen, Yue Zhou 0005
ICASSP2
2024 Cross-modal Prominent Fragments Enhancement Aligning Network for Image-text Retrieval
abstract
Image-text retrieval is a widely studied topic in the field of computer vision due to the exponential growth of multimedia data, whose core concept is to measure the similarity between images and text. However, most existing retrieval methods heavily rely on cross-attention mechanisms for cross-modal fine-grained alignment, which takes into account excessive irrelevant regions and treats prominent and non-significant words equally. This paper aims to investigate an alignment approach that reduces the involvement of non-significant fragments in images and text while enhancing the alignment of prominent fragments. For this purpose, we introduce the Cross-Modal Prominent Fragments Enhancement Aligning Network(CPFEAN). In practice, we first design a novel intra-modal fragments relationship reasoning method, and subsequently employ our proposed alignment mechanism to compute the similarity between images and text. Extensive quantitative comparative experiments on MS-COCO and Flickr30K datasets demonstrate that our approach outperforms state-of-the-art methods.
Yue Zhou 0005, Zonghao Yang, Ao Chen 0003
ICME2
2024 Transformer-Based Incomplete Multi-Modal Learning for Land Cover Classification
abstract
Land cover (LC) classification via remote sensing is crucial for ecosystem monitoring and urban planning but faces the challenge of inconsistent multimodal data availability. Current techniques often falter with incomplete modalities, resulting in reduced performance and adaptability. Addressing these issues, this study propose the Transformer-based Incomplete Multi-Modal Learning (TIMML) framework. TIMML incorporates a Bernoulli indicator module during training to facilitate adaptation to missing modalities. This module, in tandem with a fusion token, is instrumental in enabling the model to handle the random omission of modalities by selectively nullifying data streams and effectively aggregates information from the remaining available modalities. Moreover, TIMML integrates a modality-aware regularization module designed to enhance the stability of the feature extraction process, especially when perturbed by the Bernoulli indicator during training. Our comprehensive experiments demonstrate that TIMML not only proficiently manages the challenge of missing modalities but also outperforms existing methods in LC classification tasks, marking a significant advancement in the field.
Guozheng Xu, Xue Jiang 0001, Yue Zhou 0005, Xingzhao Liu
IGARSS3
2024 Enhanced Tree Branch Segmentation in Urban Environments Using a Dual-Encoder Model with Graph Reasoning Decoder Block
abstract
Urban greening trees play a significant role in optimizing the environment, and the segmentation of branches can effectively help researchers assess the growth status of trees. This paper proposed an innovative RGB image-based model to segment tree branches within urban landscapes. The proposed model employed a dual-encoder architecture, including a basic encoder and an edge encoder. These components were designed with specialized blocks adept at extracting critical semantic features and edge information, enhancing the model's ability to segment fine branches. A graph reasoning decoder block with attention-based feature fusion was proposed to capture the semantic associations between regions incorporating edge information. Moreover, the elastic interaction-based loss function, a groundbreaking loss function, was introduced to ensure that the segmentation of the fine branches should be achieved smoothly and consistently. Upon evaluation against a public urban street tree dataset, the precision, recall, IoU, and accuracy of tree branch segmentation are 94.39%, 93.16%, 88.27%, and 98.78%, respectively, achieving the best result among all the tested models. This performance demonstrates the impact of deep learning on enhancing urban greening and sustainable development in effective tree management.
Yue Zhou 0005, Hancong Wang, Yanyi Liu
SMC1
2024 Semi-Supervised Scene Classification for Optical Remote Sensing Images via Label and Embedding Consistency
abstract
The utilization of unlabeled samples has contributed significantly to the achievements of semi-supervised methods in optical remote sensing image (ORSI) scene classification. However, existing methods face the challenge of effectively integrating labeled and unlabeled data during model training. To mitigate these challenges, a semi-supervised label and embedding consistency network (SS-LEC) is proposed for OSRI scene classification. Specifically, given an image, SS-LEC enables the high-confidence prediction from a weak-augmentation view consistent with the prediction from a strong-augmentation view, while also ensuring consistency in embeddings derived from middle-augmentation views. Moreover, a soft learning schedule is proposed to strategically focus on varied consistency tasks at different stages of training. Our experiments on two ORSI datasets showcase SS-LEC’s superior classification performance over existing semi-supervised methods. Notably, under label-scarce scenarios with only four labeled images per category, SS-LEC achieves classification accuracies of 92.04% on the EuroSAT dataset and 70.19% on the NWPU-RESISC45 dataset. These results set new benchmarks and demonstrating superior classification performance in challenging conditions with limited labeled data.
Guozheng Xu, Xue Jiang 0001, Yue Zhou 0005, Xingzhao Liu
IEEE Geosci. Remote. Sens. Lett.4
2024 Robust Land Cover Classification With Multimodal Knowledge Distillation
abstract
In recent years, enormous studies have been conducted to improve the land cover (LC) classification performance of multimodal remote sensing (RS) data, which outperforms single-modal-based methods by a large margin due to information diversity. To go a step further, we develop a two-branch patch-based convolutional neural network (CNN) with an encoder–decoder (ED) module to fuse multimodal RS data information. A knowledge distillation in model (DIM) module is proposed to guild per-modality encoder learning with the final fused information to enable multimodal data fusion more effectively. Moreover, utilizing multimodal information to guide single-modal learning still remains to be explored. To this end, a knowledge distillation cross-model (DCM) module is designed to improve single-modal LC classification with multimodal knowledge distillation, which bridges the gap between single-modal-based and multimodal-based methods. In particular, the multimodal-based method is taken as a teacher to transfer knowledge to single-modal-based methods. Extensive experiments are carried out on two multimodal RS datasets, including hyperspectral (HS) and light detection and ranging (LiDAR) data, i.e., the Houston2013 dataset, and HS and synthetic aperture radar (SAR) data, i.e., the Berlin dataset. The results demonstrate the effectiveness and superiority of the proposed multimodal fusion strategy in comparison with several state-of-the-art multimodal RS data classification methods. Also, the proposed DCM module improves the LC classification performance of single-modal methods by a large margin.
Guozheng Xu, Xue Jiang 0001, Yue Zhou 0005, Shutao Li 0001, Xingzhao Liu, Peiwen Lin
IEEE Trans. Geosci. Remote. Sens.3
2024 DGA: Direction-Guided Attack Against Optical Aerial Detection in Camera Shooting Direction-Agnostic Scenarios
abstract
Patch-based adversarial attacks have increasingly aroused concerns due to their application potential in military and civilian fields. In aerial imagery, numerous targets exhibit inherent directionality, such as vehicles and ships, giving rise to the emergence of oriented object detection tasks; similarly, adversarial patches also exhibit intrinsic orientation due to their lack of perfect symmetry. Existing methods presuppose a static alignment between the adversarial patch’s orientation and the camera’s coordinate system – an assumption that is frequently violated in aerial images, whose effectiveness degrades in real-world scenarios. In this paper, we investigate the often-neglected aspect of patch orientation in adversarial attacks and its impact on camouflage effectiveness, particularly when the orientation is not congruent with the target. A new Directional Guided Attack (DGA) framework is proposed for deceiving real-world aerial detectors, which shows robust and adaptable attack performance in camera shooting direction agnostic (CSDA) scenarios. The core idea of DGA is to utilize affine transformations to constrain the relative orientation of the patch to the target and introduce three types of loss to reduce target detection confidence, make the color printable, and smooth the patch color. We introduce a direction-guided evaluation methodology to bridge the gap between patch performance in the digital domain and its actual real-world efficacy. Moreover, we establish a drone-based vehicle detection dataset (SJTU-4K), which labels the orientation of the target, to assess the robustness of patches under various shooting altitudes and views. Extensive proportionally scaled and 1:1 experiments are performed in physical scenarios, demonstrating the superiority and potential of the proposed framework for real-world attacks.
Yue Zhou 0005, Shu-Qi Sun, Xue Jiang 0001, Guozheng Xu, Fengyuan Hu, Xingzhao Liu
IEEE Trans. Geosci. Remote. Sens.1
2023 The KFIoU Loss for Rotated Object Detection
Xue Yang 0005, Yue Zhou 0005, Gefan Zhang, Jirui Yang, Wentao Wang 0009, Junchi Yan, Xiaopeng Zhang 0008, Qi Tian 0001
ICLR2
2023 H2RBox: Horizontal Box Annotation is All You Need for Oriented Object Detection
Xue Yang 0005, Gefan Zhang, Yue Zhou 0005, Junchi Yan
ICLR4
2023 SODet: A LiDAR-Based Object Detector in Bird's-Eye View
Jin Pang, Yue Zhou 0005
ICONIP (12)2
2023 Adversarial Example Generation Method for Object Detection in Remote Sensing Images
abstract
Object detection in remote sensing images is an essential application of deep learning. However, due to the vulnerability of deep learning models, they are susceptible to adversarial attacks, which can undermine their reliability and accuracy. While significant progress has been made in the field of adversarial attacks, most of the work has focused on image classification tasks due to the complexity of object detection. In this paper, we propose a target camouflage method based on adversarial attacks that can mislead detectors and hide targets with minimal pixel perturbations. Experiments on the DIOR dataset demonstrate the effectiveness of our approach. Our method generates adversarial examples that can successfully fool Faster R-CNN into failing to detect objects with minimal perturbations.
Wanghan Jiang, Yue Zhou 0005, Xue Jiang 0001
IGARSS2
2023 H2RBox-v2: Incorporating Symmetry for Boosting Horizontal Box Supervised Oriented Object Detection
abstract
With the rapidly increasing demand for oriented object detection, e.g. in autonomous driving and remote sensing, the recently proposed paradigm involving weakly-supervised detector H2RBox for learning rotated box (RBox) from the more readily-available horizontal box (HBox) has shown promise. This paper presents H2RBox-v2, to further bridge the gap between HBox-supervised and RBox-supervised oriented object detection. Specifically, we propose to leverage the reflection symmetry via flip and rotate consistencies, using a weakly-supervised network branch similar to H2RBox, together with a novel self-supervised branch that learns orientations from the symmetry inherent in visual objects. The detector is further stabilized and enhanced by practical techniques to cope with peripheral issues e.g. angular periodicity. To our best knowledge, H2RBox-v2 is the first symmetry-aware self-supervised paradigm for oriented object detection. In particular, our method shows less susceptibility to low-quality annotation and insufficient training data compared to H2RBox. Specifically, H2RBox-v2 achieves very close performance to a rotation annotation trained counterpart -- Rotated FCOS: 1) DOTA-v1.0/1.5/2.0: 72.31%/64.76%/50.33% vs. 72.44%/64.53%/51.77%; 2) HRSC: 89.66% vs. 88.99%; 3) FAIR1M: 42.27% vs. 41.25%.
Yi Yu 0010, Xue Yang 0005, Qingyun Li, Yue Zhou 0005, Feipeng Da, Junchi Yan
NeurIPS4
2023 GRD: An Ultra-Lightweight SAR Ship Detector Based on Global Relationship Distillation
abstract
Most existing works on lightweight SAR ship detectors sacrifice a lot of detection accuracy to reduce model size. In this letter, we propose an ultra-lightweight detector based on distillation technology, which can reduce the parameter quantity of the model while minimizing the damage to the model’s detection accuracy. Due to the scattering interference and speckle noise in SAR images, directly applying the existing ultra-lightweight detectors cannot achieve satisfactory performance for ship detection. As a result, we design a global relationship distillation (GRD) algorithm for the ultra-lightweight SAR ship detector. This algorithm can preserve more global relationships from the teacher and mitigate the accuracy degradation caused by the noise and interference, especially in complex inshore scenarios. Besides, the features learned by this algorithm are robust, and the pruned model is more stable. The superiority of the GRD method over several state-of-the-art distillation methods has been evaluated on the HRSID dataset.
Yue Zhou 0005, Xue Jiang 0001, Lin Chen 0037, Xingzhao Liu
IEEE Geosci. Remote. Sens. Lett.1
2023 Detecting Rotated Objects as Gaussian Distributions and its 3-D Generalization
abstract
Existing detection methods commonly use a parameterized bounding box (BBox) to model and detect (horizontal) objects and an additional rotation angle parameter is used for rotated objects. We argue that such a mechanism has fundamental limitations in building an effective regression loss for rotation detection, especially for high-precision detection with high IoU (e.g., 0.75). Instead, we propose to model the rotated objects as Gaussian distributions. A direct advantage is that our new regression loss regarding the distance between two Gaussians e.g., Kullback-Leibler Divergence (KLD), can well align the actual detection performance metric, which is not well addressed in existing methods. Moreover, the two bottlenecks i.e., boundary discontinuity and square-like problem also disappear. We also propose an efficient Gaussian metric-based label assignment strategy to further boost the performance. Interestingly, by analyzing the BBox parameters' gradients under our Gaussian-based KLD loss, we show that these parameters are dynamically updated with interpretable physical meaning, which help explain the effectiveness of our approach, especially for high-precision detection. We extend our approach from 2-D to 3-D with a tailored algorithm design to handle the heading estimation, and experimental results on twelve public datasets (2-D/3-D, aerial/text/face images) with various base detectors show its superiority.
Xue Yang 0005, Gefan Zhang, Xiaojiang Yang, Yue Zhou 0005, Wentao Wang 0009, Jin Tang 0001, Junchi Yan
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Bottom-Up Transformer Reasoning Network for Text-Image Retrieval
Zonghao Yang, Yue Zhou 0005, Ao Chen 0003
ICONIP (6)2
2022 A Three-Stage Cascade Rotating Regression Network for SAR Target Rotation Detection
abstract
With the development of deep learning, SAR target rotating detection has become a research hotspot. However, the regression process of the rotated boxes is challenging, especially for the rotating detection in target gathering areas such as docks and airports. This is because the regions of interest corresponding to adjacent targets have a relatively large overlap, which may cause data misalignment during the regression and classification process. To solve this problem, we propose a three-stage cascaded rotating regression network (3SCR2Net). Specifically, a horizontal anchor will be rotated and corrected three times continuously. 3SCR2Net overcomes the original problem of information misalignment caused by rotation detection in a complex background. In addition, two decoders are designed in this paper to decode the regression parameters. Experimental results on the SSDD+ dataset demonstrate that the proposed network outperforms five state-of-the-art methods in mean average accuracy (mAP).
Xue Jiang 0001, Yue Zhou 0005, Xingzhao Liu
IGARSS3
2022 Benchmark for Arbitrary-Oriented SAR Ship Detection
abstract
A growing number of researchers have begun to use arbitrary-oriented detectors to detect ships. Because in the remote sensing dataset, ships are shown to be narrow and dense. Most of the deep learning methods used in synthetic aperture radar (SAR) ship detection are the same as or variants of those of optical remote sensing. However, the experimental settings of different papers are different. Therefore, various arbitrary-oriented target detectors can not be compared fairly on the SAR dataset. To solve this problem, we developed an arbitrary-oriented SAR ship detection benchmark, which provides strong baselines and state-of-the-art methods in rotation detection. All benchmark methods are tested on RSSDD datasets, and the code is publicly released at https://github.com/open-mmlab/mmrotate. Meanwhile, this paper further explores the role of ImageNet pretrained weight in SAR target detection and tries to train the SAR pretrained weight through unsupervised learning.
Yue Zhou 0005, Xue Jiang 0001, Zhou Li 0002, Xingzhao Liu
IGARSS1
2022 MMRotate: A Rotated Object Detection Benchmark using PyTorch
abstract
We present an open-source toolbox, named MMRotate, which provides a coherent algorithm framework of training, inferring, and evaluation for the popular rotated object detection algorithm based on deep learning. MMRotate implements 18 state-of-the-art algorithms and supports the three most frequently used angle definition methods. To facilitate future research and industrial applications of rotated object detection-related problems, we also provide a large number of trained models and detailed benchmarks to give insights into the performance of rotated object detection. MMRotate is publicly released at https://github.com/open-mmlab/mmrotate.
Yue Zhou 0005, Xue Yang 0005, Gefan Zhang, Yanyi Liu, Liping Hou, Xue Jiang 0001, Xingzhao Liu, Junchi Yan, Chengqi Lyu, Kai Chen 0026
ACM Multimedia1
2022 Balance Scene Learning Mechanism for Offshore and Inshore Ship Detection in SAR Images
abstract
Huge imbalance of different scenes’ sample numbers seriously reduces synthetic aperture radar (SAR) ship detection accuracy. Thus, to solve this problem, this letter proposes a balance scene learning mechanism (BSLM) for offshore and inshore ship detection in SAR images. BSLM involves three steps: 1) based on unsupervised representation learning, a generative adversarial network (GAN) is used to extract the scene features of SAR images; 2) using these features, a scene binary cluster (offshore/inshore) is conducted by${K}$-means; and 3) finally, the small cluster’s samples (inshore) are augmented via replication, rotation transformation or noise addition to balance another big cluster (offshore), so as to eliminate scene learning bias and obtain balanced learning representation ability that can enhance learning benefits and improve detection accuracy. This letter applies BSLM to four widely used and open-sourced deep learning detectors, i.e., faster regions-convolutional neural network (Faster R-CNN), Cascade R-CNN, single shot multibox detector (SSD), and RetinaNet, to verify its effectiveness. Experimental results on the open SAR ship detection data set (SSDD) reveal that BSLM can greatly improve detection accuracy, especially for more complex inshore scenes.
Tianwen Zhang, Xiaoling Zhang 0002, Jun Shi 0002, Shunjun Wei, Yue Zhou 0005
IEEE Geosci. Remote. Sens. Lett.8
2022 HOG-ShipCLSNet: A Novel Deep Learning Network With HOG Feature Fusion for SAR Ship Classification
abstract
Ship classification in synthetic aperture radar (SAR) images is a fundamental and significant step in ocean surveillance. Recently, with the rise of deep learning (DL), modern abstract features from convolutional neural networks (CNNs) have hugely improved SAR ship classification accuracy. However, most existing CNN-based SAR ship classifiers overly rely on abstract features, but uncritically abandon traditional mature hand-crafted features, which may incur some challenges for further improving accuracy. Hence, this article proposes a novel DL network with histogram of oriented gradient (HOG) feature fusion (HOG-ShipCLSNet) for preferable SAR ship classification. In HOG-ShipCLSNet, four mechanisms are proposed to ensure superior classification accuracy, that is, 1) a multiscale classification mechanism (MS-CLS-Mechanism); 2) a global self-attention mechanism (GS-ATT-Mechanism); 3) a fully connected balance mechanism (FC-BAL-Mechanism); and 4) an HOG feature fusion mechanism (HOG-FF-Mechanism). We perform sufficient ablation studies to confirm the effectiveness of these four mechanisms. Finally, our experimental results on two open SAR ship datasets (OpenSARShip and FUSAR-Ship) jointly reveal that HOG-ShipCLSNet dramatically outperforms both modern CNN-based methods and traditional hand-crafted feature methods.
Tianwen Zhang, Xiaoling Zhang 0002, Xiao Ke, Xiaowo Xu, Xu Zhan, Chen Wang 0041, Yue Zhou 0005, Dece Pan, Jun Shi 0002, Shunjun Wei
IEEE Trans. Geosci. Remote. Sens.9
2021 Dense Label Encoding for Boundary Discontinuity Free Rotation Detection
abstract
Rotation detection serves as a fundamental building block in many visual applications involving aerial image, scene text, and face etc. Differing from the dominant regression-based approaches for orientation estimation, this paper explores a relatively less-studied methodology based on classification. The hope is to inherently dismiss the boundary discontinuity issue as encountered by the regression-based detectors. We propose new techniques to push its frontier in two aspects: i) new encoding mechanism: the design of two Densely Coded Labels (DCL) for angle classification, to replace the Sparsely Coded Label (SCL) in existing classification-based detectors, leading to three times training speed increase as empirically observed across benchmarks, further with notable improvement in detection accuracy; ii) loss re-weighting: we propose Angle Distance and Aspect Ratio Sensitive Weighting (ADARSW), which improves the detection accuracy especially for square-like objects, by making DCL-based detectors sensitive to angular distance and object’s aspect ratio. Extensive experiments and visual analysis on large-scale public datasets for aerial images i.e. DOTA, UCAS-AOD, HRSC2016, as well as scene text dataset ICDAR2015 and MLT, show the effectiveness of our approach. The source code is available at $\color{Red}{\text{DCL}}$ and is also integrated in our open source rotation detection benchmark:$\color{Red}{\text{RotationDetection}}$.
Xue Yang 0005, Liping Hou, Yue Zhou 0005, Wentao Wang 0009, Junchi Yan
CVPR3
2021 Bridging the GAP Between Outputs: Domain Adaptation for Lung Cancer IHC Segmentation
abstract
Lung cancer is globally the most common cause of cancer-related death and immunotherapy is one of the most promising treatments for various cancers. Since there are some differences in characteristics between lung cancer immunohistochemistry (IHC) image datasets collected by different hospitals or laboratories, and relabeling new datasets is time-consuming and laborious, we thus propose an effective domain adaptation framework for IHC image segmentation. In our method, maximum mean discrepancy of segmentation outputs is extracted and combined with generative adversarial networks to bridge the gap between the outputs from source and target domains. To make the networks more instructive, pseudo target labels generated by the source model are employed. Meanwhile, adaptive hyperparameter tuning offers much flexibility for the training process. The proposed model has been evaluated over our own datasets and other public medical benchmarks, and a comparison with other recent works demonstrates its effectiveness and robustness. Furthermore, the model is exerted to extract quantitative and spatial features of conventional lung cancer IHC images, which can be utilized for overall survival (OS) and relapse-free survival (RFS) predictions.
Li Diao, Haoyue Guo, Yue Zhou 0005, Yayi He
ICIP3
2021 Arbitrary-Oriented SAR Ship Detection Via Frequency Learning
abstract
A growing number of researchers begin to use the arbitrary-oriented detector to detect ships. Because in the remote sensing dataset, ships are shown to be narrow and dense. Most of the deep learning methods used in synthetic aperture radar (SAR) ship detection are the same as or variants of those of optical remote sensing. However, due to the relatively low resolution and signal-to-noise ratio, the amplitude information in SAR image becomes more contaminated than that in optical image. Therefore, it is very difficult to train the arbitrary-oriented target detection task based on SAR dataset. In order to solve this problem, we attempt to further extract the frequency domain features of the SAR image by introducing a frequency attention module. The discrete cosine transform (DCT) to converts SAR image from spatial domain to frequency domain. Experimental results on HRSID datasets show that the proposed method can significantly improve the performance of arbitrary-oriented SAR ship detection.
Yue Zhou 0005, Xue Jiang 0001, Zhou Li 0002, Xingzhao Liu
IGARSS1
2021 Joint Intention and Trajectory Prediction Based on Transformer
abstract
Although autonomous driving technology has made tremendous progress in recent years, it is still challenging to predict the intentions and trajectories of pedestrians. The state-of-the-art methods suffer from two problems. (1) Existing works consider these two tasks separately, ignoring the connection between them. (2) The selection and integration of inputs for these tasks are not well designed. In this paper, these two tasks are taken into consideration in a unified model. In this way, the information provided by the labels of each other is shared, improving the performance of both tasks. Besides, in addition to the bounding boxes and speeds, orientation and road semantic segmentation features are taken into consideration to show the potential intention and road context of the pedestrian. And all the inputs are weighted by an attention module before integration. Meanwhile, a Transformer encoder is applied in our method to extract the temporal information from the fused feature sequence. Our method outperforms all previous models for both trajectory prediction and intention prediction tasks on the JAAD dataset and PIE dataset.
Ze Sui, Yue Zhou 0005, Xu Zhao 0001, Ao Chen 0003, Yiyang Ni 0004
IROS2
2021 Convolutional Neural Network-Based Dictionary Learning for SAR Target Recognition
abstract
In this letter, a novel convolutional neural network (CNN)-based dictionary learning (DL) method is proposed for synthetic aperture radar (SAR) target recognition. Different from conventional target recognition schemes, which consist of the hand-crafted feature extraction followed by a classifier, the proposed scheme utilizes a well-designed ConvNet as the feature extractor, and it can automatically learn hierarchies of features from the training data set. The outputs of the ConvNet are regarded as multifeature and are used for the following multi-DL. For a classification task, we take into consideration the mean-squared error (MSE) combined with a regularization term as the loss function. As a result, the whole architecture combines the ConvNet and DL as an end-to-end framework. We show the back propagation of the loss and update the variables using the stochastic gradient descent with the momentum method. Experiments performed on the moving and stationary target automatic recognition (MSTAR) data set exhibit that the proposed method outperforms many state-of-the-art DL and CNN methods in terms of recognition performance.
Yue Zhou 0005, Xue Jiang 0001, Xingzhao Liu, Zhixin Zhou
IEEE Geosci. Remote. Sens. Lett.2
2020 SAR Target Classification with Limited Data via Data Driven Active Learning
abstract
With the rapid development of deep learning, more and more deep neural networks with strong discrimination have come up. One reason why deep learning models can achieve such good results is the abundant annotation data. However, obtaining such considerable amount of annotation data is costly, especially in the field of synthetic aperture radar (SAR). High-quality SAR dataset cannot be constructed without the support of the specialists and institutes in the related field, which leads to the limited amount of labeled training data of SAR. In order to solve the problem that deep neural networks (DNNs) may face under limited labeled training data, researchers often use data augmentation methods to increase the number of labeled samples to boost model performance. But in fact, a large number of augmented training samples not only introduce extra noise, but also additional training time. In this paper, we introduce the active learning into SAR target recognition, which used to help specialists select the samples that are most worth labeling. Moreover, we propose a data-driven active learning scheme named Ranking Loss Module (RLM), which does not rely on artificially strategies to select samples. In contrast to data augmentation, it can improve the performance of the model while reducing the number of training data samples. Experimental results based on MSTAR dataset demonstrate the superiority of the proposed scheme. Using the RLM, better performance can be achieved with only one third of the full training samples.
Yue Zhou 0005, Xue Jiang 0001, Zhou Li 0002, Xingzhao Liu
IGARSS1
2020 An Attention Enhanced Graph Convolutional Network for Semantic Segmentation
Ao Chen 0003, Yue Zhou 0005
PRCV (1)2
2018 Hierarchical Online Multi-person Pose Tracking with Multiple Cues
Chuanzhi Xu, Yue Zhou 0005
ICONIP (6)2
2018 Image Super-Resolution Based on Dense Convolutional Network
Yue Zhou 0005
PRCV (2)2
2018 Human Trajectory Prediction with Social Information Encoding
Siqi Ren, Yue Zhou 0005, Liming He
PRCV (1)2
2018 Consistent Online Multi-object Tracking with Part-Based Deep Network
Chuanzhi Xu, Yue Zhou 0005
PRCV (2)2
2017 A Point and Line Features Based Method for Disturbed Surface Motion Estimation
Xiang Li 0041, Yue Zhou 0005
ICONIP (3)2
2017 Visual Saliency Based Blind Image Quality Assessment via Convolutional Neural Network
Yue Zhou 0005
ICONIP (6)2
2017 Multi-camera Tracking Exploiting Person Re-ID Technique
Yiming Liang, Yue Zhou 0005
ICONIP (3)2
2017 Online Tracking with Convolutional Neural Networks
Yue Zhou 0005
ICONIP (3)2
2016 Bi-Lp-Norm Sparsity Pursuiting Regularization for Blind Motion Deblurring
Wanlin Gan, Yue Zhou 0005, Liming He
ICONIP (2)2
2016 A New Blind Image Quality Assessment Based on Pairwise
Jianbin Jiang, Yue Zhou 0005, Liming He
ICONIP (4)2
2016 Robust Part-Based Correlation Tracking
Yue Zhou 0005
ICONIP (2)2
2016 Fusion of Multi-view Multi-exposure Images with Delaunay Triangulation
Hanyi Yu, Yue Zhou 0005
ICONIP (2)2
2014 Online Detection of Concept Drift in Visual Tracking
Yue Zhou 0005
ICONIP (3)2
2013 Superpixel based color contrast and color distribution driven salient object detection
Keren Fu, Chen Gong 0002, Jie Yang 0002, Yue Zhou 0005, Irene Y. H. Gu
Signal Process. Image Commun.4
2012 Object Tracking within the Framework of Concept Drift
Yue Zhou 0005, Jie Yang 0002
ACCV (3)2
2012 Salient Object Detection via Color Contrast and Color Distribution
Keren Fu, Chen Gong 0002, Jie Yang 0002, Yue Zhou 0005
ACCV (1)4
2011 Robust facial feature points extraction in color images
Yue Zhou 0005, Yin Li 0003, Meilin Ge
Eng. Appl. Artif. Intell.1
2010 Optimum Subspace Learning and Error Correction for Tensors
Yin Li 0003, Junchi Yan, Yue Zhou 0005, Jie Yang 0002
ECCV (3)3
2010 Tensor error correction for corrupted values in visual data
abstract
The multi-channel image or the video clip has the natural form of tensor. The values of the tensor can be corrupted due to noise in the acquisition process. We consider the problem of recovering a tensor L of visual data from its corrupted observations X = L + S, where the corrupted entries S are unknown and unbounded, but are assumed to be sparse. Our work is built on the recent studies about the recovery of corrupted low-rank matrix via trace norm minimization. We extend the matrix case to the tensor case by the definition of tensor trace norm in [6]. Furthermore, the problem of tensor is formulated as a convex optimization, which is much harder than its matrix form. Thus, we develop a high quality algorithm to efficiently solve the problem. Our experiments show potential applications of our method and indicate a robust and reliable solution.
Yin Li 0003, Yue Zhou 0005, Junchi Yan, Jie Yang 0002, Xiangjian He
ICIP2
2010 A Novel Approach to Detect Ship-Radiated Signal Based on HMT
abstract
In the presence of non-gaussian noise, we propose a method for the detection of underwater ship-radiated signal. The wavelet decomposition of the underwater signal yields a natural tree structure, which is further modeled by the Hidden Markov Tree (HMT). Therefore, the signal is represented as the parameter of the correspondent HMT. We analysis the likelihood defined on the parameters and form the new detection criteria. Experimental results demonstrate a reliable and robust solution of our method.
Yue Zhou 0005, Zhibin Niu, Chenghao Wang 0007
ICPR1
2009 Visual Saliency Based on Conditional Entropy
Yin Li 0003, Yue Zhou 0005, Junchi Yan, Zhibin Niu, Jie Yang 0002
ACCV (1)2
2009 Incremental sparse saliency detection
abstract
By the guidance of attention, human visual system is able to locate objects of interest in complex scene. We propose a new visual saliency detection model for both image and video. Inspired by biological vision, saliency is defined locally. Lossy compression is adopted, where the saliency of a location is measured by the Incremental Coding Length(ICL). The ICL is computed by presenting the center patch as the sparsest linear representation of its surroundings. The final saliency map is generated by accumulating the coding length. The model is tested on both images and videos. The results indicate a reliable and robust saliency of our method.
Yin Li 0003, Yue Zhou 0005, Xiaochao Yang, Jie Yang 0002
ICIP2
2008 Bayesian nonstationary source separation
Qinghua Huang, Jie Yang 0002, Yue Zhou 0005
Neurocomputing3
2008 Gait recognition based on dynamic region analysis
Xiaochao Yang, Yue Zhou 0005, Tianhao Zhang 0002, Guang Shu, Jie Yang 0002
Signal Process.2
2007 Variational Bayesian method for speech enhancement
Qinghua Huang, Jie Yang 0002, Yue Zhou 0005
Neurocomputing3
2007 Subspace evolution analysis for face representation and recognition
Huahua Wang, Yue Zhou 0005, Xinliang Ge, Jie Yang 0002
Pattern Recognit.2
2007 Region partition and feature matching based color recognition of tongue image
Jie Yang 0002, Yue Zhou 0005
Pattern Recognit. Lett.3
2006 Color Texture Segmentation Based on Quaternion-Gabor Features
Yue Zhou 0005, Yong-Gang Wang
CIARP2
2006 Color Texture Segmentation Using Quaternion-Gabor Filters
abstract
A simple and fast quaternion-Gabor filter is introduced for color texture segmentation in this paper. The algorithm achieves color texture multichannel Gabor filtering via DRBFT and IDRBFT. Despite the simplicity of the whole algorithm, the segmentation results are rather encouraging. In addition, the proposed algorithm is compatible with the measurement of gray-value texture.
Hui Wang 0012, Yue Zhou 0005, Jie Yang 0002
ICIP3
2006 Variational Bayesian Method for Temporally Correlated Source Separation
Qinghua Huang, Jie Yang 0002, Yue Zhou 0005
ICONIP (1)3
2006 Flexible background mixture models for foreground segmentation
Jian Cheng 0001, Jie Yang 0002, Yue Zhou 0005, Yingying Cui
Image Vis. Comput.3
2006 GACV: Geodesic-Aided C-V method
Yue Zhou 0005, Jie Yang 0002
Pattern Recognit.2
2004 An Extended Speech De-noising Method Using GGM-Based ICA Feature Extraction
Yue Zhou 0005, Jie Yang 0002
CIARP2
2004 JSEG Based Color Separation of Tongue Image in Traditional Chinese Medicine
Yue Zhou 0005, Jie Yang 0002
CIARP2
2004 Unsupervised Segmentation on Image with JSEG Using Soft Class Map
Yuanjie Zheng, Jie Yang 0002, Yue Zhou 0005
IDEAL3
2004 Unsupervised Image Segmentation with Fuzzy Connectedness
Yuanjie Zheng, Jie Yang 0002, Yue Zhou 0005
PRICAI3