Xiaowei Hu 0001

dblp:151/5859-1 · DBLP profile ↗
← Back
58ranked-venue papers
11as first author
39since 2021 · last 2026
0000-0002-5708-7018ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 7 first-author · 27 since 2021Artificial intelligence and machine learning · 33 · 7 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI
abstract
Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a 7B-parameter general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis.
Tianbin Li, Yanzhou Su, Wei Li 0320, Zhe Chen 0017, Ziyan Huang, Guoan Wang, Chenglong Ma 0002, Yanjun Li 0007, Shixiang Tang, Xiaowei Hu 0001, Zhongying Deng, Yuanfeng Ji, Jin Ye 0002, Yu Qiao 0001, Junjun He
AAAI14
2026 SurgPub-Video: A Comprehensive Surgical Video Framework for Enhanced Surgical Intelligence in Vision-Language Model
abstract
Vision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these challenges, we make the following contributions: (i) SurgPub-Video, a comprehensive dataset of over 3,000 surgical videos and 25 million annotated frames across 11 specialities, sourced from peer-reviewed clinical journals, (ii) SurgLLaVA-Video, a specialized VLM for surgical video understanding, built upon the TinyLLaVA-Video architecture that supports both video-level and frame-level inputs, and (iii) a video-level surgical Visual Question Answering (VQA) benchmark, covering diverse 11 surgical specialities, such as vascular, cardiology, and thoracic. Extensive experiments, conducted on the proposed benchmark and three additional surgical downstream tasks (action recognition, skill assessment, and triplet recognition), show that SurgLLaVA-Video significantly outperforms both general-purpose and surgical-specific VLMs with only three billion parameters.
Yaoqian Li, Xikai Yang, Dunyuan Xu, Litao Zhao, Xiaowei Hu 0001, Jinpeng Li 0004, Pheng-Ann Heng
AAAI6
2026 IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
Guibao Shen, Quande Liu, Jialin Gao, Lan Du 0002, Cunjian Chen, Chi-Wing Fu, Xiaowei Hu 0001, Pheng-Ann Heng
AAAI10
2026 Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
Xiaowei Hu 0001, Zhenghao Xing, Tianyu Wang 0003, Chi-Wing Fu, Pheng-Ann Heng
Int. J. Comput. Vis.1
2026 Moving Beyond Functional Connectivity: Time-Series Modeling for fMRI-Based Brain Disorder Classification
abstract
Functional magnetic resonance imaging (fMRI) enables non-invasive brain disorder classification by capturing blood-oxygen-level-dependent (BOLD) signals. However, most existing methods rely on functional connectivity (FC) via Pearson correlation, which reduces 4D BOLD signals to static 2D matrices-discarding temporal dynamics and capturing only linear inter-regional relationships. In this work, we benchmark state-of-the-art temporal models (e.g., time-series models: PatchTST, TimesNet, TimeMixer) on raw BOLD signals across five public datasets. Results show these models consistently outperform traditional FC-based approaches, highlighting the value of directly modeling temporal information such as cycle-like oscillatory fluctuations and drift-like slow baseline trends. Building on this insight, we propose DeCI, a simple yet effective framework that integrates two key principles: (i) Cycle and Drift Decomposition to disentangle cycle and drift within each ROI (Region of Interest); and (ii) Channel-Independence to model each ROI separately, improving robustness and reducing overfitting. Extensive experiments demonstrate that DeCI achieves superior classification accuracy and generalization compared to both FC-based and temporal baselines. Our findings advocate for a shift toward end-to-end temporal modeling in fMRI analysis to better capture complex brain dynamics. The code is available at https://github.com/Levi-Ackman/DeCI.
Guoqi Yu, Xiaowei Hu 0001, Angelica I. Avilés-Rivero, Anqi Qiu
IEEE Trans. Medical Imaging2
2025 EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights
abstract
Traffic Anomaly Understanding (TAU) is essential for improving public safety and transportation efficiency by enabling timely detection and response to incidents. Beyond existing methods, which rely largely on visual data, we propose to consider audio cues, a valuable source that offers strong hints to anomaly scenarios such as crashes and honking. Our contributions are twofold. First, we compile AV-TAU, the first large-scale audio-visual dataset for TAU, providing 29,865 traffic anomaly videos and 149,325 Q&A pairs, while supporting five essential TAU tasks. Second, we develop EchoTraffic, a multimodal LLM that integrates audio and visual data for TAU, through our audio-insight frame selector and dynamic connector to effectively extract crucial audio cues for anomaly understanding with a two-phase training framework. Experimental results on AV-TAU manifest that EchoTraffic sets a new SOTA performance in TAU, outperforming the existing multimodal LLMs. Our contributions, including AV-TAU and EchoTraffic, pave a new direction for multimodal TAU.
Zhenghao Xing, Hao Chen 0193, Binzhu Xie, Xuemiao Xu, Jianye Hao, Chi-Wing Fu, Xiaowei Hu 0001, Pheng-Ann Heng
CVPR9
2025 Fast Image Super-Resolution via Consistency Rectified Flow
Wenbo Li 0002, Haoze Sun, Zhixin Wang, Long Peng 0003, Xiaowei Hu 0001, Renjing Pei, Pheng-Ann Heng
ICCV9
2025 MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models can generate high-quality images but lack fine-grained control of visual concepts, limiting their creativity. Thus, we introduce component-controllable personalization, a new task that enables users to customize and reconfigure individual components within concepts. This task faces two challenges: semantic pollution, where undesired elements disrupt the target concept, and semantic imbalance, which causes disproportionate learning of the target concept and component. To address these, we design MagicTailor, a framework that uses Dynamic Masked Degradation to adaptively perturb unwanted visual semantics and Dual-Stream Balancing for more balanced learning of desired visual semantics. The experimental results show that MagicTailor achieves superior performance in this task and enables more personalized and creative image generation.
Jiancheng Huang, Jinbin Bai, Hao Chen 0193, Guangyong Chen, Xiaowei Hu 0001, Pheng-Ann Heng
IJCAI7
2025 Device-Cloud Collaborative Learning Framework for Efficient Unknown Object Detection
abstract
Unknown object detection aims to build detectors capable of identifying out-of-distribution objects, a critical need for applications like autonomous driving and traffic monitoring. However, limited device resources restrict existing methods from achieving accurate detection on the device side. Addressing this gap, this paper introduces a device-cloud collaborative framework named DCCUOD that enhances device model performance through efficient cloud collaboration. Our framework employs an energy-based sampling function on devices to target samples with unknown objects, coupled with a collaborative pseudo-labeling strategy to generate accurate pseudo-labels. Additionally, a two-stage training paradigm enables continuous improvements of device models on both known and unknown objects. Our study is the first to explore device-cloud collaborative learning for UOD tasks. Experimental results show that the device model is three times smaller and seven times faster than cloud models, with minimal performance trade-offs.
Kewei Zhao, Xiaowei Hu 0001, Qinya Li
ACM Multimedia2
2025 Real-World Adverse Weather Image Restoration via Dual-Level Reinforcement Learning with High-Quality Cold Start
abstract
Adverse weather severely impairs real-world visual perception, while existing vision models trained on synthetic data with fixed parameters struggle to generalize to complex degradations. To address this, we first construct HFLS-Weather, a physics-driven, high-fidelity dataset that simulates diverse weather phenomena, and then design a dual-level reinforcement learning framework initialized with HFLS-Weather for cold-start training. Within this framework, at the local level, weather-specific restoration models are refined through perturbation-driven image quality optimization, enabling reward-based learning without paired supervision; at the global level, a meta-controller dynamically orchestrates model selection and execution order according to scene degradation. This framework enables continuous adaptation to real-world conditions and achieves state-of-the-art performance across a wide range of adverse weather scenarios. Code is available at https://github.com/xxclfy/AgentRL-Real-Weather
Fuyang Liu, Xiaowei Hu 0001
NeurIPS3
2025 SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
abstract
Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativity in practice. This paper introduces scene-oriented story generation, addressing two key challenges: (i) scene planning, where current methods fail to ensure scene-level narrative coherence by relying solely on text descriptions, and (ii) scene consistency, which remains largely unexplored in terms of maintaining scene consistency across multiple stories. We propose SceneDecorator, a training-free framework that employs VLM-Guided Scene Planning to ensure narrative coherence across different scenes in a ``global-to-local'' manner, and Long-Term Scene-Sharing Attention to maintain long-term scene consistency and subject diversity across generated stories. Extensive experiments demonstrate the superior performance of SceneDecorator, highlighting its potential to unleash creativity in the fields of arts, films, and games.
Quanjian Song, Fei Shen 0004, Xiaowei Hu 0001, Cunjian Chen, Pheng-Ann Heng
NeurIPS6
2025 Demystify Transformers & Convolutions in Modern Image Deep Networks
abstract
Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these advancements are not solely attributable to novel feature transformation designs; certain benefits also arise from advanced network-level and block-level architectures. This paper aims to identify the real gains of popular convolution and attention operators through a detailed study. We find that the key difference among these feature transformation modules, such as attention or convolution, lies in their spatial feature aggregation approach, known as the "spatial token mixer" (STM). To facilitate an impartial comparison, we introduce a unified architecture to neutralize the impact of divergent network-level and block-level designs. Subsequently, various STMs are integrated into this unified framework for comprehensive comparative analysis. Our experiments on various tasks and an analysis of inductive bias show a significant performance boost due to advanced network-level and block-level designs, but performance differences persist among different STMs. Our detailed analysis also reveals various findings about different STMs, including effective receptive fields, invariance, and adversarial robustness tests.
Xiaowei Hu 0001, Min Shi 0004, Weiyun Wang, Sitong Wu, Linjie Xing, Wenhai Wang, Xizhou Zhou, Lewei Lu, Jie Zhou 0001, Xiaogang Wang 0005, Yu Qiao 0001, Jifeng Dai
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Unifying Physically-Informed Weather Priors in a Single Model for Image Restoration Across Multiple Adverse Weather Conditions
abstract
Image restoration under multiple adverse weather conditions aims to develop a single model to recover the underlying scene with high visibility. Weather-related artifacts vary with the particle’s distance to the camera according to the established scene visibility analysis, where close and faraway regions are more affected by falling drops and fog effects, respectively. In challenging weather conditions, existing image restoration methods fall short by not accounting for the varying impact of adverse weather on different scene regions. We develop a novel unified imaging model combined with a weather-prior-based network that directly incorporates weather-specific physical imaging processes into the restoration process. This approach not only enhances visibility in both near and distant regions affected by drops but also outperforms current state-of-the-art methods by effectively mitigating artifacts such as fog. Our contributions include a comprehensive analysis of weather-related visual factors and the development of an innovative network architecture that leverages estimated occlusion and transmission to restore scene details. Experimental results on three synthetic benchmarks, including our Weather30K dataset, along with two all-weather datasets, and a real-world benchmark with challenging mixed weather conditions, show the superiority of our method against state-of-the-art methods.
Xiaowei Hu 0001, Lei Zhu 0003, Pheng-Ann Heng
IEEE Trans. Circuits Syst. Video Technol.2
2024 Semi-supervised TEE Segmentation via Interacting with SAM Equipped with Noise-Resilient Prompting
abstract
Semi-supervised learning (SSL) is a powerful tool to address the challenge of insufficient annotated data in medical segmentation problems. However, existing semi-supervised methods mainly rely on internal knowledge for pseudo labeling, which is biased due to the distribution mismatch between the highly imbalanced labeled and unlabeled data. Segmenting left atrial appendage (LAA) from transesophageal echocardiogram (TEE) images is a typical medical image segmentation task featured by scarcity of professional annotations and diverse data distributions, for which existing SSL models cannot achieve satisfactory performance. In this paper, we propose a novel strategy to mitigate the inherent challenge of distribution mismatch in SSL by, for the first time, incorporating a large foundation model (i.e. SAM in our implementation) into an SSL model to improve the quality of pseudo labels. We further propose a new self-reconstruction mechanism to generate both noise-resilient prompts to demonically improve SAM’s generalization capability over TEE images and self-perturbations to stabilize the training process and reduce the impact of noisy labels. We conduct extensive experiments on an in-house TEE dataset; experimental results demonstrate that our method achieves better performance than state-of-the-art SSL models.
Yidan Feng, Haoneng Lin, Yiting Fan, Alex Pui-Wai Lee, Xiaowei Hu 0001, Harry Qin
AAAI6
2024 Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling
abstract
Predicting multivariate time series is crucial, demanding precise modeling of intricate patterns, including inter-series dependencies and intra-series variations. Distinctive trend characteristics in each time series pose challenges, and existing methods, relying on basic moving average kernels, may struggle with the non-linear structure and complex trends in real-world data. Given that, we introduce a learnable decomposition strategy to capture dynamic trend information more reasonably. Additionally, we propose a dual attention module tailored to capture inter-series dependencies and intra-series variations simultaneously for better time series forecasting, which is implemented by channel-wise self-attention and autoregressive self-attention. To evaluate the effectiveness of our method, we conducted experiments across eight open-source datasets and compared it with the state-of-the-art methods. Through the comparison results, our $\textbf{Leddam}$ ($\textbf{LE}arnable$ $\textbf{D}ecomposition$ and $\textbf{D}ual $ $\textbf{A}ttention$ $\textbf{M}odule$) not only demonstrates significant advancements in predictive performance but also the proposed decomposition strategy can be plugged into other methods with a large performance-boosting, from 11.87% to 48.56% MSE error degradation. Code is available at this link: https://github.com/Levi-Ackman/Leddam.
Guoqi Yu, Xiaowei Hu 0001, Angelica I. Avilés-Rivero, Harry Qin
ICML3
2024 TrafficMOT: A Challenging Dataset for Multi-Object Tracking in Complex Traffic Scenarios
abstract
ACM Multimedia 2024, Melbourne, Australia, Oct 28 - Nov 1, 2024
Yanqi Cheng, Zhongying Deng, Dongdong Chen 0001, Xiaowei Hu 0001, Pietro Liò, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
ACM Multimedia6
2024 Video Instance Shadow Detection Under the Sun and Sky
abstract
Instance shadow detection, crucial for applications such as photo editing and light direction estimation, has undergone significant advancements in predicting shadow instances, object instances, and their associations. The extension of this task to videos presents challenges in annotating diverse video data and addressing complexities arising from occlusion and temporary disappearances within associations. In response to these challenges, we introduce ViShadow, a semi-supervised video instance shadow detection framework that leverages both labeled image data and unlabeled video data for training. ViShadow features a two-stage training pipeline: the first stage, utilizing labeled image data, identifies shadow and object instances through contrastive learning for cross-frame pairing. The second stage employs unlabeled videos, incorporating an associated cycle consistency loss to enhance tracking ability. A retrieval mechanism is introduced to manage temporary disappearances, ensuring tracking continuity. The SOBA-VID dataset, comprising unlabeled training videos and labeled testing videos, along with the SOAP-VID metric, is introduced for the quantitative evaluation of VISD solutions. The effectiveness of ViShadow is further demonstrated through various video-level applications such as video inpainting, instance cloning, shadow editing, and text-instructed shadow-object manipulation.
Zhenghao Xing, Tianyu Wang 0003, Xiaowei Hu 0001, Chi-Wing Fu, Pheng-Ann Heng
IEEE Trans. Image Process.3
2024 Dynamic Message Propagation Network for RGB-D and Video Salient Object Detection
abstract
Exploiting long-range semantic contexts and geometric information is crucial to infer salient objects from RGB and depth features. However, existing methods mainly focus on excavating local features within fixed regions by continuously feeding forward networks. In this article, we introduce Dynamic Message Propagation (DMP) to dynamically learn context information within more flexible regions. We integrate DMP into a Siamese-based network to process the RGB image and depth map separately and design a multi-level feature fusion module to explore cross-level information between refined RGB and depth features. Extensive experiments show clear improvements of our method over 17 methods on six benchmark datasets for RGB-D salient object detection (SOD). Additionally, our method outperforms its competitors for the video SOD task. Code is available at https://github.com/chenbaian-cs/DMPNet .
Baian Chen, Zhilei Chen, Xiaowei Hu 0001, Jun Xu 0019, Haoran Xie 0001, Harry Qin, Mingqiang Wei
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Matching Is Not Enough: A Two-Stage Framework for Category-Agnostic Pose Estimation
abstract
Category-agnostic pose estimation (CAPE) aims to predict keypoints for arbitrary categories given support images with keypoint annotations. Existing approaches match the keypoints across the image for localization. However, such a one-stage matching paradigm shows inferior accuracy: the prediction heavily relies on the matching results, which can be noisy due to the open set nature in CAPE. For example, two mirror-symmetric keypoints (e.g., left and right eyes) in the query image can both trigger high similarity on certain support keypoints (eyes), which leads to duplicated or opposite predictions. To calibrate the inaccurate matching results, we introduce a two-stage framework, where matched keypoints from the first stage are viewed as similarity-aware position proposals. Then, the model learns to fetch relevant features to correct the initial proposals in the second stage. We instantiate the framework with a transformer model tailored for CAPE. The transformer encoder incorporates specific designs to improve the representation and similarity modeling in the first matching stage. In the second stage, similarity-aware proposals are packed as queries in the decoder for refinement via cross-attention. Our method surpasses the previous best approach by large margins on CAPE benchmark MP-100 on both accuracy and efficiency. Code available at github.com/flyinglynx/CapeFormer
Min Shi 0004, Zihao Huang 0001, Xianzheng Ma, Xiaowei Hu 0001, Zhiguo Cao 0001
CVPR4
2023 InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
abstract
Compared to the great progress of large-scale vision transformers (ViTs) in recent years, large-scale models based on convolutional neural networks (CNNs) are still in an early state. This work presents a new large-scale CNN-based foundation model, termed InternImage, which can obtain the gain from increasing parameters and training data like ViTs. Different from the recent CNNs that focus on large dense kernels, InternImage takes deformable convolution as the core operator, so that our model not only has the large effective receptive field required for downstream tasks such as detection and segmentation, but also has the adaptive spatial aggregation conditioned by input and task information. As a result, the proposed InternImage reduces the strict inductive bias of traditional CNNs and makes it possible to learn stronger and more robust patterns with large-scale parameters from massive data like ViTs. The effectiveness of our model is proven on challenging benchmarks including ImageNet, COCO, andADE20K. It is worth mentioning that InternImage-H achieved a new record 65.4 mAP on COCO test-dev and 62.9 mIoU on ADE20K, outperforming current leading CNNs and ViTs.
Wenhai Wang, Jifeng Dai, Zhe Chen 0017, Zhenhang Huang, Xizhou Zhu, Xiaowei Hu 0001, Tong Lu 0002, Lewei Lu, Hongsheng Li 0001, Xiaogang Wang 0001, Yu Qiao 0001
CVPR7
2023 Video Dehazing via a Multi-Range Temporal Alignment Network with Physical Prior
abstract
Video dehazing aims to recover haze-free frames with high visibility and contrast. This paper presents a novel framework to effectively explore the physical haze priors and aggregate temporal information. Specifically, we design a memory-based physical prior guidance module to encode the prior-related features into long-range memory. Besides, we formulate a multi-range scene radiance recovery module to capture space-time dependencies in multiple space-time ranges, which helps to effectively aggregate temporal information from adjacent frames. Moreover, we construct the first large-scale outdoor video dehazing benchmark dataset, which contains videos in various real-world scenarios. Experimental results on both synthetic and real conditions show the superiority of our proposed method.
Xiaowei Hu 0001, Lei Zhu 0003, Qi Dou 0001, Jifeng Dai, Yu Qiao 0001, Pheng-Ann Heng
CVPR2
2023 Learning Weather-General and Weather-Specific Features for Image Restoration Under Multiple Adverse Weather Conditions
abstract
Image restoration under multiple adverse weather conditions aims to remove weather-related artifacts by using a single set of network parameters. In this paper, we find that image degradations under different weather conditions contain general characteristics as well as their specific characteristics. Inspired by this observation, we design an efficient unified framework with a two-stage training strategy to explore the weather-general and weather-specific features. The first training stage aims to learn the weather-general features by taking the images under various weather conditions as inputs and outputting the coarsely restored results. The second training stage aims to learn to adaptively expand the specific parameters for each weather type in the deep model, where the requisite positions for expanding weather-specific parameters are automatically learned. Hence, we can obtain an efficient and unified model for image restoration under multiple adverse weather conditions. Moreover, we build the first real-world benchmark dataset with multiple weather conditions to better deal with realworld weather scenarios. Experimental results show that our method achieves superior performance on all the synthetic and real-world benchmarks. Codes and datasets are available at this repository.
Yurui Zhu, Tianyu Wang 0003, Xueyang Fu, Xuanyu Yang, Xin Guo 0018, Jifeng Dai, Yu Qiao 0001, Xiaowei Hu 0001
CVPR8
2023 SILT: Shadow-aware Iterative Label Tuning for Learning to Detect Shadows from Noisy Labels
abstract
Existing shadow detection datasets often contain missing or mislabeled shadows, which can hinder the performance of deep learning models trained directly on such data. To address this issue, we propose SILT, the Shadow-aware Iterative Label Tuning framework, which explicitly considers noise in shadow labels and trains the deep model in a self-training manner. Specifically, we incorporate strong data augmentations with shadow counterfeiting to help the network better recognize non-shadow regions and alleviate overfitting. We also devise a simple yet effective label tuning strategy with global-local fusion and shadow-aware filtering to encourage the network to make significant refinements on the noisy labels. We evaluate the performance of SILT by relabeling the test set of the SBU [55] dataset and conducting various experiments. Our results show that even a simple U-Net [42] trained with SILT can outperform all state-of-the-art methods by a large margin. When trained on SBU / UCF [78] / ISTD [56], our network can successfully reduce the Balanced Error Rate by 25.2% / 36.9% / 21.3% over the best state-of-the-art method.
Tianyu Wang 0003, Xiaowei Hu 0001, Chi-Wing Fu
ICCV3
2023 IDRNet: Intervention-Driven Relation Network for Semantic Segmentation
abstract
Co-occurrent visual patterns suggest that pixel relation modeling facilitates dense prediction tasks, which inspires the development of numerous context modeling paradigms, \emph{e.g.}, multi-scale-driven and similarity-driven context schemes. Despite the impressive results, these existing paradigms often suffer from inadequate or ineffective contextual information aggregation due to reliance on large amounts of predetermined priors. To alleviate the issues, we propose a novel \textbf{I}ntervention-\textbf{D}riven \textbf{R}elation \textbf{Net}work (\textbf{IDRNet}), which leverages a deletion diagnostics procedure to guide the modeling of contextual relations among different pixels. Specifically, we first group pixel-level representations into semantic-level representations with the guidance of pseudo labels and further improve the distinguishability of the grouped representations with a feature enhancement module. Next, a deletion diagnostics procedure is conducted to model relations of these semantic-level representations via perceiving the network outputs and the extracted relations are utilized to guide the semantic-level representations to interact with each other. Finally, the interacted representations are utilized to augment original pixel-level representations for final predictions. Extensive experiments are conducted to validate the effectiveness of IDRNet quantitatively and qualitatively. Notably, our intervention-driven context scheme brings consistent performance improvements to state-of-the-art segmentation frameworks and achieves competitive results on popular benchmark datasets, including ADE20K, COCO-Stuff, PASCAL-Context, LIP, and Cityscapes.
Zhenchao Jin, Xiaowei Hu 0001, Lingting Zhu, Luchuan Song, Lequan Yu
NeurIPS2
2023 Instance Shadow Detection With a Single-Stage Detector
abstract
This article formulates a new problem, instance shadow detection, which aims to detect shadow instance and the associated object instance that cast each shadow in the input image. To approach this task, we first compile a new dataset with the masks for shadow instances, object instances, and shadow-object associations. We then design an evaluation metric for quantitative evaluation of the performance of instance shadow detection. Further, we design a single-stage detector to perform instance shadow detection in an end-to-end manner, where the bidirectional relation learning module and the deformable maskIoU head are proposed in the detector to directly learn the relation between shadow instances and object instances and to improve the accuracy of the predicted masks. Finally, we quantitatively and qualitatively evaluate our method on the benchmark dataset of instance shadow detection and show the applicability of our method on light direction estimation and photo editing.
Tianyu Wang 0003, Xiaowei Hu 0001, Pheng-Ann Heng, Chi-Wing Fu
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Deep Texture-Aware Features for Camouflaged Object Detection
abstract
Camouflaged object detection is a challenging task that aims to identify objects having similar texture to the surroundings. This paper presents to amplify the subtle texture difference between camouflaged objects and the background for camouflaged object detection by formulating multiple texture-aware refinement modules to learn the texture-aware features in a deep convolutional neural network. The texture-aware refinement module computes the biased co-variance matrices of feature responses to extract the texture information, adopts an affinity loss to learn a set of parameter maps that help to separate the texture between camouflaged objects and the background, and leverages a boundary-consistency loss to explore the structures of object details. We evaluate our network on the benchmark datasets for camouflaged object detection both qualitatively and quantitatively. Experimental results show that our approach outperforms various state-of-the-art methods by a large margin.
Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Yangyang Xu 0003, Weiming Wang 0002, Zijun Deng, Pheng-Ann Heng
IEEE Trans. Circuits Syst. Video Technol.2
2023 Representative Feature Alignment for Adaptive Object Detection
abstract
Unsupervised domain adaptation for object detection aims to generalize the object detector trained on the label-rich source domain to the unlabeled target domain. Recently, existing works adopt the instance-level alignment or pixel-level alignment to perform domain transfer, which can effectively avoid the negative transfer due to the diverse background between domains. However, we find that they treat all the regions of an instance feature equally without suppressing background area. They do not segment the specific texture and discriminative regions of objects, which are transferable during adaptation. We call the features that combine the local structure feature and semantic discriminant features as representative features. We propose a novel Representative Feature Alignment (RFA) model to align the features extracted from representative patterns of objects, i.e. representative features, for domain adaptation. Specifically, the representative features are extracted by the Representative Feature Extraction (RFE) submodules. The RFE submodules take the features extracted from different intermediate layers of the detector as input, and filter out the representative features layer-by-layer via integrating class weighting generator, category selection and class activation mapping. Then the representative features from multi-layers are further adaptively aggregated to obtain the final representative features, which are utilized to conduct feature alignment in a class-aware manner. Our representative features are free of untransferable regions and background areas, which leads to better feature alignment. Extensive experimental results show that the proposed model outperforms state-of-the-art methods on a few benchmark datasets.
Shan Xu 0006, Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Yangyang Xu 0003, Liangui Dai, Kup-Sze Choi, Pheng-Ann Heng
IEEE Trans. Circuits Syst. Video Technol.4
2022 Enhancing Pseudo Label Quality for Semi-supervised Domain-Generalized Medical Image Segmentation
abstract
Generalizing the medical image segmentation algorithms to unseen domains is an important research topic for computer-aided diagnosis and surgery. Most existing methods require a fully labeled dataset in each source domain. Although some researchers developed a semi-supervised domain generalized method, it still requires the domain labels. This paper presents a novel confidence-aware cross pseudo supervision algorithm for semi-supervised domain generalized medical image segmentation. The main goal is to enhance the pseudo label quality for unlabeled images from unknown distributions. To achieve it, we perform the Fourier transformation to learn low-level statistic information across domains and augment the images to incorporate cross-domain information. With these augmentations as perturbations, we feed the input to a confidence-aware cross pseudo supervision network to measure the variance of pseudo labels and regularize the network to learn with more confident pseudo labels. Our method sets new records on public datasets, i.e., M&Ms and SCGM. Notably, without using domain labels, our method surpasses the prior art that even uses domain labels by 11.67% on Dice on M&Ms dataset with 2% labeled data. Code is available at https://github.com/XMed-Lab/EPL SemiDG.
Huifeng Yao, Xiaowei Hu 0001, Xiaomeng Li 0001
AAAI2
2022 Learning Shadow Correspondence for Video Shadow Detection
Xinpeng Ding, Xiaowei Hu 0001, Xiaomeng Li 0001
ECCV (17)3
2022 Sparse2Dense: Learning to Densify 3D Features for 3D Object Detection
abstract
LiDAR-produced point clouds are the major source for most state-of-the-art 3D object detectors. Yet, small, distant, and incomplete objects with sparse or few points are often hard to detect. We present Sparse2Dense, a new framework to efficiently boost 3D detection performance by learning to densify point clouds in latent space. Specifically, we first train a dense point 3D detector (DDet) with a dense point cloud as input and design a sparse point 3D detector (SDet) with a regular point cloud as input. Importantly, we formulate the lightweight plug-in S2D module and the point cloud reconstruction module in SDet to densify 3D features and train SDet to produce 3D features, following the dense 3D features in DDet. So, in inference, SDet can simulate dense 3D features from regular (sparse) point cloud inputs without requiring dense inputs. We evaluate our method on the large-scale Waymo Open Dataset and the Waymo Domain Adaptation Dataset, showing its high performance and efficiency over the state of the arts.
Tianyu Wang 0003, Xiaowei Hu 0001, Zhengzhe Liu, Chi-Wing Fu
NeurIPS2
2021 Learning Semantic Context from Normal Samples for Unsupervised Anomaly Detection
abstract
Unsupervised anomaly detection aims to identify data samples that have low probability density from a set of input samples, and only the normal samples are provided for model training. The inference of abnormal regions on the input image requires an understanding of the surrounding semantic context. This work presents a Semantic Context based Anomaly Detection Network, SCADN, for unsupervised anomaly detection by learning the semantic context from the normal samples. To achieve this, we first generate multi-scale striped masks to remove a part of regions from the normal samples, and then train a generative adversarial network to reconstruct the unseen regions. Note that the masks are designed in multiple scales and stripe directions, and various training examples are generated to obtain the rich semantic context . In testing, we obtain an error map by computing the difference between the reconstructed image and the input image for all samples, and infer the abnormal samples based on the error maps. Finally, we perform various experiments on three public benchmark datasets and a new dataset LaceAD collected by us, and show that our method clearly outperforms the current state-of-the-art methods.
Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Pheng-Ann Heng
AAAI4
2021 Single-Stage Instance Shadow Detection With Bidirectional Relation Learning
abstract
Instance shadow detection aims to find shadow instances paired with the objects that cast the shadows. The previous work adopts a two-stage framework to first predict shadow instances, object instances, and shadow-object associations from the region proposals, then leverage a post-processing to match the predictions to form the final shadow-object pairs. In this paper, we present a new single-stage fully-convolutional network architecture with a bidirectional relation learning module to directly learn the relations of shadow and object instances in an end-to-end manner. Compared with the prior work, our method actively explores the internal relationship between shadows and objects to learn a better pairing between them, thus improving the overall performance for instance shadow detection. We evaluate our method on the benchmark dataset for instance shadow detection, both quantitatively and visually. The experimental results demonstrate that our method clearly outperforms the state-of-the-art method.
Tianyu Wang 0003, Xiaowei Hu 0001, Chi-Wing Fu, Pheng-Ann Heng
CVPR2
2021 Global guidance network for breast lesion segmentation in ultrasound images
Cheng Xue 0003, Lei Zhu 0003, Huazhu Fu, Xiaowei Hu 0001, Xiaomeng Li 0001, Pheng-Ann Heng
Medical Image Anal.4
2021 SAC-Net: Spatial Attenuation Context for Salient Object Detection
abstract
This paper presents a new deep neural network design for salient object detection by maximizing the integration of local and global image context within, around, and beyond the salient objects. Our key idea is to adaptively propagate and aggregate the image context features with variable attenuation over the entire feature maps. To achieve this, we design the spatial attenuation context (SAC) module to recurrently translate and aggregate the context features independently with different attenuation factors and then to attentively learn the weights to adaptively integrate the aggregated context features. By further embedding the module to process individual layers in a deep network, namely SAC-Net, we can train the network end-To-end and optimize the context features for detecting salient objects. Compared with 29 state-of-The-Art methods, experimental results show that our method performs favorably over all the others on six common benchmark data, both quantitatively and visually.
Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Tianyu Wang 0003, Pheng-Ann Heng
IEEE Trans. Circuits Syst. Video Technol.1
2021 Learning Gated Non-Local Residual for Single-Image Rain Streak Removal
abstract
This work presents a gated non-local deep residual learning framework for image deraining. It can avoid the over-deraining or under-deraining caused by the global residual learning in existing deraining networks, since the learned soft gate in our method adaptively adjusts the amount of global residual to be passed for generating the final derained result. To generate feature maps for global residual prediction, we develop a non-local guided attention module (NLAM), which first obtains non-local features by exploiting spatial inter-dependencies among all the feature positions of local features produced by convolutional neural network (CNN), and then leverages the attention mechanism to merge the local and non-local features based on their complementary relation. Moreover, we develop a channel-wise gated prediction module to learn a soft gate on the global residual by explicitly modelling channel inter-dependencies of the feature maps obtained from NLAM. Experiments on four deraining benchmark datasets and real-world rainy images show that our network has a quantitative and qualitative improvement over state-of-the-arts.
Lei Zhu 0003, Zijun Deng, Xiaowei Hu 0001, Haoran Xie 0001, Xuemiao Xu, Harry Qin, Pheng-Ann Heng
IEEE Trans. Circuits Syst. Video Technol.3
2021 Revisiting Shadow Detection: A New Benchmark Dataset for Complex World
abstract
Shadow detection in general photos is a nontrivial problem, due to the complexity of the real world. Though recent shadow detectors have already achieved remarkable performance on various benchmark data, their performance is still limited for general real-world situations. In this work, we collected shadow images for multiple scenarios and compiled a new dataset of 10,500 shadow images, each with labeled ground-truth mask, for supporting shadow detection in the complex world. Our dataset covers a rich variety of scene categories, with diverse shadow sizes, locations, contrasts, and types. Further, we comprehensively analyze the complexity of the dataset, present a fast shadow detection network with a detail enhancement module to harvest shadow details, and demonstrate the effectiveness of our method to detect shadows in general situations.
Xiaowei Hu 0001, Tianyu Wang 0003, Chi-Wing Fu, Yitong Jiang, Qiong Wang 0001, Pheng-Ann Heng
IEEE Trans. Image Process.1
2021 Single-Image Real-Time Rain Removal Based on Depth-Guided Non-Local Features
abstract
Rain is a common weather phenomenon that affects environmental monitoring and surveillance systems. According to an established rain model (Garg and Nayar, 2007), the scene visibility in the rain varies with the depth from the camera, where objects faraway are visually blocked more by the fog than by the rain streaks. However, existing datasets and methods for rain removal ignore these physical properties, thus limiting the rain removal efficiency on real photos. In this work, we analyze the visual effects of rain subject to scene depth and formulate a rain imaging model that collectively considers rain streaks and fog. Also, we prepare a dataset called RainCityscapes on real outdoor photos. Furthermore, we design a novel real-time end-to-end deep neural network, for which we train to learn the depth-guided non-local features and to regress a residual map to produce a rain-free output image. We performed various experiments to visually and quantitatively compare our method with several state-of-the-art methods to show its superiority over others.
Xiaowei Hu 0001, Lei Zhu 0003, Tianyu Wang 0003, Chi-Wing Fu, Pheng-Ann Heng
IEEE Trans. Image Process.1
2021 SALMNet: A Structure-Aware Lane Marking Detection Network
abstract
Lane marking detection is a fundamental task, which serves as an important prerequisite for automatic driving or driver-assistance systems. However, the complex and uncontrollable driving road environment as well as the discontinuous lane marking appearance make this task challenging. In this work, a novel deep neural network architecture is presented to detect lane markings in a complex environment by analyzing their structure information. There are two contributions to the network design. Firstly, a semantic-guided channel attention (SGCA) module is developed to select the low-level features of a deep convolutional neural network by taking the high-level features as the guidance. Secondly, a pyramid deformable convolution (PDC) module is formulated to enlarge the receptive fields and to capture the complex structures of lane markings by applying deformable convolutions on multiple feature maps with different scales. Hence, our network can better reduce false detection and enhance lane marking structures simultaneously. The experimental results on three benchmark datasets for lane marking detection show that our method outperforms other methods on all the benchmark datasets.
Xuemiao Xu, Tianfei Yu, Xiaowei Hu 0001, Wing W. Y. Ng, Pheng-Ann Heng
IEEE Trans. Intell. Transp. Syst.3
2021 Rotation-Oriented Collaborative Self-Supervised Learning for Retinal Disease Diagnosis
abstract
The automatic diagnosis of various conventional ophthalmic diseases from fundus images is important in clinical practice. However, developing such automatic solutions is challenging due to the requirement of a large amount of training data and the expensive annotations for medical images. This paper presents a novel self-supervised learning framework for retinal disease diagnosis to reduce the annotation efforts by learning the visual features from the unlabeled images. To achieve this, we present a rotation-oriented collaborative method that explores rotation-related and rotation-invariant features, which capture discriminative structures from fundus images and also explore the invariant property used for retinal disease classification. We evaluate the proposed method on two public benchmark datasets for retinal disease classification. The experimental results demonstrate that our method outperforms other self-supervised feature learning methods (around 4.2% area under the curve (AUC)). With a large amount of unlabeled data available, our method can surpass the supervised baseline for pathologic myopia (PM) and is very close to the supervised baseline for age-related macular degeneration (AMD), showing the potential benefit of our method in clinical practice.
Xiaomeng Li 0001, Xiaowei Hu 0001, Xiaojuan Qi 0001, Lequan Yu, Wei Zhao 0029, Pheng-Ann Heng, Lei Xing 0001
IEEE Trans. Medical Imaging2
2020 Instance Shadow Detection
abstract
Instance shadow detection is a brand new problem, aiming to find shadow instances paired with object instances. To approach it, we first prepare a new dataset called SOBA, named after Shadow-OBject Association, with 3,623 pairs of shadow and object instances in 1,000 photos, each with individual labeled masks. Second, we design LISA, named after Light-guided Instance Shadow-object Association, an end-to-end framework to automatically predict the shadow and object instances, together with the shadow-object associations and light direction. Then, we pair up the predicted shadow and object instances, and match them with the predicted shadow-object associations to generate the final results. In our evaluations, we formulate a new metric named the shadow-object average precision to measure the performance of our results. Further, we conducted various experiments and demonstrate our method's applicability on light direction estimation and photo editing.
Tianyu Wang 0003, Xiaowei Hu 0001, Qiong Wang 0001, Pheng-Ann Heng, Chi-Wing Fu
CVPR2
2020 GrabAR: Occlusion-aware Grabbing Virtual Objects in AR
abstract
Existing augmented reality (AR) applications often ignore the occlusion between real hands and virtual objects when incorporating virtual objects in user's views. The challenges come from the lack of accurate depth and mismatch between real and virtual depth. This paper presents GrabAR1, a new approach that directly predicts the real-and-virtual occlusion and bypasses the depth acquisition and inference. Our goal is to enhance AR applications with interactions between hand (real) and grabbable objects (virtual). With paired images of hand and object as inputs, we formulate a compact deep neural network that learns to generate the occlusion mask. To train the network, we compile a large dataset, including synthetic data and real data. We then embed the trained network in a prototyping AR system to support real-time grabbing of virtual objects. Further, we demonstrate the performance of our method on various virtual objects, compare our method with others through two user studies, and showcase a rich variety of interaction scenarios, in which we can use bare hand to grab virtual objects and directly manipulate them.
Xiao Tang 0005, Xiaowei Hu 0001, Chi-Wing Fu, Daniel Cohen-Or
UIST2
2020 Direction-Aware Spatial Context Features for Shadow Detection and Removal
abstract
Shadow detection and shadow removal are fundamental and challenging tasks, requiring an understanding of the global image semantics. This paper presents a novel deep neural network design for shadow detection and removal by analyzing the spatial image context in a direction-aware manner. To achieve this, we first formulate the direction-aware attention mechanism in a spatial recurrent neural network (RNN) by introducing attention weights when aggregating spatial context features in the RNN. By learning these weights through training, we can recover direction-aware spatial context (DSC) for detecting and removing shadows. This design is developed into the DSC module and embedded in a convolutional neural network (CNN) to learn the DSC features at different levels. Moreover, we design a weighted cross entropy loss to make effective the training for shadow detection and further adopt the network for shadow removal by using a euclidean loss function and formulating a color transfer function to address the color and luminosity inconsistencies in the training pairs. We employed two shadow detection benchmark datasets and two shadow removal benchmark datasets, and performed various experiments to evaluate our method. Experimental results show that our method performs favorably against the state-of-the-art methods for both shadow detection and shadow removal.
Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Harry Qin, Pheng-Ann Heng
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Aggregating Attentional Dilated Features for Salient Object Detection
abstract
This paper presents a novel deep learning model to aggregate the attentional dilated features for salient object detection by exploring the complementary information between the global and local context in a convolutional neural network. There are two technical contributions to our network design. First, we develop an attentional dense atrous (dilated) spatial pyramid pooling (AD-ASPP) module to selectively use the local saliency cues captured by dilated convolutions with a small rate and the global saliency cues captured by dilated convolutions with a large rate. Second, taking the feature pyramid network as the backbone, we develop an aggregation network to integrate the refined features by formulating two consecutive chains of residual learning based modules: one chain from deep to shallow layers while another chain from shallow to deep layers. We evaluate our network on seven widely-used saliency detection benchmarks by comparing it against 21 state-of-the-art methods. Experimental results show that our network outperforms others on all the seven benchmark datasets.
Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Harry Qin, Pheng-Ann Heng
IEEE Trans. Circuits Syst. Video Technol.3
2020 CANet: Cross-Disease Attention Network for Joint Diabetic Retinopathy and Diabetic Macular Edema Grading
abstract
Diabetic retinopathy (DR) and diabetic macular edema (DME) are the leading causes of permanent blindness in the working-age population. Automatic grading of DR and DME helps ophthalmologists design tailored treatments to patients, thus is of vital importance in the clinical practice. However, prior works either grade DR or DME, and ignore the correlation between DR and its complication, i.e., DME. Moreover, the location information, e.g., macula and soft hard exhaust annotations, are widely used as a prior for grading. Such annotations are costly to obtain, hence it is desirable to develop automatic grading methods with only image-level supervision. In this article, we present a novel cross-disease attention network (CANet) to jointly grade DR and DME by exploring the internal relationship between the diseases with only image-level supervision. Our key contributions include the disease-specific attention module to selectively learn useful features for individual diseases, and the disease-dependent attention module to further capture the internal relationship between the two diseases. We integrate these two attention modules in a deep network to produce disease-specific and disease-dependent features, and to maximize the overall performance jointly for grading DR and DME. We evaluate our network on two public benchmark datasets, i.e., ISBI 2018 IDRiD challenge dataset and Messidor dataset. Our method achieves the best result on the ISBI 2018 IDRiD challenge dataset and outperforms other methods on the Messidor dataset. Our code is publicly available at https://github.com/xmengli999/CANet.
Xiaomeng Li 0001, Xiaowei Hu 0001, Lequan Yu, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng
IEEE Trans. Medical Imaging2
2020 ψ-Net: Stacking Densely Convolutional LSTMs for Sub-Cortical Brain Structure Segmentation
abstract
Sub-cortical brain structure segmentation is of great importance for diagnosing neuropsychiatric disorders. However, developing an automatic approach to segmenting sub-cortical brain structures remains very challenging due to the ambiguous boundaries, complex anatomical structures, and large variance of shapes. This paper presents a novel deep network architecture, namely Ψ -Net, for sub-cortical brain structure segmentation, aiming at selectively aggregating features and boosting the information propagation in a deep convolutional neural network (CNN). To achieve this, we first formulate a densely convolutional LSTM module (DC-LSTM) to selectively aggregate the convolutional features with the same spatial resolution at the same stage of a CNN. This helps to promote the discriminativeness of features at each CNN stage. Second, we stack multiple DC-LSTMs from the deepest stage to the shallowest stage to progressively enrich low-level feature maps with high-level context. We employ two benchmark datasets on sub-cortical brain structure segmentation, and perform various experiments to evaluate the proposed Ψ -Net. The experimental results show that our network performs favorably against the state-of-the-art methods on both benchmark datasets.
Xiaowei Hu 0001, Lei Zhu 0003, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng
IEEE Trans. Medical Imaging2
2020 Saliency-Aware Texture Smoothing
abstract
Texture smoothing aims to smooth out textures in images, while retaining the prominent structures. This paper presents a saliency-aware approach to the problem with two key contributions. First, we design a deep saliency network with guided non-local blocks (GNLBs) for learning long-range pixel dependencies by taking the predicted saliency map at former layer as the guidance image to help suppress the non-saliency regions in the shallow layer. The GNLB computes the saliency response at a position by a weighted sum of features at all positions, and enables us to produce results that outperform existing deep saliency models. Second, we formulate a joint optimization framework to take saliency information when iteratively separating textures from structures: on the texture layer, we smooth out structures with the help of the saliency information and migrate structures from the texture to structure layer, while on the structure layer, we adopt another deep model to detect edges and simultaneous sparse coding to push textures back to the texture layer. We tested our method on a rich variety of images and compared it with several state-of-the-art methods. Both visual and quantitative comparison results show that our method better preserves structures while removing the texture components.
Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng
IEEE Trans. Vis. Comput. Graph.2
2019 Depth-Attentional Features for Single-Image Rain Removal
abstract
Rain is a common weather phenomenon, where object visibility varies with depth from the camera and objects faraway are visually blocked more by fog than by rain streaks. Existing methods and datasets for rain removal, however, ignore these physical properties, thereby limiting the rain removal efficiency on real photos. In this work, we first analyze the visual effects of rain subject to scene depth and formulate a rain imaging model collectively with rain streaks and fog; by then, we prepare a new dataset called RainCityscapes with rain streaks and fog on real outdoor photos. Furthermore, we design an end-to-end deep neural network, where we train it to learn depth-attentional features via a depth-guided attention mechanism, and regress a residual map to produce the rain-free image output. We performed various experiments to visually and quantitatively compare our method with several state-of-the-art methods to demonstrate its superiority over the others.
Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Pheng-Ann Heng
CVPR1
2019 Deep Multi-Model Fusion for Single-Image Dehazing
abstract
This paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature integration module to maximize the integration of the convolutional neural network (CNN) features at different CNN layers and generate the attentional multi-level integrated features (AMLIF). Then, from the AMLIF, we further predict a haze-free result for an atmospheric scattering model, as well as for four haze-layer separation models, and then fuse the results together to produce the final haze-free image. To evaluate the effectiveness of our method, we compare our network with several state-of-the-art methods on two widely-used dehazing benchmark datasets, as well as on two sets of real-world hazy images. Experimental results demonstrate clear quantitative and qualitative improvements of our method over the state-of-the-arts.
Zijun Deng, Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Qing Zhang 0006, Harry Qin, Pheng-Ann Heng
ICCV3
2019 Mask-ShadowGAN: Learning to Remove Shadows From Unpaired Data
abstract
This paper presents a new method for shadow removal using unpaired data, enabling us to avoid tedious annotations and obtain more diverse training samples. However, directly employing adversarial learning and cycle-consistency constraints is insufficient to learn the underlying relationship between the shadow and shadow-free domains, since the mapping between shadow and shadow-free images is not simply one-to-one. To address the problem, we formulate Mask-ShadowGAN, a new deep framework that automatically learns to produce a shadow mask from the input shadow image and then takes the mask to guide the shadow generation via re-formulated cycle-consistency constraints. Particularly, the framework simultaneously learns to produce shadow masks and learns to remove shadows, to maximize the overall performance. Also, we prepared an unpaired dataset for shadow removal and demonstrated the effectiveness of Mask-ShadowGAN on various experiments, even it was trained on unpaired data.
Xiaowei Hu 0001, Yitong Jiang, Chi-Wing Fu, Pheng-Ann Heng
ICCV1
2019 Probabilistic Multilayer Regularization Network for Unsupervised 3D Brain Image Registration
Xiaowei Hu 0001, Lei Zhu 0003, Pheng-Ann Heng
MICCAI (2)2
2019 CATARACTS: Challenge on automatic tool annotation for cataRACT surgery
Hassan Al Hajj, Mathieu Lamard, Pierre-Henri Conze, Soumali Roychowdhury, Xiaowei Hu 0001, Gabija Marsalkaite, Odysseas Zisimopoulos, Muneer Ahmad Dedmari, Fenqiang Zhao, Jonas Prellberg, Manish Sahu, Adrian Galdran, Teresa Araujo, Duc My Vo, Chandan Panda, Navdeep Dahiya, Satoshi Kondo, Zhengbing Bian, Gwenolé Quellec
Medical Image Anal.5
2019 SINet: A Scale-Insensitive Convolutional Neural Network for Fast Vehicle Detection
abstract
Vision-based vehicle detection approaches achieve incredible success in recent years with the development of deep convolutional neural network (CNN). However, existing CNN-based algorithms suffer from the problem that the convolutional features are scale-sensitive in object detection task but it is common that traffic images and videos contain vehicles with a large variance of scales. In this paper, we delve into the source of scale sensitivity, and reveal two key issues: 1) existing RoI pooling destroys the structure of small scale objects and 2) the large intra-class distance for a large variance of scales exceeds the representation capability of a single network. Based on these findings, we present a scale-insensitive convolutional neural network (SINet) for fast detecting vehicles with a large variance of scales. First, we present a context-aware RoI pooling to maintain the contextual information and original structure of small scale objects. Second, we present a multi-branch decision network to minimize the intra-class distance of features. These lightweight techniques bring zero extra time complexity but prominent detection accuracy improvement. The proposed techniques can be equipped with any deep network architectures and keep them trained end-to-end. Our SINet achieves state-of-the-art performance in terms of accuracy and speed (up to 37 FPS) on the KITTI benchmark and a new highway dataset, which contains a large variance of scales and extremely small objects.
Xiaowei Hu 0001, Xuemiao Xu, Yongjie Xiao, Hao Chen 0011, Shengfeng He, Harry Qin, Pheng-Ann Heng
IEEE Trans. Intell. Transp. Syst.1
2019 Deep Attentive Features for Prostate Segmentation in 3D Transrectal Ultrasound
abstract
Automatic prostate segmentation in transrectal ultrasound (TRUS) images is of essential importance for image-guided prostate interventions and treatment planning. However, developing such automatic solutions remains very challenging due to the missing/ambiguous boundary and inhomogeneous intensity distribution of the prostate in TRUS, as well as the large variability in prostate shapes. This paper develops a novel 3D deep neural network equipped with attention modules for better prostate segmentation in TRUS by fully exploiting the complementary information encoded in different layers of the convolutional neural network (CNN). Our attention module utilizes the attention mechanism to selectively leverage the multi-level features integrated from different layers to refine the features at each individual layer, suppressing the non-prostate noise at shallow layers of the CNN and increasing more prostate details into features at deep layers. Experimental results on challenging 3D TRUS volumes show that our method attains satisfactory segmentation performance. The proposed attention mechanism is a general strategy to aggregate multi-level deep features and has the potential to be used for other medical image segmentation tasks. The code is publicly available at https://github.com/wulalago/DAF3D.
Yi Wang 0031, Dong Ni 0001, Haoran Dou, Xiaowei Hu 0001, Lei Zhu 0003, Xin Yang 0009, Harry Qin, Pheng-Ann Heng, Tianfu Wang 0001
IEEE Trans. Medical Imaging4
2018 Recurrently Aggregating Deep Features for Salient Object Detection
abstract
Salient object detection is a fundamental yet challenging problem in computer vision, aiming to highlight the most visually distinctive objects or regions in an image. Recent works benefit from the development of fully convolutional neural networks (FCNs) and achieve great success by integrating features from multiple layers of FCNs. However, the integrated features tend to include non-salient regions (due to low level features of the FCN) or lost details of salient objects (due to high level features of the FCN) when producing the saliency maps. In this paper, we develop a novel deep saliency network equipped with recurrently aggregated deep features (RADF) to more accurately detect salient objects from an image by fully exploiting the complementary saliency information captured in different layers. The RADF utilizes the multi-level features integrated from different layers of a FCN to recurrently refine the features at each layer, suppressing the non-salient noise at low-level of the FCN and increasing more salient details into features at high layers. We perform experiments to evaluate the effectiveness of the proposed network on 5 famous saliency detection benchmarks and compare it with 15 state-of-the-art methods. Our method ranks first in 4 of the 5 datasets and second in the left dataset.
Xiaowei Hu 0001, Lei Zhu 0003, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng
AAAI1
2018 Direction-Aware Spatial Context Features for Shadow Detection
abstract
Shadow detection is a fundamental and challenging task, since it requires an understanding of global image semantics and there are various backgrounds around shadows. This paper presents a novel network for shadow detection by analyzing image context in a direction-aware manner. To achieve this, we first formulate the direction-aware attention mechanism in a spatial recurrent neural network (RNN) by introducing attention weights when aggregating spatial context features in the RNN. By learning these weights through training, we can recover direction-aware spatial context (DSC) for detecting shadows. This design is developed into the DSC module and embedded in a CNN to learn DSC features at different levels. Moreover, a weighted cross entropy loss is designed to make the training more effective. We employ two common shadow detection benchmark datasets and perform various experiments to evaluate our network. Experimental results show that our network outperforms state-of-the-art methods and achieves 97% accuracy and 38% reduction on balance error rate.
Xiaowei Hu 0001, Lei Zhu 0003, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng
CVPR1
2018 Bidirectional Feature Pyramid Network with Recurrent Attention Residual Modules for Shadow Detection
Lei Zhu 0003, Zijun Deng, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Harry Qin, Pheng-Ann Heng
ECCV (6)3
2018 R³Net: Recurrent Residual Refinement Network for Saliency Detection
abstract
Saliency detection is a fundamental yet challenging task in computer vision, aiming at highlighting the most visually distinctive objects in an image. We propose a novel recurrent residual refinement network (R^3Net) equipped with residual refinement blocks (RRBs) to more accurately detect salient regions of an input image. Our RRBs learn the residual between the intermediate saliency prediction and the ground truth by alternatively leveraging the low-level integrated features and the high-level integrated features of a fully convolutional network (FCN). While the low-level integrated features are capable of capturing more saliency details, the high-level integrated features can reduce non-salient regions in the intermediate prediction. Furthermore, the RRBs can obtain complementary saliency information of the intermediate prediction, and add the residual into the intermediate prediction to refine the saliency maps. We evaluate the proposed R^3Net on five widely-used saliency detection benchmarks by comparing it with 16 state-of-the-art saliency detectors. Experimental results show that our network outperforms our competitors in all the benchmark datasets.
Zijun Deng, Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Harry Qin, Guoqiang Han 0002, Pheng-Ann Heng
IJCAI2
2018 Deep Attentional Features for Prostate Segmentation in Ultrasound
Yi Wang 0031, Zijun Deng, Xiaowei Hu 0001, Lei Zhu 0003, Xin Yang 0009, Xuemiao Xu, Pheng-Ann Heng, Dong Ni 0001
MICCAI (4)3