EDBT 2026 Demo / reviewers in the wild / expert
Rui Xu 0002
dblp:00/4859-2
· DBLP profile ↗
48ranked-venue papers
14as first author
29since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 10 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mining Scene Structural Guidance for Thermal Images in Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation from RGB images has seen significant advancements recently, primarily because it eliminates the need for ground truth data during training. However, applying this technique to thermal images remains challenging due to their inherent characteristics, such as low contrast, low texture, and low signal-to-noise ratio, which impede accurate self-supervision. In this paper, we propose leveraging reliable and distinct scene structural information from thermal images to enhance self-supervised signals. We introduce structural losses, including explicit structural loss in the image space and implicit structural loss in the feature space, to improve self-supervised depth estimation. This approach mitigates the interference caused by the degraded characteristics of thermal images. Our method demonstrates superior performance compared to previous state-of-the-art approaches on the ViViD benchmark dataset, both quantitatively and qualitatively. Xinchen Ye, Xia Mao, Rui Xu 0002 |
ICASSP | 3 |
| 2025 | Delving into Transformer-based Network Architecture for Guided Depth Super-ResolutionabstractGuided Depth Super-Resolution (GDSR) enhances low-resolution (LR) depth maps by leveraging high-resolution (HR) color images. The primary challenges involve achieving effective cross-modal data alignment and fusion, as well as incorporating multi-scale information within the Transformer architecture. To address these challenges, we propose a novel network architecture named DRMPNet which integrates two key components: Offset-based Detail Refinement (ODR) and Structure-guided Multi-scale Perception (SMP). ODR leverages offset calibration and window cross-attention to align and fuse LR depth maps with color images, effectively recovering local depth details. Meanwhile, SMP employs a structure generator and multi-scale cross-attention to capture scene details and structures at multiple scales, thereby enhancing the network’s contextual understanding. Extensive experiments on various benchmark datasets demonstrate the effectiveness of our method. Xinchen Ye, Aokai Zhang, Rui Xu 0002 |
ICASSP | 3 |
| 2025 | TextBraTS: Text-Guided Volumetric Brain Tumor Segmentation with Innovative Dataset Development and Fusion Module Exploration
Rahul Kumar Jain 0001, Yinhao Li 0002, Ruibo Hou, Jingliang Cheng, Guohua Zhao, Lanfen Lin, Rui Xu 0002, Yen-Wei Chen 0001 |
MICCAI (6) | 9 |
| 2025 | EchoCardMAE: Video Masked Auto-Encoders Customized for Echocardiography
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Miao Zhang 0004, Yi Wang 0037, Xin Fan 0001, Hongkai Wang 0002, Qingxiong Yue, Xiangjian He, Yen-Wei Chen 0001 |
MICCAI (13) | 2 |
| 2025 | Guided Infrared Image Super-Resolution via Cross-modal Progressive GuidanceabstractGuided Infrared image Super-Resolution (GISR) aims to reconstruct low-resolution infrared images by leveraging high-resolution visible images that provide rich geometric and high-frequency details. The primary challenge lies in establishing cross-modal association to effectively extract and fuse complementary features while mitigating redundant information. Therefore, we propose a novel network named CPGNet, which integrates two key components: the Cross-modal Gating Module (CGM) and the Cross-modal Collaborative Module (CCM). CGM employs cross-gating mechanism combined with asymmetric convolutions to dynamically enhance the salient features from both modalities and filter out irrelevant information. Meanwhile, CCM utilizes spatial and channel collaborative importance mapping along with masking mechanism to effectively explore and combine relevant details from two modalities and generate guidance information for infrared reconstruction. Additionally, we design a hierarchical architecture for progressive guidance, which fuses infrared features with cross-modal guidance cues by progressively integrating guidance information. Extensive experiments on various datasets demonstrate the effectiveness of our proposed method. Xinchen Ye, Rui Xu 0002 |
ICMR | 3 |
| 2025 | Dynamic Motion Modeling for Enhanced Visual-Inertial OdometryabstractVisual-inertial odometry (VIO) faces a key challenge in accurately capturing the dynamics of camera motion across different trajectories, which is essential for reliable pose estimation. In this paper, we propose a novel network architecture, named DMMNet, equipped with two pivotal modules: the Attention-Driven Motion Modeling (AMM) module and the Dynamic Motion Adaptation (DMA) module. AMM enhances motion feature extraction by modeling dynamic motion between frames. DMA improves the model's adaptability to varying motion states by adaptively adjusting translation and rotation weights, thereby enhancing the modeling of complex dynamic trajectories. Experimental results show that DMMNet outperforms existing VIO methods on the KITTI dataset, demonstrating its strong generalization and adaptability in dynamic environments. Xinchen Ye, Rui Xu 0002 |
ICMR | 3 |
| 2025 | Semantics-Driven Contrastive Learning for Real-World Depth Super ResolutionabstractLow-resolution (LR) depth maps captured by depth sensors often suffer from structural distortions, noise, and blurring, limiting their practical usability. While most existing depth super-resolution (DSR) methods rely on synthetic datasets, they fail to accurately model real-world degradations, leading to poor performance on real-world data. To address this limitation, we identify two key challenges in the real-world DSR task: structural contour inconsistency and regional degradation inconsistency. The former arises from structural distortions in LR depth maps, while the latter stems from varying degradation levels in smooth regions. Upon this, we propose a Semantics-Driven Contrastive Learning (SDCL) pipeline for real-world DSR, leveraging semantic priors from the SAM model to enhance structural contour reconstruction and region-wise degradation handling. We introduce two novel contrastive loss functions: Structural Contour Alignment (SCA) loss, which aligns depth contours with semantic boundaries, and Regional Degradation Discrimination (RDD) loss, which optimizes smooth region restoration through region-level contrastive learning. Our approach is model-agnostic and can be seamlessly integrated into existing DSR frameworks. Experiments demonstrate that our method significantly enhances DSR performance on real-world DSR datasets. Xinchen Ye, Aokai Zhang, Rui Xu 0002 |
ACM Multimedia | 3 |
| 2025 | Structure-preserving dental plaque segmentation via dynamically complementary information interaction
Rui Xu 0002, Baoli Sun, Tiantian Yan, Zhihui Wang 0001 |
Multim. Syst. | 2 |
| 2025 | Self-Supervised Monocular Depth Estimation From Videos via Adaptive Reconstruction ConstraintsabstractTo estimate depth maps from monocular videos in a self-supervised way, existing methods simultaneously predict the pose changes between adjacent frames and the depth maps of each frame, and then reconstruct the forward or backward frames using them, thereby casting depth estimation as a frame reconstruction problem. The corresponding reconstruction loss, which serves as a key supervision signal for training the whole network, can adversely affect the depth estimation accuracy if it is not properly established. In this paper, we propose a novel self-supervised monocular depth estimation method from videos via adaptive reconstruction constraints, i.e., designing the loss functions by establishing more accurate reconstruction constraints. Specifically, we first propose a pose-adaptive reconstruction loss to adaptively select the optimal pose parameterizations that yield the minimum reconstruction errors, reducing the impact of inaccurate posture on frame reconstruction. Then, we propose a region-sensitive reconstruction loss that fully utilizes the pretrained image reconstruction model to adaptively identify the poorly reconstructed regions and characterize the deviation of these regions on feature space. Finally, we additionally construct a multi-frame depth estimation network and design a reconstruction-guided bidirectional distillation loss to adaptively adjust the direction of distillation between networks of multi-frame and monocular depth estimation based on their current reconstruction quality, which encourages them to learn from each other and benefits the core task of monocular depth estimation. With our proposed losses, we achieve superior performance in comparison with state-of-the-art methods on benchmark datasets. Xinchen Ye, Yuxiang Ou, Rui Xu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Feature Separation in Diffuse Lung Disease Image Classification by Using Evolutionary Algorithm-Based NASabstractIn the field of diagnosing lung diseases, the application of neural networks (NNs) in image classification exhibits significant potential. However, NNs are considered "black boxes," making it difficult to discern their decision-making processes, thereby leading to skepticism and concern regarding NNs. This compromises model reliability and hampers intelligent medicine's development. To tackle this issue, we introduce the Evolutionary Neural Architecture Search (EvoNAS). In image classification tasks, EvoNAS initially utilizes an Evolutionary Algorithm to explore various Convolutional Neural Networks, ultimately yielding an optimized network that excels at separating between redundant texture features and the most discriminative ones. Retaining the most discriminative features improves classification accuracy, particularly in distinguishing similar features. This approach illuminates the intrinsic mechanics of classification, thereby enhancing the accuracy of the results. Subsequently, we incorporate a Differential Evolution algorithm based on distribution estimation, significantly enhancing search efficiency. Employing visualization techniques, we demonstrate the effectiveness of EvoNAS, endowing the model with interpretability. Finally, we conduct experiments on the diffuse lung disease texture dataset using EvoNAS. Compared to the original network, the classification accuracy increases by 0.56%. Moreover, our EvoNAS approach demonstrates significant advantages over existing methods in the same dataset. Dan Shao, Lin Lin 0008, Guoliang Gong, Rui Xu 0002, Shoji Kido, HongWei Cui |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | UW-Adapter: Adapting Monocular Depth Estimation Model in Underwater ScenesabstractEstimating depth maps from monocular underwater images poses one of the most challenging problems in underwater applications. Due to the lack of large-scale paired underwater color-depth datasets for effective training, existing style transfer-based and self-supervision-based approaches can improve the performance of depth estimation to some extent, but they remain unsatisfactory. Leveraging the power of massive training datasets, foundation models designed for terrestrial monocular depth estimation have demonstrated superior performance across various scenes. These models provide rich prior knowledge of 3D perception, which can be valuable for underwater depth estimation. Upon this, we introduce tunable adapters (UW-Adapter) that tailor a pre-trained foundation model specifically for underwater depth estimation, customizing it to the unique characteristics of underwater imagery. Our approach involves freezing the parameters of the pre-trained model and updating only the adapters through self-supervision. To address the complex degradation of underwater images, we propose two adapters: the transmission adapter and the high-frequency adapter. These adapters incorporate depth clues and high-frequency information as prior knowledge, thereby enhancing the performance of pre-trained model in underwater depth estimation. Experimental results demonstrate that by integrating lightweight adapters into off-the-shelf depth estimation foundation models, our method achieves superior performance across multiple datasets. Xinchen Ye, Rui Xu 0002 |
IEEE Trans. Multim. | 3 |
| 2025 | Is a Pure Transformer Effective for Separated and Online Multi-Object Tracking?abstractRecent advances in multi-object tracking (MOT) have demonstrated significant success in short-term association within the separated tracking-by-detection online paradigm. However, long-term tracking remains challenging. While graph-based approaches address this by modeling trajectories as global graphs, these methods are unsuitable for real-time applications due to their non-online nature. In this article, we review the concept of trajectory graphs and propose a novel perspective by representing them as directed acyclic graphs. This representation can be described using frame-ordered object sequences and binary adjacency matrices. We observe that this structure naturally aligns with Transformer attention mechanisms, enabling us to model the association problem using a classic Transformer architecture. Based on this insight, we introduce a concise pure transformer (PuTR) to validate the effectiveness of Transformer in unifying short- and long-term tracking for separated online MOT. Extensive experiments on four diverse datasets (SportsMOT, DanceTrack, MOT17, and MOT20) demonstrate that PuTR effectively establishes a solid baseline compared to existing foundational online methods while exhibiting superior domain adaptation capabilities. Furthermore, the separated nature enables efficient training and inference, making it suitable for practical applications. Implementation code and trained models are available at https://github.com/chongweiliu/PuTR . Chongwei Liu, Zhihui Wang 0001, Rui Xu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Novelty Detection Based Discriminative Multiple Instance Feature Mining to Classify NSCLC PD-L1 Status on HE-Stained Histopathological Images
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yi Wang 0037, Hongkai Wang 0002, Dingpin Huang, Fangyi Xu, Yi Gan, Yuan Tu, Hongjie Hu |
MICCAI (4) | 1 |
| 2024 | Low-resolution few-shot learning via multi-space knowledge distillation
Xinchen Ye, Baoli Sun, Hairui Yang, Rui Xu 0002, Zhihui Wang 0001 |
Inf. Sci. | 6 |
| 2024 | Reconciling global and local optimal label assignments for heavily occluded pedestrian detection
Chongwei Liu, Zhihui Wang 0001, Rui Xu 0002 |
Multim. Syst. | 4 |
| 2024 | Addressing Challenges of Incorporating Appearance Cues Into Heuristic Multi-Object Tracker via a Novel Feature ParadigmabstractIn the field of Multi-Object Tracking (MOT), the incorporation of appearance cues into tracking-by-detection heuristic trackers using re-identification (ReID) features has posed limitations on its advancement. The existing ReID paradigm involves the extraction of coarse-grained object-level feature vectors from cropped objects at a fixed input size using a ReID model, and similarity computation through a simple normalized inner product. However, MOT requires fine-grained features from different object regions and more accurate similarity measurements to identify individuals, especially in the presence of occlusion. To address these limitations, we propose a novel feature paradigm. In this paradigm, we extract the feature map from the entire frame image to preserve object sizes and represent objects using a set of fine-grained features from different object regions. These features are sampled from adaptive patches within the object bounding box on the feature map to effectively capture local appearance cues. We introduce Mutual Ratio Similarity (MRS) to accurately measure the similarity of the most discriminative region between two objects based on the sampled patches, which proves effective in handling occlusion. Moreover, we propose absolute Intersection over Union (AIoU) to consider object sizes in feature cost computation. We integrate our paradigm with advanced motion techniques to develop a heuristic Motion-Feature joint multi-object tracker, MoFe. Within it, we reformulate the track state transition of tracklets to better model their life cycle, and firstly introduce a runtime recorder after MoFe to refine trajectories. Extensive experiments on five benchmarks, i.e., GMOT-40, BDD100k, DanceTrack, MOT17, and MOT20, demonstrate that MoFe achieves state-of-the-art performance in robustness and generalizability without any fine-tuning, and even surpasses the performance of fine-tuned ReID features. Chongwei Liu, Zhihui Wang 0001, Rui Xu 0002 |
IEEE Trans. Image Process. | 4 |
| 2024 | C2ANet: Cross-Scale and Cross-Modality Aggregation Network for Scene Depth Super-ResolutionabstractExisting depth super-resolution (DSR) methods typically utilize an additional high-resolution (HR) color image of the same scene as assistance to recover the low-resolution (LR) depth map. Although these color-guided methods have achieved impressive progress, they easily face with color image under-utilization and mis-utilization issues. In this article, we deeply investigate the above problems and further propose a novel DSR framework to alleviate them. Specifically, we propose a Cross-scale and Cross-modality Aggregation Network(C$^{2}$ANet)to learn abundant and accurate complementarity from color images to help recover the degraded depth map. Our C$^{2}$ANet can simultaneously extract multi-scale representations from color images with parallel network hierarchies, and effectively aggregate cross-scale and cross-modality contexts to boost HR representations in each hierarchy. Then, to appropriately use the guided color image, we further design a Feature Aggregation Module (FAM) to adaptively select and fuse task-relevant features, which consists of (1) afeature alignment blockto learn transformation offsets and align upsampled features with targeted HR features, and (2) afeature fusion blockbased on cross-attention mechanism to maintain strong structural context and suppress texture distraction. Experimental results on synthetic and real-world benchmark datasets demonstrate the superiority of our proposed method in comparison with other state-of-the-art DSR methods. Xinchen Ye, Baoli Sun, Rui Xu 0002, Zhihui Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Cross-Modality depth Estimation via Unsupervised Stereo RGB-to-infrared TranslationabstractExisting depth estimation methods infer scene depth only from stereo visible light (RGB) images. Since RGB imaging is sensitive to changes in light, it’s difficult to estimate depth information accurately in some degraded visibility conditions. In contrast, infrared (IR) imaging captures thermal radiation and is not affected by brightness changing, providing extra clues for depth estimation. However, most datasets used for training in depth estimation do not have IR images paired with RGB-D data. Therefore, how to obtain the paired IR images and exploit the respective advantages of RGB and IR images to improve the performance of depth estimation, is of vital importance. Our core idea is to first develop an unsupervised RGB-to-IR translation (RIT) network with proposed Fourier domain adaptation and multi-space warping regularization to synthesize stereo IR images from their corresponding stereo RGB images. And then modified depth estimation backbones can be used as the cross-modality depth estimation (CDE) network to infer disparity maps from cross-modal RGB-IR stereo pairs. Assisted by the synthetic stereo IR images, we obtain superior performance just by flexibly deploying our framework to several off-the-shelf depth estimation backbones of single-modality (RGB) based methods. Shi Tang, Xinchen Ye, Rui Xu 0002 |
ICASSP | 4 |
| 2023 | Underwater Depth Estimation via Stereo Adaptation NetworksabstractWith the fast development and wide application of stereo depth estimation, adequate high-quality stereo training data with groundtruth depth information plays an important role, but is not easily acquired in underwater environments. Therefore, satisfactory performance of depth estimation is difficult to achieve in underwater environments. In addition, the domain gap also leads to the failure of directly applying existing models of terrestrial scene to underwater scene. Therefore, this paper proposes a novel underwater depth estimation network which can infer depth maps from real underwater stereo images in an adaptation manner. The proposed learning pipeline mainly contains three different adaptation modules, i.e., style adaptation, semantic adaptation and disparity range adaptation, to progressively adapt a terrestrial depth estimation model to the underwater domain. Specifically, due to the lack of underwater training data, we first propose a depth-aware stereo image translation network to synthesize stylized underwater stereo images from terrestrial dataset, thus benefiting the effective training of depth estimation network. Then, considering the weak generalization to the real underwater data when only trained on the above synthetic data, we present a self-ensembling semantic adaptation for depth estimation network to minimize the semantic domain discrepancy between synthetic and real underwater data. Meanwhile, we design a disparity range adaptation module to address the problem of disparity range miss-match between both data, thus obtaining more accurate depth predictions for large-disparity-span underwater images. Experimental results show that by integrating the proposed adaptation modules into the off-the-shelf depth estimation backbones, our method successfully achieves superior performance of underwater depth estimation compared to other state-of-the-art methods. Xinchen Ye, Yazhi Yuan, Rui Xu 0002, Zhihui Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 LesionsabstractAutomatic segmentation of COVID-19 lesions is essential for computer-aided diagnosis. However, this task remains challenging because widely-used supervised based methods require large-scale annotated data that is difficult to obtain. Although an unsupervised method based on anomaly detection has shown promising results in [1], its performance is relatively poor. We address this problem by proposing a pixel-level and affinity-level knowledge distillation method. It obtains a pre-trained teacher network with rich semantic knowledge of CT images by constructing and training an auto-encoder at first, and then trains a student network with the same architecture as the teacher by distilling the teacher’s knowledge only from normal CT images, and finally localizes COVID-19 lesions using the feature discrepancy between the teacher and the student networks. Besides, except for the traditional pixel-level distillation, we design the affinity-level distillation that takes into account the pairwise relationship of features to fully distill effective knowledge. We evaluate this method by using three different COVID-19 datasets and the experimental results show that the segmentation performance is largely improved when it is compared with the other existing unsupervised anomaly detection methods. Rui Xu 0002, Xinchen Ye, Yen-Wei Chen 0001, Fangyi Xu, Wenchao Zhu, Hongjie Hu, Xiaofeng Qu, Shoji Kido, Noriyuki Tomiyama |
ICASSP | 1 |
| 2022 | Underwater Stereo Matching Via Unsupervised Appearance And Feature Adaptation NetworksabstractStereo matching has been widely used to estimate depth maps in terrestrial environments. However, it is difficult to achieve appealing performance in underwater environments, since adequate underwater stereo data with groundtruth depth information is not easily available for training an underwater depth estimation model. In addition, the domain gap also leads to the failure of directly applying existing models of terrestrial scenes to underwater scenes. Therefore, this paper proposes a novel underwater depth estimation network which can infer depth maps from real underwater stereo images in an unsupervised adaptation manner. The proposed learning pipeline contains style adaptation (SA) in appearance space and feature adaptation (FA) in semantic space to progressively adapt the depth estimation models to underwater domain. Experimental results show that by integrating the proposed adaptation modules into the off-the-shelf stereo matching backbones, our method achieves a superior performance of underwater depth estimation compared to other state-of-the-art methods. Yazhi Yuan, Xinchen Ye, Dian Zheng, Rui Xu 0002 |
ICASSP | 5 |
| 2022 | Learning Data Hallucination and Reciprocal Guidance for Underwater Depth Estimation and Color CorrectionabstractUnderwater vision is typically more difficult to tackle than open-air vision due to the degraded visibility and geometrical distortion, which impedes the development of underwater machine vision. Hence, we propose a joint depth estimation and color correction framework for underwater monocular images via data hallucination and reciprocal guidance learning. Specifically, due to the lack of labeled underwater data, we first design a data hallucination network to translate terrestrial images to multi-style synthetic underwater images while retaining the scene structure of terrestrial images from a single-source-multi-target perspective, benefiting the effective training of the joint tasks. Then, considering the strong connection between both tasks, we design a collaborative network to learn the reciprocal guidance between tasks from a multi-task perspective, thus improving the performance of each task. The whole framework can be trained end-to-end, and performs favorably against state-of-the-art methods in both depth estimation and color correction tasks. Xinchen Ye, Rui Xu 0002 |
ICME | 4 |
| 2022 | Local-Region and Cross-Dataset Contrastive Learning for Retinal Vessel Segmentation
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yen-Wei Chen 0001 |
MICCAI (2) | 1 |
| 2022 | Low-Dose CT Reconstruction via Dual-Domain Learning and Controllable Modulation
Xinchen Ye, Rui Xu 0002, Zhihui Wang 0001 |
MICCAI (6) | 3 |
| 2021 | Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-ResolutionabstractExisting color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training samples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected in actual testing environment. Therefore, we explore for the first time to learn the cross-modality knowledge at training stage, where both RGB and depth modalities are available, but test on the target dataset, where only single depth modality exists. Our key idea is to distill the knowledge of scene structural guidance from RGB modality to the single DSR task without changing its network architecture. Specifically, we construct an auxiliary depth estimation (DE) task that takes an RGB image as input to estimate a depth map, and train both DSR task and DE task collaboratively to boost the performance of DSR. Upon this, a cross-task interaction module is proposed to realize bilateral cross-task knowledge transfer. First, we design a cross-task distillation scheme that encourages DSR and DE networks to learn from each other in a teacher-student role-exchanging fashion. Then, we advance a structure prediction (SP) task that provides extra structure regularization to help both DSR and DE networks learn more informative structure representations for depth recovery. Extensive experiments demonstrate that our scheme achieves superior performance in comparison with other DSR methods. Baoli Sun, Xinchen Ye, Baopu Li, Zhihui Wang 0001, Rui Xu 0002 |
CVPR | 6 |
| 2021 | DPNet: Detail-preserving network for high quality monocular depth estimation
Xinchen Ye, Shude Chen, Rui Xu 0002 |
Pattern Recognit. | 3 |
| 2021 | VolumeNet: A Lightweight Parallel Network for Super-Resolution of MR and CT Volumetric DataabstractDeep learning-based super-resolution (SR) techniques have generally achieved excellent performance in the computer vision field. Recently, it has been proven that three-dimensional (3D) SR for medical volumetric data delivers better visual results than conventional two-dimensional (2D) processing. However, deepening and widening 3D networks increases training difficulty significantly due to the large number of parameters and small number of training samples. Thus, we propose a 3D convolutional neural network (CNN) for SR of magnetic resonance (MR) and computer tomography (CT) volumetric data called ParallelNet using parallel connections. We construct a parallel connection structure based on the group convolution and feature aggregation to build a 3D CNN that is as wide as possible with a few parameters. As a result, the model thoroughly learns more feature maps with larger receptive fields. In addition, to further improve accuracy, we present an efficient version of ParallelNet (called VolumeNet), which reduces the number of parameters and deepens ParallelNet using a proposed lightweight building block module called the Queue module. Unlike most lightweight CNNs based on depthwise convolutions, the Queue module is primarily constructed using separable 2D cross-channel convolutions. As a result, the number of network parameters and computational complexity can be reduced significantly while maintaining accuracy due to full channel fusion. Experimental results demonstrate that the proposed VolumeNet significantly reduces the number of model parameters and achieves high precision results compared to state-of-the-art methods in tasks of brain MR image SR, abdomen CT image SR, and reconstruction of super-resolution 7T-like images from their 3T counterparts. Yinhao Li 0002, Yutaro Iwamoto, Lanfen Lin, Rui Xu 0002, Ruofeng Tong 0001, Yen-Wei Chen 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Unsupervised Monocular Depth Estimation via Recursive Stereo DistillationabstractExisting unsupervised monocular depth estimation methods resort to stereo image pairs instead of ground-truth depth maps as supervision to predict scene depth. Constrained by the type of monocular input in testing phase, they fail to fully exploit the stereo information through the network during training, leading to the unsatisfactory performance of depth estimation. Therefore, we propose a novel architecture which consists of a monocular network (Mono-Net) that infers depth maps from monocular inputs, and a stereo network (Stereo-Net) that further excavates the stereo information by taking stereo pairs as input. During training, the sophisticated Stereo-Net guides the learning of Mono-Net and devotes to enhance the performance of Mono-Net without changing its network structure and increasing its computational burden. Thus, monocular depth estimation with superior performance and fast runtime can be achieved in testing phase by only using the lightweight Mono-Net. For the proposed framework, our core idea lies in: 1) how to design the Stereo-Net so that it can accurately estimate depth maps by fully exploiting the stereo information; 2) how to use the sophisticated Stereo-Net to improve the performance of Mono-Net. To this end, we propose a recursive estimation and refinement strategy for Stereo-Net to boost its performance of depth estimation. Meanwhile, a multi-space knowledge distillation scheme is designed to help Mono-Net amalgamate the knowledge and master the expertise from Stereo-Net in a multi-scale fashion. Experiments demonstrate that our method achieves the superior performance of monocular depth estimation in comparison with other state-of-the-art methods. Xinchen Ye, Xin Fan 0001, Mingliang Zhang 0002, Rui Xu 0002 |
IEEE Trans. Image Process. | 4 |
| 2021 | Joint Extraction of Retinal Vessels and Centerlines Based on Deep Semantics and Multi-Scaled Cross-Task AggregationabstractRetinal vessel segmentation and centerline extraction are crucial steps in building a computer-aided diagnosis system on retinal images. Previous works treat them as two isolated tasks, while ignoring their tight association. In this paper, we propose a deep semantics and multi-scaled cross-task aggregation network that takes advantage of the association to jointly improve their performances. Our network is featured by two sub-networks. The forepart is a deep semantics aggregation sub-network that aggregates strong semantic information to produce more powerful features for both tasks, and the tail is a multi-scaled cross-task aggregation sub-network that explores complementary information to refine the results. We evaluate the proposed method on three public databases, which are DRIVE, STARE and CHASE_DB1. Experimental results show that our method can not only simultaneously extract retinal vessels and their centerlines but also achieve the state-of-the-art performances on both tasks. Rui Xu 0002, Xinchen Ye, Lin Lin 0008, Liang Li 0002, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Unsupervised Content-Preserved Adaptation Network for Classification of Pulmonary Textures from Different CT ScannersabstractDeep network based methods have been proposed for accurate classification of pulmonary textures on CT images. However, such methods well-trained on CT data from one scanner cannot perform well when they are directly applied to the data from other scanners. This domain shift problem is caused by different physical components and scanning protocols of different CT scanners. In this paper, we propose an unsupervised content-preserved adaptation network to address this problem. Our method can make a previously well-trained deep network to be adapted for the data of a new CT scanner and does not require the laboring annotation to delineate pulmonary texture regions on the new CT data. Extensive evaluations have been carried on images collected from GE and Toshiba CT scanners and show that the proposed method can alleviate the performance degradation problem of classifying pulmonary textures from different CT scanners. Rui Xu 0002, Zhen Cong, Xinchen Ye, Shoji Kido, Noriyuki Tomiyama |
ICASSP | 1 |
| 2020 | Retinal Vessel Segmentation via a Semantics and Multi-Scale Aggregation NetworkabstractPrecise segmentation of retinal vessels is crucial for a computer-aided diagnosis system of retinal fundus images. However, this task remains challenging due to large variations in scales and poor segmentation of capillary vessels. In this paper, we propose a semantics and multi-scale aggregation network to address these difficulties. It includes semantics aggregation blocks that are designed for aggregating stronger high-level semantic information. These carefully designed blocks produce more semantic feature representation that is helpful for capillary vessel identification and vessel connection. Besides, a multi-scale aggregation block is designed by employing parallel dilated convolutional filters with different dilation rates to fully exploit the multi-scale information. We evaluate the network by using two public databases of retinal vessel segmentation and compare its performance with several leading methods published in the past several years. Extensive evaluations show that the proposed network has achieved the state-of-the-art performance on the public CHASE DB1 and HRF datasets. Rui Xu 0002, Xinchen Ye, Guiliang Jiang, Liang Li 0002 |
ICASSP | 1 |
| 2020 | Cascaded Detail-Aware Network for Unsupervised Monocular Depth EstimationabstractExisting unsupervised learning methods usually reformulate the depth estimation into the image reconstruction problem by training on stereo image pairs to circumvent the need of dense labeled ground truth depth information. Most of them are designed based on a simple encoder-decoder backbone architecture, which has limited expression for context information and suffers from the loss of depth details. In this paper, we propose a cascaded detail-aware network which contains a contextual network (CN) followed by consecutive spatial networks (SNs) to make an unsupervised coarse-to-fine prediction. CN aims to provide good initialized depth estimation results by introducing a multi-scale attention fusion module to enhance the ability of feature representation. Then, SN is progressively applied on the coarse depth map to produce refined depth outputs by exploiting abundant spatial details from input color image. Moreover, we design a robust loss function that further considers the penalty of photometric errors and the occlusion, and strengthens the recovery of spatial details for better depth estimation. Experimental results show that the proposed method achieves promising performance. Xinchen Ye, Mingliang Zhang 0002, Xin Fan 0001, Rui Xu 0002, Juncheng Pu, Ruoke Yan |
ICME | 4 |
| 2020 | Unsupervised Detection of Pulmonary Opacities for Computer-Aided Diagnosis of COVID-19 on CT ImagesabstractCOVID-19 emerged towards the end of 2019 which was identified as a global pandemic by the world heath organization (WHO). With the rapid spread of COVID-19, the number of infected and suspected patients has increased dramatically. Chest computed tomography (CT) has been recognized as an efficient tool for the diagnosis of COVID-19. However, the huge CT data make it difficult for radiologist to fully exploit them on the diagnosis. In this paper, we propose a computer-aided diagnosis system that can automatically analyze CT images to distinguish the COVID-19 against to community-acquired pneumonia (CAP). The proposed system is based on an unsupervised pulmonary opacity detection method that locates opacity regions by a detector unsupervisedly trained from CT images with normal lung tissues. Radiomics based features are extracted insides the opacity regions, and fed into classifiers for classification. We evaluate the proposed CAD system by using 200 CT images collected from different patients in several hospitals. The accuracy, precision, recall, f1-score and AUC achieved are 95.5%, 100%, 91%, 95.1% and 95.9% respectively, exhibiting the promising capacity on the differential diagnosis of COVID-19 from CT images. Rui Xu 0002, Xiao Cao, Yen-Wei Chen 0001, Xinchen Ye, Lin Lin 0008, Wenchao Zhu, Fangyi Xu, Hongjie Hu, Shoji Kido, Noriyuki Tomiyama |
ICPR | 1 |
| 2020 | BG-Net: Boundary-Guided Network for Lung Segmentation on Clinical CT ImagesabstractLung segmentation on CT images is a crucial step for a computer-aided diagnosis system of lung diseases. The existing deep learning based lung segmentation methods are less efficient to segment lungs on clinical CT images, especially that the segmentation on lung boundaries is not accurate enough due to complex pulmonary opacities in practical clinics. In this paper, we propose a boundary-guided network (BG-Net) to address this problem. It contains two auxiliary branches that seperately segment lungs and extract the lung boundaries, and an aggregation branch that efficiently exploits lung boundary cues to guide the network for more accurate lung segmentation on clinical CT images. We evaluate the proposed method on a private dataset collected from the Osaka university hospital and four public datasets including StructSeg [1], HUG [2], VESSEL12 [3], and a Novel Coronavirus 2019 (COVID-19) dataset [4]. Experimental results show that the proposed method can segment lungs more accurately and outperform several other deep learning based methods. Rui Xu 0002, Yi Wang 0037, Xinchen Ye, Lin Lin 0008, Yen-Wei Chen 0001, Shoji Kido, Noriyuki Tomiyama |
ICPR | 1 |
| 2020 | Detail- Revealing Deep Low-Dose CT ReconstructionabstractLow-dose CT imaging emerges with low radiation risk due to the reduction of radiation dose, but brings negative impact on the imaging quality. This paper addresses the problem of low-dose CT reconstruction. Previous methods are unsatisfactory due to the inaccurate recovery of image details under the strong noise generated by the reduction of radiation dose, which directly affects the final diagnosis. To suppress the noise effectively while retain the structures well, we propose a detail-revealing dual-branch aggregation network to effectively reconstruct the degraded CT image. Specifically, the main reconstruction branch iteratively exploits and compensates the reconstruction errors to gradually refine the CT image, while the prior branch is to learn the structure details as prior knowledge to help recover the CT image. A sophisticated detail-revealing loss is designed to fuse the information from both branches and guide the learning to obtain better performance from pixel-wise and holistic perspectives respectively. Experimental results show that our method outperforms the state-of-art methods in both PSNR and SSIM metrics. Xinchen Ye, Yuyao Xu, Rui Xu 0002, Shoji Kido, Noriyuki Tomiyama |
ICPR | 3 |
| 2020 | Evolutionary Neural Network and Visualization for CNN-based Pulmonary Textures ClassificationabstractAccurate classification and comprehensive explanation is crucial to build a computer aided diagnosis (CAD) system of diffuse lung disease (DLD). Although deep neural networks (DNNs) have been applied to this task, the classification performance and reliability are not satisfied for medical clinical requirements. Specifically, DNNs are regarded as unexplainable “black-box” in general, and, thus, are not deemed reliable by expects. In this paper, we propose a neural network structure search approach based on evolutionary algorithm to improve the DNN's effectiveness and interpretability, and applied to the pulmonary textures classification problem. Through this network structure search approach, we find out how a DNN's subnet recognize the pulmonary textures features, then filter out the redundant subnets, and retain the most distinctive feature subnets. Besides, we utilize the method of feature visualization and the fine-grained heat map of the activation to interpret network's decision-making process. Finally, through quantitatively and qualitatively evaluate on a real dataset of diffuse lung disease, we verify the effectiveness of this neural network structure search approach on VggNet and ResNet, and achieve the state-of-the-art performance. We can classify the pulmonary textures on high-resolution computed tomography (HRCT) images. Guoliang Gong, Lin Lin 0008, Zhaoyang Wu, Rui Xu 0002, Shoji Kido |
ICTAI | 4 |
| 2020 | Boosting Connectivity in Retinal Vessel Segmentation via a Recursive Semantics-Guided Network
Rui Xu 0002, Xinchen Ye, Lin Lin 0008, Yen-Wei Chen 0001 |
MICCAI (5) | 1 |
| 2020 | Depth Super-Resolution via Deep Controllable Slicing NetworkabstractDue to the imaging limitation of depth sensors, high-resolution (HR) depth maps are often difficult to be acquired directly, thus effective depth super-resolution (DSR) algorithms are needed to generate HR output from its low-resolution (LR) counterpart. Previous methods treat all depth regions equally without considering different extents of degradation at region-level, and regard DSR under different scales as independent tasks without considering the modeling of different scales, which impede further performance improvement and practical use of DSR. To alleviate these problems, we propose a deep controllable slicing network from a novel perspective. Specifically, our model is to learn a set of slicing branches in a divide-and-conquer manner, parameterized by a distance-aware weighting scheme to adaptively aggregate different depths in an ensemble. Each branch that specifies a depth slice (e.g., the region in some depth range) tends to yield accurate depth recovery. Meanwhile, a scale-controllable module that extracts depth features under different scales is proposed and inserted into the front of slicing network, and enables finely-grained control of the depth restoration results of slicing network with a scale hyper-parameter. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our method achieves superior performance. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li |
ACM Multimedia | 5 |
| 2020 | Depth upsampling based on deep edge-aware learning
Zhihui Wang 0001, Xinchen Ye, Baoli Sun, Jing-Yu Yang 0002, Rui Xu 0002 |
Pattern Recognit. | 5 |
| 2020 | Deep Joint Depth Estimation and Color Correction From Monocular Underwater Images Based on Unsupervised Adaptation NetworksabstractDegraded visibility and geometrical distortion typically make the underwater vision more intractable than open air vision, which impedes the development of underwater-related machine vision and robotic perception. Therefore, this paper addresses the problem of joint underwater depth estimation and color correction from monocular underwater images, which aims at enjoying the mutual benefits between these two related tasks from a multi-task perspective. Our core ideas lie in our new deep learning architecture. Due to the lack of effective underwater training data, and the weak generalization to the real-world underwater images trained on synthetic data, we consider the problem from a novel perspective of style-level and feature-level adaptation, and propose an unsupervised adaptation network to deal with the joint learning problem. Specifically, a style adaptation network (SAN) is first proposed to learn a style-level transformation to adapt in-air images to the style of underwater domain. Then, we formulate a task network (TN) to jointly estimate the scene depth and correct the color from a single underwater image by learning domain-invariant representations. The whole framework can be trained end-to-end in an adversarial learning manner. Extensive experiments are conducted under air-to-water domain adaptation settings. We show that the proposed method performs favorably against state-of-the-art methods in both depth estimation and color correction tasks. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Rui Xu 0002, Xin Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | PMBANet: Progressive Multi-Branch Aggregation Network for Scene Depth Super-ResolutionabstractDepth map super-resolution is an ill-posed inverse problem with many challenges. First, depth boundaries are generally hard to reconstruct particularly at large magnification factors. Second, depth regions on fine structures and tiny objects in the scene are destroyed seriously by downsampling degradation. To tackle these difficulties, we propose a progressive multi-branch aggregation network (PMBANet), which consists of stacked MBA blocks to fully address the above problems and progressively recover the degraded depth map. Specifically, each MBA block has multiple parallel branches: 1) The reconstruction branch is proposed based on the designed attention-based error feed-forward/-back modules, which iteratively exploits and compensates the downsampling errors to refine the depth map by imposing the attention mechanism on the module to gradually highlight the informative features at depth boundaries. 2) We formulate a separate guidance branch as prior knowledge to help to recover the depth details, in which the multi-scale branch is to learn a multi-scale representation that pays close attention at objects of different scales, while the color branch regularizes the depth map by using auxiliary color information. Then, a fusion block is introduced to adaptively fuse and select the discriminative features from all the branches. The design methodology of our whole network is well-founded, and extensive experiments on benchmark datasets demonstrate that our method achieves superior performance in comparison with the state-of-the-art methods. Our code and models are available athttps://github.com/Sunbaoli/PMBANet_DSR/. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li |
IEEE Trans. Image Process. | 5 |
| 2020 | Pulmonary Textures Classification via a Multi-Scale Attention NetworkabstractPrecise classification of pulmonary textures is crucial to develop a computer aided diagnosis (CAD) system of diffuse lung diseases (DLDs). Although deep learning techniques have been applied to this task, the classification performance is not satisfied for clinical requirements, since commonly-used deep networks built by stacking convolutional blocks are not able to learn discriminative feature representation to distinguish complex pulmonary textures. For addressing this problem, we design a multi-scale attention network (MSAN) architecture comprised by several stacked residual attention modules followed by a multi-scale fusion module. Our deep network can not only exploit powerful information on different scales but also automatically select optimal features for more discriminative feature representation. Besides, we develop visualization techniques to make the proposed deep model transparent for humans. The proposed method is evaluated by using a large dataset. Experimental results show that our method has achieved the average classification accuracy of 94.78% and the average f-value of 0.9475 in the classification of 7 categories of pulmonary textures. Besides, visualization results intuitively explain the working behavior of the deep network. The proposed method has achieved the state-of-the-art performance to classify pulmonary textures on high resolution CT images. Rui Xu 0002, Zhen Cong, Xinchen Ye, Yasushi Hirano, Shoji Kido, Tomoko Gyobu, Yutaka Kawata, Osamu Honda, Noriyuki Tomiyama |
IEEE J. Biomed. Health Informatics | 1 |
| 2019 | Unsupervised Monocular Depth Estimation Based on Dual Attention Mechanism and Depth-Aware LossabstractMost existing monocular depth estimation approaches are su- pervised, but enough quantities of ground truth depth data are required during training. To cope with this, recent techniques deal with the depth estimation task in an unsupervised man- ner, i.e., replacing the use of depth data with easily obtained stereo images for training. Based on this, we propose a nov- el unsupervised learning architecture, which integrates dual attention mechanism into the framework and designs a depth- aware loss for better depth estimation. Specifically, to en- hance the ability of feature representations, we introduce a d- ual attention module to capture global feature dependencies in spatial and channel dimensions for scene understanding and depth estimation. Meanwhile, we propose a depth-aware loss that fully addresses the occlusion problem in brightness con- stancy assumption, the intrinsic characteristics of depth map, and the left-right consistency problem, respectively. Besides, an adversarial loss is employed to discriminate synthetic or realistic depth maps by training a discriminator so as to pro- duce better results. Extensive experiments on KITTI dataset show that our approach achieves state-of-the-art performance compared with other monocular depth estimation methods. Xinchen Ye, Mingliang Zhang 0002, Rui Xu 0002, Xin Fan 0001, Zhu Liu 0004, Jiaao Zhang |
ICME | 3 |
| 2018 | Pulmonary Textures Classification Using A Deep Neural Network with Appearance and Geometry CuesabstractClassification of pulmonary textures on CT images is essential for the development of a computer-aided diagnosis system of diffuse lung diseases. In this paper, we propose a novel method to classify pulmonary textures by using a deep neural network, which can make full use of appearance and geometry cues of textures via a dual-branch architecture. The proposed method has been evaluated by a dataset that includes seven kinds of typical pulmonary textures. Experimental results show that our method outperforms the state-of-the-art methods including feature engineering based method and convolutional neural network based method. Rui Xu 0002, Zhen Cong, Xinchen Ye, Yasushi Hirano, Shoji Kido |
ICASSP | 1 |
| 2011 | Classification of Diffuse Lung Disease Patterns on High-Resolution Computed Tomography by a Bag of Words Approach
Rui Xu 0002, Yasushi Hirano, Rie Tachibana, Shoji Kido |
MICCAI (3) | 1 |
| 2009 | Generalized N-dimensional principal component analysis (GND-PCA) and its application on construction of statistical appearance models for medical volumes with fewer samples
Rui Xu 0002, Yen-Wei Chen 0001 |
Neurocomputing | 1 |
| 2008 | Semiautomatic non-rigid 3-D image registration for MR-Guided Liver Cancer SurgeryabstractRecently a growing interest has been seen in minimally invasive treatments with open configuration magnetic resonance (Open-MR) scanners. In this paper, we proposed a semi-automatic non-rigid 3D MR-CT image registration technique for MR-Guided Liver Cancer Surgery in which cancer tissues are coagulated by microwave ablation. Because of the lower magnetic field (0.5 T) and various different surgical conditions, sometimes tumors can not be visualized clearly on Open-MR volumes. Combining of CT volumes acquired before surgery, it is possible to identify the tumor's location by application of registration techniques. Since such a registration problem belongs to a non-rigid one considering the easy deformation of livers, free-form deformation (FFD) based registration method is applied. Similarity measurement in registration is normalized mutual information (NMI). Both phantom and clinical experiments show that the registration is accurate enough (ap1.45 mm) for liver cancer surgery given some proper processing steps. Yen-Wei Chen 0001, Katsumi Tsubokawa, Rui Xu 0002, Shigehiro Morikawa, Yoshimasa Kurumi |
ICIP | 3 |
| 2007 | Appearance Models for Medical Volumes with Few Samples by Generalized 3D-PCA
Rui Xu 0002, Yen-Wei Chen 0001 |
ICONIP (1) | 1 |