Xinchen Ye

dblp:119/1472 · DBLP profile ↗
← Back
74ranked-venue papers
25as first author
41since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 56 · 20 first-author · 31 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Color-guided depth super-resolution with structure selection and modulation
Xinchen Ye
Multim. Syst.2
2026 Harnessing the power of local representations for few-shot classification
Shi Tang, Guiming Luo, Xinchen Ye, Zhiyi Xia
Pattern Recognit.3
2025 Few-shot Image Classification based on Attribute Prediction and Selection
abstract
Few-shot learning addresses the challenges of image classification with limited samples, but current methods often fail to fully utilize sample correlations and external semantic information, leading to low accuracy. To overcome these limitations, we propose a few-shot image classification method based on attribute prediction and selection. In addition to the image branch that extracts image features, an extra attribute branch employs an attribute prediction network to predict attribute features of the query set samples. This ensures that both the support set and the query set can utilize attribute information as assistance. Additionally, we propose an attribute selection network to distinguish and select discriminative attribute features, thereby increasing the utilization of task-related attribute information. Finally, the processed attribute features and the image features are merged together as the ultimate features for final classification. Extensive experiments on mainstream benchmarks demonstrate the superiority of our method.
Boqian Liu, Xinchen Ye, Guanqiao Chen
ICASSP3
2025 Self-Supervised Monocular Depth Estimation from Videos via Pose-Adaptive Reconstruction
abstract
Self-supervised depth estimation from videos involves predicting the depth map of a target frame and the pose changes between source and target frames. The reconstructed source frame is aligned with the target view using the predicted pose and depth information. Precise pose estimation significantly impacts frame reconstruction, which, in turn, affects depth estimation performance. In our approach, we emphasize the importance of pose estimation, often overlooked by existing methods. We propose a self-supervised loss function that improves pose-adaptive reconstruction. Specifically, we decompose the conventional pose network into three parallel branches, each estimating pure translation, pure rotation, and full 6-DoF pose components independently. Our pose-adaptive reconstruction loss selects optimal pose parameterizations to minimize reconstruction errors, mitigating the impact of inaccurate posture. Our proposed pose estimation framework outperforms state-of-the-art methods on benchmark datasets.
Boqian Liu, Xinchen Ye
ICASSP3
2025 2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT Images
abstract
Classifying the status of NSCLC PD-L1 on chest CT is a cost-effective and non-invasive method. The existing multiple instance learning (MIL) methods are not effective for this task, due to the lack of an efficient feature encoder for 3D instances and ignoring the importance of representative instance selection. Thus, they cannot capture weak visual cues related to PD-L1 status on CT images. To address this, we propose a 2.5D top-K ranked multiple instance learning method. We design a 2.5D instance feature encoder, which takes advantage of knowledge from a 2D pre-trained model and has trainable parameters to learn information for 3D instances. In addition, we design a top-K ranked multiple instance learning strategy, which fully exploits the bag-level labels to select representative instances to eliminate the effect of atypical instances and guide the network to learn effective information. We demonstrate that our method can not only outperform state-of-the-art MIL methods on the PD-L1 status classification but also generalize well on a COVID-19 classification task.
Huadong Liu, Yongcen Li, Xinchen Ye, Hongkai Wang 0002, Yi Wang 0037, Dingpin Huang, Fangyi Xu, Yi Gan, Yuan Tu, Hongjie Hu
ICASSP4
2025 Mining Scene Structural Guidance for Thermal Images in Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation from RGB images has seen significant advancements recently, primarily because it eliminates the need for ground truth data during training. However, applying this technique to thermal images remains challenging due to their inherent characteristics, such as low contrast, low texture, and low signal-to-noise ratio, which impede accurate self-supervision. In this paper, we propose leveraging reliable and distinct scene structural information from thermal images to enhance self-supervised signals. We introduce structural losses, including explicit structural loss in the image space and implicit structural loss in the feature space, to improve self-supervised depth estimation. This approach mitigates the interference caused by the degraded characteristics of thermal images. Our method demonstrates superior performance compared to previous state-of-the-art approaches on the ViViD benchmark dataset, both quantitatively and qualitatively.
Xinchen Ye, Xia Mao, Rui Xu 0002
ICASSP1
2025 Delving into Transformer-based Network Architecture for Guided Depth Super-Resolution
abstract
Guided Depth Super-Resolution (GDSR) enhances low-resolution (LR) depth maps by leveraging high-resolution (HR) color images. The primary challenges involve achieving effective cross-modal data alignment and fusion, as well as incorporating multi-scale information within the Transformer architecture. To address these challenges, we propose a novel network architecture named DRMPNet which integrates two key components: Offset-based Detail Refinement (ODR) and Structure-guided Multi-scale Perception (SMP). ODR leverages offset calibration and window cross-attention to align and fuse LR depth maps with color images, effectively recovering local depth details. Meanwhile, SMP employs a structure generator and multi-scale cross-attention to capture scene details and structures at multiple scales, thereby enhancing the network’s contextual understanding. Extensive experiments on various benchmark datasets demonstrate the effectiveness of our method.
Xinchen Ye, Aokai Zhang, Rui Xu 0002
ICASSP1
2025 EchoCardMAE: Video Masked Auto-Encoders Customized for Echocardiography
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Miao Zhang 0004, Yi Wang 0037, Xin Fan 0001, Hongkai Wang 0002, Qingxiong Yue, Xiangjian He, Yen-Wei Chen 0001
MICCAI (13)3
2025 Guided Infrared Image Super-Resolution via Cross-modal Progressive Guidance
abstract
Guided Infrared image Super-Resolution (GISR) aims to reconstruct low-resolution infrared images by leveraging high-resolution visible images that provide rich geometric and high-frequency details. The primary challenge lies in establishing cross-modal association to effectively extract and fuse complementary features while mitigating redundant information. Therefore, we propose a novel network named CPGNet, which integrates two key components: the Cross-modal Gating Module (CGM) and the Cross-modal Collaborative Module (CCM). CGM employs cross-gating mechanism combined with asymmetric convolutions to dynamically enhance the salient features from both modalities and filter out irrelevant information. Meanwhile, CCM utilizes spatial and channel collaborative importance mapping along with masking mechanism to effectively explore and combine relevant details from two modalities and generate guidance information for infrared reconstruction. Additionally, we design a hierarchical architecture for progressive guidance, which fuses infrared features with cross-modal guidance cues by progressively integrating guidance information. Extensive experiments on various datasets demonstrate the effectiveness of our proposed method.
Xinchen Ye, Rui Xu 0002
ICMR1
2025 Dynamic Motion Modeling for Enhanced Visual-Inertial Odometry
abstract
Visual-inertial odometry (VIO) faces a key challenge in accurately capturing the dynamics of camera motion across different trajectories, which is essential for reliable pose estimation. In this paper, we propose a novel network architecture, named DMMNet, equipped with two pivotal modules: the Attention-Driven Motion Modeling (AMM) module and the Dynamic Motion Adaptation (DMA) module. AMM enhances motion feature extraction by modeling dynamic motion between frames. DMA improves the model's adaptability to varying motion states by adaptively adjusting translation and rotation weights, thereby enhancing the modeling of complex dynamic trajectories. Experimental results show that DMMNet outperforms existing VIO methods on the KITTI dataset, demonstrating its strong generalization and adaptability in dynamic environments.
Xinchen Ye, Rui Xu 0002
ICMR1
2025 Semantics-Driven Contrastive Learning for Real-World Depth Super Resolution
abstract
Low-resolution (LR) depth maps captured by depth sensors often suffer from structural distortions, noise, and blurring, limiting their practical usability. While most existing depth super-resolution (DSR) methods rely on synthetic datasets, they fail to accurately model real-world degradations, leading to poor performance on real-world data. To address this limitation, we identify two key challenges in the real-world DSR task: structural contour inconsistency and regional degradation inconsistency. The former arises from structural distortions in LR depth maps, while the latter stems from varying degradation levels in smooth regions. Upon this, we propose a Semantics-Driven Contrastive Learning (SDCL) pipeline for real-world DSR, leveraging semantic priors from the SAM model to enhance structural contour reconstruction and region-wise degradation handling. We introduce two novel contrastive loss functions: Structural Contour Alignment (SCA) loss, which aligns depth contours with semantic boundaries, and Regional Degradation Discrimination (RDD) loss, which optimizes smooth region restoration through region-level contrastive learning. Our approach is model-agnostic and can be seamlessly integrated into existing DSR frameworks. Experiments demonstrate that our method significantly enhances DSR performance on real-world DSR datasets.
Xinchen Ye, Aokai Zhang, Rui Xu 0002
ACM Multimedia1
2025 Self-Supervised Monocular Depth Estimation From Videos via Adaptive Reconstruction Constraints
abstract
To estimate depth maps from monocular videos in a self-supervised way, existing methods simultaneously predict the pose changes between adjacent frames and the depth maps of each frame, and then reconstruct the forward or backward frames using them, thereby casting depth estimation as a frame reconstruction problem. The corresponding reconstruction loss, which serves as a key supervision signal for training the whole network, can adversely affect the depth estimation accuracy if it is not properly established. In this paper, we propose a novel self-supervised monocular depth estimation method from videos via adaptive reconstruction constraints, i.e., designing the loss functions by establishing more accurate reconstruction constraints. Specifically, we first propose a pose-adaptive reconstruction loss to adaptively select the optimal pose parameterizations that yield the minimum reconstruction errors, reducing the impact of inaccurate posture on frame reconstruction. Then, we propose a region-sensitive reconstruction loss that fully utilizes the pretrained image reconstruction model to adaptively identify the poorly reconstructed regions and characterize the deviation of these regions on feature space. Finally, we additionally construct a multi-frame depth estimation network and design a reconstruction-guided bidirectional distillation loss to adaptively adjust the direction of distillation between networks of multi-frame and monocular depth estimation based on their current reconstruction quality, which encourages them to learn from each other and benefits the core task of monocular depth estimation. With our proposed losses, we achieve superior performance in comparison with state-of-the-art methods on benchmark datasets.
Xinchen Ye, Yuxiang Ou, Rui Xu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2025 Readiness Evaluation of Freeways for Lane-Detection Performance of Lidar-Based Automated Vehicles: A Field Test Analysis
abstract
Advanced driver-assistance systems with lane-detection functions are increasingly deployed in automated vehicles (AVs), but current road infrastructures may not accommodate to AVs’ safe operations. Freeways, where reliable perception is crucial for safe navigation, present higher safety risks. Lidar’s ability to provide precise depth information and function effectively in challenging scenarios becomes essential. However, fewer studies have explored roadway readiness for lidar-based performance. This study identified the effects of freeway design and condition on lane-detection performance of lidar-based AVs through a field test on two typical freeways in Shanghai, China, and the freeway readiness was evaluated. The Tongji University Road and Traffic Data Acquisition System was used for data collection. The lane-detection failure was treated as the label. Ten variables of five feature types were considered as the parameters, including road geometry, road segment, road marking, vehicle operation, and environment. The XGBoost ensemble machine learning algorithm and the SHapley Additive exPlanations (SHAP) were used for modeling and interpretation, respectively. Results show: 1) all features except for light condition were strongly correlated to lidar-based lane-detection failures; 2) higher failure probability was observed under circumstances like the presence of larger change rate of vertical curves, right-most lane with entrance or exit, taper extensions, special markings like words, worn-out markings, higher speed, and closer leading large-vehicle distance; and 3) several interaction effects were discovered. Results provide three contributions to AV safety: 1) make freeway design better accommodate to AVs; 2) provide ODD management references; and 3) offer technical improvement focuses to manufacturers.
Xinchen Ye, Xuesong Wang 0006, Salvatore Cafiso, Alfredo García 0003
IEEE Trans. Intell. Transp. Syst.1
2025 High-Quality Reconstruction of Depth Maps From Graph-Based Non-Uniform Sampling
abstract
Depth sensing is essential for intelligent computer vision applications, but it often suffers from low range precision and spatial resolution. To address this problem, we propose a novel framework that combines non-uniform sampling and reconstruction based on graph theory. Our framework consists of two main components: (1) a graph Laplacian induced non-uniform sampling (GLINUS) scheme that samples depth signals more densely around edges and contours than in smooth regions, and (2) an ensemble of priors (EoP) model that reconstructs the high-quality depth map using adaptive dual-tree discrete wavelet packets (ADDWP) transform, graph total variation regularizer, and graph Laplacian regularizer with color guidance. We solve the reconstruction problem using the alternating direction method of multipliers (ADMM). Our experiments demonstrate that our framework can capture fine structures and global information in depth signals and produce superior depth reconstruction results.
Jing-Yu Yang 0002, Yusen Hou, Xinchen Ye, Pascal Frossard, Kun Li 0001
IEEE Trans. Multim.4
2025 UW-Adapter: Adapting Monocular Depth Estimation Model in Underwater Scenes
abstract
Estimating depth maps from monocular underwater images poses one of the most challenging problems in underwater applications. Due to the lack of large-scale paired underwater color-depth datasets for effective training, existing style transfer-based and self-supervision-based approaches can improve the performance of depth estimation to some extent, but they remain unsatisfactory. Leveraging the power of massive training datasets, foundation models designed for terrestrial monocular depth estimation have demonstrated superior performance across various scenes. These models provide rich prior knowledge of 3D perception, which can be valuable for underwater depth estimation. Upon this, we introduce tunable adapters (UW-Adapter) that tailor a pre-trained foundation model specifically for underwater depth estimation, customizing it to the unique characteristics of underwater imagery. Our approach involves freezing the parameters of the pre-trained model and updating only the adapters through self-supervision. To address the complex degradation of underwater images, we propose two adapters: the transmission adapter and the high-frequency adapter. These adapters incorporate depth clues and high-frequency information as prior knowledge, thereby enhancing the performance of pre-trained model in underwater depth estimation. Experimental results demonstrate that by integrating lightweight adapters into off-the-shelf depth estimation foundation models, our method achieves superior performance across multiple datasets.
Xinchen Ye, Rui Xu 0002
IEEE Trans. Multim.1
2024 Novelty Detection Based Discriminative Multiple Instance Feature Mining to Classify NSCLC PD-L1 Status on HE-Stained Histopathological Images
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yi Wang 0037, Hongkai Wang 0002, Dingpin Huang, Fangyi Xu, Yi Gan, Yuan Tu, Hongjie Hu
MICCAI (4)4
2024 Low-resolution few-shot learning via multi-space knowledge distillation
Xinchen Ye, Baoli Sun, Hairui Yang, Rui Xu 0002, Zhihui Wang 0001
Inf. Sci.2
2024 GAN-BodyPose: Real-time 3D human body pose data key point detection and quality assessment assisted by generative adversarial network
Xicheng Zhu, Xinchen Ye
Image Vis. Comput.2
2024 CustomDepth: Customizing point-wise depth categories for depth completion
Shenglun Chen, Xinchen Ye, Hong Zhang 0011, Zhihui Wang 0001
Pattern Recognit. Lett.2
2024 C2ANet: Cross-Scale and Cross-Modality Aggregation Network for Scene Depth Super-Resolution
abstract
Existing depth super-resolution (DSR) methods typically utilize an additional high-resolution (HR) color image of the same scene as assistance to recover the low-resolution (LR) depth map. Although these color-guided methods have achieved impressive progress, they easily face with color image under-utilization and mis-utilization issues. In this article, we deeply investigate the above problems and further propose a novel DSR framework to alleviate them. Specifically, we propose a Cross-scale and Cross-modality Aggregation Network(C$^{2}$ANet)to learn abundant and accurate complementarity from color images to help recover the degraded depth map. Our C$^{2}$ANet can simultaneously extract multi-scale representations from color images with parallel network hierarchies, and effectively aggregate cross-scale and cross-modality contexts to boost HR representations in each hierarchy. Then, to appropriately use the guided color image, we further design a Feature Aggregation Module (FAM) to adaptively select and fuse task-relevant features, which consists of (1) afeature alignment blockto learn transformation offsets and align upsampled features with targeted HR features, and (2) afeature fusion blockbased on cross-attention mechanism to maintain strong structural context and suppress texture distraction. Experimental results on synthetic and real-world benchmark datasets demonstrate the superiority of our proposed method in comparison with other state-of-the-art DSR methods.
Xinchen Ye, Baoli Sun, Rui Xu 0002, Zhihui Wang 0001
IEEE Trans. Multim.1
2024 Digging into Depth and Color Spaces: A Mapping Constraint Network for Depth Super-Resolution
abstract
Scene depth super-resolution (DSR) poses an inherently ill-posed problem due to the extremely large space of one-to-many mapping functions from a given low-resolution (LR) depth map, which possesses limited depth information, to multiple plausible high-resolution (HR) depth maps. This characteristic renders the task highly challenging, as identifying an optimal solution becomes significantly intricate amidst this multitude of potential mappings. While simplistic constraints have been proposed to address the DSR task, the relationship between LR and HR depth maps and the color image has not been thoroughly investigated. In this paper, we introduce a novel mapping constraint network (MCNet) that incorporates additional constraints derived from both LR depth maps and color images. This integration aims to optimize the space of mapping functions and enhance the performance of DSR. Specifically, alongside the primary DSR network (DSRNet) dedicated to learning LR-to-HR mapping, we have developed an auxiliary degradation network (ADNet) that operates in reverse, generating the LR depth map from the reconstructed HR depth map to obtain depth features in LR space. To enhance the learning process of DSRNet in LR-to-HR mapping, we introduce two mapping constraints in LR space: (1) the cycle-consistent constraint, which offers additional supervision by establishing a closed loop between LR-to-HR and HR-to-LR mappings, and (2) the region-level contrastive constraint, aimed at reinforcing region-specific HR representations by explicitly modeling the consistency between LR and HR spaces. To leverage the color image effectively, we introduce a feature screening module to adaptively fuse color features at different layers, which can simultaneously maintain strong structural context and suppress texture distraction through subspace generation and image projection. Comprehensive experimental results across synthetic and real-world benchmark datasets unequivocally demonstrate the superiority of our proposed method over state-of-the-art DSR methods. Our MCNet achieves an average MAD reduction of 3.7% and 7.5% over state-of-the-art DSR method for ×8 and ×16 cases on Milddleburry dataset, respectively, without incurring additional costs during inference.
Baoli Sun, Tiantian Yan, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Discriminative Segment Focus Network for Fine-grained Video Action Recognition
abstract
Fine-grained video action recognition aims at identifying minor and discriminative variations among fine categories of actions. While many recent action recognition methods have been proposed to better model spatio-temporal representations, how to model the interactions among discriminative atomic actions to effectively characterize inter-class and intra-class variations has been neglected, which is vital for understanding fine-grained actions. In this work, we devise a Discriminative Segment Focus Network (DSFNet) to mine the discriminability of segment correlations and localize discriminative action-relevant segments for fine-grained video action recognition. Firstly, we propose a hierarchic correlation reasoning (HCR) module which explicitly establishes correlations between different segments at multiple temporal scales and enhances each segment by exploiting the correlations with other segments. Secondly, a discriminative segment focus (DSF) module is devised to localize the most action-relevant segments from the enhanced representations of HCR by enforcing the consistency between the discriminability and the classification confidence of a given segment with a consistency constraint. Finally, these localized segment representations are combined with the global action representation of the whole video for boosting final recognition. Extensive experimental results on two fine-grained action recognition datasets, i.e., FineGym and Diving48, and two action recognition datasets, i.e., Kinetics400 and Something-Something, demonstrate the effectiveness of our approach compared with the state-of-the-art methods.
Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Cross-Modality depth Estimation via Unsupervised Stereo RGB-to-infrared Translation
abstract
Existing depth estimation methods infer scene depth only from stereo visible light (RGB) images. Since RGB imaging is sensitive to changes in light, it’s difficult to estimate depth information accurately in some degraded visibility conditions. In contrast, infrared (IR) imaging captures thermal radiation and is not affected by brightness changing, providing extra clues for depth estimation. However, most datasets used for training in depth estimation do not have IR images paired with RGB-D data. Therefore, how to obtain the paired IR images and exploit the respective advantages of RGB and IR images to improve the performance of depth estimation, is of vital importance. Our core idea is to first develop an unsupervised RGB-to-IR translation (RIT) network with proposed Fourier domain adaptation and multi-space warping regularization to synthesize stereo IR images from their corresponding stereo RGB images. And then modified depth estimation backbones can be used as the cross-modality depth estimation (CDE) network to infer disparity maps from cross-modal RGB-IR stereo pairs. Assisted by the synthetic stereo IR images, we obtain superior performance just by flexibly deploying our framework to several off-the-shelf depth estimation backbones of single-modality (RGB) based methods.
Shi Tang, Xinchen Ye, Rui Xu 0002
ICASSP2
2023 Exploring Coarse-to-Fine Action Token Localization and Interaction for Fine-grained Video Action Recognition
abstract
Vision transformers have achieved impressive performance for video action recognition due to their strong capability of modeling long-range dependencies among spatio-temporal tokens. However, as for fine-grained actions, subtle and discriminative differences mainly exist in the regions of actors, directly utilizing vision transformers without removing irrelevant tokens will compromise recognition performance and lead to high computational costs. In this paper, we propose a coarse-to-fine action token localization and interaction network, namely C2F-ALIN, that dynamically localizes the most informative tokens at a coarse granularity and then partitions these located tokens to a fine granularity for sufficient fine-grained spatio-temporal interaction. Specifically, in the coarse stage, we devise a discriminative token localization module to accurately identify informative tokens and to discard irrelevant tokens, where each localized token corresponds to a large spatial region, thus effectively preserving the continuity of action regions.In the fine stage, we only further partition the localized tokens obtained in the coarse stage into a finer granularity and then characterize fine-grained token interactions in two aspects: (1) first using vanilla transformers to learn compact dependencies among all discriminative tokens; and (2) proposing a global contextual interaction module which enables each fine-grained tokens to communicate with all the spatio-temporal tokens and to embed the global context. As a result, our coarse-to-fine strategy is able to identify more relevant tokens and integrate global context for high recognition accuracy while maintaining high efficiency.Comprehensive experimental results on four widely used action recognition benchmarks, including FineGym, Diving48, Kinetics and Something-Something, clearly demonstrate the advantages of our proposed method in comparison with other state-of-the-art ones.
Baoli Sun, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Multimedia2
2023 Feature enhancement network for stereo matching
Shenglun Chen, Hong Zhang 0011, Baoli Sun, Xinchen Ye, Zhihui Wang 0001
Image Vis. Comput.5
2023 Underwater Depth Estimation via Stereo Adaptation Networks
abstract
With the fast development and wide application of stereo depth estimation, adequate high-quality stereo training data with groundtruth depth information plays an important role, but is not easily acquired in underwater environments. Therefore, satisfactory performance of depth estimation is difficult to achieve in underwater environments. In addition, the domain gap also leads to the failure of directly applying existing models of terrestrial scene to underwater scene. Therefore, this paper proposes a novel underwater depth estimation network which can infer depth maps from real underwater stereo images in an adaptation manner. The proposed learning pipeline mainly contains three different adaptation modules, i.e., style adaptation, semantic adaptation and disparity range adaptation, to progressively adapt a terrestrial depth estimation model to the underwater domain. Specifically, due to the lack of underwater training data, we first propose a depth-aware stereo image translation network to synthesize stylized underwater stereo images from terrestrial dataset, thus benefiting the effective training of depth estimation network. Then, considering the weak generalization to the real underwater data when only trained on the above synthetic data, we present a self-ensembling semantic adaptation for depth estimation network to minimize the semantic domain discrepancy between synthetic and real underwater data. Meanwhile, we design a disparity range adaptation module to address the problem of disparity range miss-match between both data, thus obtaining more accurate depth predictions for large-disparity-span underwater images. Experimental results show that by integrating the proposed adaptation modules into the off-the-shelf depth estimation backbones, our method successfully achieves superior performance of underwater depth estimation compared to other state-of-the-art methods.
Xinchen Ye, Yazhi Yuan, Rui Xu 0002, Zhihui Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Semantic Decomposition Network With Contrastive and Structural Constraints for Dental Plaque Segmentation
abstract
Segmenting dental plaque from images of medical reagent staining provides valuable information for diagnosis and the determination of follow-up treatment plan. However, accurate dental plaque segmentation is a challenging task that requires identifying teeth and dental plaque subjected to semantic-blur regions (i.e., confused boundaries in border regions between teeth and dental plaque) and complex variations of instance shapes, which are not fully addressed by existing methods. Therefore, we propose a semantic decomposition network (SDNet) that introduces two single-task branches to separately address the segmentation of teeth and dental plaque and designs additional constraints to learn category-specific features for each branch, thus facilitating the semantic decomposition and improving the performance of dental plaque segmentation. Specifically, SDNet learns two separate segmentation branches for teeth and dental plaque in a divide-and-conquer manner to decouple the entangled relation between them. Each branch that specifies a category tends to yield accurate segmentation. To help these two branches better focus on category-specific features, two constraint modules are further proposed: 1) contrastive constraint module (CCM) to learn discriminative feature representations by maximizing the distance between different category representations, so as to reduce the negative impact of semantic-blur regions on feature extraction; 2) structural constraint module (SCM) to provide complete structural information for dental plaque of various shapes by the supervision of an boundary-aware geometric constraint. Besides, we construct a large-scale open-source Stained Dental Plaque Segmentation dataset (SDPSeg), which provides high-quality annotations for teeth and dental plaque. Experimental results on SDPSeg datasets show SDNet achieves state-of-the-art performance.
Baoli Sun, Xinchen Ye, Zhihui Wang 0001, Xiaolong Luo, Heli Gao
IEEE Trans. Medical Imaging3
2022 Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 Lesions
abstract
Automatic segmentation of COVID-19 lesions is essential for computer-aided diagnosis. However, this task remains challenging because widely-used supervised based methods require large-scale annotated data that is difficult to obtain. Although an unsupervised method based on anomaly detection has shown promising results in [1], its performance is relatively poor. We address this problem by proposing a pixel-level and affinity-level knowledge distillation method. It obtains a pre-trained teacher network with rich semantic knowledge of CT images by constructing and training an auto-encoder at first, and then trains a student network with the same architecture as the teacher by distilling the teacher’s knowledge only from normal CT images, and finally localizes COVID-19 lesions using the feature discrepancy between the teacher and the student networks. Besides, except for the traditional pixel-level distillation, we design the affinity-level distillation that takes into account the pairwise relationship of features to fully distill effective knowledge. We evaluate this method by using three different COVID-19 datasets and the experimental results show that the segmentation performance is largely improved when it is compared with the other existing unsupervised anomaly detection methods.
Rui Xu 0002, Xinchen Ye, Yen-Wei Chen 0001, Fangyi Xu, Wenchao Zhu, Hongjie Hu, Xiaofeng Qu, Shoji Kido, Noriyuki Tomiyama
ICASSP3
2022 Underwater Stereo Matching Via Unsupervised Appearance And Feature Adaptation Networks
abstract
Stereo matching has been widely used to estimate depth maps in terrestrial environments. However, it is difficult to achieve appealing performance in underwater environments, since adequate underwater stereo data with groundtruth depth information is not easily available for training an underwater depth estimation model. In addition, the domain gap also leads to the failure of directly applying existing models of terrestrial scenes to underwater scenes. Therefore, this paper proposes a novel underwater depth estimation network which can infer depth maps from real underwater stereo images in an unsupervised adaptation manner. The proposed learning pipeline contains style adaptation (SA) in appearance space and feature adaptation (FA) in semantic space to progressively adapt the depth estimation models to underwater domain. Experimental results show that by integrating the proposed adaptation modules into the off-the-shelf stereo matching backbones, our method achieves a superior performance of underwater depth estimation compared to other state-of-the-art methods.
Yazhi Yuan, Xinchen Ye, Dian Zheng, Rui Xu 0002
ICASSP3
2022 Learning Data Hallucination and Reciprocal Guidance for Underwater Depth Estimation and Color Correction
abstract
Underwater vision is typically more difficult to tackle than open-air vision due to the degraded visibility and geometrical distortion, which impedes the development of underwater machine vision. Hence, we propose a joint depth estimation and color correction framework for underwater monocular images via data hallucination and reciprocal guidance learning. Specifically, due to the lack of labeled underwater data, we first design a data hallucination network to translate terrestrial images to multi-style synthetic underwater images while retaining the scene structure of terrestrial images from a single-source-multi-target perspective, benefiting the effective training of the joint tasks. Then, considering the strong connection between both tasks, we design a collaborative network to learn the reciprocal guidance between tasks from a multi-task perspective, thus improving the performance of each task. The whole framework can be trained end-to-end, and performs favorably against state-of-the-art methods in both depth estimation and color correction tasks.
Xinchen Ye, Rui Xu 0002
ICME2
2022 Local-Region and Cross-Dataset Contrastive Learning for Retinal Vessel Segmentation
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yen-Wei Chen 0001
MICCAI (2)3
2022 Low-Dose CT Reconstruction via Dual-Domain Learning and Controllable Modulation
Xinchen Ye, Rui Xu 0002, Zhihui Wang 0001
MICCAI (6)1
2022 Fine-grained Action Recognition with Robust Motion Representation Decoupling and Concentration
abstract
Fine-grained action recognition is a challenging task that requires identifying discriminative and subtle motion variations among fine-grained action classes. Existing methods typically focus on spatio-temporal feature extraction and long-temporal modeling to characterize complex spatio-temporal patterns of fine-grained actions. However, the learned spatio-temporal features without explicit motion modeling may emphasize more on visual appearance than on motion, which could compromise the learning of effective motion features required for fine-grained temporal reasoning. Therefore, how to decouple robust motion representations from the spatio-temporal features and further effectively leverage them to enhance the learning of discriminative features still remains less explored, which is crucial for fine-grained action recognition. In this paper, we propose a motion representation decoupling and concentration network (MDCNet) to address these two key issues. First, we devise a motion representation decoupling (MRD) module to disentangle the spatio-temporal representation into appearance and motion features through contrastive learning from video and segment views. Next, in the proposed motion representation concentration (MRC) module, the decoupled motion representations are further leveraged to learn a universal motion prototype shared across all the instances of each action class. Finally, we project the decoupled motion features onto all the motion prototypes through semantic relations to obtain the concentrated action-relevant features for each action class, which can effectively characterize the temporal distinctions of fine-grained actions for improved recognition performance. Comprehensive experimental results on four widely used action recognition benchmarks, i.e., FineGym, Diving48, Kinetics400 and Something-Something, clearly demonstrate the superiority of our proposed method in comparison with other state-of-the-art ones.
Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Multimedia2
2022 Brain gray matter nuclei segmentation on quantitative susceptibility mapping using dual-branch convolutional neural network
Chao Chai, Pengchong Qiao, Bin Zhao 0007, Huiying Wang, Wen Shen 0008, Chen Cao 0007, Xinchen Ye
Artif. Intell. Medicine9
2022 The Farther the Better: Balanced Stereo Matching via Depth-Based Sampling and Adaptive Feature Refinement
abstract
Existing stereo matching methods achieve satisfactory average accuracy on a whole predicted disparity map under common global metrics, but ignore the fine-grained performance at region level, especially for far regions in the scene, which is more crucial in actual auto-driving scenarios. There are two factors accounting for this problem: 1) Depth resolution. Existing methods use disparity-based sampling to extract matching candidates uniformly according to the disparity range, but leads to sparser sampling density at far regions than that of close regions in terms of depth range, resulting in low depth resolution in far regions. 2) Feature discriminability. Limited image resolution and inferior feature extraction at far regions result in the obtained features with low discrimination, which influences the subsequent matching process between stereo images. To improve the estimation accuracy of far regions and thus achieve a balanced performance at region level, we design a novel two-stage Balanced Stereo Matching Network (BSMNet) to address the above problems. The coarse stage of BSMNet introduces a direct depth-based sampling strategy, which generates matching candidates according to scene depth instead of disparity, thus improving the depth resolution and obtaining initial depth map with more balanced accuracy. Then, a depth refinement stage is proposed to solve the problem of low feature discriminability and further optimizes the initial depth map obtained from the coarse stage. It selects matching candidates and computes their similarity scores from a carefully designed adaptive feature volume guided by a learnable scale map, thus making the final estimation more accurate. Different from the existing methods that construct stereo matching based on disparity prediction, our proposed pipeline is to directly optimize on depth information. Experiments show that our BSMNet can obtain an obvious performance improvement at far regions without discarding that at close regions, so as to largely outperform existing state-of-the-art methods.
Hong Zhang 0011, Xinchen Ye, Shenglun Chen, Zhihui Wang 0001, Wanli Ouyang
IEEE Trans. Circuits Syst. Video Technol.2
2021 Leveraging Line-Point Consistence To Preserve Structures for Wide Parallax Image Stitching
abstract
Generating high-quality stitched images with natural structures is a challenging task in computer vision. In this paper, we succeed in preserving both local and global geometric structures for wide parallax images, while reducing artifacts and distortions. A projective invariant, Characteristic Number, is used to match co-planar local sub-regions for input images. The homography between these well-matched sub-regions produces consistent line and point pairs, suppressing artifacts in overlapping areas. We explore and introduce global collinear structures into an objective function to specify and balance the desired characters for image warping, which can preserve both local and global structures while alleviating distortions. We also develop comprehensive measures for stitching quality to quantify the collinearity of points and the discrepancy of matched line pairs by considering the sensitivity to linear structures for human vision. Extensive experiments demonstrate the superior performance of the proposed method over the state-of-the-art by presenting sharp textures and preserving prominent natural structures in stitched images. Especially, our method not only exhibits lower errors but also the least divergence across all test images. Code is available at https://github.com/dut-media-lab/Image-Stitching.
Qi Jia 0001, Zhengjun Li, Xin Fan 0001, Shiyu Teng, Xinchen Ye, Longin Jan Latecki
CVPR6
2021 Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution
abstract
Existing color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training samples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected in actual testing environment. Therefore, we explore for the first time to learn the cross-modality knowledge at training stage, where both RGB and depth modalities are available, but test on the target dataset, where only single depth modality exists. Our key idea is to distill the knowledge of scene structural guidance from RGB modality to the single DSR task without changing its network architecture. Specifically, we construct an auxiliary depth estimation (DE) task that takes an RGB image as input to estimate a depth map, and train both DSR task and DE task collaboratively to boost the performance of DSR. Upon this, a cross-task interaction module is proposed to realize bilateral cross-task knowledge transfer. First, we design a cross-task distillation scheme that encourages DSR and DE networks to learn from each other in a teacher-student role-exchanging fashion. Then, we advance a structure prediction (SP) task that provides extra structure regularization to help both DSR and DE networks learn more informative structure representations for depth recovery. Extensive experiments demonstrate that our scheme achieves superior performance in comparison with other DSR methods.
Baoli Sun, Xinchen Ye, Baopu Li, Zhihui Wang 0001, Rui Xu 0002
CVPR2
2021 DPNet: Detail-preserving network for high quality monocular depth estimation
Xinchen Ye, Shude Chen, Rui Xu 0002
Pattern Recognit.1
2021 Sparse intrinsic decomposition and applications
Kun Li 0001, Xinchen Ye, Chenggang Yan 0001, Jing-Yu Yang 0002
Signal Process. Image Commun.3
2021 Unsupervised Monocular Depth Estimation via Recursive Stereo Distillation
abstract
Existing unsupervised monocular depth estimation methods resort to stereo image pairs instead of ground-truth depth maps as supervision to predict scene depth. Constrained by the type of monocular input in testing phase, they fail to fully exploit the stereo information through the network during training, leading to the unsatisfactory performance of depth estimation. Therefore, we propose a novel architecture which consists of a monocular network (Mono-Net) that infers depth maps from monocular inputs, and a stereo network (Stereo-Net) that further excavates the stereo information by taking stereo pairs as input. During training, the sophisticated Stereo-Net guides the learning of Mono-Net and devotes to enhance the performance of Mono-Net without changing its network structure and increasing its computational burden. Thus, monocular depth estimation with superior performance and fast runtime can be achieved in testing phase by only using the lightweight Mono-Net. For the proposed framework, our core idea lies in: 1) how to design the Stereo-Net so that it can accurately estimate depth maps by fully exploiting the stereo information; 2) how to use the sophisticated Stereo-Net to improve the performance of Mono-Net. To this end, we propose a recursive estimation and refinement strategy for Stereo-Net to boost its performance of depth estimation. Meanwhile, a multi-space knowledge distillation scheme is designed to help Mono-Net amalgamate the knowledge and master the expertise from Stereo-Net in a multi-scale fashion. Experiments demonstrate that our method achieves the superior performance of monocular depth estimation in comparison with other state-of-the-art methods.
Xinchen Ye, Xin Fan 0001, Mingliang Zhang 0002, Rui Xu 0002
IEEE Trans. Image Process.1
2021 Joint Extraction of Retinal Vessels and Centerlines Based on Deep Semantics and Multi-Scaled Cross-Task Aggregation
abstract
Retinal vessel segmentation and centerline extraction are crucial steps in building a computer-aided diagnosis system on retinal images. Previous works treat them as two isolated tasks, while ignoring their tight association. In this paper, we propose a deep semantics and multi-scaled cross-task aggregation network that takes advantage of the association to jointly improve their performances. Our network is featured by two sub-networks. The forepart is a deep semantics aggregation sub-network that aggregates strong semantic information to produce more powerful features for both tasks, and the tail is a multi-scaled cross-task aggregation sub-network that explores complementary information to refine the results. We evaluate the proposed method on three public databases, which are DRIVE, STARE and CHASE_DB1. Experimental results show that our method can not only simultaneously extract retinal vessels and their centerlines but also achieve the state-of-the-art performances on both tasks.
Rui Xu 0002, Xinchen Ye, Lin Lin 0008, Liang Li 0002, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics3
2020 Unsupervised Content-Preserved Adaptation Network for Classification of Pulmonary Textures from Different CT Scanners
abstract
Deep network based methods have been proposed for accurate classification of pulmonary textures on CT images. However, such methods well-trained on CT data from one scanner cannot perform well when they are directly applied to the data from other scanners. This domain shift problem is caused by different physical components and scanning protocols of different CT scanners. In this paper, we propose an unsupervised content-preserved adaptation network to address this problem. Our method can make a previously well-trained deep network to be adapted for the data of a new CT scanner and does not require the laboring annotation to delineate pulmonary texture regions on the new CT data. Extensive evaluations have been carried on images collected from GE and Toshiba CT scanners and show that the proposed method can alleviate the performance degradation problem of classifying pulmonary textures from different CT scanners.
Rui Xu 0002, Zhen Cong, Xinchen Ye, Shoji Kido, Noriyuki Tomiyama
ICASSP3
2020 Retinal Vessel Segmentation via a Semantics and Multi-Scale Aggregation Network
abstract
Precise segmentation of retinal vessels is crucial for a computer-aided diagnosis system of retinal fundus images. However, this task remains challenging due to large variations in scales and poor segmentation of capillary vessels. In this paper, we propose a semantics and multi-scale aggregation network to address these difficulties. It includes semantics aggregation blocks that are designed for aggregating stronger high-level semantic information. These carefully designed blocks produce more semantic feature representation that is helpful for capillary vessel identification and vessel connection. Besides, a multi-scale aggregation block is designed by employing parallel dilated convolutional filters with different dilation rates to fully exploit the multi-scale information. We evaluate the network by using two public databases of retinal vessel segmentation and compare its performance with several leading methods published in the past several years. Extensive evaluations show that the proposed network has achieved the state-of-the-art performance on the public CHASE DB1 and HRF datasets.
Rui Xu 0002, Xinchen Ye, Guiliang Jiang, Liang Li 0002
ICASSP2
2020 Cascaded Detail-Aware Network for Unsupervised Monocular Depth Estimation
abstract
Existing unsupervised learning methods usually reformulate the depth estimation into the image reconstruction problem by training on stereo image pairs to circumvent the need of dense labeled ground truth depth information. Most of them are designed based on a simple encoder-decoder backbone architecture, which has limited expression for context information and suffers from the loss of depth details. In this paper, we propose a cascaded detail-aware network which contains a contextual network (CN) followed by consecutive spatial networks (SNs) to make an unsupervised coarse-to-fine prediction. CN aims to provide good initialized depth estimation results by introducing a multi-scale attention fusion module to enhance the ability of feature representation. Then, SN is progressively applied on the coarse depth map to produce refined depth outputs by exploiting abundant spatial details from input color image. Moreover, we design a robust loss function that further considers the penalty of photometric errors and the occlusion, and strengthens the recovery of spatial details for better depth estimation. Experimental results show that the proposed method achieves promising performance.
Xinchen Ye, Mingliang Zhang 0002, Xin Fan 0001, Rui Xu 0002, Juncheng Pu, Ruoke Yan
ICME1
2020 Unsupervised Detection of Pulmonary Opacities for Computer-Aided Diagnosis of COVID-19 on CT Images
abstract
COVID-19 emerged towards the end of 2019 which was identified as a global pandemic by the world heath organization (WHO). With the rapid spread of COVID-19, the number of infected and suspected patients has increased dramatically. Chest computed tomography (CT) has been recognized as an efficient tool for the diagnosis of COVID-19. However, the huge CT data make it difficult for radiologist to fully exploit them on the diagnosis. In this paper, we propose a computer-aided diagnosis system that can automatically analyze CT images to distinguish the COVID-19 against to community-acquired pneumonia (CAP). The proposed system is based on an unsupervised pulmonary opacity detection method that locates opacity regions by a detector unsupervisedly trained from CT images with normal lung tissues. Radiomics based features are extracted insides the opacity regions, and fed into classifiers for classification. We evaluate the proposed CAD system by using 200 CT images collected from different patients in several hospitals. The accuracy, precision, recall, f1-score and AUC achieved are 95.5%, 100%, 91%, 95.1% and 95.9% respectively, exhibiting the promising capacity on the differential diagnosis of COVID-19 from CT images.
Rui Xu 0002, Xiao Cao, Yen-Wei Chen 0001, Xinchen Ye, Lin Lin 0008, Wenchao Zhu, Fangyi Xu, Hongjie Hu, Shoji Kido, Noriyuki Tomiyama
ICPR5
2020 BG-Net: Boundary-Guided Network for Lung Segmentation on Clinical CT Images
abstract
Lung segmentation on CT images is a crucial step for a computer-aided diagnosis system of lung diseases. The existing deep learning based lung segmentation methods are less efficient to segment lungs on clinical CT images, especially that the segmentation on lung boundaries is not accurate enough due to complex pulmonary opacities in practical clinics. In this paper, we propose a boundary-guided network (BG-Net) to address this problem. It contains two auxiliary branches that seperately segment lungs and extract the lung boundaries, and an aggregation branch that efficiently exploits lung boundary cues to guide the network for more accurate lung segmentation on clinical CT images. We evaluate the proposed method on a private dataset collected from the Osaka university hospital and four public datasets including StructSeg [1], HUG [2], VESSEL12 [3], and a Novel Coronavirus 2019 (COVID-19) dataset [4]. Experimental results show that the proposed method can segment lungs more accurately and outperform several other deep learning based methods.
Rui Xu 0002, Yi Wang 0037, Xinchen Ye, Lin Lin 0008, Yen-Wei Chen 0001, Shoji Kido, Noriyuki Tomiyama
ICPR4
2020 Detail- Revealing Deep Low-Dose CT Reconstruction
abstract
Low-dose CT imaging emerges with low radiation risk due to the reduction of radiation dose, but brings negative impact on the imaging quality. This paper addresses the problem of low-dose CT reconstruction. Previous methods are unsatisfactory due to the inaccurate recovery of image details under the strong noise generated by the reduction of radiation dose, which directly affects the final diagnosis. To suppress the noise effectively while retain the structures well, we propose a detail-revealing dual-branch aggregation network to effectively reconstruct the degraded CT image. Specifically, the main reconstruction branch iteratively exploits and compensates the reconstruction errors to gradually refine the CT image, while the prior branch is to learn the structure details as prior knowledge to help recover the CT image. A sophisticated detail-revealing loss is designed to fuse the information from both branches and guide the learning to obtain better performance from pixel-wise and holistic perspectives respectively. Experimental results show that our method outperforms the state-of-art methods in both PSNR and SSIM metrics.
Xinchen Ye, Yuyao Xu, Rui Xu 0002, Shoji Kido, Noriyuki Tomiyama
ICPR1
2020 Boosting Connectivity in Retinal Vessel Segmentation via a Recursive Semantics-Guided Network
Rui Xu 0002, Xinchen Ye, Lin Lin 0008, Yen-Wei Chen 0001
MICCAI (5)3
2020 Depth Super-Resolution via Deep Controllable Slicing Network
abstract
Due to the imaging limitation of depth sensors, high-resolution (HR) depth maps are often difficult to be acquired directly, thus effective depth super-resolution (DSR) algorithms are needed to generate HR output from its low-resolution (LR) counterpart. Previous methods treat all depth regions equally without considering different extents of degradation at region-level, and regard DSR under different scales as independent tasks without considering the modeling of different scales, which impede further performance improvement and practical use of DSR. To alleviate these problems, we propose a deep controllable slicing network from a novel perspective. Specifically, our model is to learn a set of slicing branches in a divide-and-conquer manner, parameterized by a distance-aware weighting scheme to adaptively aggregate different depths in an ensemble. Each branch that specifies a depth slice (e.g., the region in some depth range) tends to yield accurate depth recovery. Meanwhile, a scale-controllable module that extracts depth features under different scales is proposed and inserted into the front of slicing network, and enables finely-grained control of the depth restoration results of slicing network with a scale hyper-parameter. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our method achieves superior performance.
Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li
ACM Multimedia1
2020 DRM-SLAM: Towards dense reconstruction of monocular SLAM with scene depth fusion
Xinchen Ye, Xiang Ji 0005, Baoli Sun, Shenglun Chen, Zhihui Wang 0001
Neurocomputing1
2020 Unsupervised detail-preserving network for high quality monocular depth estimation
Mingliang Zhang 0002, Xinchen Ye, Xin Fan 0001
Neurocomputing2
2020 Unsupervised depth estimation from monocular videos with hybrid geometric-refined loss and contextual attention
Mingliang Zhang 0002, Xinchen Ye, Xin Fan 0001
Neurocomputing2
2020 Depth upsampling based on deep edge-aware learning
Zhihui Wang 0001, Xinchen Ye, Baoli Sun, Jing-Yu Yang 0002, Rui Xu 0002
Pattern Recognit.2
2020 A sparsity-promoting image decomposition model for depth recovery
Xinchen Ye, Mingliang Zhang 0002, Jing-Yu Yang 0002, Xin Fan 0001, Fangfang Guo
Pattern Recognit.1
2020 Deep Joint Depth Estimation and Color Correction From Monocular Underwater Images Based on Unsupervised Adaptation Networks
abstract
Degraded visibility and geometrical distortion typically make the underwater vision more intractable than open air vision, which impedes the development of underwater-related machine vision and robotic perception. Therefore, this paper addresses the problem of joint underwater depth estimation and color correction from monocular underwater images, which aims at enjoying the mutual benefits between these two related tasks from a multi-task perspective. Our core ideas lie in our new deep learning architecture. Due to the lack of effective underwater training data, and the weak generalization to the real-world underwater images trained on synthetic data, we consider the problem from a novel perspective of style-level and feature-level adaptation, and propose an unsupervised adaptation network to deal with the joint learning problem. Specifically, a style adaptation network (SAN) is first proposed to learn a style-level transformation to adapt in-air images to the style of underwater domain. Then, we formulate a task network (TN) to jointly estimate the scene depth and correct the color from a single underwater image by learning domain-invariant representations. The whole framework can be trained end-to-end in an adversarial learning manner. Extensive experiments are conducted under air-to-water domain adaptation settings. We show that the proposed method performs favorably against state-of-the-art methods in both depth estimation and color correction tasks.
Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Rui Xu 0002, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 PMBANet: Progressive Multi-Branch Aggregation Network for Scene Depth Super-Resolution
abstract
Depth map super-resolution is an ill-posed inverse problem with many challenges. First, depth boundaries are generally hard to reconstruct particularly at large magnification factors. Second, depth regions on fine structures and tiny objects in the scene are destroyed seriously by downsampling degradation. To tackle these difficulties, we propose a progressive multi-branch aggregation network (PMBANet), which consists of stacked MBA blocks to fully address the above problems and progressively recover the degraded depth map. Specifically, each MBA block has multiple parallel branches: 1) The reconstruction branch is proposed based on the designed attention-based error feed-forward/-back modules, which iteratively exploits and compensates the downsampling errors to refine the depth map by imposing the attention mechanism on the module to gradually highlight the informative features at depth boundaries. 2) We formulate a separate guidance branch as prior knowledge to help to recover the depth details, in which the multi-scale branch is to learn a multi-scale representation that pays close attention at objects of different scales, while the color branch regularizes the depth map by using auxiliary color information. Then, a fusion block is introduced to adaptively fuse and select the discriminative features from all the branches. The design methodology of our whole network is well-founded, and extensive experiments on benchmark datasets demonstrate that our method achieves superior performance in comparison with the state-of-the-art methods. Our code and models are available athttps://github.com/Sunbaoli/PMBANet_DSR/.
Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li
IEEE Trans. Image Process.1
2020 Pulmonary Textures Classification via a Multi-Scale Attention Network
abstract
Precise classification of pulmonary textures is crucial to develop a computer aided diagnosis (CAD) system of diffuse lung diseases (DLDs). Although deep learning techniques have been applied to this task, the classification performance is not satisfied for clinical requirements, since commonly-used deep networks built by stacking convolutional blocks are not able to learn discriminative feature representation to distinguish complex pulmonary textures. For addressing this problem, we design a multi-scale attention network (MSAN) architecture comprised by several stacked residual attention modules followed by a multi-scale fusion module. Our deep network can not only exploit powerful information on different scales but also automatically select optimal features for more discriminative feature representation. Besides, we develop visualization techniques to make the proposed deep model transparent for humans. The proposed method is evaluated by using a large dataset. Experimental results show that our method has achieved the average classification accuracy of 94.78% and the average f-value of 0.9475 in the classification of 7 categories of pulmonary textures. Besides, visualization results intuitively explain the working behavior of the deep network. The proposed method has achieved the state-of-the-art performance to classify pulmonary textures on high resolution CT images.
Rui Xu 0002, Zhen Cong, Xinchen Ye, Yasushi Hirano, Shoji Kido, Tomoko Gyobu, Yutaka Kawata, Osamu Honda, Noriyuki Tomiyama
IEEE J. Biomed. Health Informatics3
2019 Does Video Content Facilitate or Impair Comprehension of Documentaries? The Effect of Cognitive Abilities and Eye Movement Strategy
Yueyuan Zheng, Xinchen Ye, Janet Hui-wen Hsiao
CogSci2
2019 Graph Based Non-Uniform Sampling and Reconstruction of Depth Maps
abstract
High-quality depth sensing is highly demanded in intelligent computer vision, 3DTV, and many other related fields. However, prevalent time-of-fly (ToF) depth sensors are of low resolution as the number of pixel-level demodulators is limited. Moreover, the rectangular sampling does not consider the signal characteristics of depth maps. Being a departure of previous resolution enhancement on rectangular sampling, this paper investigates the non-uniform sampling of depth maps, and the high-resolution depth reconstruction from limited non-uniformly distributed samples. The proposed depth sampling and reconstruction schemes are developed based on graph signal processing. We first propose a graph-based non-uniform sampling (GNS) scheme, where depth signals are sampled based on the response of a high-pass graph filter, which results in denser sampling around discontinuities such as edges and contours than in smooth regions. We then propose a graph-based depth reconstruction (GDR) framework,where a graph Laplacian regularizer is designed to fully exploit structural correlation between the depth and photometric images. To solve the reconstruction problem, we derive an efficient algorithm based on the alternating direction method of multipliers (ADMM). Experimental results show that the GNS-GDR non-uniform sampling and reconstruction method achieves high-quality depth sensing, outperforming several state-of-the-art schemes.
Jing-Yu Yang 0002, Xinchen Ye, Pascal Frossard, Kun Li 0001
ICIP3
2019 Unsupervised Monocular Depth Estimation Based on Dual Attention Mechanism and Depth-Aware Loss
abstract
Most existing monocular depth estimation approaches are su- pervised, but enough quantities of ground truth depth data are required during training. To cope with this, recent techniques deal with the depth estimation task in an unsupervised man- ner, i.e., replacing the use of depth data with easily obtained stereo images for training. Based on this, we propose a nov- el unsupervised learning architecture, which integrates dual attention mechanism into the framework and designs a depth- aware loss for better depth estimation. Specifically, to en- hance the ability of feature representations, we introduce a d- ual attention module to capture global feature dependencies in spatial and channel dimensions for scene understanding and depth estimation. Meanwhile, we propose a depth-aware loss that fully addresses the occlusion problem in brightness con- stancy assumption, the intrinsic characteristics of depth map, and the left-right consistency problem, respectively. Besides, an adversarial loss is employed to discriminate synthetic or realistic depth maps by training a discriminator so as to pro- duce better results. Extensive experiments on KITTI dataset show that our approach achieves state-of-the-art performance compared with other monocular depth estimation methods.
Xinchen Ye, Mingliang Zhang 0002, Rui Xu 0002, Xin Fan 0001, Zhu Liu 0004, Jiaao Zhang
ICME1
2018 Pulmonary Textures Classification Using A Deep Neural Network with Appearance and Geometry Cues
abstract
Classification of pulmonary textures on CT images is essential for the development of a computer-aided diagnosis system of diffuse lung diseases. In this paper, we propose a novel method to classify pulmonary textures by using a deep neural network, which can make full use of appearance and geometry cues of textures via a dual-branch architecture. The proposed method has been evaluated by a dataset that includes seven kinds of typical pulmonary textures. Experimental results show that our method outperforms the state-of-the-art methods including feature engineering based method and convolutional neural network based method.
Rui Xu 0002, Zhen Cong, Xinchen Ye, Yasushi Hirano, Shoji Kido
ICASSP3
2018 Depth Super-Resolution with Deep Edge-Inference Network and Edge-Guided Depth Filling
abstract
In this paper, we propose a novel depth super-resolution framework with deep edge-inference network and edge-guided depth filling. We first construct a convolutional neural network (CNN) architecture to learn a binary map of depth edge location from low resolution depth map and corresponding color image. Then, a fast edge-guided depth filling strategy is proposed to interpolate the missing depth constrained by the acquired edges to prevent predicting across the depth boundaries. Experimental results show that our method outperforms the state-of-art methods in both the edges inference and the final results of depth super-resolution, and generalizes well for handling depth data captured in different scenes.
Xinchen Ye, Xiangyue Duan
ICASSP1
2018 High Quality Depth Estimation from Monocular Images Based on Depth Prediction and Enhancement Sub-Networks
abstract
This paper addresses the problem of depth estimation from a single RGB image. Previous methods mainly focus on the problems of depth prediction accuracy and output depth resolution, but seldom of them can tackle these two problems well. Here, we present a novel depth estimation framework based on deep convolutional neural network (CNN) to learn the mapping between monocular images and depth maps. The proposed architecture can be divided into two components, i.e., depth prediction and depth enhancement sub-networks. We first design a depth prediction network based on the ResNet architecture to infer the scene depth from color image. Then, a depth enhancement network is concatenated to the end of the depth prediction network to obtain a high resolution depth map. Experimental results show that the proposed method outperforms other methods on benchmark RGB-D datasets and achieves state-of-the-art performance.
Xiangyue Duan, Xinchen Ye
ICME2
2018 Dense Reconstruction from Monocular Slam with Fusion of Sparse Map-Points and Cnn-Inferred Depth
abstract
Real-time monocular visual SLAM approaches relying on building sparse correspondences between two or multiple views of the scene, are capable of accurately tracking camera pose and inferring structure of the environment. However, these methods have the common problem, i.e., the reconstructed 3D map is extremely sparse. Recently, convolutional neural network (CNN) is widely used for estimating scene depth from monocular color images. As we observe, sparse map-points generated from epipolar geometry are locally accurate, while CNN-inferred depth map contains high-level global context but generates blurry depth boundaries. Therefore, we propose a depth fusion framework to yield a dense monocular reconstruction that fully exploits the sparse depth samples and the CNN-inferred depth. Color key-frames are employed to guide the depth reconstruction process, avoiding smoothing over depth boundaries. Experimental results on benchmark datasets show the robustness and accuracy of our method.
Xiang Ji 0005, Xinchen Ye, Hongcan Xu
ICME2
2017 Intrinsic decomposition from a single RGB-D image with sparse and non-local priors
abstract
This paper proposes a new intrinsic image decomposition method that decomposes a single RGB-D image into reflectance and shading components. We observe and verify that, a shading image mainly contains smooth regions separated by curves, and its gradient distribution is sparse. We therefore use ℓ1-norm to model the direct irradiance component - the main sub-component extracted from shading component. Moreover, a non-local prior weighted by a bilateral kernel on a larger neighborhood is designed to fully exploit structural correlation in the reflectance component to improve the decomposition performance. The model is solved by the alternating direction method under the augmented Lagrangian multiplier (ADM-ALM) framework. Experimental results on both synthetic and real datasets demonstrate that the proposed method yields better results and enjoys lower complexity compared with two state-of-the-art methods.
Kun Li 0001, Jing-Yu Yang 0002, Xinchen Ye
ICME4
2017 Reconstruction of Structurally-Incomplete Matrices With Reweighted Low-Rank and Sparsity Priors
abstract
Most matrix reconstruction methods assume that missing entries randomly distribute in the incomplete matrix, and the low-rank prior or its variants are used to well pose the problem. However, in practical applications, missing entries are structurally rather than randomly distributed, and cannot be handled by the rank minimization prior individually. To remedy this, this paper introduces new matrix reconstruction models using double priors on the latent matrix, named Reweighted Low-rank and Sparsity Priors (ReLaSP). In the proposed ReLaSP models, the matrix is regularized by a low-rank prior to exploit the inter-column and inter-row correlations, and its columns (rows) are regularized by a sparsity prior under a dictionary to exploit intra-column (-row) correlations. Both the low-rank and sparse priors are reweighted on the fly to promote low-rankness and sparsity, respectively. Numerical algorithms to solve our ReLaSP models are derived via the alternating direction method under the augmented Lagrangian multiplier framework. Results on synthetic data, image restoration tasks, and seismic data interpolation show that the proposed ReLaSP models are quite effective in recovering matrices degraded by highly structural missing and various types of noise, complementing the classic matrix reconstruction models that handle random missing only.
Jing-Yu Yang 0002, Xuemeng Yang, Xinchen Ye, Chunping Hou
IEEE Trans. Image Process.3
2016 Completion of structurally-incomplete matrices with reweighted low-rank and sparsity priors
abstract
Most matrix completion methods impose a low-rank prior or its variants to well pose the problem. However, the rank minimization is problematic to handle matrices with structural missing. To remedy this, this paper introduces a new matrix completion method using double priors on the latent matrix, named Reweighted Low-rank and Sparsity Priors. In the proposed model, the matrix is regularized by a low-rank prior to exploit the inter-column (row) correlations, and its columns (rows) are regularized by a sparsity prior under a dictionary to exploit intra-column (row) correlations. Both the low-rank and sparse priors are reweighted on the fly to promote low-rankness and sparsity, respectively. Numerical algorithm to solve our model is derived via the alternating direction method under the augmented Lagrangian multiplier framework. Experimental results show that our model is quite effective in recovering matrices with highly-structural missing, complementing the classic matrix completion models that handle random missing only.
Jing-Yu Yang 0002, Xuemeng Yang, Xinchen Ye
ICASSP3
2016 Energy-efficient strategy for cloud storage based on the characteristics of remote sensing image data
abstract
To the problem that random placement in distributed file system leads to the low utilization of servers, this article models according to the characteristics of remote sensing image data blocks, clusters data by setting storage centre with access frequency so that the adjacent remote sensing image data blocks in spatial position are near each other in physical storage as well. It promotes the response speed of the system, places data blocks according to data block groups, reshuffle the non-grouped data blocks at the system low load and turns off dispensable datanodes to achieve energy saving. The experiments suggest that in the inquiry of remote sensing image data, it is more efficient to place data blocks according to their own characteristics than random placement. Compared with common dynamic data placement strategies, this strategy performs better in energy saving when the system is at moderate load.
Xinchen Ye, Yurong Qian, Jiaji Wu, Hailong Zhang 0002
IGARSS1
2016 Depth refinement for binocular kinect RGB-D cameras
abstract
This paper presents a novel depth refinement framework for binocular Kinect RGB-D cameras for obtaining high quality depth map. Firstly, we build a binocular depth sensing system using two Kinect v2 cameras, and analyze the systematic error of the system from two aspects, i.e., camera interaction and intrinsic characteristics. Then, the captured depth maps from different views are fused to fully exploit the inter-view correlations, and an error compensation method is proposed to remove the systematic errors from the fused depth map. Finally, an edge-guided depth propagation scheme is used to refine the depth map from binocular depth map. Experimental results show that the proposed framework is able to substantially improve the quality of depth image.
Jinghui Bai, Jing-Yu Yang 0002, Xinchen Ye, Chunping Hou
VCIP3
2016 Depth recovery via decomposition of polynomial and piece-wise constant signals
abstract
This paper proposes a novel decomposition model for high-quality depth recovery (DMDR) from low quality depth measurement accompanied by high-resolution RGB image. We observe that depth patches extracted from the depth map containing smooth regions separated by curves, can be decomposed simultaneously by a low-order polynomial surface and a piece-wise constant signal. In our model, the polynomial surface component is regularized by least-square polynomial smoothing, while the piece-wise constant component is constrained by total variation filtering. The model is effectively solved by the alternating direction method under the augmented Lagrangian multiplier (ALM-ADM) algorithm. Experimental results show that our method is able to handle various types of depth degradation under the designed signal decomposition model, and produces high-quality depth recovery results.
Xinchen Ye, Jing-Yu Yang 0002, Chunping Hou, Yao Wang 0001
VCIP1
2015 Foreground-Background Separation From Video Clips via Motion-Assisted Matrix Restoration
abstract
Separation of video clips into foreground and background components is a useful and important technique, making recognition, classification, and scene analysis more efficient. In this paper, we propose a motion-assisted matrix restoration (MAMR) model for foreground-background separation in video clips. In the proposed MAMR model, the backgrounds across frames are modeled by a low-rank matrix, while the foreground objects are modeled by a sparse matrix. To facilitate efficient foreground-background separation, a dense motion field is estimated for each frame, and mapped into a weighting matrix which indicates the likelihood that each pixel belongs to the background. Anchor frames are selected in the dense motion estimation to overcome the difficulty of detecting slowly moving objects and camouflages. In addition, we extend our model to a robust MAMR model against noise for practical applications. Evaluations on challenging datasets demonstrate that our method outperforms many other state-of-the-art methods, and is versatile for a wide range of surveillance videos.
Xinchen Ye, Jing-Yu Yang 0002, Kun Li 0001, Chunping Hou, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2014 Background extraction from video sequences via motion-assisted matrix completion
abstract
Background extraction from video sequences is a useful and important technique in video surveillance. This paper proposes a motion-assisted matrix completion model for background extraction from video sequences. A binary motion map is first calculated for each frame by optical flow. By excluding areas associated with moving objects with the binary motion maps, the background extraction is formulated into a motion-assisted matrix completion (MAMC) problem. Experimental results show that our method not only extracts promising backgrounds but also outperforms many state-of-the-art methods in distinguishing moving objects on challenging datasets.
Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001
ICIP3
2014 Color-Guided Depth Recovery From RGB-D Data Using an Adaptive Autoregressive Model
abstract
This paper proposes an adaptive color-guided autoregressive (AR) model for high quality depth recovery from low quality measurements captured by depth cameras. We observe and verify that the AR model tightly fits depth maps of generic scenes. The depth recovery task is formulated into a minimization of AR prediction errors subject to measurement consistency. The AR predictor for each pixel is constructed according to both the local correlation in the initial depth map and the nonlocal similarity in the accompanied high quality color image. We analyze the stability of our method from a linear system point of view, and design a parameter adaptation scheme to achieve stable and accurate depth recovery. Quantitative and qualitative evaluation compared with ten state-of-the-art schemes show the effectiveness and superiority of our method. Being able to handle various types of depth degradations, the proposed method is versatile for mainstream depth sensors, time-of-flight camera, and Kinect, as demonstrated by experiments on real systems.
Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou, Yao Wang 0001
IEEE Trans. Image Process.2
2012 Depth Recovery Using an Adaptive Color-Guided Auto-Regressive Model
Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou
ECCV (5)2