VLDB 2026 Research / reviewers in the wild / expert
Baoli Sun
dblp:264/9035
· DBLP profile ↗
24ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-2861-4288ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniAlign: A Universal Cross-Modality Knowledge Alignment Framework for Fine-Grained Action RecognitionabstractThe key to fine-grained video action recognition is identifying subtle differences between action categories. Relying solely on visual features supervised by action labels makes it challenging to characterize robust and discriminative action dynamics from videos. With significant advancements in human pose estimation and the powerful capabilities of Vision-Language Models (VLMs), obtaining reliable and cost-free human pose data and textual semantics has become increasingly feasible, enabling their effective use in fine-grained action recognition. However, the inherent disparities in feature representations across different modalities necessitate a robust alignment strategy to achieve opti mal fusion. To address this, we propose a universal cross-modality knowledge alignment framework, namely UniAlign, to transfer the knowledge from such pre-trained multi-modal models into action recognition models. Specifically, UniAlign introduces two additional branches to extract pose features and textual semantics with the pre-trained pose encoder and VLM. To align the action relevant cues among video features, pose features, and textual semantics, we propose a Cross-Modality Similarity Aggregation module (CMSA) that utilizes the importance of different modal cues while aggregating cross-modal similarities. Additionally, we adopt a fine-tuning mechanism similar to Exponential Moving Average (EMA) to refine the textual semantics, ensuring that the semantic representations encoded by VLMs are preserved while being optimized towards the specific task preferences. Extensive experiments on widely used fine-grained action recognition benchmarks (e.g., FineGym, NTURGB-D, Diving48) and coarse-grained K400 dataset demonstrate the effectiveness of the proposed UniAlign method. Yihan Wang 0011, Baoli Sun, Xinzhu Ma, Zhihui Wang 0001, Zhiyong Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2026 | Toward an Effective Action-Region Tracking Framework for Fine-Grained Video Action RecognitionabstractFine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture coarse-grained motion patterns but struggle to identify subtle details in local regions evolving over time. In this work, we introduce the action-region tracking (ART) framework, a novel solution leveraging a query-response mechanism to discover and track the dynamics of distinctive local details, enabling distinguishing similar actions effectively. Specifically, we propose a region-specific semantic activation module that employs discriminative and text-constrained semantics serve as queries to capture the most action-related region responses in each video frame, facilitating interaction among spatial and temporal dimensions with corresponding video features. The captured region responses are then organized into action tracklets, which characterize the region-based action dynamics by linking related responses across different video frames in a coherent sequence. The text-constrained queries are designed to expressly encode nuanced semantic representations derived from the textual descriptions of action labels, as extracted by the language branches within visual language models. To optimize generated action tracklets, we design a multilevel tracklet contrastive constraint among multiple region responses at spatial and temporal levels, which can effectively distinguish individual region responses in each video frame (spatial level) and establish the correlation of similar region responses between adjacent video frames (temporal level). In addition, we implement a task-specific fine-tuning mechanism to refine textual semantics during training. This ensures that the semantic representations encoded by vision language models (VLMs) are not only preserved but also optimized for specific task preferences. Comprehensive experiments on several widely used action recognition benchmarks, i.e., FineGym, Diving48, NTURGB-D, Kinetics, and Something-Something, clearly demonstrate the superiority to previous state-of-the-art baselines. Baoli Sun, Yihan Wang 0011, Xinzhu Ma, Zhihui Wang 0001, Kun Lu 0003, Zhiyong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | RobAVA: A Large-Scale Dataset and Baseline Towards Video Based Robotic Arm Action Understanding
Baoli Sun, Xinzhu Ma, Anqi Zou, Chuixuan Fan, Zhihui Wang 0001, Kun Lu 0003, Zhiyong Wang 0001 |
ICCV | 1 |
| 2025 | Structure-preserving dental plaque segmentation via dynamically complementary information interaction
Rui Xu 0002, Baoli Sun, Tiantian Yan, Zhihui Wang 0001 |
Multim. Syst. | 3 |
| 2025 | Referring Video Object Segmentation With Cross-Modality Proxy QueriesabstractReferring video object segmentation (RVOS) is an emerging cross-modality task that aims to generate pixel-level maps of the target objects referred by given textual expressions. The main concept involves learning an accurate alignment visual elements and language expressions within a semantic space. Recent approaches address cross-modality alignment through conditional queries, tracking the target object using a queryresponse based mechanism built upon transformer structure. However, they exhibit two limitations: (1) these conditional queries, identifying the same object across different frames through the same query, lack inter-frame dependency and variation modeling, making accurate target tracking challenging amid significant frame-to-frame variations; and (2) they handle the temporal feature of a video and build visual-language interaction sequentially, integrating textual constraints belatedly, which may cause the video features potentially focus on the non-referred objects. Therefore, we propose a novel RVOS architecture called ProxyFormer, which introduces a set of proxy queries to integrate visual and text semantics and facilitate the flow of semantics between them. By progressively updating and propagating proxy queries across multiple stages of video feature encoder, ProxyFormer ensures that the video features are as focused as much possible on the object of interest. This dynamic evolution of the queries across video also enables the proxy queries to establish inter-frame dependencies, enhancing the accuracy and coherence of object tracking throughout the video sequence. To mitigate the high computational costs associated with full spatio-temporal interactions between video and proxy queries, we propose to decouple cross-modality interactions into their temporal and spatial dimensions, respectively. Additionally, we design a Joint Semantic Consistency (JSC) training strategy to align semantic consensus between the proxy queries and the combined videotext pairs. Comprehensive experiments on four widely used RVOS benchmarks, i.e., Ref-Youtube-VOS, Ref-DAVIS17, A2D-Sentences and JHMDB-Sentences, clearly demonstrate the superiority of our ProxyFormer to the state-of-the-art methods Baoli Sun, Xinzhu Ma, Zhihui Wang 0001, Zhiyong Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | P$^{2}$M: Progressive Perspective Mining for Referring Video Object Segmentation
Yihan Wang 0011, Baoli Sun, Xinzhu Ma, Hong-Wei Ge, Jiulin Fan |
IEEE Trans. Multim. | 2 |
| 2025 | Real-Time Depth Completion With Multimodal Feature AlignmentabstractAs a key problem in computer vision, depth completion aims to recover dense depth maps from sparse ones [generally derived from light detection and ranging (LiDAR)]. Most methods introduce synchronous RGB images and leverage multimodal fusion to integrate multimodal features from these modalities to describe the complete scene. However, their different natural characteristics lead to inconsistency in features, potentially impacting the effectiveness of multimodal feature fusion. To address this issue, we propose a feature alignment network (FANet) that introduces an alignment scheme to enhance the consistency between multimodal features. This scheme aligns the modality-invariant semantic context, which is invariant to changes in modality and represents the correlation between a pixel and its surroundings. Specifically, we first design an asymmetric context extraction (ACE) module to extract modality-invariant semantic contexts from multimodal features within limited GPU memory, and then pull them closer to improve consistency. Crucially, our alignment scheme is only applied during the training phase, and no additional computation cost is incurred in the inference phase. Moreover, we introduce a simple yet effective refinement module to refine estimated results via residual learning based on intermediate depth maps and sparse depth maps. Extensive experiments on KITTI and VOID datasets demonstrate that our method achieves competitive performance against typical real-time methods. In addition, we embed the proposed alignment scheme and refinement module into other methods to demonstrate their effectiveness. Shenglun Chen, Xinzhu Ma, Hong Zhang 0011, Baoli Sun, Zhihui Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Fine-grained Video Semantic Distillation for Video-Text Retrieval
Zuyi Pei, Baoli Sun, Zhihui Wang 0001 |
MMAsia | 2 |
| 2024 | Low-resolution few-shot learning via multi-space knowledge distillation
Xinchen Ye, Baoli Sun, Hairui Yang, Rui Xu 0002, Zhihui Wang 0001 |
Inf. Sci. | 3 |
| 2024 | C2ANet: Cross-Scale and Cross-Modality Aggregation Network for Scene Depth Super-ResolutionabstractExisting depth super-resolution (DSR) methods typically utilize an additional high-resolution (HR) color image of the same scene as assistance to recover the low-resolution (LR) depth map. Although these color-guided methods have achieved impressive progress, they easily face with color image under-utilization and mis-utilization issues. In this article, we deeply investigate the above problems and further propose a novel DSR framework to alleviate them. Specifically, we propose a Cross-scale and Cross-modality Aggregation Network(C$^{2}$ANet)to learn abundant and accurate complementarity from color images to help recover the degraded depth map. Our C$^{2}$ANet can simultaneously extract multi-scale representations from color images with parallel network hierarchies, and effectively aggregate cross-scale and cross-modality contexts to boost HR representations in each hierarchy. Then, to appropriately use the guided color image, we further design a Feature Aggregation Module (FAM) to adaptively select and fuse task-relevant features, which consists of (1) afeature alignment blockto learn transformation offsets and align upsampled features with targeted HR features, and (2) afeature fusion blockbased on cross-attention mechanism to maintain strong structural context and suppress texture distraction. Experimental results on synthetic and real-world benchmark datasets demonstrate the superiority of our proposed method in comparison with other state-of-the-art DSR methods. Xinchen Ye, Baoli Sun, Rui Xu 0002, Zhihui Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Digging into Depth and Color Spaces: A Mapping Constraint Network for Depth Super-ResolutionabstractScene depth super-resolution (DSR) poses an inherently ill-posed problem due to the extremely large space of one-to-many mapping functions from a given low-resolution (LR) depth map, which possesses limited depth information, to multiple plausible high-resolution (HR) depth maps. This characteristic renders the task highly challenging, as identifying an optimal solution becomes significantly intricate amidst this multitude of potential mappings. While simplistic constraints have been proposed to address the DSR task, the relationship between LR and HR depth maps and the color image has not been thoroughly investigated. In this paper, we introduce a novel mapping constraint network (MCNet) that incorporates additional constraints derived from both LR depth maps and color images. This integration aims to optimize the space of mapping functions and enhance the performance of DSR. Specifically, alongside the primary DSR network (DSRNet) dedicated to learning LR-to-HR mapping, we have developed an auxiliary degradation network (ADNet) that operates in reverse, generating the LR depth map from the reconstructed HR depth map to obtain depth features in LR space. To enhance the learning process of DSRNet in LR-to-HR mapping, we introduce two mapping constraints in LR space: (1) the cycle-consistent constraint, which offers additional supervision by establishing a closed loop between LR-to-HR and HR-to-LR mappings, and (2) the region-level contrastive constraint, aimed at reinforcing region-specific HR representations by explicitly modeling the consistency between LR and HR spaces. To leverage the color image effectively, we introduce a feature screening module to adaptively fuse color features at different layers, which can simultaneously maintain strong structural context and suppress texture distraction through subspace generation and image projection. Comprehensive experimental results across synthetic and real-world benchmark datasets unequivocally demonstrate the superiority of our proposed method over state-of-the-art DSR methods. Our MCNet achieves an average MAD reduction of 3.7% and 7.5% over state-of-the-art DSR method for ×8 and ×16 cases on Milddleburry dataset, respectively, without incurring additional costs during inference. Baoli Sun, Tiantian Yan, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Discriminative Segment Focus Network for Fine-grained Video Action RecognitionabstractFine-grained video action recognition aims at identifying minor and discriminative variations among fine categories of actions. While many recent action recognition methods have been proposed to better model spatio-temporal representations, how to model the interactions among discriminative atomic actions to effectively characterize inter-class and intra-class variations has been neglected, which is vital for understanding fine-grained actions. In this work, we devise a Discriminative Segment Focus Network (DSFNet) to mine the discriminability of segment correlations and localize discriminative action-relevant segments for fine-grained video action recognition. Firstly, we propose a hierarchic correlation reasoning (HCR) module which explicitly establishes correlations between different segments at multiple temporal scales and enhances each segment by exploiting the correlations with other segments. Secondly, a discriminative segment focus (DSF) module is devised to localize the most action-relevant segments from the enhanced representations of HCR by enforcing the consistency between the discriminability and the classification confidence of a given segment with a consistency constraint. Finally, these localized segment representations are combined with the global action representation of the whole video for boosting final recognition. Extensive experimental results on two fine-grained action recognition datasets, i.e., FineGym and Diving48, and two action recognition datasets, i.e., Kinetics400 and Something-Something, demonstrate the effectiveness of our approach compared with the state-of-the-art methods. Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Exploring Coarse-to-Fine Action Token Localization and Interaction for Fine-grained Video Action RecognitionabstractVision transformers have achieved impressive performance for video action recognition due to their strong capability of modeling long-range dependencies among spatio-temporal tokens. However, as for fine-grained actions, subtle and discriminative differences mainly exist in the regions of actors, directly utilizing vision transformers without removing irrelevant tokens will compromise recognition performance and lead to high computational costs. In this paper, we propose a coarse-to-fine action token localization and interaction network, namely C2F-ALIN, that dynamically localizes the most informative tokens at a coarse granularity and then partitions these located tokens to a fine granularity for sufficient fine-grained spatio-temporal interaction. Specifically, in the coarse stage, we devise a discriminative token localization module to accurately identify informative tokens and to discard irrelevant tokens, where each localized token corresponds to a large spatial region, thus effectively preserving the continuity of action regions.In the fine stage, we only further partition the localized tokens obtained in the coarse stage into a finer granularity and then characterize fine-grained token interactions in two aspects: (1) first using vanilla transformers to learn compact dependencies among all discriminative tokens; and (2) proposing a global contextual interaction module which enables each fine-grained tokens to communicate with all the spatio-temporal tokens and to embed the global context. As a result, our coarse-to-fine strategy is able to identify more relevant tokens and integrate global context for high recognition accuracy while maintaining high efficiency.Comprehensive experimental results on four widely used action recognition benchmarks, including FineGym, Diving48, Kinetics and Something-Something, clearly demonstrate the advantages of our proposed method in comparison with other state-of-the-art ones. Baoli Sun, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001 |
ACM Multimedia | 1 |
| 2023 | Feature enhancement network for stereo matching
Shenglun Chen, Hong Zhang 0011, Baoli Sun, Xinchen Ye, Zhihui Wang 0001 |
Image Vis. Comput. | 3 |
| 2023 | Iterative Class Prototype Calibration for Transductive Zero-Shot LearningabstractZero-shot learning (ZSL) typically suffers from the domain shift issue since the projected feature embedding of unseen samples mismatch with the corresponding class semantic prototypes, making it very challenging to fine-tune an optimal visual-semantic mapping for the unseen domain. Some existing transductive ZSL methods solve this problem by introducing unlabeled samples of the unseen domain, in which the projected features of unseen samples are still not discriminative and tend to be distributed around prototypes of seen classes. Therefore, how to effectively align the projection features of samples in unseen classes with corresponding predefined class prototypes is crucial for promoting the generalization of ZSL models. In this paper, we propose a novel Iterative Class Prototype Calibration (ICPC) framework for transductive ZSL which consists of a pseudo-labeling stage and a model retraining stage to address the above key issue. First, in the labeling stage, we devise a Class Prototype Calibration (CPC) module to calibrate the predefined class prototypes of the unseen domain by estimating the real center of projected feature distribution, which achieves better matching of sample points and class prototypes. Next, in the retraining stage, we devise a Certain Samples Screening (CSS) module to select relatively certain unseen samples with high confidence and align them with predefined class prototypes in the embedding space. A progressive training strategy is adopted to select more certain samples and update the proposed model with augmented training data. Extensive experiments on AwA2, CUB, and SUN datasets demonstrate that the proposed scheme achieves new state-of-the-art in the conventional setting under both standard split (SS) and proposed split (PS). Hairui Yang, Baoli Sun, Baopu Li, Caifei Yang, Zhihui Wang 0001, Jenhui Chen, Lei Wang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Semantic Decomposition Network With Contrastive and Structural Constraints for Dental Plaque SegmentationabstractSegmenting dental plaque from images of medical reagent staining provides valuable information for diagnosis and the determination of follow-up treatment plan. However, accurate dental plaque segmentation is a challenging task that requires identifying teeth and dental plaque subjected to semantic-blur regions (i.e., confused boundaries in border regions between teeth and dental plaque) and complex variations of instance shapes, which are not fully addressed by existing methods. Therefore, we propose a semantic decomposition network (SDNet) that introduces two single-task branches to separately address the segmentation of teeth and dental plaque and designs additional constraints to learn category-specific features for each branch, thus facilitating the semantic decomposition and improving the performance of dental plaque segmentation. Specifically, SDNet learns two separate segmentation branches for teeth and dental plaque in a divide-and-conquer manner to decouple the entangled relation between them. Each branch that specifies a category tends to yield accurate segmentation. To help these two branches better focus on category-specific features, two constraint modules are further proposed: 1) contrastive constraint module (CCM) to learn discriminative feature representations by maximizing the distance between different category representations, so as to reduce the negative impact of semantic-blur regions on feature extraction; 2) structural constraint module (SCM) to provide complete structural information for dental plaque of various shapes by the supervision of an boundary-aware geometric constraint. Besides, we construct a large-scale open-source Stained Dental Plaque Segmentation dataset (SDPSeg), which provides high-quality annotations for teeth and dental plaque. Experimental results on SDPSeg datasets show SDNet achieves state-of-the-art performance. Baoli Sun, Xinchen Ye, Zhihui Wang 0001, Xiaolong Luo, Heli Gao |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Fine-grained Action Recognition with Robust Motion Representation Decoupling and ConcentrationabstractFine-grained action recognition is a challenging task that requires identifying discriminative and subtle motion variations among fine-grained action classes. Existing methods typically focus on spatio-temporal feature extraction and long-temporal modeling to characterize complex spatio-temporal patterns of fine-grained actions. However, the learned spatio-temporal features without explicit motion modeling may emphasize more on visual appearance than on motion, which could compromise the learning of effective motion features required for fine-grained temporal reasoning. Therefore, how to decouple robust motion representations from the spatio-temporal features and further effectively leverage them to enhance the learning of discriminative features still remains less explored, which is crucial for fine-grained action recognition. In this paper, we propose a motion representation decoupling and concentration network (MDCNet) to address these two key issues. First, we devise a motion representation decoupling (MRD) module to disentangle the spatio-temporal representation into appearance and motion features through contrastive learning from video and segment views. Next, in the proposed motion representation concentration (MRC) module, the decoupled motion representations are further leveraged to learn a universal motion prototype shared across all the instances of each action class. Finally, we project the decoupled motion features onto all the motion prototypes through semantic relations to obtain the concentrated action-relevant features for each action class, which can effectively characterize the temporal distinctions of fine-grained actions for improved recognition performance. Comprehensive experimental results on four widely used action recognition benchmarks, i.e., FineGym, Diving48, Kinetics400 and Something-Something, clearly demonstrate the superiority of our proposed method in comparison with other state-of-the-art ones. Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001 |
ACM Multimedia | 1 |
| 2022 | Discriminative Feature Mining and Enhancement Network for Low-Resolution Fine-Grained Image RecognitionabstractExisting fine-grained image recognition methods are difficult to learn complete discriminative features from low-resolution (LR) data, because the original subtle inter-class distinctions become slimmer with the reduction of the image resolution. Besides, existing methods of LR fine-grained image recognition and general LR image recognition only consider the restoration and extraction of global discriminative features, ignoring unreliable local fine-grained details can be detrimental to final recognition. To address the above problems, we propose a multi-tasking framework, discriminative feature mining and enhancement network (DME-Net), for the LR fine-grained image recognition task, which aims to capture the reliable object descriptions from macro and micro perspectives, respectively. Macroscopically, we train the framework’s ability to recover and extract global discriminative features based on the whole images. Microscopically, we purposefully reinforce the framework’s ability to repair and capture the local discriminative details on the mined informative parts. To precisely excavate the most potential parts, we design an informative part mining (IPM) module, in which we firstly employ a part generation layer to predict several part masks that focus on different discriminative parts under the guidance of discrepancy loss and discriminant loss. Then we introduce a part selection (PS) submodule to further screen out a group of most informative parts from the predicted part masks according to their corresponding scores, which measure the semantic correlation degree of each part to the others. Experimental results on three benchmark datasets and one retail product dataset consistently show that our proposed framework can significantly boost the performance of the baseline model. Besides, extensive ablation studies are conducted, which further prove the effectiveness of each component of our designs. Tiantian Yan, Baoli Sun, Zhihui Wang 0001, Zhongxuan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-ResolutionabstractExisting color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training samples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected in actual testing environment. Therefore, we explore for the first time to learn the cross-modality knowledge at training stage, where both RGB and depth modalities are available, but test on the target dataset, where only single depth modality exists. Our key idea is to distill the knowledge of scene structural guidance from RGB modality to the single DSR task without changing its network architecture. Specifically, we construct an auxiliary depth estimation (DE) task that takes an RGB image as input to estimate a depth map, and train both DSR task and DE task collaboratively to boost the performance of DSR. Upon this, a cross-task interaction module is proposed to realize bilateral cross-task knowledge transfer. First, we design a cross-task distillation scheme that encourages DSR and DE networks to learn from each other in a teacher-student role-exchanging fashion. Then, we advance a structure prediction (SP) task that provides extra structure regularization to help both DSR and DE networks learn more informative structure representations for depth recovery. Extensive experiments demonstrate that our scheme achieves superior performance in comparison with other DSR methods. Baoli Sun, Xinchen Ye, Baopu Li, Zhihui Wang 0001, Rui Xu 0002 |
CVPR | 1 |
| 2020 | Depth Super-Resolution via Deep Controllable Slicing NetworkabstractDue to the imaging limitation of depth sensors, high-resolution (HR) depth maps are often difficult to be acquired directly, thus effective depth super-resolution (DSR) algorithms are needed to generate HR output from its low-resolution (LR) counterpart. Previous methods treat all depth regions equally without considering different extents of degradation at region-level, and regard DSR under different scales as independent tasks without considering the modeling of different scales, which impede further performance improvement and practical use of DSR. To alleviate these problems, we propose a deep controllable slicing network from a novel perspective. Specifically, our model is to learn a set of slicing branches in a divide-and-conquer manner, parameterized by a distance-aware weighting scheme to adaptively aggregate different depths in an ensemble. Each branch that specifies a depth slice (e.g., the region in some depth range) tends to yield accurate depth recovery. Meanwhile, a scale-controllable module that extracts depth features under different scales is proposed and inserted into the front of slicing network, and enables finely-grained control of the depth restoration results of slicing network with a scale hyper-parameter. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our method achieves superior performance. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li |
ACM Multimedia | 2 |
| 2020 | DRM-SLAM: Towards dense reconstruction of monocular SLAM with scene depth fusion
Xinchen Ye, Xiang Ji 0005, Baoli Sun, Shenglun Chen, Zhihui Wang 0001 |
Neurocomputing | 3 |
| 2020 | Depth upsampling based on deep edge-aware learning
Zhihui Wang 0001, Xinchen Ye, Baoli Sun, Jing-Yu Yang 0002, Rui Xu 0002 |
Pattern Recognit. | 3 |
| 2020 | Deep Joint Depth Estimation and Color Correction From Monocular Underwater Images Based on Unsupervised Adaptation NetworksabstractDegraded visibility and geometrical distortion typically make the underwater vision more intractable than open air vision, which impedes the development of underwater-related machine vision and robotic perception. Therefore, this paper addresses the problem of joint underwater depth estimation and color correction from monocular underwater images, which aims at enjoying the mutual benefits between these two related tasks from a multi-task perspective. Our core ideas lie in our new deep learning architecture. Due to the lack of effective underwater training data, and the weak generalization to the real-world underwater images trained on synthetic data, we consider the problem from a novel perspective of style-level and feature-level adaptation, and propose an unsupervised adaptation network to deal with the joint learning problem. Specifically, a style adaptation network (SAN) is first proposed to learn a style-level transformation to adapt in-air images to the style of underwater domain. Then, we formulate a task network (TN) to jointly estimate the scene depth and correct the color from a single underwater image by learning domain-invariant representations. The whole framework can be trained end-to-end in an adversarial learning manner. Extensive experiments are conducted under air-to-water domain adaptation settings. We show that the proposed method performs favorably against state-of-the-art methods in both depth estimation and color correction tasks. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Rui Xu 0002, Xin Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | PMBANet: Progressive Multi-Branch Aggregation Network for Scene Depth Super-ResolutionabstractDepth map super-resolution is an ill-posed inverse problem with many challenges. First, depth boundaries are generally hard to reconstruct particularly at large magnification factors. Second, depth regions on fine structures and tiny objects in the scene are destroyed seriously by downsampling degradation. To tackle these difficulties, we propose a progressive multi-branch aggregation network (PMBANet), which consists of stacked MBA blocks to fully address the above problems and progressively recover the degraded depth map. Specifically, each MBA block has multiple parallel branches: 1) The reconstruction branch is proposed based on the designed attention-based error feed-forward/-back modules, which iteratively exploits and compensates the downsampling errors to refine the depth map by imposing the attention mechanism on the module to gradually highlight the informative features at depth boundaries. 2) We formulate a separate guidance branch as prior knowledge to help to recover the depth details, in which the multi-scale branch is to learn a multi-scale representation that pays close attention at objects of different scales, while the color branch regularizes the depth map by using auxiliary color information. Then, a fusion block is introduced to adaptively fuse and select the discriminative features from all the branches. The design methodology of our whole network is well-founded, and extensive experiments on benchmark datasets demonstrate that our method achieves superior performance in comparison with the state-of-the-art methods. Our code and models are available athttps://github.com/Sunbaoli/PMBANet_DSR/. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li |
IEEE Trans. Image Process. | 2 |