VLDB 2026 Research / reviewers in the wild / expert
Qiudan Zhang
dblp:190/4336
· DBLP profile ↗
25ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0001-6067-8188ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DBGroup: Dual-Branch Point Grouping for Weakly Supervised 3D Semantic Instance SegmentationabstractWeakly supervised 3D instance segmentation is essential for 3D scene understanding, especially as the growing scale of data and high annotation costs associated with fully supervised approaches. Existing methods primarily rely on two forms of weak supervision: one-thing-one-click annotations and bounding box annotations, both of which aim to reduce labeling efforts. However, these approaches still encounter limitations, including labor-intensive annotation processes, high complexity, and reliance on expert annotators. To address these challenges, we propose DBGroup, a two-stage weakly supervised 3D instance segmentation framework that leverages scene-level annotations as a more efficient and scalable alternative. In the first stage, we introduce a Dual-Branch Point Grouping module to generate pseudo labels guided by semantic and mask cues extracted from multi-view images. To further improve label quality, we develop two refinement strategies: Granularity-Aware Instance Merging and Semantic Selection and Propagation. The second stage involves multi-round self-training on an end-to-end instance segmentation network using the refined pseudo-labels. Additionally, we introduce an Instance Mask Filter strategy to address inconsistencies within the pseudo labels. Extensive experiments demonstrate that DBGroup achieves competitive performance compared to sparse-point-level supervised 3D instance segmentation methods, while surpassing state-of-the-art scene-level supervised 3D semantic segmentation approaches. Xuexun Liu, Xiaoxu Xu, Qiudan Zhang, Lin Ma 0002, Xu Wang 0006 |
AAAI | 3 |
| 2026 | A Visual-Linguistic Approach for Robust RGB-Thermal Tracking With Dynamic Template AdaptationabstractRGB-T tracking seeks to improve tracking robust ness in complex environments by exploiting the complementary advantages of RGB and thermal infrared (TIR) modalities. Nevertheless, current RGB-T tracking methods encounter two critical limitations stemming from their reliance on fixed template images for search region matching. First, the resemblance between the initial fixed template and the target in later frames diminishes as time goes on, causing tracking reliability to decline. Second, background clutter within template bounding boxes introduces disruptive noise that compromises tracking accuracy. To address these challenges, we propose a new referring RGB-T tracking approach that integrates visual-linguistic cues to enhance multimodal data, enabling background noise elimination in the template and dynamic template image updating through adaptive mechanisms. Additionally, to facilitate research in language guided multimodal tracking, we construct a large-scale dataset named Refer-RGBT, which comprises 1508 multimodal video pairs with synchronized RGB/TIR sequences of diverse objects. Each sequence is annotated by descriptive textual captions that detail their appearance and actions. Extensive evaluations on the Refer-RGBT dataset demonstrate that our proposed approach achieves cutting-edge performance compared to the state-of-the art (SOTA) methods. Xu Wang 0006, Huanxin Zheng, Haohong Liao, Qiudan Zhang, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 4 |
| 2025 | CWC-DNERF: Compact Dynamic Neural Radiance Field VIA Discrete Wavelet Transform And Learnable CodebooksabstractNeural radiance fields have significantly advanced dynamic scene reconstruction and novel view synthesis. However, relying on multiple implicit multi-layer perceptrons for reconstructing dynamic scenes is computationally expensive. Recent methods have alleviated this challenge by introducing explicit data structures, such as voxel grids and feature planes, but these significantly increase storage demands and complicate network transmission. We propose Cwc-DNeRF, a compact dynamic NeRF representation that leverages discrete wavelet transform (DWT) and learnable codebooks to achieve superior storage efficiency while maintaining competitive rendering quality compared to K-Planes. In Stage I, DWT and trainable masks are employed to optimize parameter efficiency, resulting in sparse space planes. In Stage II, learnable codebooks are introduced for the space-time planes to merge redundant spatio-temporal features further reducing the storage demand. Additionally, a data compression pipeline is applied to compress both sparse space plane parameters and codebooks. Experimental results on D-NeRF and DyNeRF datasets show that our method achieves state-of-the-art rendering quality within a 10MB storage budget while retaining the benefits of explicit feature planes. Yaojian Xu, Qiudan Zhang, Longhao Zou, Qiong Liu 0001, Xu Wang 0006 |
ICIP | 3 |
| 2025 | Semantic driven ViT and dual-level token fusion for weakly supervised semantic segmentation
Qiudan Zhang, Jianmin Jiang |
Neurocomputing | 2 |
| 2025 | Geometry-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Asymmetric Dual-Domain Collaborative LearningabstractBeing able to estimate monocular depth for spherical panoramas is of fundamental importance in 3D scene perception. However, spherical distortion severely limits the effectiveness of vanilla convolutions. To push the envelope of accuracy, recent approaches attempt to utilize Tangent projection (TP) to estimate the depth of$360 ^{\circ }$images. Yet, these methods still suffer from discrepancies and inconsistencies among patch-wise tangent images, as well as the lack of accurate ground truth depth maps under a supervised fashion. In this paper, we propose a geometry-aware self-supervised$360 ^{\circ }$image depth estimation methodology that explores the complementary advantages of TP and Equirectangular projection (ERP) by an asymmetric dual-domain collaborative learning strategy. Especially, we first develop a lightweight asymmetric dual-domain depth estimation network, which enables to aggregate depth-related features from a single TP domain, and then produce depth distributions of the TP and ERP domains via collaborative learning. This effectively mitigates stitching artifacts and preserves fine details in depth inference without overspending model parameters. In addition, a frequent-spatial feature concentration module is devised to simultaneously capture non-local Fourier features and local spatial features, such that facilitating the efficient exploration of monocular depth cues. Moreover, we introduce a geometric structural alignment module to further improve geometric structural consistency among tangent images. Extensive experiments illustrate that our designed approach outperforms existing self-supervised$360 ^{\circ }$depth estimation methods on three publicly available benchmark datasets. Xu Wang 0006, Ziyan He, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 3 |
| 2025 | Weakly-Supervised 3D Visual Grounding Based on Visual Language AlignmentabstractLearning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual grounding approaches require a substantial number of bounding box annotations for text queries, which is time-consuming and labor-intensive to obtain. In this paper, we propose3D-VLA, a weakly supervised approach for3Dvisual grounding based onVisualLanguageAlignment. Our 3D-VLA exploits the superior ability of current large-scale vision-language models (VLMs) on aligning the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds with no need for fine-grained box annotations in the training procedure. During the inference stage, the learned text-3D correspondence will help us ground the text queries to the 3D target objects even without 2D images. To the best of our knowledge, this is the first work to investigate 3D visual grounding in a weakly supervised manner by involving large scale vision-language models, and extensive experiments on ReferIt3D and ScanRefer datasets demonstrate that our 3D-VLA achieves comparable and even superior results over the fully supervised methods. Xiaoxu Xu, Yitian Yuan, Qiudan Zhang, Wenhui Wu 0001, Zequn Jie, Lin Ma 0002, Xu Wang 0006 |
IEEE Trans. Multim. | 3 |
| 2025 | Hierarchical Uncertainty-Aware Salient Object Detection for $360 ^{\circ }$ Images via Bi-Projection Collaborative Learningabstract$360^{\circ }$salient object detection has recently received much attention for 3D scene perception owing to its omnidirectional field of view (FoV). The capability of recognizing salient objects of$360^{\circ }$images remains technically challenging due to severe spherical distortion. In this paper, we develop a hierarchical uncertainty-aware$360^{\circ }$image salient object detection methodology that explicitly explores the geometric and spatial complementary coherence of Tangent projection (TP) and Equirectangular projection (ERP) by a collaborative learning strategy. Concretely, to mitigate spherical distortion, we first intend to learn saliency-related features from less-distorted tangent images, in which a deformation-aware attention block is introduced to mitigate the geometric distortion caused by projecting a$360^{\circ }$image onto a 2D plane. However, the discrepancies among tangent images pose a new challenge to$360^{\circ }$image salient object detection. To tackle this issue and achieve accurate localization for salient objects of all sizes, we design a spatial-frequency saliency feature aggregation module to leverage fast Fourier convolution to capture global contextual information from ERP images, such that obtaining more representative saliency features. Moreover, a hierarchical uncertainty-aware bi-projection consistency learning module with strong local-global information embedding capabilities is constructed, which learns the geometric and spatial correlations between tangent images and ERP images via a collaborative learning strategy. Ultimately, salient object maps are produced for$360^{\circ }$images on the basis of the merged saliency features driven by the uncertainty. Extensive experiments show that our developed method improves${\mathrm{F}}_\beta ^{\sigma }$by an average of 31.67% compared to twenty existing advanced methods on the publicly available 360-SOD dataset. Qiudan Zhang, Kaiyu Ji, Xu Wang 0006, Zhaoqing Pan, Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2025 | RGB-D Data Compression via Bi-Directional Cross-Modal Prior Transfer and Enhanced Entropy ModelingabstractRGB-D data, being homogeneous cross-modal data, demonstrates significant correlations among data elements. However, current research focuses only on a uni-directional pattern of cross-modal contextual information, neglecting the exploration of bi-directional relationships in the compression field. Thus, we propose a joint RGB-D compression scheme, which is combined with Bi-Directional Cross-Modal Prior Transfer (Bi-CPT) modules and a Bi-Directional Cross-Modal Enhanced Entropy (Bi-CEE) model. The Bi-CPT module is designed for compact representations of cross-modal features, effectively eliminating spatial and modality redundancies at different granularity levels. In contrast to the traditional entropy models, our proposed Bi-CEE model not only achieves spatial-channel contextual adaptation through partitioning RGB and depth features but also incorporates information from other modalities as prior to enhance the accuracy of probability estimation for latent variables. Furthermore, this model enables parallel multi-stage processing to accelerate coding. Experimental results demonstrate the superiority of our proposed framework over the current compression scheme, outperforming both rate-distortion performance and downstream tasks, including surface reconstruction and semantic segmentation. The source code will be available at https://github.com/xyy7/Learning-based-RGB-D-Image-Compression . Yuyu Xu, Qiudan Zhang, Wenhui Wu 0001, Yun Zhang 0002, Xu Wang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | 3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
Xiaoxu Xu, Yitian Yuan, Jinlong Li 0003, Qiudan Zhang, Zequn Jie, Lin Ma 0002, Hao Tang 0005, Nicu Sebe, Xu Wang 0006 |
ECCV (73) | 4 |
| 2024 | LESS: Label-Efficient and Single-Stage Referring 3D SegmentationabstractReferring 3D Segmentation is a visual-language task that segments all points of the specified object from a 3D point cloud described by a sentence of query. Previous works perform a two-stage paradigm, first conducting language-agnostic instance segmentation then matching with given text query. However, the semantic concepts from text query and visual cues are separately interacted during the training, and both instance and semantic labels for each object are required, which is time consuming and human-labor intensive. To mitigate these issues, we propose a novel Referring 3D Segmentation pipeline, Label-Efficient and Single-Stage, dubbed LESS, which is only under the supervision of efficient binary mask. Specifically, we design a Point-Word Cross-Modal Alignment module for aligning the fine-grained features of points and textual embedding. Query Mask Predictor module and Query-Sentence Alignment module are introduced for coarse-grained alignment between masks and query. Furthermore, we propose an area regularization loss, which coarsely reduces irrelevant background predictions on a large scale. Besides, a point-to-point contrastive loss is proposed concentrating on distinguishing points with subtly similar features. Through extensive experiments, we achieve state-of-the-art performance on ScanRefer dataset by surpassing the previous methods about 3.7% mIoU using only binary labels. Code is available at https://github.com/mellody11/LESS. Xuexun Liu, Xiaoxu Xu, Jinlong Li 0003, Qiudan Zhang, Xu Wang 0006, Nicu Sebe, Lin Ma 0002 |
NeurIPS | 4 |
| 2024 | Multi-exposure embeddings for graph learning: Towards high dynamic range image saliency predictionabstractAbstract Identifying saliency in high dynamic range (HDR) images is a fundamentally important issue in HDR imaging, and plays critical roles towards comprehensive scene understanding. Most of existing studies leverage hand‐crafted features for HDR image saliency prediction, lacking the capabilities of fully exploiting the characteristics of HDR image (i.e. wider luminance range and richer colour gamut). Here, systematical studies are carried out on HDR image saliency prediction by proposing a new framework to single out the contributions from multi‐exposure images. Specifically, inspired by the mechanism of HDR imaging, the method first utilizes graph neural networks to model the relations among multi‐exposure images and the tone‐mapped image obtained from an HDR image, enabling more discriminative saliency‐related feature representations. Subsequently, the saliency features driven by global semantic knowledge are aggregated from the tone‐mapped image through enhancing global context‐aware semantic information. Finally, a fusion module is designed to integrate saliency‐oriented feature representations originated from multi‐exposure images and the tone‐mapped image, producing the saliency maps of HDR images. Moreover, a new challenging HDR eye fixation database (HDR‐EYEFix) is created, expecting to further contribute the research on HDR image saliency prediction. Experiment results show that the method obtains superior performance compared to the state‐of‐the‐art methods. Jun Xing, Qiudan Zhang, Xuelin Shen, Xu Wang 0006 |
IET Image Process. | 2 |
| 2024 | Towards 360$^{\circ }$ image compression for machines via modulating pixel significance
Silin Zheng, Xuelin Shen, Qiudan Zhang, Zhuo Chen 0006, Wenhan Yang, Xu Wang 0006 |
Multim. Tools Appl. | 3 |
| 2024 | Distortion-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Hybrid Projection Fusion and Structural RegularitiesabstractOwing to the rapid development of emerging 360$^{\circ }$panoramic imaging techniques, indoor 360$^{\circ }$depth estimation has aroused extensive attention in the community. Due to the lack of available ground truth depth data, it is extremely urgent to model indoor 360$^{\circ }$depth estimation in self-supervised mode. However, self-supervised 360$^{\circ }$depth estimation suffers from two major limitations. One is the distortion and network training problems caused by Equirectangular projection (ERP), and the other is that texture-less regions are quite difficult to back-propagate in self-supervised mode. Hence, to address the above issues, we introduce spherical view synthesis for learning self-supervised 360$^{\circ }$depth estimation. Specifically, to alleviate the ERP-related problems, we first propose a dual-branch distortion-aware network to produce the coarse depth map, including a distortion-aware module and a hybrid projection fusion module. Subsequently, the coarse depth map is utilized for spherical view synthesis, in which a spherically weighted loss function for view reconstruction and depth smoothing is investigated to optimize the projection distribution problem of 360$^{\circ }$images. In addition, two structural regularities of indoor 360$^{\circ }$scenes are devised as two additional supervisory signals to efficiently optimize our self-supervised 360$^{\circ }$depth estimation model, containing the principal-direction normal constraint and the co-planar depth constraint. The principal-direction normal constraint is designed to align the normal of the 360$^{\circ }$image with the direction of the vanishing points. Meanwhile, we employ the co-planar depth constraint to fit the estimated depth of each pixel through its 3D plane. Finally, a depth map is obtained for the 360$^{\circ }$image. Experimental results illustrate that our proposed method achieves superior performance than the current advanced depth estimation methods on four publicly available datasets. Xu Wang 0006, Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 3 |
| 2024 | Weakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-LabelingabstractLearning to build 3D scene graphs is essential for real-world perception in a structured and rich fashion. However, previous 3D scene graph generation methods utilize a fully supervised learning manner and require a large amount of entity-level annotation data of objects and relations, which is extremely resource-consuming and tedious to obtain. To tackle this problem, we propose 3D-VLAP, a weakly-supervised 3D scene graph generation method via Visual-Linguistic Assisted Pseudo-labeling. Specifically, our 3D-VLAP exploits the superior ability of current large-scale visual-linguistic models to align the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds. First, we establish the positional correspondence from 3D point clouds to 2D images via camera intrinsic and extrinsic parameters, thereby achieving alignment of 3D point clouds and 2D images. Subsequently, a large-scale cross-modal visual-linguistic model is employed to indirectly align 3D instances with the textual category labels of objects by matching 2D images with object category labels. The pseudo labels for objects and relations are then produced for 3D-VLAP model training by calculating the similarity between visual embeddings and textual category embeddings of objects and relations encoded by the visual-linguistic model, respectively. Ultimately, we design an edge self-attention based graph neural network to generate scene graphs of 3D point clouds. Experiments demonstrate that our 3D-VLAP achieves comparable results with current fully supervised methods, meanwhile alleviating the data annotation pressure. Xu Wang 0006, Qiudan Zhang, Wenhui Wu 0001, Mark Junjie Li, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 3 |
| 2023 | Hybrid Prior-Based Diminished Reality for Indoor Panoramic Images
Jiashu Liu, Qiudan Zhang, Xuelin Shen, Wenhui Wu 0001, Xu Wang 0006 |
CGI (3) | 2 |
| 2023 | Salient Object Detection on 360° Omnidirectional Image with Bi-Branch Hybrid Projection NetworkabstractWith the advent of panoramic cameras, modeling saliency in 360° omnidirectional images becomes very urgent and challenging. However, severe distortions limit the prediction accuracy of 360° saliency model. In this paper, we devise a bi-branch hybrid projection network (HPNet), which exploits characteristics of equirectangular projection (ERP) and cubic map projection (CMP) formats to predict salient objects in 360° omnidirectional images. Specifically, an ERP image and a CMP image are first fed into a bi-branch network to aggregate the comprehensive features of the omnidirectional image. Subsequently, to explore the coherence among ERP and CMP images, we design a hybrid projection feature fusion module to efficiently combine CMP and ERP features extracted from different layers. Ultimately, a progressive prediction module is developed to refine the features and locate salient objects incrementally, and then produce the final saliency map for the 360° omnidirectional image. Experimental results illustrate that our model is superior to the existing advanced methods in two publicly available datasets. Qiudan Zhang, Xuelin Shen, Xu Wang 0006 |
MMSP | 2 |
| 2022 | A Learning-based Framework for Multi-View Instance Segmentation in PanoramaabstractThe application of panoramas in computer vision has received a lot of attention due to its ability to represent information about the surrounding environment. Instance segmentation on panoramas can make machines better understand 3D scenes. However, there are few efforts have been made on detecting the instance for panoramas. The main challenge is that in a panorama, objects are subject to three variations: geometric distortion, edge discontinuity and minification of objects. To address the above issues, we propose an instance segmentation method for panoramas based upon the multi-view fusion. First, a set of sub-views are sampled from a panorama by a rectilinear projection, and then the Cascade Mask R-CNN is employed to perform instance segmentation on each sub-view. Subsequently, we reverse mapping the obtained instance segmentation results of sub-views back to the panorama. Finally, we combine the complementary advantages of large and small fields of view, and merge the instance segmentation results of the entire panorama with the reorganized instance segmentation results to obtain high-quality instance segmentation results for panoramas. The quantitative experiments illustrate that our proposed method obtains higher Mean Average Precision (0.28) than existing benchmark methods. Weihao Ye, Ziyang Mai, Qiudan Zhang, Xu Wang 0006 |
DSAA | 3 |
| 2022 | Self-supervised Indoor 360-Degree Depth Estimation via Structural Regularization
Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Wenhui Wu 0001, Xu Wang 0006 |
PRICAI (3) | 2 |
| 2022 | Deep stereoscopic image saliency inspired stereoscopic image thumbnail generation
Yu Zhou 0027, Xiaotong Xiao, Qiudan Zhang, Xu Wang 0006, Jianmin Jiang |
Multim. Tools Appl. | 3 |
| 2022 | Adaptive Viewpoint Feature Enhancement-Based Binocular Stereoscopic Image Saliency DetectionabstractModeling 3D visual saliency has received great attention due to the development of emerging 3D display technologies. Traditional methods relying on low-level features may not be efficient in interpreting 3D visual content from high-level semantic perspective. Despite numerous efforts dedicated to this area, existing 3D visual saliency detection methods do not necessarily excel in exploring the stereoscopic image saliency driven by the intra-view and inter-view dependencies among left and right views. In this paper, we propose a visual saliency detection method for stereoscopic images grounded on adaptive viewpoint feature enhancement via binocular vision. More specifically, the correlation among left and right views is investigated through a delicately designed binocular stereoscopic saliency feature aggregation module, enabling the generation of more representative saliency features towards binocular vision. Subsequently, to further aggregate the saliency features in multiple scales, we design a progressive attention-based saliency feature pyramid extraction module to effectively integrate the features from top-level to down-level based on the network hierarchy mechanism. The saliency maps are ultimately produced for stereoscopic images by evaluating the obtained saliency features. In addition, we create a stereoscopic image saliency dataset (SIS-3D) that includes 1086 stereoscopic image pairs with various content and their corresponding human eye fixation annotations, aiming to further facilitate the research on visual saliency detection for stereoscopic images. Extensive experiments demonstrate that our proposed method improves CC by an average of 4.02% compared to representative counterparts on the newly built saliency dataset and another publicly available dataset. Qiudan Zhang, Xiaotong Xiao, Xu Wang 0006, Shiqi Wang 0001, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | A Multi-Task Collaborative Network for Light Field Salient Object DetectionabstractBeing able to predict the salient object is of fundamental importance in image processing and computer vision. With numerous approaches proposed for automatic image and video salient object detection, much less work has been dedicated to detecting and segmenting salient objects from light fields. In this article, based on the intrinsic characteristics of light fields, we carefully explore the complementary coherence among multiple cues including spatial, edge and depth information, and elaborately design a multi-task collaborative network for light field salient object detection. More specifically, the correlation mechanisms among edge detection, depth inference and salient object detection are carefully investigated to facilitate the representative saliency features. We first model the coherence among low-level features and heuristic semantic priors, as well as the edge information. Subsequently, the depth-oriented saliency features are derived from the geometry of light fields, in which the 3D convolution operation is leveraged with powerful representation capability to model the disparity correlations among multiple viewpoint images. Finally, a feature-enhanced salient object generator is developed to integrate these complementary saliency features, leading to the final salient object predictions for light fields. Quantitative and qualitative experiments demonstrate the superiority of our proposed model against the state-of-the-art methods over the public light field salient object detection datasets. Qiudan Zhang, Shiqi Wang 0001, Xu Wang 0006, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Geometry Auxiliary Salient Object Detection for Light Fields via Graph Neural NetworksabstractLight field imaging, originated from the availability of light field capture technology, offers a wide range of applications in the field of computational vision. The capability of predicting salient objects of light fields remains technologically challenging due to its complicated geometry structure. In this paper, we propose a light field salient object detection approach that formulates the geometric coherence among multiple views of light fields as graphs, where the angular/central views represent the nodes and their relations compose the edges. The spatial and disparity correlations between multiple views are effectively explored through multi-scale graph neural networks, enabling the more comprehensive understanding of light field content and more representative and discriminative saliency features generation. Moreover, a multi-scale saliency feature consistency learning module is embedded to enhance the saliency features. Finally, an accurate salient object map is produced for the light field based upon the extracted features. In addition, we establish a new light field salient object detection dataset (CITYU-Lytro) that contains 817 light fields with diverse contents and their corresponding annotations, aiming to further promote the research on light field salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably compared with the state-of-the-art methods on the benchmark datasets. Qiudan Zhang, Shiqi Wang 0001, Xu Wang 0006, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Image Process. | 1 |
| 2020 | Multi-Exposure Decomposition-Fusion Model for High Dynamic Range Image Saliency DetectionabstractHigh dynamic range (HDR) imaging techniques have witnessed a great improvement in the past few decades. However, saliency detection task on HDR content is still far from well explored. In this paper, we introduce a multi-exposure decomposition-fusion model for HDR image saliency detection inspired by the brightness adaption mechanism. The proposed model is composed of three modules. Firstly, a decomposition module converts the input raw HDR image into a stack of LDR images by uniformly sampling the exposure time range. Secondly, a saliency region proposal network is employed to generate the candidate saliency maps for each LDR image in the exposure stack. Finally, an uncertainty weighting based fusion algorithm is applied to generate the overall saliency map for the input HDR image by merging the obtained LDR saliency maps. Extensive experiments show that our proposed model achieves superior performance compared with the state-of-the-art methods on the existing HDR eye fixation databases. The source code of the proposed model are made publicly available at https://github.com/sunnycia/DFHSal. Xu Wang 0006, Zhenhao Sun, Qiudan Zhang, Yuming Fang 0001, Lin Ma 0002, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Learning to Explore Saliency for Stereoscopic Videos Via Component-Based InteractionabstractIn this paper, we devise a saliency prediction model for stereoscopic videos that learns to explore saliency inspired by the component-based interactions including spatial, temporal, as well as depth cues. The model first takes advantage of specific structure of 3D residual network (3D-ResNet) to model the saliency driven by spatio-temporal coherence from consecutive frames. Subsequently, the saliency inferred by implicit-depth is automatically derived based on the displacement correlation between left and right views by leveraging a deep convolutional network (ConvNet). Finally, a component-wise refinement network is devised to produce final saliency maps over time by aggregating saliency distributions obtained from multiple components. In order to further facilitate research towards stereoscopic video saliency, we create a new dataset including 175 stereoscopic video sequences with diverse content, as well as their dense eye fixation annotations. Extensive experiments support that our proposed model can achieve superior performance compared to the state-of-the-art methods on all publicly available eye fixation datasets. Qiudan Zhang, Xu Wang 0006, Shiqi Wang 0001, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Image Process. | 1 |
| 2019 | Learning to Explore Intrinsic Saliency for Stereoscopic VideoabstractThe human visual system excels at biasing the stereoscopic visual signals by the attention mechanisms. Traditional methods relying on the low-level features and depth relevant information for stereoscopic video saliency prediction have fundamental limitations. For example, it is cumbersome to model the interactions between multiple visual cues including spatial, temporal, and depth information as a result of the sophistication. In this paper, we argue that the high-level features are crucial and resort to the deep learning framework to learn the saliency map of stereoscopic videos. Driven by spatio-temporal coherence from consecutive frames, the model first imitates the mechanism of saliency by taking advantage of the 3D convolutional neural network. Subsequently, the saliency originated from the intrinsic depth is derived based on the correlations between left and right views in a data-driven manner. Finally, a Convolutional Long Short-Term Memory (Conv-LSTM) based fusion network is developed to model the instantaneous interactions between spatio-temporal and depth attributes, such that the ultimate stereoscopic saliency maps over time are produced. Moreover, we establish a new large-scale stereoscopic video saliency dataset (SVS) including 175 stereoscopic video sequences and their fixation density annotations, aiming to comprehensively study the intrinsic attributes for stereoscopic video saliency detection. Extensive experiments show that our proposed model can achieve superior performance compared to the state-of-the-art methods on the newly built dataset for stereoscopic videos. Qiudan Zhang, Xu Wang 0006, Shiqi Wang 0001, Shikai Li, Sam Kwong, Jianmin Jiang |
CVPR | 1 |