EDBT 2026 Demo / reviewers in the wild / expert
Anjie Wang
dblp:221/9646
· DBLP profile ↗
16ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Entity completion for industrial knowledge graph based on zero-shot learning
Yin Cai, Zhijun Fang 0001, Anjie Wang, Zheyi Cheng |
Data Min. Knowl. Discov. | 5 |
| 2026 | Learning Monocular Depth via Cascaded Iterative Refinement in Visual-Echo ScenesabstractIn recent years, integrating multimodal information, particularly visual and echo data, has shown great promise for improving depth estimation performance. While existing works demonstrate that combining binaural echo features with image attributes can enhance depth estimation, they often use rudimentary feature alignment and fusion methods, failing to fully exploit the complementary nature of cross-modal information and limiting integration effectiveness. To address these challenges, this paper introduces an innovative multimodal fusion framework. First, the framework incorporates a combination of multi-scale self-attention and cross-attention mechanisms, establishing correlations between features and facilitating cohesive interactions between the visual and echo domains. Furthermore, we propose an incremental feature updating mechanism based on Convolutional Gated Recurrent Units (ConvGRU), which implements cascaded iterative optimization, integrating contextual features with the multi-scale fused features from both echo and image modalities. In each iteration, the framework preserves contextual information from previous steps while employing a multi-level loss function to guide result updates. This approach effectively captures spatial structural information and progressively enhances depth estimation accuracy. Comprehensive experimental evaluations on the Replica, Matterport3D and BatVision (BV1) datasets validate the effectiveness of the proposed method. Comparative analyses with state-of-the-art monocular plus echo methods underscore the superior performance achievable through this novel framework. Anjie Wang, Zhijun Fang 0001, Leidong Fan, Guibiao Liao, Siwei Ma 0001, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | A Unified Inverse-Tone-Mapped HDR Video Quality Assessment Method across Two HDR FormatsabstractHigh Dynamic Range Video Quality Assessment (HDR VQA) plays a pivotal role in Inverse Tone Mapping (ITM) research. Existing HDR VQA datasets and models mainly focus on a single HDR format, leading to poor generalization and limited application scope. To address this problem, this paper proposes a reference-free CONTrastive ITM-VQA (CONT-ITM-VQA) model via format transformation-based data augmentation and contrastive learning. Specifically, the format transformation-based data augmentation improves the model generalization, via applying Opto-electronic Transfer Function (OETF) transformations between the HDR formats; while the contrastive learning-based quality-related feature alignment aligns the quality features from different HDR formats of the same video to obtain more effective quality representations. It is worth noting that our method can be extended to other HDR-related quality assessment, not limited to ITM-HDR VQA. Experimental results demonstrate that our model closely mimics subjective judgments. Leidong Fan, Xiongkuo Min, Qing Li 0029, Anjie Wang |
ICME | 4 |
| 2025 | Sparse-view 3D Open-vocabulary Gaussian Splatting via Collaborative Contrastive Learningabstract3D Gaussian Splatting-based Open-vocabulary 3D segmentation has shown impressive performance with dense input images. However, existing methods exhibit poor results when confronted with sparse inputs, primarily due to limited overlap among input views and insufficient view supervision provided. To tackle these challenges, we propose SpContrast, a novel framework that creates additional semantic constraints to enhance sparse-view 3D open-vocabulary segmentation. First, we introduce Collaborative Contrastive Learning (CCL), which creates instructive multi-view semantic constraints by collaboratively mining semantic interactions between training and online-rendered novel views. Motivated by the principle that semantically consistent features should converge and divergent ones separate, CCL establishes cross-view contrastive constraints to enhance semantic coherence. Second, to alleviate the adverse impact of false negative samples caused by semantic inconsistencies within the same object, we present Region-aware Negative Sampling (RNS). RNS rectifies these false negative samples, and treats them as hard samples during our contrastive optimization, leading to improved object completeness and more accurate segmentation. Extensive experiments on challenging sparse-input datasets, including Replica and ScanNet, demonstrate the superiority of SpContrast, achieving 7.6% and 8.3% mIoU improvements for 3D open-vocabulary segmentation. Guibiao Liao, Anjie Wang, Mingxuan Chen, Zhijun Fang 0001 |
ICME | 2 |
| 2025 | MSCC-RetNet: a multi-scale color corrected retinex network for underwater image enhancement
Benxue Sun, Mingxuan Chen, Liming Hu, Anjie Wang, Zhijun Fang 0001 |
Multim. Syst. | 4 |
| 2025 | Self-Attention Sliding Window Enhanced Canonical Correlation Analysis for Incipient Fault Detection in Dynamic Industrial ProcessesabstractThanks to its excellent nonlinear representation ability, deep neural network (DNN) has been designed to aid canonical correlation analysis (CCA) for stable kernel representation. These DNN-aided CCA methods make use of correlation as the optimization objective and exploit its change to distinguish system states. However, the maximum correlation-based training strategy lacks robustness to effectively tackle system false alarms caused by changes in operating conditions. To this end, this work proposes a self-attention sliding window enhanced CCA (SaECCA) for incipient fault detection in dynamic industrial processes. The main novelties of this work include the following: first, with the aid of DNNs, an adaptive weighting mechanism with self-attention is developed to amplify the incipient fault information; second, residual estimation-based, a novel and robust optimization objective for nonlinear CCA is formulated; third, an SaECCA-based fault detection algorithm is designed, whose convergence and detectability are illustrated via theoretical analysis. Studies on a three-tank system simulation and a multiphase flow industrial process are presented to verify the effectiveness of the proposed SaECCA method. Anjie Wang, Guang Wang 0002, Jianfang Jiao, Shen Yin |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Multi-modal Scene Global Fusion Framework for Enhanced Depth Estimation
Anjie Wang, Xujun Wei, Mingxuan Chen, Yongbin Gao, Zhijun Fang 0001, Siwei Ma 0001 |
ICONIP (9) | 1 |
| 2024 | Dual-stream multi-label image classification model enhanced by feature reconstruction
Liming Hu, Mingxuan Chen, Anjie Wang, Zhijun Fang 0001 |
Multim. Syst. | 3 |
| 2023 | Depth Estimation of Multi-Modal Scene Based on Multi-Scale ModulationabstractAs multimodal information is complementary, effectively utilizing scene multimodal information has become an increasingly important research topic for many scholars. This paper proposes a novel multi-scale global learning strategy that utilizes both echo and visual modal data as inputs to estimate scene depth. The framework involves constructing a multi-scale feature extraction method using pyramid pooling modules to aggregate contextual information from different regions and improve global information acquisition ability. Furthermore, a recurrent multi-scale feature modulation module is introduced to generate more semantic and accurate spatial representations in each iteration update process. Additionally, a multi-scale fusion method is constructed for the fusion of echo and visual modalities. The proposed method's superior performance is demonstrated through sufficient experiments conducted on the Replica dataset. Anjie Wang, Zhijun Fang 0001, Yongbin Gao, Gaofeng Cao, Siwei Ma 0001 |
ICIP | 1 |
| 2023 | A Decoupled Kernel Prediction Network Guided by Soft Mask for Single Image HDR ReconstructionabstractRecent works on single image high dynamic range (HDR) reconstruction fail to hallucinate plausible textures, resulting in information missing and artifacts in large-scale under/over-exposed regions. In this article, a decoupled kernel prediction network is proposed to infer an HDR image from a low dynamic range (LDR) image. Specifically, we first adopt a simple module to generate a preliminary result, which can precisely estimate well-exposed HDR regions. Meanwhile, an encoder-decoder backbone network with a soft mask guidance module is presented to predict pixel-wise kernels, which is further convolved with the preliminary result to obtain the final HDR output. Instead of traditional kernels, our predicted kernels are decoupled along the spatial and channel dimensions. The advantages of our method are threefold at least. First, our model is guided by the soft mask so that it can focus on the most relevant information for under/over-exposed regions. Second, pixel-wise kernels are able to adaptively solve the different degradations for differently exposed regions. Third, decoupled kernels can avoid information redundancy across channels and reduce the solution space of our model. Thus, our method is able to hallucinate fine details in the under/over-exposed regions and renders visually pleasing results. Extensive experiments demonstrate that our model outperforms state-of-the-art ones. Gaofeng Cao, Fei Zhou 0001, Kanglin Liu, Anjie Wang, Leidong Fan |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | KPN-MFI: A Kernel Prediction Network with Multi-frame Interaction for Video Inverse Tone MappingabstractUp to now, the image-based inverse tone mapping (iTM) models have been widely investigated, while there is little research on video-based iTM methods. It would be interesting to make use of these existing image-based models in the video iTM task. However, directly transferring the imagebased iTM models to video data without modeling spatial-temporal information remains nontrivial and challenging. Considering both the intra-frame quality and the inter-frame consistency of a video, this article presents a new video iTM method based on a kernel prediction network (KPN), which takes advantage of multi-frame interaction (MFI) module to capture temporal-spatial information for video data. Specifically, a basic encoder-decoder KPN, essentially designed for image iTM, is trained to guarantee the mapping quality within each frame. More importantly, the MFI module is incorporated to capture temporal-spatial context information and preserve the inter-frame consistency by exploiting the correction between adjacent frames. Notably, we can readily extend any existing image iTM models to video iTM ones by involving the proposed MFI module. Furthermore, we propose an inter-frame brightness consistency loss function based on the Gaussian pyramid to reduce the video temporal inconsistency. Extensive experiments demonstrate that our model outperforms state-ofthe-art image and video-based methods. The code is available at https://github.com/caogaofeng/KPNMFI. Gaofeng Cao, Fei Zhou 0001, Han Yan 0003, Anjie Wang, Leidong Fan |
IJCAI | 4 |
| 2021 | Point AE-DCGAN: A deep learning model for 3D point cloud lossy geometry compressionabstract3D point cloud has been widely applied in virtual reality and augmented reality. A complex 3D scene always needs a large number of the point cloud to represent and demands a lot of space to store. Thus, point cloud compression becomes a crucial issue to research. In this paper, we propose a novel lossy geometric compression method of autoencoder based on DCGAN optimization. This method can reconstruct a high-quality point cloud and solves a large area of missing points in the process of compression and decompression. To improve the point cloud codec performance, we propose a multi-scale 3D deconvolution hopping connection structure to obtain a better-quality reconstructed point cloud under low bit rates. Our approach is the first GAN-based point cloud compression algorithm to our knowledge. Compared with state-of-the-art methods on the MVUB dataset, our approach achieves a better rate-distortion performance and visual quality. Zhijun Fang 0001, Yongbin Gao, Siwei Ma 0001, Yaochu Jin, Anjie Wang |
DCC | 7 |
| 2020 | Self-Supervised Learning of Depth and Pose Using Cycle Generative Adversarial NetworkabstractIn recent years, large amount of ground truth data is typically required to feed into the supervised depth estimation models to produce satisfactory performance, while it is usually costly and impracticable to acquire depth ground truth. In this paper, we propose a self-supervised joint deep learning pipeline for depth and pose estimation of monocular video sequences, which uses the cycle generation adversarial network structure to extend the existing reconstruction loss function based on photometric consistency. The generation function of the algorithm learns to synthesize the adjacent image to predict the depth map and the relative target pose, and the discriminant function learns the dispersion of the monocular images to correctly classify the realism of the composite image. At the same time, a reconstruction loss function based on pose consistency is used to assist the generator function in training. Extensive experimental results on the KITTI dataset show superior performance of the proposed method. Yunhe Tong, Anjie Wang, Songchao Tan, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 2 |
| 2020 | Adversarial Learning for Joint Optimization of Depth and Ego-MotionabstractIn recent years, supervised deep learning methods have shown a great promise in dense depth estimation. However, massive high-quality training data are expensive and impractical to acquire. Alternatively, self-supervised learning-based depth estimators can learn the latent transformation from monocular or binocular video sequences by minimizing the photometric warp error between consecutive frames, but they suffer from the scale ambiguity problem or have difficulty in estimating precise pose changes between frames. In this paper, we propose a joint self-supervised deep learning pipeline for depth and ego-motion estimation by employing the advantages of adversarial learning and joint optimization with spatial-temporal geometrical constraints. The stereo reconstruction error provides the spatial geometric constraint to estimate the absolute scale depth. Meanwhile, the depth map with an absolute scale and a pre-trained pose network serves as a good starting point for direct visual odometry (DVO). DVO optimization based on spatial geometric constraints can result in a fine-grained ego-motion estimation with the additional backpropagation signals provided to the depth estimation network. Finally, the spatial and temporal domain-based reconstructed views are concatenated, and the iterative coupling optimization process is implemented in combination with the adversarial learning for accurate depth and precise ego-motion estimation. The experimental results show superior performance compared with state-of-the-art methods for monocular depth and ego-motion estimation on the KITTI dataset and a great generalization ability of the proposed approach. Anjie Wang, Zhijun Fang 0001, Yongbin Gao, Songchao Tan, Shanshe Wang, Siwei Ma 0001, Jenq-Neng Hwang |
IEEE Trans. Image Process. | 1 |
| 2019 | Unsupervised Learning of Depth and Ego-Motion with Spatial-Temporal Geometric ConstraintsabstractIn this paper, we propose an unsupervised joint deep learning pipeline for depth and ego-motion estimation that explicitly incorporated with traditional spatial-temporal geometric constraints. The stereo reconstruction error provides the spatial geometric constraint to estimate the absolute scale depth. Meanwhile, the depth map with absolute scale and a pre-trained pose network serve as a good starting point for direct visual odometry (DVO), resulting in a fine-grained ego-motion estimation with the additional back-propagation signals provided to the depth estimation network. The proposed joint training pipeline enables an iterative coupling optimization process for accurate depth and precise ego-motion estimation. The experimental results show the state-of-the-art performance for monocular depth and ego-motion estimation on the KITTI dataset and a great generalization ability of the proposed approach. Anjie Wang, Yongbin Gao, Zhijun Fang 0001, Shanshe Wang, Siwei Ma 0001, Jenq-Neng Hwang |
ICME | 1 |
| 2019 | Unsupervised learning of depth estimation based on attention model and global pose optimization
Renyue Dai, Yongbin Gao, Zhijun Fang 0001, Anjie Wang, Juan Zhang 0001, Cengsi Zhong |
Signal Process. Image Commun. | 5 |