VLDB 2026 Research / reviewers in the wild / expert
Zhiheng Fu
dblp:214/9102
· DBLP profile ↗
26ranked-venue papers
8as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made in recent years, existing methods assume that all samples are correctly matched. However, in real-world scenarios, due to high triplet annotation costs, CIR datasets inevitably contain annotation errors, resulting in incorrectly matched triplets. To address this issue, the problem of Noisy Triplet Correspondence (NTC) has attracted growing attention. We argue that noise in CIR can be categorized into two types: cross-modal correspondence noise and modality-inherent noise. The former arises from mismatches across modalities, whereas the latter originates from intra-modal background interference or visual factors irrelevant to the coarse-grained modification annotations. However, modality-inherent noise is often overlooked, and research on cross-modal correspondence noise remains nascent. To tackle above issues, we propose the Invariance and discrimiNaTion-awarE Noise neTwork (INTENT), comprising two components: Visual Invariant Composition and Bi-Objective Discriminative Learning, specifically designed to handle the two-aspect noise. The former applies causal intervention on the visual side via Fast Fourier Transform (FFT) to generate intervened composed features, enforcing visual invariance and enabling the model to ignore modality-inherent noise during composition. The latter adopts collaborative optimization with both positive and negative samples, and constructs a scalable decision boundary that dynamically adjusts decisions based on the loyalty degree, enabling robust correspondence discrimination. Extensive experiments on two widely used benchmark datasets demonstrate the superiority and robustness of INTENT. Zhiwei Chen 0003, Yupeng Hu 0003, Zhiheng Fu, Zixu Li 0001, Qinlei Huang, Yinwei Wei |
AAAI | 3 |
| 2026 | ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalabstractWith the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing attention from researchers. Unlike unimodal video retrieval methods, the CVR task takes a multi-modal query consisting of a reference video and a piece of modification text as input. The modification text conveys the user's intended alterations to the reference video. Based on this input, the model aims to retrieve the most relevant target video. In the CVR task, there exists a substantial discrepancy in information density between video and text modalities. Traditional composition methods tend to bias the composed feature toward the reference video, which leads to suboptimal retrieval performance. This limitation is significant due to the presence of three core challenges: (1) modal contribution entanglement, (2) explicit optimization of composed features, and (3) retrieval uncertainty. To address these challenges, we propose the evidence-dRivEn dual-sTream diRectionAl anChor calibration networK (ReTrack). ReTrack is the first CVR framework that improves multi-modal query understanding by calibrating directional bias in composed features. It consists of three key modules: Semantic Contribution Disentanglement, Composition Geometry Calibration, and Reliable Evidence-driven Alignment. Specifically, ReTrack estimates the semantic contribution of each modality to calibrate the directional bias of the composed feature. It then uses the calibrated directional anchors to compute bidirectional evidence that drives reliable composed-to-target similarity estimation. Moreover, ReTrack exhibits strong generalization to the Composed Image Retrieval (CIR) task, achieving SOTA performance across three benchmark datasets in both CVR and CIR scenarios. Zixu Li 0001, Yupeng Hu 0003, Zhiwei Chen 0003, Qinlei Huang, Guozhi Qiu, Zhiheng Fu, Meng Liu 0006 |
AAAI | 6 |
| 2026 | HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has demonstrated promising applications in personalized search and recommendation systems, it encounters a severe challenge in practical scenarios known as the Noise Triplet Correspondence (NTC) problem. This issue primarily arises from the high cost and subjectivity involved in annotating triplet data. To address this problem, we identify two central challenges: the precise estimation of composed semantic discrepancy and the insufficient progressive adaptation to modification discrepancy. To tackle these challenges, we propose a cHrono-synergiA roBust progressIve learning framework for composed image reTrieval (HABIT), which consists of two core modules. First, the Mutual Knowledge Estimation Module quantifies sample cleanliness by calculating the Transition Rate of mutual information between the composed feature and the target image, thereby effectively identifying clean samples that align with the intended modification semantics. Second, the Dual-consistency Progressive Learning Module introduces a collaborative mechanism between the historical and current models, simulating human habit formation to retain good habits and calibrate bad habits, ultimately enabling robust learning under the presence of NTC. Extensive experiments conducted on two standard CIR datasets demonstrate that HABIT significantly outperforms most methods under various noise ratios, exhibiting superior robustness and retrieval performance. Zixu Li 0001, Yupeng Hu 0003, Zhiwei Chen 0003, Qinlei Huang, Zhiheng Fu, Yinwei Wei |
AAAI | 6 |
| 2026 | TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image RetrievalabstractComposed Image Retrieval (CIR) is an important image retrieval paradigm that enables users to retrieve a target image using a multimodal query that consists of a reference image and modification text. Although research on CIR has made significant progress, prevailing setups still rely simple modification texts that typically cover only a limited range of salient changes, which induces two limitations highly relevant to practical applications, namely Insufficient Entity Coverage and Clause-Entity Misalignment. In order to address these issues and bring CIR closer to real-world use, we construct two instruction-rich multi-modification datasets, M-FashionIQ and M-CIRR. In addition, we propose TEMA, the Text-oriented Entity Mapping Architecture, which is the first CIR framework designed for multi-modification while also accommodating simple modifications. Extensive experiments on four benchmark datasets demonstrate that TEMA’s superiority in both original and multi-modification scenarios, while maintaining an optimal balance between retrieval accuracy and computational efficiency. Our codes and constructed multi-modification dataset (M-FashionIQ and M-CIRR) are available at https://github.com/lee-zixu/ACL26-TEMA/ Zixu Li 0001, Yupeng Hu 0003, Zhiheng Fu, Zhiwei Chen 0003, Yongqi Li 0001, Liqiang Nie |
ACL (1) | 3 |
| 2026 | VLGDiff: A Vision-Language Guided Diffusion Model for Image Inpainting
Shuling Zheng, Zhan Li 0004, Zhanglu Chen, Yinglue Yang, Zhiheng Fu, Feng Zhang 0047 |
ICIC (21) | 5 |
| 2026 | IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video RetrievalabstractComposed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal correspondences, they often assume modified objects appear directly in videos. However, modification texts frequently describe concepts not explicitly presented but implicitly expressed through semantically related visual cues (e.g., “cake” implying “birthday party”). Current approaches typically rely on aligning explicit feature representations within the concrete space, neglecting critical latent associations. To address this, we propose an adaptIve scheMa-ImAGery enhanced composItional NEtwork (IMAGINE). Unlike standard explicit matching, IMAGINE materializes implicit semantics (termed schema imagery) via dynamic multimodal prototypes. These prototypes capture shared latent concepts to adaptively modulate visual features, effectively injecting implicit guidance into the retrieval process. By bridging the gap between explicit visual contents and implicit retrieval intentions, IMAGINE achieves state-of-the-art performance in both CVR and Composed Image Retrieval (CIR) across three widely used benchmarks. Zixu Li 0001, Zhiwei Chen 0003, Zhiheng Fu, Yupeng Hu 0003 |
ICMR | 4 |
| 2026 | RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image RetrievalabstractComposed Image Retrieval (CIR) constitutes a pivotal paradigm requiring models to perform joint reasoning on reference images and modification texts. However, the prevalence of Noisy Triplet Correspondence (NTC) in large-scale datasets severely constrains model performance. Existing denoising methods either target binary mismatches or rely on scalar-based point-wise estimation, neglecting rich global structural correlations among sample populations and dynamic value variations during training, thereby yielding suboptimal results. This paper identifies two critical unresolved challenges: Global Structural Inconsistency of Semantic Correlations and Hard Sample Discrimination Uncertainty. To address these, we propose RankVR, a framework designed to construct a robust CIR model via global structure consistency and dynamic value perception. Specifically, we introduce the Global Structure Consistency Perception (GSCP) module, which utilizes the Effective Rank of the Correlation Matrix to decouple clean samples from structural noise. By measuring rank difference, GSCP identifies samples disrupting macroscopic semantic symmetry. Furthermore, we develop the Adaptive Semantic Value Calibration (ASVC) module to distinguish high-value hard clean samples. By integrating training potential and reliability, it dynamically quantifies the semantic value of each triplet, ensuring effective utilization of hard samples while suppressing noise characterized by logical conflicts. Extensive experiments on the FashionIQ and CIRR benchmark datasets demonstrate that RankVR significantly outperforms existing state-of-the-art methods, validating its superior robustness in noisy environments. Zixu Li 0001, Zhiheng Fu, Zhiwei Chen 0003, Qinlei Huang, Yupeng Hu 0003 |
ICMR | 3 |
| 2026 | STABLE: Efficient Hybrid Nearest Neighbor Search via Magnitude-Uniformity and Cardinality-RobustnessabstractHybrid Approximate Nearest Neighbor Search (Hybrid ANNS) is a foundational search technology for large-scale heterogeneous data and has gained significant attention in both academia and industry. However, current approaches overlook the heterogeneity in data distribution, thus ignoring two major challenges: the Compatibility Barrier for Similarity Magnitude Heterogeneity and the Tolerance Bottleneck to Attribute Cardinality. To overcome these issues, we propose the robuSt he Terogeneity-Aware hyBrid retrievaL framEwork, STABLE, designed for accurate, efficient, and robust hybrid ANNS under datasets with various distributions. Specifically, we introduce an enhAnced heterogeneoUs semanTic perceptiOn (AUTO) metric to achieve a joint measurement of feature similarity and attribute consistency, addressing similarity magnitude heterogeneity and improving robustness to datasets with various attribute cardinalities. Thereafter, we construct our Heterogeneous Emanticre Lation graPh (HELP) index based on AUTO to organize heterogeneous semantic relations. Finally, we employ a novel Dynamic Heterogeneity Routing method to ensure an efficient search. Extensive experiments on five feature vector benchmarks with various attribute cardinalities demonstrate the superior performance of STABLE. Qianyun Yang, Zhiwei Chen 0003, Yupeng Hu 0003, Zixu Li 0001, Zhiheng Fu, Liqiang Nie |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2026 | REFINE: Composed Video Retrieval via Shared and Differential Semantics EnhancementabstractComposed Video Retrieval (CVR) is a novel video retrieval paradigm. Unlike traditional single-modal video retrieval paradigms (e.g., text to video or video to video), CVR employs multi-modal queries (including both a reference video and a natural language modification) to retrieve the target video that best matches the modified reference video. Existing CVR methods primarily rely on generalized knowledge from vision-language pretrained models or utilize caption expansions to enhance video comprehension. However, these approaches overlook the benefits offered by the shareability and variability of videos for multi-modal query understanding. To overcome this limitation, we introduce a novel CVR framework named shaREd and diFferential semantIcs eNhancement nEtwork ( REFINE ). REFINE is the first framework to exploit the shareability and variability of videos to improve multi-modal query comprehension. Specifically, REFINE leverages learnable tokens to achieve enhanced shared feature representation. Moreover, it introduces a carefully designed Differential Block to disentangle differential semantics between frames and employs modification associations to guide multi-modal query feature fusion. Additionally, REFINE has been extended to the Composed Image Retrieval task, making it effectively generalize across existing composed multi-modal retrieval scenarios and outperform existing methods. Extensive qualitative and quantitative evaluations on four benchmark datasets validate the superiority of the proposed REFINE framework. Yupeng Hu 0003, Zixu Li 0001, Zhiwei Chen 0003, Qinlei Huang, Zhiheng Fu, Liqiang Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalabstractThe objective of Composed Image Retrieval (CIR) is to identify a target image that meets the requirement based on a multimodal query (including the reference image and the modification text) provided by the user. Despite the notable success of existing approaches, they fail to adequately address the modification relation between visual entities and modification actions. This limitation is non-trivial due to three challenges: 1) irrelevant factor perturbation, 2) vague semantic boundaries, and 3) implicit modification relations. To address the above challenges, we propose an Entity miNing and modifiCation relatiOn binDing nEtwoRk (ENCODER), which has been designed to mine visual entities and modification actions, and then bind modification relations. Among the various components of the proposed ENCODER, we have initially designed the Latent Factor Filter (LFF) module to filter visual and textual latent factors related to modification semantics based on a threshold gating mechanism. Secondly, we propose Entity-Action Binding (EAB), which comprises modality-shared Learnable Relation Queries (LRQ) that are capable of mining visual entities and modification actions, as well as learning implicit modification relations for entity-action binding. Finally, the Multi-scale Composition module is introduced to achieve multi-scale feature composition, with guidance provided by entity-action binding. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed method. Zixu Li 0001, Zhiwei Chen 0003, Haokun Wen, Zhiheng Fu, Yupeng Hu 0003, Weili Guan |
AAAI | 4 |
| 2025 | PAIR: Complementarity-guided Disentanglement for Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a novel image retrieval paradigm that aims at searching for the target images via the multimodal query including a reference image and a modification text. Although existing works have made significant progress, they overlook the inter-modal coherence and incoherence relations modeling, hindering the retrieval accuracy of CIR models. This limitation is non-trivial due to the following two challenges: 1) inter-modal incoherence and 2) intra-modal entanglement. To address the above challenges, we propose a comPlementArity-guided dIsentanglement netwoRk (PAIR), which can disentangle the features of multimodal queries from a semantic coherence perspective, thereby facilitating the identification of both complementary coherent and incoherent features. Furthermore, based on disentangled features, PAIR develops an asymmetric feature composition module, which is designed to enhance the retrieval performance of the model. Extensive experiments on three benchmark datasets demonstrate the superiority of PAIR. The code is available at https://zhihfu.github.io/PAIR.github.io/. Zhiheng Fu, Zixu Li 0001, Zhiwei Chen 0003, Xuemeng Song, Yupeng Hu 0003, Liqiang Nie |
ICASSP | 1 |
| 2025 | Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment RetrievalabstractCurrent text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieves the target moment using purified multimodal representations. Following this paradigm, we introduce the Denoise-then-Retrieve Network (DRNet), comprising Text-Conditioned Denoising (TCD) and Text-Reconstruction Feedback (TRF) modules. TCD integrates cross-attention and structured state space blocks to dynamically identify noisy clips and produce a noise mask to purify multimodal video representations. TRF further distills a single query embedding from purified video representations and aligns it with the text embedding, serving as auxiliary supervision for denoising during training. Finally, we perform conditional retrieval using text embeddings on purified video representations for accurate VMR. Experiments on Charades-STA and QVHighlights demonstrate that our approach surpasses state-of-the-art methods on all metrics. Furthermore, our denoise-then-retrieve paradigm is adaptable and can be seamlessly integrated into advanced VMR models to boost performance. Jiuxin Cao, Bo Miao, Zhiheng Fu, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian |
IJCAI | 4 |
| 2025 | OFFSET: Segmentation-based Focus Shift Revision for Composed Image RetrievalabstractComposed Image Retrieval (CIR) represents a novel retrieval paradigm that is capable of expressing users' intricate retrieval requirements flexibly. It enables the user to give a multimodal query, comprising a reference image and a modification text, and subsequently retrieve the target image. Notwithstanding the considerable advances made by prevailing methodologies, CIR remains in its nascent stages due to two limitations: 1) inhomogeneity between dominant and noisy portions in visual data is ignored, leading to query feature degradation, and 2) the priority of textual data in the image modification process is overlooked, which leads to a visual focus bias. To address these two limitations, this work presents a focus mapping-based feature extractor, which consists of two modules: dominant portion segmentation and dual focus mapping. It is designed to identify significant dominant portions in images and guide the extraction of visual and textual data features, thereby reducing the impact of noise interference. Subsequently, we propose a textually guided focus revision module, which can utilize the modification requirements implied in the text to perform adaptive focus revision on the reference image, thereby enhancing the perception of the modification focus on the composed features. The aforementioned modules collectively constitute the segmentatiOn-based Focus shiFt reviSion nETwork (OFFSET), and comprehensive experiments on four benchmark datasets substantiate the superiority of our proposed method. The codes and data are available on https://zivchen-ty.github.io/OFFSET.github.io/. Zhiwei Chen 0003, Yupeng Hu 0003, Zixu Li 0001, Zhiheng Fu, Xuemeng Song, Liqiang Nie |
ACM Multimedia | 4 |
| 2025 | HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video RetrievalabstractComposed Video Retrieval (CVR) is a challenging video retrieval task that utilizes multi-modal queries, consisting of a reference video and modification text, to retrieve the desired target video. The core of this task lies in understanding the multi-modal composed query and achieving accurate composed feature learning. Within multi-modal queries, the video modality typically carries richer semantic content compared to the textual modality. However, previous works have largely overlooked the disparity in information density between these two modalities. This limitation can lead to two critical issues: 1) modification subject referring ambiguity and 2) limited detailed semantic focus, both of which degrade the performance of CVR models. To address the aforementioned issues, we propose a novel CVR framework, namely the Hierarchical Uncertainty-aware Disambiguation network (HUD). HUD is the first framework that leverages the disparity in information density between video and text to enhance multi-modal query understanding. It comprises three key components: (a) Holistic Pronoun Disambiguation, (b) Atomistic Uncertainty Modeling, and (c) Holistic-to-Atomistic Alignment. By exploiting overlapping semantics through holistic cross-modal interaction and fine-grained semantic alignment via atomistic-level cross-modal interaction, HUD enables effective object disambiguation and enhances the focus on detailed semantics, thereby achieving precise composed feature learning. Moreover, our proposed HUD is also applicable to the Composed Image Retrieval (CIR) task and achieves state-of-the-art performance across three benchmark datasets for both CVR and CIR tasks. The codes are available on https://zivchen-ty.github.io/HUD.github.io/. Zhiwei Chen 0003, Yupeng Hu 0003, Zixu Li 0001, Zhiheng Fu, Haokun Wen, Weili Guan |
ACM Multimedia | 4 |
| 2025 | Memory guided representation learning for cross-domain face anti-spoofing
Pengchao Deng, Zhiheng Fu, Shengjun Xu, Chenyang Ge, Farid Boussaïd, Mohammed Bennamoun |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | WSSIC-Net: Weakly-Supervised Semantic Instance Completion of 3D Point Cloud ScenesabstractSemantic instance completion aims to recover the complete 3D shapes of foreground objects together with their labels from a partial 2.5D scan of a scene. Previous works have relied on full supervision, which requires ground-truth annotations, in the form of bounding boxes and complete 3D objects. This has greatly limited their real-world application because the acquisition of ground-truth data is very costly and time-consuming. To address this bottleneck, we propose a Weakly-Supervised Semantic Instance Completion Network (WSSIC-Net), which learns real-world partial point cloud object completion without requiring the ground truth of complete 3D objects. Instead, WSSIC-Net leverages 3D ground-truth bounding boxes, partial objects of a raw scene, and unpaired synthetic 3D point clouds. More specifically, a 3D detector is used to encode partial point clouds into proposal features, which are then fed into two branches. The first branch uses fully supervised box prediction based on proposal features. The second branch, hereinafter called instance completion, leverages the proposal features as partial object features to achieve weakly-supervised instance completion. A Generative Adversarial Network (GAN) completes the partial features of the 2.5D foreground objects of real-world scenes using only unpaired but semantically-consistent complete synthetic point clouds. In our experiments, we demonstrate that the fully-supervised 3D detection and the weakly-supervised instance completion complement one another. The qualitative and quantitative evaluations on the ScanNet v2 dataset demonstrate that the proposed "weakly-supervised" approach consistently achieves comparable performance to the state-of-the-art "fully supervised" methods. Zhiheng Fu, Yulan Guo, Minglin Chen, Qingyong Hu, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun |
IEEE Trans. Image Process. | 1 |
| 2025 | CompletionMamba: Taming State Space Model for Point Cloud CompletionabstractPoint cloud completion aims to reconstruct complete 3D shapes from partial scans. The long-range dependencies between points and shape perception are crucial for this task. While Transformers are effective due to their global processing ability, the quadratic complexity of their attention mechanism makes them unsuitable for long sequences when computational resources are constrained. As an alternative, State Space Models (SSMs) provide a memory-efficient solution for handling long-range dependencies, yet applying them directly to unordered point clouds presents challenges because of their intrinsic causality requirements. Existing methods attempt to address this by sorting points along a single axis. This, however, often overlooks complex causal relationships in 3D space since adjacency relationships based on Euclidean distance between points in the 3D space may not be preserved by this linear arrangement. To overcome this issue, we introduce CompletionMamba, a novel SSM-based network designed to harness SSMs for capturing both global and local dependencies within a point cloud. Initially, the input point cloud is causally structured by rearranging its coordinates. Then, a local SSM framework is proposed that defines neighborhood spaces around each point based on Euclidean distance, enhancing the causal structure. Although local SSM enhances relationships in short and long distance sequences, it still lacks full shape modeling of point cloud. To address this, we propose a novel shape-aware Mamba by integrating the shape code of each 3D shape into the model, enabling shape information propagation to all points. Our experiments show that CompletionMamba achieves state-of-the-art performance on both the MVP and PCN datasets. Zhiheng Fu, Longguang Wang, Lian Xu, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun |
IEEE Trans. Image Process. | 1 |
| 2024 | AEDNet: Adaptive Embedding and Multiview-Aware Disentanglement for Point Cloud Completion
Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun |
ECCV (11) | 1 |
| 2023 | VAPCNet: Viewpoint-Aware 3D Point Cloud CompletionabstractMost existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating the viewpoint of each incomplete object is usually time-consuming and leads to huge annotation cost. In this paper, we thus propose an unsupervised viewpoint representation learning scheme for 3D point cloud completion without explicit viewpoint estimation. To be specific, we learn abstract representations of partial scans to distinguish various viewpoints in the representation space rather than the explicit estimation in the 3D space. We also introduce a Viewpoint-Aware Point cloud Completion Network (VAPCNet) with flexible adaption to various viewpoints based on the learned representations. The proposed viewpoint representation learning scheme can extract discriminative representations to obtain accurate viewpoint information. Reported experiments on two popular public datasets show that our VAPCNet achieves state-of-the-art performance for the point cloud completion task. Source code is available at https://github.com/FZH92128/VAPCNet. Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun |
ICCV | 1 |
| 2023 | PMNet: A Point-to-Mesh Network for 3-D Semantic Instance ReconstructionabstractSemantic instance reconstruction attracts increasing attention in several areas such as mobile mapping, scene reconstruction, and robot navigation. Although much progresses have been made in recent years, the reconstruction performance is highly sensitive to occlusions and noises. To address these issues, we incorporate point cloud completion into a novel semantic instance reconstruction network PMNet, which consists of a 3-D object detection module, a point cloud completion module, and a mesh generation module. Based on the candidate instance proposals and their proposal features obtained in the object detection module, a point encoder layer is proposed to learn the local geometric features from the point cloud belonging to the detected instances, and a feature transformation layer is utilized to align the proposal features with the local geometric features. These two types of features are then fused and fed into the point cloud decoder to predict the complete point cloud of each instance. The mesh is finally reconstructed for each instance by the mesh generation module. Quantitative and qualitative experiments conducted on the ScanNetv2 dataset demonstrate that the proposed PMNet achieves the best reconstruction performance on real-world point clouds. Junhui Wan, Zhiheng Fu, Minglin Chen, Peng Zhang 0079, Hanyun Wang, Yulan Guo |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Multi-stage information diffusion for joint depth and surface normal estimation
Zhiheng Fu, Siyu Hong, Hamid Laga, Mohammed Bennamoun, Farid Boussaïd, Yulan Guo |
Pattern Recognit. | 1 |
| 2023 | Local-to-Global Cost Aggregation for Semantic CorrespondenceabstractEstablishing visual correspondences across semantically similar images is challenging due to intra-class variations, viewpoint changes, repetitive patterns, and background clutter. Recent approaches focus on cost aggregation to achieve promising performance. However, these methods fail to jointly utilize local and global cues to suppress unreliable matches. In this paper, we propose a cost aggregation network with convolutions and transformers, dubbed CACT. Different from existing methods, CACT refines the correlation map in a local-to-global manner by utilizing the strengths of convolutions and transformers in different stages. Additionally, considering the bidirectional nature of the correlation map, we propose a dual-path learning framework to work parallelly. Benefiting from the proposed framework, we can use 2D blocks to construct a cost aggregator to improve the efficiency of our model. Experimental results on the SPair-71k, PF-PASCAL, and PF-WILLOW datasets show that the proposed method outperforms the most state-of-the-art methods. Zi Wang 0008, Zhiheng Fu, Yulan Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Distortion-Aware Monocular Depth Estimation for Omnidirectional ImagesabstractImage distortion is a main challenge for tasks on panoramas. In this work, we propose a Distortion-Aware Monocular Omnidirectional (DAMO) network to estimate dense depth maps from indoor panoramas. First, we introduce a distortion-aware module to extract semantic features from omnidirectional images. Specifically, we exploit deformable convolution to adjust its sampling grids to geometric distortions on panoramas. We also utilize a strip pooling module to sample against horizontal distortion introduced by inverse gnomonic projection. Second, we introduce a plug-and-play spherical-aware weight matrix for our loss function to handle the uneven distribution of areas projected from a sphere. Experiments on the 360D dataset show that the proposed method can effectively extract semantic features from distorted panoramas and alleviate the supervision bias caused by distortion. It achieves the state-of-the-art performance on the 360D dataset with high efficiency. Hong-Xiang Chen, Kunhong Li 0001, Zhiheng Fu, Zonghao Chen, Yulan Guo |
IEEE Signal Process. Lett. | 3 |
| 2021 | Adv-Depth: Self-Supervised Monocular Depth Estimation With an Adversarial LossabstractLoss function plays a key role in self-supervised monocular depth estimation methods. Current reprojection loss functions are hand-designed and mainly focus on local patch similarity but overlook the global distribution differences between a synthetic image and a target image. In this paper, we leverage global distribution differences by introducing an adversarial loss into the training stage of self-supervised depth estimation. Specifically, we formulate this task as a novel view synthesis problem. We use a depth estimation module and a pose estimation module to form a generator, and then design a discriminator to learn the global distribution differences between real and synthetic images. With the learned global distribution differences, the adversarial loss can be back-propagated to the depth estimation module to improve its performance. Experiments on the KITTI dataset have demonstrated the effectiveness of the adversarial loss. The adversarial loss is further combined with the reprojection loss to achieve the state-of-the-art performance on the KITTI dataset. Kunhong Li 0001, Zhiheng Fu, Hanyun Wang, Zonghao Chen, Yulan Guo |
IEEE Signal Process. Lett. | 2 |
| 2018 | Simultaneous Context Feature Learning and Hashing for Large Scale Loop Closure DetectionabstractVisual loop closure is important in pose tracking and relocalization in many robotics and Argument Reality (AR) systems. For large and highly repetitive environments, sparse keypoint-based methods face several challenges, especially the discriminability of descriptors. In this paper, we propose an augmented descriptor by combining ORB feature and the context descriptor to increase its discriminability and matching performance. An end-to-end network is adopted to perform simultaneous feature learning and code hashing for the context. In addition, feature position clustering is used to reduce the number of contexts. Besides, hash mapping is adopted to reduce the dimensionality of ORB features. Finally, the context descriptors and ORB features with dimensionality reduction are stacked. Experimental results on the NewCollege and TUM datasets demonstrate that our algorithm achieves higher precision/recall and faster speed than the original algorithm proposed by Antonio et al. [1]. Zhiheng Fu, Yulan Guo, Wei An 0003 |
ICPR | 1 |
| 2017 | FSVO: Semi-direct monocular visual odometry using fixed mapsabstractWe propose a fixed-map semi-direct visual odometry (FSVO) algorithm for Micro Aerial Vehicles (MAVs). The proposed approach does not need computationally expensive feature extraction and matching techniques for motion estimation at each frame. Instead, we extract and match ORiented Brief (ORB) features between keyframes and assist-frames. We replace the incremental map generation step in traditional algorithms with fixed map generation at keyframe and assistframe only in our algorithm, resulting in reduced storage memory and higher flexibility for relocalization. Based on the fixed-map, we design a new keyframe selection criterion and a relocalization step. Our algorithm has no limit on the orientation of the camera and reduces drifting effectively. Experimental results on the EuRoC and KITTI datasets show that our algorithm achieves higher precision and robustness than the SVO algorithm. Zhiheng Fu, Yulan Guo, Zaiping Lin, Wei An 0003 |
ICIP | 1 |