EDBT 2026 Demo / reviewers in the wild / expert
Min Wang 0019
dblp:181/2695-19
· DBLP profile ↗
40ranked-venue papers
10as first author
35since 2021 · last 2026
0000-0003-3048-6980ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-author · 27 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Object Detection with Active Exploration and Spatiotemporal AggregationabstractClassic object detectors are fundamentally limited by their reliance on a single, static viewpoint, which often suffers from occlusions, challenging scales, and ambiguous perspectives. Video-based object detectors can partially mitigate this problem by aggregating information from multiple views, but they usually work with passively gathered sequences and are in no means guaranteed to provide sufficient information for all objects of interest. On the other hand, active vision methods can seek out better views, yet existing active object detectors typically only try to search for a single optimal perspective and discard valuable information gathered along their path. In this article, we introduce a new paradigm for object detection that unifies active exploration with cumulative spatiotemporal aggregation. We train an embodied agent to intelligently explore its environment, guided by a novel, detection-aware reward function that directly encourages seeking out views that resolve visual ambiguities. To leverage this active exploration, we introduce a robust aggregation pipeline that adeptly harnesses SAM 2’s temporal reasoning capabilities to fuse information from the agent’s entire trajectory into a single, coherent set of detections. Through extensive experiments on AI2-THOR, we demonstrate that our framework provides a consistent and substantial performance uplift when applied to a wide range of state-of-the-art detectors, establishing a strong and versatile new baseline for the next generation of active vision systems. Peiwei Li, Min Wang 0019, Wengang Zhou 0001, Yufei Yin, Yebo Bao, Guodong Shen, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex Environments
Min Wang 0019, Peiwei Li, Wengang Zhou 0001, Houqiang Li |
ICCV | 2 |
| 2025 | From History to Goal: Enhanced Vision-and-Language Navigation with Historical TraceabilityabstractIn Vision-and-Language Navigation (VLN), most methods ignore the navigation history after each episode, which is unrealistic for navigation in a persistent environment. Recent methods concatenate historical trajectories with the current one for history awareness, but also introduce redundant observations that offer little information gain and even harmful noise. To this end, we propose a history-traceable framework for VLN named TraceNav, which selects relevant historical information during persistent navigation. Specifically, it employs a multi-granularity matching strategy that consists of image-text matching and instruction-trajectory matching. In image-text matching, we accurately identify target candidates from historical observations using vision-language models and object detection models. Instruction-trajectory matching employs a cross-modal transformer to infer the degree of match between candidates’ trajectories and text instructions. We validate our method in similar pre-exploration and iterative setups, achieving performance that exceeds existing methods. We also validate our method on multiple datasets and achieve significant performance improvements over several state-of-the-art VLN methods. Our code is released on https://github.com/zhuxinguang33/TraceNav. Xinguang Zhu, Min Wang 0019, Li Li 0040, Wengang Zhou 0001, Houqiang Li |
ICME | 2 |
| 2025 | DP-Habitat: Bridging the Gap Between Simulation and Reality for Visual Navigation in Dynamic Pedestrian EnvironmentsabstractVisual navigation in dynamic environments poses a considerable challenge, particularly in scenarios with diverse pedestrian behaviors. Traditional simulators primarily focus on static scenes, while existing dynamic pedestrian simulators often suffer limitations such as monotonous pedestrian models, lack of interaction with the environment, and constrained scenarios. These deficiencies lead to notable discrepancies from real-world dynamic pedestrian environments. To bridge this gap, we introduce DP-Habitat, a dynamic pedestrian simulator developed on the Habitat platform. DP-Habitat efficiently simulates a wide range of complex and realistic human behaviors, with flexible interactions between pedestrian models and environments. It also supports rapid deployment of pedestrian models across various scenes, thereby more accurately replicating the complexities of real-world dynamic pedestrian settings. Additionally, we present Adaptive Object Navigation with Dynamic Mapping (AON-DM), a novel baseline method specifically designed for dynamic pedestrian settings. AON-DM integrates real-time pedestrian tracking and predictive modeling with a hybrid path planning strategy, markedly improving navigation efficiency and success rates. Our experimental results reveal that dynamic pedestrians significantly affect visual navigation performance within DP-Habitat, with AON-DM achieving superior effectiveness compared to existing methods under these challenging conditions. Furthermore, our approach maintains high performance in real-world scenarios, highlighting its practical applicability and robustness. The code and data are available at https://github.com/qinliangql/DP-Habitat.git. Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
ICRA | 2 |
| 2025 | Single-Source Dual-Stream Representation Learning for DNA Sequence ClassificationabstractDNA sequence classification is pivotal in genomics and bioinformatics for elucidating biological functions and diseases. Traditional sequence alignment methods, while precise, face significant challenges when applied to extensive, diverse datasets due to their computational intensity and limitations in scalability. To this end, digital encoding techniques transform DNA into vectors suitable for machine learning. However, these approaches often lose essential sequential or structural information, affecting the classification accuracy. In this work, we propose a novel approach, Single-Source Dual-stream representation learning (SSD), to enhance DNA sequence classification. SSD achieves this by extracting two pseudo-modalities from single-source data and integrating them into a dual-stream representation. Specifically, SSD regards DNA sequence as text and its Frequency Chaos Game Representation (FCGR) as image, effectively reframing DNA classification as a multi-modal learning task to capture diverse feature perspectives. We use BERT for DNA sequence text and Vision Transformer (ViT) for FCGR image encoding, with pre-training to capture information from both scales and obtain more generalized representations. Subsequently, an adaptive fusion module is designed to fuse the dual-stream representations before hierarchical classification to fully exploit the strengths of both modalities. Extensive experiments on three datasets reveal that SSD outperforms existing methods, highlighting its robust generalization and potential for novel genomic discoveries. Code is available at https://github.com/jiaruizhou/SSD. Zongmeng Zhang, Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
ICMR | 3 |
| 2025 | HandNeRF++: Modeling Animatable Interacting Hands With Neural Radiance FieldsabstractIn this work, we explore the rendering of photo-realistic free-viewpoint hand pose animation. We present HandNeRF, the first NeRF-based framework to reconstruct accurate appearance and geometry for interacting hands. To overcome the texture contamination and shape artifact problems when dealing with complex interacting scenarios, we further introduce HandNeRF++ to achieve better performance. In our advanced framework, a pose-driven deformation field is designed to establish correspondence from diverse poses to a canonical space, where the pose- and shape-disentangled NeRFs are optimized. To enhance the geometry and texture cues in rarely-observed areas for interacting hands, we establish a connection between the interacting hands by proposing the adaptive hand-sharing technique for cross-hand augmentation. Meanwhile, we further leverage the hand poses to generate fine-grained density priors, serving as valuable guidance for occlusion-aware geometry learning. Furthermore, a neural feature distillation method and a neural refiner are proposed to facilitate color optimization and further polish the renderings. With the collaboration of all the modules and strategies, our HandNeRF++ significantly advances the capabilities of NeRF-based 3D reconstruction in the context of interacting hands. Extensive experiments are conducted to validate the merits of the proposed frameworks. We report a series of state-of-the-art results both qualitatively and quantitatively. Zhiyang Guo, Wengang Zhou 0001, Min Wang 0019, Li Li 0040, Houqiang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | GaussNav: Gaussian Splatting for Visual NavigationabstractIn embodied vision, Instance ImageGoal Navigation (IIN) requires an agent to locate a specific object depicted in a goal image within an unexplored environment. The primary challenge of IIN arises from the need to recognize the target object across varying viewpoints while ignoring potential distractors. Existing map-based navigation methods typically use Bird's Eye View (BEV) maps, which lack detailed texture representation of a scene. Consequently, while BEV maps are effective for semantic-level visual navigation, they are struggling for instance-level tasks. To this end, we propose a new framework for IIN, Gaussian Splatting for Visual Navigation (GaussNav), which constructs a novel map representation based on 3D Gaussian Splatting (3DGS). The GaussNav framework enables the agent to memorize both the geometry and semantic information of the scene, as well as retain the textural features of objects. By matching renderings of similar objects with the target, the agent can accurately identify, ground, and navigate to the specified object. Our GaussNav framework demonstrates a significant performance improvement, with Success weighted by Path Length (SPL) increasing from 0.347 to 0.578 on the challenging Habitat-Matterport 3D (HM3D) dataset. Xiaohan Lei, Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Revisit Weakly Supervised Hashing With Deep Multi-Modal Foundation ModelsabstractVision-Language Pretraining (VLP) has developed a series of fancy foundation models, which continuously advance the state-of-the-art on various multimodal tasks. However, there has been limited exploration of their potential for large-scale image retrieval. In a real-world image retrieval system, images are collected together with user-annotated tags from the web. These tags contain various information about the corresponding image and could be used as weak supervision for image representation learning. In this paper, we seek to harness the powerful image-and-text alignment ability of VLP foundation models to enhance compact image representation. Specifically, we propose a new weakly supervised hashing framework, which learns a deep hashing network and enhances weak supervision alternatively. First, we extract the image and tag representation from VLP foundation models, and learn the deep hashing network with a policy gradient process, which directly optimizes the retrieval performance, i.e., mAP. Then given the learned deep hashing network, we further enhance the weak supervision with a separate probabilistic decision process. This process also optimizes the retrieval performance by the ground-truth defined with the learned hashing network. These two processes are alternatively repeated until a fixed number of steps. Experiments on public image datasets prove the effectiveness of our method. Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Motion-Aware 3D Gaussian Splatting for Efficient Dynamic Scene Reconstructionabstract3D Gaussian Splatting (3DGS) has become an emerging tool for dynamic scene reconstruction. However, existing methods mainly focus on developing various strategies to extend static 3DGS into a time-variant representation, while overlooking the rich motion information implicitly carried by 2D observations, thus suffering from performance degradation and model redundancy. To address the above problem, we propose a novel motion-aware enhancement framework for dynamic scene reconstruction, which mines useful motion cues from optical flow to improve different paradigms of dynamic 3DGS. Specifically, we first step beyond the vanilla render-based cross-dimensional supervision that suffers from ambiguity and instability, and establish a more robust and effective dense correspondence between 3D Gaussian movements and pixel-level flows. Then a novel flow augmentation method is introduced with additional insights into uncertainty and loss collaboration. Furthermore, for the prevalent deformation-based paradigm that presents a harder optimization problem, a transient-aware deformation auxiliary module is proposed. We conduct extensive experiments on both multi-view and monocular scenes to verify the merits of our work. Compared with the baselines, our method shows significant superiority in both rendering quality and efficiency. The code will be publicly available athttps://github.com/jasongzy/MAGS. Zhiyang Guo, Wengang Zhou 0001, Li Li 0040, Min Wang 0019, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Adaptive Bit Selection for Scalable Deep HashingabstractDeep Hashing is one of the most important methods for generating compact feature representation in content-based image retrieval. However, in various application scenarios, it requires training different models with diversified memory and computational resource costs. To address this problem, in this paper, we propose a new scalable deep hashing framework, which aims to generate binary codes with different code lengths by adaptive bit selection. Specifically, the proposed framework consists of two alternative steps, i.e., bit pool generation and adaptive bit selection. In the first step, a deep feature extraction model is trained to output binary codes by optimizing retrieval performance and bit properties. In the second step, we select informative bits from the generated bit pool with reinforcement learning algorithm, in which the same retrieval performance and bit properties are directly used in computing reward. The bit pool can be further updated by fine-tuning the deep feature extraction model with more attention on the selected bits. Hence, these two steps are alternatively iterated until convergence is achieved. Notably, most existing binary hashing methods can be readily integrated into our framework to generate scalable binary codes. Experiments on four public image datasets prove the effectiveness of the proposed framework for image retrieval tasks. Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Image Process. | 1 |
| 2025 | LayoutEnc: Leveraging Enhanced Layout Representations for Transformer-based Complex Scene SynthesisabstractIn complex scene synthesis, the effective representation of layouts is paramount. This paper introduces LayoutEnc, an advanced approach specifically designed to enhance layout representation by improving interpretability, robustness, and expressiveness, thereby facilitating more efficient image transformation. Distinct from conventional approaches that homogenize layout and image data, LayoutEnc distinctively processes various data modalities, enhancing the fidelity and interpretability of the layout representation. We apply stochastic noise injection to image tokens to align training and inference conditions, thereby fortifying the robustness of the layout representation. Additionally, LayoutEnc employs a two-stage multi-scale guidance learning strategy, to meticulously extract and refine semantic and textural features from training images. This enriched layout representation is then adeptly integrated into a transformer-based image generation framework, facilitating controlled and nuanced scene synthesis. Experimental results on the COCO-stuff and Visual Genome datasets demonstrate that LayoutEnc outperforms prior works in metrics such as FID and Scene-FID scores. The code and demo are available on https://github.com/qsun1/LayoutEnc . Qi Sun 0005, Min Wang 0019, Li Li 0040, Wengang Zhou 0001, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Instance-Aware Exploration-Verification-Exploitation for Instance ImageGoal NavigationabstractAs a new embodied vision task, Instance ImageGoal Navigation (IIN) aims to navigate to a specified object depicted by a goal image in an unexplored environment. The main challenge of this task lies in identifying the target object from different viewpoints while rejecting similar distractors. Existing ImageGoal Navigation methods usually adopt the simple Exploration-Exploitation framework and ignore the identification of specific instance during navigation. In this work, we propose to imitate the human behaviour of “getting closer to confirm” when distinguishing objects from a distance. Specifically, we design a new modular navigation framework named Instance-aware Exploration-Verification-Exploitation (IEVE) for instancelevel image goal navigation. Our method allows for active switching among the exploration, verification, and exploitation actions, thereby facilitating the agent in making reasonable decisions under different situations. On the challenging HabitatMatterport 3D semantic (HM3D-SEM) dataset, our method surpasses previous state-of-the-art work, with a classical segmentation model (0.684 vs. 0.561 success) or a robust model (0.702 vs. 0.561 success). Our code will be made publicly available at https://github.com/XiaohanLei/IEVE. Xiaohan Lei, Min Wang 0019, Wengang Zhou 0001, Li Li 0040, Houqiang Li |
CVPR | 2 |
| 2024 | Image2Sentence based Asymmetrical Zero-shot Composed Image RetrievalabstractThe task of composed image retrieval (CIR) aims to retrieve images based on the query image and the text describing the users' intent.
Existing methods have made great progress with the advanced large vision-language (VL) model in CIR task, however, they generally suffer from two main issues: lack of labeled triplets for model training and difficulty of deployment on resource-restricted environments when deploying the large vision-language model. To tackle the above problems, we propose Image2Sentence based Asymmetric zero-shot composed image retrieval (ISA), which takes advantage of the VL model and only relies on unlabeled images for composition learning. In the framework, we propose a new adaptive token learner that maps an image to a sentence in the word embedding space of VL model. The sentence adaptively captures discriminative visual information and is further integrated with the text modifier. An asymmetric structure is devised for flexible deployment, in which the lightweight model is adopted for the query side while the large VL model is deployed on the gallery side. The global contrastive distillation and the local alignment regularization are adopted for the alignment between the light model and the VL model for CIR task. Our experiments demonstrate that the proposed ISA could better cope with the real retrieval scenarios and further improve retrieval accuracy and efficiency. Yongchao Du, Min Wang 0019, Wengang Zhou 0001, Shuping Hui, Houqiang Li |
ICLR | 2 |
| 2024 | SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalabstractDifferent from traditional video retrieval, sign language retrieval is more biased towards understanding the semantic information of human actions contained in video clips. Previous works typically only encode RGB videos to obtain high-level semantic features, resulting in local action details drowned in a large amount of visual information redundancy. Furthermore, existing RGB-based sign retrieval works suffer from the huge memory cost of dense visual data embedding in end-to-end training, and adopt offline RGB encoder instead, leading to suboptimal feature representation. To address these issues, we propose a novel sign language representation framework called Semantically Enhanced Dual-Stream Encoder (SEDS), which integrates Pose and RGB modalities to represent the local and global information of sign language videos. Specifically, the Pose encoder embeds the coordinates of keypoints corresponding to human joints, effectively capturing detailed action features. For better context-aware fusion of two video modalities, we propose a Cross Gloss Attention Fusion (CGAF) module to aggregate the adjacent clip features with similar semantic information from intra-modality and inter-modality. Moreover, a Pose-RGB Fine-grained Matching Objective is developed to enhance the aggregated fusion feature by contextual matching of fine-grained dual-stream features. Besides the offline RGB encoder, the whole framework only contains learnable lightweight networks, which can be trained end-to-end. Extensive experiments demonstrate that our framework significantly outperforms state-of-the-art methods on various datasets. Code will be available at https://github.com/longtaojiang/SEDS. Longtao Jiang, Min Wang 0019, Zecheng Li 0002, Yao Fang, Wengang Zhou 0001, Houqiang Li |
ACM Multimedia | 2 |
| 2024 | P-RAG: Progressive Retrieval Augmented Generation For Planning on Embodied Everyday Task
Weiye Xu 0003, Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
ACM Multimedia | 2 |
| 2024 | Towards Codebook-Free Deep Probabilistic Quantization for Image RetrievalabstractAs a classical feature compression technique, quantization is usually coupled with inverted indices for scalable image retrieval. Most quantization methods explicitly divide feature space into Voronoi cells, and quantize feature vectors in each cell into the centroids learned from data distribution. However, Voronoi decomposition is difficult to achieve discriminative space partition for semantic image retrieval. In this paper, we explore semantic-aware feature space partition by deep neural network instead of Voronoi cells. To this end, we propose a new deep probabilistic quantization method, abbreviated as DeepIndex, which constructs inverted indices without explicit centroid learning. In our method, the deep neural network takes an image as input and outputs its probability of being put into each inverted index list. During training, we progressively quantize each image into the inverted lists with the top- T maximal probabilities, and calculate the reward of each trial based on retrieval accuracy. We optimize the deep neural network to maximize the probability of the inverted list with maximal reward. In this way, the retrieval performance is directly optimized, leading to a more semantically discriminative space partition than other quantization methods. The experiments on public image datasets demonstrate the effectiveness of our DeepIndex method on semantic image retrieval. Min Wang 0019, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | MASA: Motion-Aware Masked Autoencoder With Semantic Alignment for Sign Language RecognitionabstractSign language recognition (SLR) has long been plagued by insufficient model representation capabilities. Although current pre-training approaches have alleviated this dilemma to some extent and yielded promising performance by employing various pretext tasks on sign pose data, these methods still suffer from two primary limitations: i) Explicit motion information is usually disregarded in previous pretext tasks, leading to partial information loss and limited representation capability. ii) Previous methods focus on the local context of a sign pose sequence, without incorporating the guidance of the global meaning of lexical signs. To this end, we propose a Motion-Aware masked autoencoder with Semantic Alignment (MASA) that integrates rich motion cues and global semantic information in a self-supervised learning paradigm for SLR. Our framework contains two crucial components, i.e., a motion-aware masked autoencoder (MA) and a momentum semantic alignment module (SA). Specifically, in MA, we introduce an autoencoder architecture with a motion-aware masked strategy to reconstruct motion residuals of masked frames, thereby explicitly exploring dynamic motion cues among sign pose sequences. Moreover, in SA, we embed our framework with global semantic awareness by aligning the embeddings of different augmented samples from the input sequence in the shared latent space. In this way, our framework can simultaneously learn local motion cues and global semantic features for comprehensive sign language representation. Furthermore, we conduct extensive experiments to validate the effectiveness of our method, achieving new state-of-the-art performance on four public benchmarks. The source code are publicly available athttps://github.com/sakura/MASA. Weichao Zhao, Hezhen Hu, Wengang Zhou 0001, Yunyao Mao, Min Wang 0019, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Self-Supervised Representation Learning With Spatial-Temporal Consistency for Sign Language RecognitionabstractRecently, there have been efforts to improve the performance in sign language recognition by designing self-supervised learning methods. However, these methods capture limited information from sign pose data in a frame-wise learning manner, leading to sub-optimal solutions. To this end, we propose a simple yet effective self-supervised contrastive learning framework to excavate rich context via spatial-temporal consistency from two distinct perspectives and learn instance discriminative representation for sign language recognition. On one hand, since the semantics of sign language are expressed by the cooperation of fine-grained hands and coarse-grained trunks, we utilize both granularity information and encode them into latent spaces. The consistency between hand and trunk features is constrained to encourage learning consistent representation of instance samples. On the other hand, inspired by the complementary property of motion and joint modalities, we first introduce first-order motion information into sign language modeling. Additionally, we further bridge the interaction between the embedding spaces of both modalities, facilitating bidirectional knowledge transfer to enhance sign language representation. Our method is evaluated with extensive experiments on four public benchmarks, and achieves new state-of-the-art performance with a notable margin. The source code is publicly available at https://github.com/sakura/Code. Weichao Zhao, Wengang Zhou 0001, Hezhen Hu, Min Wang 0019, Houqiang Li |
IEEE Trans. Image Process. | 4 |
| 2024 | Progressive Similarity Preservation Learning for Deep Scalable Product QuantizationabstractProduct quantization is an effective strategy for compact feature learning in image retrieval, which generates compact quantization codes of different lengths for varying scenarios. However, existing deep quantization methods obtain quantization codes with different lengths by training multiple models separately for each code length, which brings about large training time cost and degrades deployment flexibility. To this end, we propose a new deep scalable Progressive Similarity Preservation Product Quantization (PSPPQ) framework, which enables us to train the quantized features in different code lengths simultaneously and imposes no additional cost during inference. By progressively approximating the ground truth similarity of image pairs, we achieve direct optimization of similarity ranking, which improves the retrieval accuracy and generates sequential quantization codes with more efficiency. Besides, by combining the advantages of classification loss and hinge loss, we design a semantic ArcFace loss to optimize our network architecture. Experiments on three datasets demonstrate the effectiveness of our proposed method with variable code lengths for scalable image retrieval. Yongchao Du, Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Multim. | 2 |
| 2024 | Structure Similarity Preservation Learning for Asymmetric Image RetrievalabstractAsymmetric image retrieval is a task that seeks to balance retrieval accuracy and efficiency by leveraging lightweight and large models for the query and gallery sides, respectively. The key to asymmetric image retrieval is realizing feature compatibility between different models. Despite the great progress, most existing approaches either rely on classifiers inherited from gallery models or simply impose constraints at the instance level, ignoring the structure of embedding space. In this work, we propose a simple yet effective structure similarity preserving method to achieve feature compatibility between query and gallery models. Specifically, we first train a product quantizer offline with the image features embedded by the gallery model. The centroid vectors in the quantizer serve as anchor points in the embedding space of the gallery model to characterize its structure. During the training of the query model, anchor points are shared by the query and gallery models. The relationships between image features and centroid vectors are considered as structure similarities and constrained to be consistent. Moreover, our approach makes no assumption about the existence of any labeled training data and thus can be extended to an unlimited amount of data. Comprehensive experiments on large-scale landmark retrieval demonstrate the effectiveness of our approach. Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Multim. | 2 |
| 2023 | HandNeRF: Neural Radiance Fields for Animatable Interacting HandsabstractWe propose a novel framework to reconstruct accurate appearance and geometry with neural radiance fields (NeRF) for interacting hands, enabling the rendering of photo-realistic images and videos for gesture animation from arbitrary views. Given multi-view images of a single hand or interacting hands, an off-the-shelf skeleton estimator is first employed to parameterize the hand poses. Then we design a pose-driven deformation field to establish correspondence from those different poses to a shared canonical space, where a pose-disentangled NeRF for one hand is optimized. Such unified modeling efficiently complements the geometry and texture cues in rarely-observed areas for both hands. Meanwhile, we further leverage the pose priors to generate pseudo depth maps as guidance for occlusion aware density learning. Moreover, a neural feature distillation method is proposed to achieve cross-domain alignment for color optimization. We conduct extensive experiments to verify the merits of our proposed HandNeRF and report a series of state-of-the-art results both qualitatively and quantitatively on the large-scale InterHand2.6M dataset. Zhiyang Guo, Wengang Zhou 0001, Min Wang 0019, Li Li 0040, Houqiang Li |
CVPR | 3 |
| 2023 | Asymmetric Feature Fusion for Image RetrievalabstractIn asymmetric retrieval systems, models with different capacities are deployed on platforms with different computational and storage resources. Despite the great progress, existing approaches still suffer from a dilemma between retrieval efficiency and asymmetric accuracy due to the limited capacity of the lightweight query model. In this work, we propose an Asymmetric Feature Fusion (AFF) paradigm, which advances existing asymmetric retrieval systems by considering the complementarity among different features just at the gallery side. Specifically, it first embeds each gallery image into various features, e.g., local features and global features. Then, a dynamic mixer is introduced to aggregate these features into compact embedding for efficient search. On the query side, only a single lightweight model is deployed for feature extraction. The query model and dynamic mixer are jointly trained by sharing a momentum-updated classifier. Notably, the proposed paradigm boosts the accuracy of asymmetric retrieval without introducing any extra overhead to the query side. Exhaustive experiments on various landmark retrieval datasets demonstrate the superiority of our paradigm. Min Wang 0019, Wengang Zhou 0001, Zhenbo Lu, Houqiang Li |
CVPR | 2 |
| 2023 | A General Rank Preserving Framework for Asymmetric Image Retrieval
Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
ICLR | 2 |
| 2023 | Deep Graph Convolutional Quantization Networks for Image RetrievalabstractTo achieve real-time online search, most image retrieval methods aim to learn compact feature representation while keeping their semantic information or intra-class relevance. In this paper, we propose a new compact feature learning method to embed the underlying manifold information from database. It integrates deep convolutional neural network (CNN) and graph convolutional neural networks (GCN) into a unified end-to-end learning framework. In the proposed method, the deep feature extracted by CNN is automatically embedded with the information from its neighbors by GCN, which possesses the ability of exploring the semantic relevance on the database manifold. Since constructing a graph over the whole database costs unaffordable memory, we build a landmark graph as database sketch. The landmark graph contains two kinds of nodes, including codewords and memory bank samples. Given an image, the deep architecture outputs the discriminative feature and its similarity with all the graph nodes. We directly use the indices of the most similar codeword nodes as the compact feature representation. To make the proposed method scalable to large datasets, a multi-graph strategy is adopted to generate compact features with adaptable code length. The experiments on two benchmark datasets demonstrate the effectiveness of the proposed method. Min Wang 0019, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li |
IEEE Trans. Multim. | 1 |
| 2023 | Hash Bit Selection With Reinforcement Learning for Image RetrievalabstractIn recent years, binary hashing methods have been widely used in large-scale multimedia retrieval because of the low computational complexity and memory cost. Generally, better retrieval accuracy can be achieved with a longer hash code, which, however, may suffer redundancy. In this paper, we propose a novel hash bit selection method, called Hash Bit Selection with Reinforcement Learning (HBS-RL), which aims to adaptively select the most informative bits from the database binary codes. In our approach, the hash bit selection problem is firstly modeled as a Markov Decision Process (MDP), which is solved with reinforcement learning. HBS-RL learns a policy for bit selection, which effectively identifies the most informative bits by directly maximizing mean Average Precision (mAP) during training. Specially, given a generated bit pool, our HBS-RL can sequentially select bits with different code lengths with a very lightweight fully-connected policy network. The proposed method is evaluated on the MNIST, CIFAR-10, ImageNet and NUS-WIDE datasets, and the results show that it significantly improves the retrieval performance of the existing unsupervised and deep supervised hashing methods. It also outperforms the state-of-the-art bit selection methods. For convenience of repeating our results, we release our source code at:https://github.com/xyez/HBS-RL. Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Multim. | 2 |
| 2023 | Weakly Supervised Hashing with Reconstructive Cross-modal AttentionabstractOn many popular social websites, images are usually associated with some meta-data such as textual tags, which involve semantic information relevant to the image and can be used to supervise the representation learning for image retrieval. However, these user-provided tags are usually polluted by noise, therefore the main challenge lies in mining the potential useful information from those noisy tags. Many previous works simply treat different tags equally to generate supervision, which will inevitably distract the network learning. To this end, we propose a new framework, termed as Weakly Supervised Hashing with Reconstructive Cross-modal Attention (WSHRCA), to learn compact visual-semantic representation with more reliable supervision for retrieval task. Specifically, for each image-tag pair, the weak supervision from tags is refined by cross-modal attention, which takes image feature as query to aggregate the most content-relevant tags. Therefore, tags with relevant content will be more prominent while noisy tags will be suppressed, which provides more accurate supervisory information. To improve the effectiveness of hash learning, the image embedding in WSHRCA is reconstructed from hash code, which is further optimized by cross-modal constraint and explicitly improves hash learning. The experiments on two widely-used datasets demonstrate the effectiveness of our proposed method for weakly-supervised image retrieval. The code is available at https://github.com/duyc168/weakly-supervised-hashing . Yongchao Du, Min Wang 0019, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Learning Token-Based Representation for Image RetrievalabstractIn image retrieval, deep local features learned in a data-driven manner have been demonstrated effective to improve retrieval performance. To realize efficient retrieval on large image database, some approaches quantize deep local features with a large codebook and match images with aggregated match kernel. However, the complexity of these approaches is non-trivial with large memory footprint, which limits their capability to jointly perform feature learning and aggregation. To generate compact global representations while maintaining regional matching capability, we propose a unified framework to jointly learn local feature representation and aggregation. In our framework, we first extract local features using CNNs. Then, we design a tokenizer module to aggregate them into a few visual tokens, each corresponding to a specific visual pattern. This helps to remove background noise, and capture more discriminative regions in the image. Next, a refinement block is introduced to enhance the visual tokens with self-attention and cross-attention. Finally, different visual tokens are concatenated to generate a compact global representation. The whole framework is trained end-to-end with image-level labels. Extensive experiments are conducted to evaluate our approach, which outperforms the state-of-the-art methods on the Revisited Oxford and Paris datasets. Min Wang 0019, Wengang Zhou 0001, Yang Hu 0006, Houqiang Li |
AAAI | 2 |
| 2022 | Contextual Similarity Distillation for Asymmetric Image RetrievalabstractAsymmetric image retrieval, which typically uses small model for query side and large model for database server, is an effective solution for resource-constrained scenarios. However, existing approaches either fail to achieve feature coherence or make strong assumptions, e.g., requiring labeled datasets or classifiers from large model, etc., which limits their practical application. To this end, we propose a flexible contextual similarity distillation framework to enhance the small query model and keep its output feature compatible with that of the large gallery model, which is crucial with asymmetric retrieval. In our approach, we learn the small model with a new contextual similarity consistency constraint without any data label. During the small model learning, it preserves the contextual similarity among each training image and its neighbors with the features extracted by the large model. Note that this simple constraint is consistent with simultaneous first-order feature vector preserving and second-order ranking list preserving. Extensive experiments show that the proposed method outperforms the state-of-the-art methods on the Revisited Oxford and Paris datasets. Min Wang 0019, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001 |
CVPR | 2 |
| 2022 | CMT: Context-Matching-Guided Transformer for 3D Tracking in Point Clouds
Zhiyang Guo, Yunyao Mao, Wengang Zhou 0001, Min Wang 0019, Houqiang Li |
ECCV (22) | 4 |
| 2022 | Deep Enhanced Weakly-Supervised Hashing With Iterative Tag RefinementabstractOn image-sharing websites, images are usually associated with user-generated tags which contain semantic information and are more easily accessible than accurate labels. It is beneficial to utilize such tags as supervised information to learn image feature representation. However, some tags are not related with the image content and disturb the feature learning process. In this paper, we are dedicated to refining such noisy tags and upgrading the image feature learning. To this end, we propose a novel deep enhanced weakly-supervised hashing method, in which tags are adaptively refined according to image content. In our approach, we first map the deep image feature representation into the tag embedding space, and learn the discriminative as well as compact feature representations with the corresponding tags. After that, by referring to the learned feature representation in the first step, we refine the tags to become consistent with image content. The above two steps are alternated until convergence. Finally, we can obtain more content-relevant tags, better image features and binary hashing functions. The experiments on two image datasets prove that the proposed method outperforms the state-of-the-art weakly-supervised deep hashing methods on image retrieval task. Min Wang 0019, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li |
IEEE Trans. Multim. | 1 |
| 2021 | Learning Deep Local Features with Multiple Dynamic Attentions for Large-Scale Image RetrievalabstractIn image retrieval, learning local features with deep convolutional networks has been demonstrated effective to improve the performance. To discriminate deep local features, some research efforts turn to attention learning. However, existing attention-based methods only generate a single attention map for each image, which limits the exploration of diverse visual patterns. To this end, we propose a novel deep local feature learning architecture to simultaneously focus on multiple discriminative local patterns in an image. In our framework, we first adaptively reorganize the channels of activation maps for multiple heads. For each head, a new dynamic attention module is designed to learn the potential attentions. The whole architecture is trained as metric learning of weighted-sum-pooled global image features, with only image-level relevance label. After the architecture training, for each database image, we select local features based on their multi-head dynamic attentions, which are further indexed for efficient retrieval. Extensive experiments show the proposed method outperforms the state-of-the-art methods on the Revisited Oxford and Paris datasets. Besides, it typically achieves competitive results even using local features with lower dimensions. Code will be released at https://github.com/CHANWH/MDA. Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
ICCV | 2 |
| 2021 | Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalabstractIn this paper, we focus on the composed query image retrieval task, namely retrieving the target images that are similar to a composed query, in which a modification text is combined with a query image to describe a user's accurate search intention. Previous methods usually focus on learning the joint image-text representations, but rarely consider the intrinsic relationship among the query image, the target image and the modification text. To address this problem, we propose a new cross-modal joint prediction and alignment framework for composed query image retrieval. In our framework, the modification text is regarded as an implicit transformation between the query image and the target image. Motivated by that, not only the combination of the query image and modification text should be similar to the target image, but also the modification text should be predicted according to the query image and the target image. We devote to aligning this relationship by a novel Joint Prediction Module (JPM). Our proposed framework can seamlessly incorporate the JPM into the existing methods to effectively improve the discrimination and robustness of visual and textual representations. The experiments on three public datasets demonstrate the effectiveness of our proposed framework, proving that our proposed JPM can be simply incorporated with the existing methods while effectively improving the performance. Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
ACM Multimedia | 2 |
| 2021 | Contextual Similarity Aggregation with Self-attention for Visual Re-rankingabstractIn content-based image retrieval, the first-round retrieval result by simple visual feature comparison may be unsatisfactory, which can be refined by visual re-ranking techniques. In image retrieval, it is observed that the contextual similarity among the top-ranked images is an important clue to distinguish the semantic relevance. Inspired by this observation, in this paper, we propose a visual re-ranking method by contextual similarity aggregation with self-attention. In our approach, for each image in the top-K ranking list, we represent it into an affinity feature vector by comparing it with a set of anchor images. Then, the affinity features of the top-K images are refined by aggregating the contextual information with a transformer encoder. Finally, the affinity features are used to recalculate the similarity scores between the query and the top-K images for re-ranking of the latter. To further improve the robustness of our re-ranking model and enhance the performance of our method, a new data augmentation scheme is designed. Since our re-ranking model is not directly involved with the visual feature used in the initial retrieval, it is ready to be applied to retrieval result lists obtained from various retrieval algorithms. We conduct comprehensive experiments on four benchmark datasets to demonstrate the generality and effectiveness of our proposed visual re-ranking method. Jianbo Ouyang, Min Wang 0019, Wengang Zhou 0001, Houqiang Li |
NeurIPS | 3 |
| 2021 | Deep Relation Embedding for Cross-Modal RetrievalabstractCross-modal retrieval aims to identify relevant data across different modalities. In this work, we are dedicated to cross-modal retrieval between images and text sentences, which is formulated into similarity measurement for each image-text pair. To this end, we propose a Cross-modal Relation Guided Network (CRGN) to embed image and text into a latent feature space. The CRGN model uses GRU to extract text feature and ResNet model to learn the globally guided image feature. Based on the global feature guiding and sentence generation learning, the relation between image regions can be modeled. The final image embedding is generated by a relation embedding module with an attention mechanism. With the image embeddings and text embeddings, we conduct cross-modal retrieval based on the cosine similarity. The learned embedding space well captures the inherent relevance between image and text. We evaluate our approach with extensive experiments on two public benchmark datasets, i.e., MS-COCO and Flickr30K. Experimental results demonstrate that our approach achieves better or comparable performance with the state-of-the-art methods with notable efficiency. Yifan Zhang 0011, Wengang Zhou 0001, Min Wang 0019, Qi Tian 0001, Houqiang Li |
IEEE Trans. Image Process. | 3 |
| 2021 | Collaborative Image Relevance Learning for Visual Re-RankingabstractIn content-based image retrieval, the initial retrieval result may be unsatisfactory, which can be refined with visual re-ranking techniques, such as query expansion, geometric verification,etc. In this work, we approach visual re-ranking from a novel perspective. Observing that the contextual similarity of images from a retrieval result list exhibits strong visual relevance, we propose to collaboratively learn the semantic relevance among images for visual re-ranking. In our approach, we represent the image set of a fixed-length retrieval list into a correlation matrix, and learn the relevance of all image pairs simultaneously with a lightweight CNN model. To optimize the CNN model, a weighted MSE loss is defined, which takes into account the sparsity of labels. To find the optimal length of retrieval result list for different queries, we present a query sensitive selection method. We conduct comprehensive experiments on five benchmark datasets, and demonstrate the generality, and effectiveness of the proposed visual re-ranking method. Jianbo Ouyang, Wengang Zhou 0001, Min Wang 0019, Qi Tian 0001, Houqiang Li |
IEEE Trans. Multim. | 3 |
| 2020 | Neighborhood Pyramid Preserving HashingabstractIn this paper, we devote our efforts to the approximate nearest neighbour (ANN) search problem and propose a new unsupervised binary hashing method, i.e., Neighbourhood Pyramid preserving Hashing (NPH). We represent the nearest neighbours of each data point in a pyramid, and as the learning objective, we impose that the pyramid neighbourhood in each level is consistently preserved across the original Euclidean space and the transformed Hamming space. The neighbourhood is quantitatively characterized by its size, defined as the average distance from the involved nearest neighbours to the referred data point. Our approach is consistent with the distance-preserving principle of binary hashing and achieves stricter neighbourhood structure preserving over previous graph hashing algorithms. The experiments on several large-scale benchmark datasets demonstrate that NPH achieves promising performances compared with those of the existing state-of-the-art unsupervised binary hashing methods. Min Wang 0019, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li |
IEEE Trans. Multim. | 1 |
| 2019 | Deep Scalable Supervised Quantization by Self-Organizing MapabstractApproximate Nearest Neighbor (ANN) search is an important research topic in multimedia and computer vision fields. In this article, we propose a new deep supervised quantization method by Self-Organizing Map to address this problem. Our method integrates the Convolutional Neural Networks and Self-Organizing Map into a unified deep architecture. The overall training objective optimizes supervised quantization loss as well as classification loss. With the supervised quantization objective, we minimize the differences on the maps between similar image pairs and maximize the differences on the maps between dissimilar image pairs. By optimization, the deep architecture can simultaneously extract deep features and quantize the features into suitable nodes in self-organizing map. To make the proposed deep supervised quantization method scalable for large datasets, instead of constructing a larger self-organizing map, we propose to divide the input space into several subspaces and construct self-organizing map in each subspace. The self-organizing maps in all the subspaces implicitly construct a large self-organizing map, which costs less memory and training time than directly constructing a self-organizing map with equal size. The experiments on several public standard datasets prove the superiority of our approaches over the existing ANN search methods. Besides, as a by-product, our deep architecture can be directly applied to visualization with little modification, and promising performance is demonstrated in the experiments. Min Wang 0019, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | A General Framework for Linear Distance Preserving HashingabstractBinary hashing approaches the approximate nearest neighbor search problem by transferring the data to Hamming space with explicit or implicit distance preserving constraint. With compact data representation, binary hashing identifies the approximate nearest neighbors via very efficient Hamming distance computation. In this paper, we propose a generic hashing framework with a new linear pairwise distance preserving objective and pointwise constraint. In our framework, the direct distance preserving objective aims to keep the linear relationship between the Euclidean distance and the Hamming distance of data points. On the other hand, to impose the pointwise constraint, we instantiate the framework from three different perspectives with pseudo-supervised, unsupervised, and supervised clues and obtain three different hashing methods. The first one is a pseudo-supervised hashing method, which adopts a certain existing unsupervised hashing method to generate binary codes as pseudo-supervised information. For the second one, we get an unsupervised hashing method by considering the quantization loss. The third one, as a supervised hashing method, learns the hash functions in a two-step paradigm. Furthermore, we improve the above-mentioned framework by constraining the global scope of the proposed linear distance preserving objective to a local range. We validate our framework on four large-scale benchmark data sets. The experiments demonstrate that our pseudo-supervised method achieves consistent improvement over the state-of-the-art unsupervised hashing methods, while our unsupervised and supervised methods achieve promising performance compared with the state-of-the-art algorithms. Min Wang 0019, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li |
IEEE Trans. Image Process. | 1 |
| 2017 | Deep Supervised Quantization by Self-Organizing MapabstractApproximate Nearest Neighbour (ANN) search is an important research topic in multimedia and computer vision fields. In this paper, we propose a new deep supervised quantization method by Self-Organizing Map (SOM) to address this problem. Our method integrates the Convolutional Neural Networks (CNN) and Self-Organizing Map into a unified deep architecture. The overall training objective includes supervised quantization loss and classification loss. With the supervised quantization loss, we minimize the differences on the maps between similar image pairs, and maximize the differences on the maps between dissimilar image pairs. By optimization, the deep architecture can simultaneously extract deep features and quantize the features into the suitable nodes in the Self-Organizing Map. The experiments on several public standard datasets prove the superiority of our approach over the existing ANN search methods. Besides, as a byproduct, our deep architecture can be directly applied to classification task and visualization with little modification, and promising performances are demonstrated on these tasks in the experiments. Min Wang 0019, Wengang Zhou 0001, Qi Tian 0001, Junfu Pu, Houqiang Li |
ACM Multimedia | 1 |
| 2016 | Linear Distance Preserving Pseudo-Supervised and Unsupervised HashingabstractWith the advantage in compact representation and efficient comparison, binary hashing has been extensively investigated for approximate nearest neighbor search. In this paper, we propose a novel and general hashing framework, which simultaneously considers a new linear pair-wise distance preserving objective and point-wise constraint. The direct distance preserving objective aims to keep the linear relationships between the Euclidean distance and the Hamming distance of data points. Based on different point-wise constraints, we propose two methods to instantiate this framework. The first one is a pseudo-supervised hashing method, which uses existing unsupervised hashing methods to generate binary codes as pseudo-supervised information. The second one is an unsupervised hashing method, in which quantization loss is considered. We validate our framework on two large-scale datasets. The experiments demonstrate that our pseudo-supervised method achieves consistent improvement for the state-of-the-art unsupervised hashing methods, while our unsupervised method outperforms the state-of-the-art methods. Min Wang 0019, Wengang Zhou 0001, Qi Tian 0001, Zhengjun Zha, Houqiang Li |
ACM Multimedia | 1 |