EDBT 2026 Demo / reviewers in the wild / expert
Zizhuo Li
dblp:304/4312
· DBLP profile ↗
18ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0003-0986-4924ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SigMa: Semantic Similarity-Guided Semi-Dense Feature MatchingabstractRecent advancements have led the image matching community to increasingly focus on obtaining subpixel-level correspondences in a detector-free manner, i.e., semi-dense feature matching. Existing methods tend to overfocus on low-level local features while ignoring equally important high-level semantic information. To tackle these shortcomings, we propose SigMa, a semantic similarity-guided semi-dense feature matching method, which leverages the strengths of both local features and high-level semantic features. First, we design a dual-branch feature extractor, comprising a convolutional network and a vision foundation model, to extract low-level local features and high-level semantic features, respectively. To fully retain the advantages of these two features and effectively integrate them, we also introduce a cross-domain feature adapter, which could overcome their spatial resolution mismatches, channel dimensionality variations, and inter-domain gaps. Furthermore, we observe that performing the transformer on the whole feature map is unnecessary because of the similarity of local representations. We design a guided pooling method based on semantic similarity. This strategy performs attention computation by selecting highly semantically similar regions, aiming to minimize information loss while maintaining computational efficiency. Extensive experiments on multiple datasets demonstrate that our method achieves a competitive accuracy-efficiency trade-off across various tasks and exhibits strong generalization capabilities across different datasets. Additionally, we conduct a series of ablation studies and analysis experiments to validate the effectiveness and rationality of our method's design. Our code is publicly available at https://github.com/ShineFox/SigMa. Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | DeMo: Deep Motion Field Consensus with Learnable Kernels for Two-view Correspondence LearningabstractAs a long-range prior, motion consensus essentially forces the overall spatial transformation between a pair of images to be smooth and consistent, which is naturally well-suited for two-view correspondence learning. However, such precious property remains under-explored by most existing studies due to the modeling challenges posed by the sparsity and uneven distributions of putative correspondences. In this paper, we propose DeMo, a novel and cutting-edge network for outlier rejection, which possesses the capacity to fully capture global motion consensus clues by way of consensus interpolation over the entire high-dimensional motion field generated by putative correspondences. Specifically, through incorporating regularization techniques into a Reproducing Kernel Hilbert Space (RKHS), a concise interpolation formula can be derived for the high-dimensional motion field, which inherently allows a closed-form solution. Subsequently, learnable deep kernels are collaboratively used to flexibly and efficiently capture the relationships between global inputs, thus maintaining the entire motion field consensus. In addition, to remedy the cubic computational overhead of explicit interpolation, a scene-adaptive sampling strategy is introduced, which implicitly selects the more scene-representative motions, reducing the computational complexity of motion consensus interpolation to be approximately linear while maintaining the accuracy. Moreover, to deal with underlying depth discontinuities caused by complicated scene variations, a local consensus complementation block is designed, which maintains local bilateral consensus across both feature and spatial channels. Without bells and whistles, DeMo achieves superior performance in various geometric tasks, including relative pose estimation, homography estimation, and visual localization. Jiajun Le, Zizhuo Li, Yixuan Yuan, Jiayi Ma 0001 |
AAAI | 3 |
| 2025 | Matching While Perceiving: Enhance Image Feature Matching with Applicable Semantic AmalgamationabstractImage feature matching is a cardinal problem in computer vision, aiming to establish accurate correspondences between two-view images. Existing methods are constrained by the performance of feature extractors and struggle to capture local information affected by sparse texture or occlusions. Recognizing that human eyes consider not only similar local geometric features but also high-level semantic information of scene objects when matching images, this paper introduces SemaGlue. This novel algorithm perceives and incorporates semantic information into the matching process. In contrast to recent approaches that leverage semantic consistency to narrow the scope of matching areas, SemaGlue achieves semantic amalgamation with the designed Semantic-Aware Fusion (SAF) Block by injecting abundant semantic features from the pre-trained segmentation model. Moreover, the Cross-Domain Alignment (CDA) Block is proposed to address domain alignment issues, bridging the gaps between semantic and geometric domains to ensure applicable semantic amalgamation. Extensive experiments demonstrate that SemaGlue outperforms state-of-the-art methods across various applications such as homography estimation, relative pose estimation, and visual localization. Zhenjie Zhu, Zizhuo Li, Tao Lu 0001, Jiayi Ma 0001 |
AAAI | 3 |
| 2025 | MINIMA: Modality Invariant Image MatchingabstractImage matching for both cross-view and cross-modality plays a critical role in multimodal perception. In practice, the modality gap caused by different imaging systems/styles poses great challenges to the matching task. Existing works try to extract invariant features for specific modalities and train on limited datasets, showing poor generalization. In this paper, we present MINIMA, a unified image matching framework for multiple cross-modal cases. Without pursuing fancy modules, our MINIMA aims to enhance universal performance from the perspective of data scaling up. For such purpose, we propose a simple yet effective data engine that can freely produce a large dataset containing multiple modalities, rich scenarios, and accurate matching labels. Specifically, we scale up the modalities from cheap but rich RGB-only matching data, by means of generative models. Under this setting, the matching labels and rich diversity of the RGB dataset are well inherited by the generated multimodal data. Benefiting from this, we construct MD-syn, a new comprehensive dataset that fills the data gap for general multimodal image matching. With MD-syn, we can directly train any advanced matching pipeline on randomly selected modality pairs to obtain cross-modal ability. Extensive experiments on in-domain and zero-shot matching tasks, including 19 cross-modal cases, demonstrate that our MINIMA can significantly outperform the baselines and even surpass modality-specific methods. The dataset and code are available at https://github.com/LSXI7/MINIMA. Jiangwei Ren, Xingyu Jiang 0005, Zizhuo Li, Dingkang Liang, Xin Zhou 0013, Xiang Bai |
CVPR | 3 |
| 2025 | CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image MatchingabstractThis prospective study proposes CoMatch, a novel semi-dense image matcher with dynamic covisibility awareness and bilateral subpixel accuracy. Firstly, observing that modeling context interaction over the entire coarse feature map elicits highly redundant computation due to the neighboring representation similarity of tokens, a covisibility-guided token condenser is introduced to adaptively aggregate tokens in light of their covisibility scores that are dynamically estimated, thereby ensuring computational efficiency while improving the representational capacity of aggregated tokens simultaneously. Secondly, considering that feature interaction with massive non-covisible areas is distracting, which may degrade feature distinctiveness, a covisibility-assisted attention mechanism is deployed to selectively suppress irrelevant message broadcast from non-covisible reduced tokens, resulting in robust and compact attention to relevant rather than all ones. Thirdly, we find that at the fine-level stage, current methods adjust only the target view's keypoints to subpixel level, while those in the source view remain restricted at the coarse level and thus not informative enough, detrimental to keypoint location-sensitive usages. A simple yet potent fine correlation module is developed to refine the matching candidates in both source and target views to subpixel level, attaining attractive performance improvement. Thorough experimentation across an array of public benchmarks affirms CoMatch's promising accuracy, efficiency, and generalizability. Zizhuo Li, Linfeng Tang, Jiayi Ma 0001 |
ICCV | 1 |
| 2025 | CorrNeXt: Making the ConvNet-Style Correspondence Pruner Stronger for Two-View GeometryabstractThe uproar over two-view correspondence pruning stems from the advent of the ConvNet-style paradigm, which showcases intrinsic proficiency in local context aggregation, tackling the context-agnostic deficiency of MLP-based methods fundamentally and delivering impressive pruning capability. To further unlock the potential of such a paradigm, this perspective study revisits its design decisions and introduces CorrNeXt, a cutting-edge ConvNet-style pruner that incorporates multiple simple but effective improvements. Firstly, we explicitly integrate 2D relative spatial knowledge into motion field modeling, arming the interconversion between unordered sparse motion vectors and ordered image-structured ones with positional awareness. Secondly, considering that existing methods struggle with perceiving global context due to limited receptive field of small-kernel convolution, we devise a context-orthogonal aggregation module that decomposes computationally expensive large-kernel depthwise convolution along channel dimension into a small square kernel, two orthogonal band kernels, and an identity mapping, enjoying large receptive field while maintaining efficiency. Thirdly, we deploy a motion field pyramid architecture that obtains and fuses multi-level motion fields, thereby facilitating the handling of the motion field's discontinuities in case of large scene disparity. Ultimately, we propose an elastic inference strategy that allows the model to introspect the confidence of its predictions at each layer, through which CorrNeXt is endowed with the flexibility of adaptively determining inference termination according to the difficulty of each image pair. Thorough experimentation affirms CorrNeXt's remarkable capabilities. Zizhuo Li, Chunbao Su, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
ACM Multimedia | 1 |
| 2025 | DeMatch++: Two-View Correspondence Learning via Deep Motion Field Decomposition and Respective Local-Context AggregationabstractTwo-view correspondence learning has increasingly focused on the coherence and smoothness of motion fields between image pairs. Conventional methods either regularize the complexity of the field function at substantial computational expense, or apply local filters that prove ineffective for large scene disparities. In this paper, we present DeMatch++, a novel network drawing inspiration from Fourier decomposition principles that decomposes the motion field to retain its primary "low-frequency" and smooth components. This approach achieves implicit regularization with lower computational overhead while exhibiting inherent piecewise smoothness. Specifically, our method decomposes the noise-contaminated motion field into multiple linearly independent basis vectors, generating smooth sub-fields that preserve the main energy of the original field. These sub-fields facilitate the recovery of a cleaner motion field for precise vector derivation. Within this framework, we aggregate local context within each sub-field while enhancing global information across all sub-fields. We also employ a masked decomposition strategy that mitigates the influence of false matches, and construct a compact representation to suppress redundant sub-fields. The complete pipeline is formulated as a discrete learnable architecture, circumventing the need for dense field computation. Extensive experiments demonstrate that DeMatch++ outperforms state-of-the-art methods while maintaining computational efficiency and piecewise smoothness. Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Learning Feature Matching via Matchable Keypoint-Assisted Graph Neural NetworkabstractAccurately matching local features between a pair of images corresponding to the same 3D scene is a challenging computer vision task. Previous studies typically utilize attention-based graph neural networks (GNNs) with fully-connected graphs over keypoints within/across images for visual and geometric information reasoning. However, in the background of local feature matching, a significant number of keypoints are non-repeatable due to factors like occlusion and failure of the detector, and thus irrelevant for message passing. The connectivity with non-repeatable keypoints not only introduces redundancy, resulting in limited efficiency (quadratic computational complexity w.r.t. the keypoint number), but also interferes with the representation aggregation process, leading to limited accuracy. Aiming at the best of both worlds on accuracy and efficiency, we propose MaKeGNN, a sparse attention-based GNN architecture which bypasses non-repeatable keypoints and leverages matchable ones to guide compact and meaningful message passing. More specifically, our Bilateral Context-Aware Sampling (BCAS) Module first dynamically samples two small sets of well-distributed keypoints with high matchability scores from the image pair. Then, our Matchable Keypoint-Assisted Context Aggregation (MKACA) Module regards sampled informative keypoints as message bottlenecks and thus constrains each keypoint only to retrieve favorable contextual information from intra- and inter-matchable keypoints, evading the interference of irrelevant and redundant connectivity with non-repeatable ones. Furthermore, considering the potential noise in initial keypoints and sampled matchable ones, the MKACA module adopts a matchability-guided attentional aggregation operation for purer data-dependent context propagation. By these means, MaKeGNN outperforms the state-of-the-arts on multiple highly challenging benchmarks, while significantly reducing computational and memory complexity compared to typical attentional GNNs. Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Seed to Prune: A Seeded Graph Neural Network for Two-View Correspondence LearningabstractWe present a simple yet tough-to-beat method dubbed SGNNet, for correspondence learning. Instead of focusing on devising sophisticated geometric extractors to explore the global or local contextual information involving all sparse correspondences as most existing studies have done, which may be biased by heavy outliers, we propose to first delve into elaborate contextual information encoded in several specific reliable correspondences, and later leverage it to achieve per-correspondence representation updating. To this end, the proposed network contains three pivotal modules: 1) dynamic seeding module, which aims to dynamically sample a set of reliable matches from the putative set as seeds to guide the network learning; 2) intraseed attention module (ISAM), which intends to capture the geometrical relations among seed matches and further leverage them to enhance seed features; and 3) dynamic unseeding module, which is designed to sufficiently aggregate favorable contextual information from seed matches and broadcast it back to features of original matches. With all the aforementioned components, the proposed SGNNet is capable of rejecting outliers from putative correspondences effectively. Extensive experiments indicate that our method beats current solid baselines and sets new SOTA scores across multiple domains and datasets. Notably, SGNNet attains an AUC@5° of 56.43% on YFCC100M without RANSAC, surpassing the most cutting-edge model by 4.51 absolute percentage points and exceeding the 55% AUC@5° bar for the first time. Project page: https://github.com/ZizhuoLi/SGNNet. Zizhuo Li, Jie Jiang 0015, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | DeMatch: Deep Decomposition of Motion Field for Two-View Correspondence LearningabstractTwo-view correspondence learning has recently focused on considering the coherence and smoothness of the motion field between an image pair. Dominant schemes include controlling the complexity of the field function with regularization or smoothing the field with local filters, but the former suffers from heavy computational burden, and the latter fails to accommodate discontinuities in the case of large scene disparities. In this paper, inspired by Fourier expansion, we propose a novel network called DeMatch, which decomposes the motion field to retain its main “low-frequency” and smooth part. This achieves implicit regularization with lower computational cost and generates piece-wise smoothness naturally. Specifically, we first decompose the rough motion field that is contaminated by false matches into several different sub-fields, which are highly smooth and contain the main energy of the original field. Then, with these smooth sub-fields, we recover a cleaner motion field from which correct motion vectors are subsequently derived. We also design a special masked decomposition strategy to further mitigate the negative influence of false matches. All the mentioned processes are finally implemented in a discrete and learnable manner, avoiding the difficulty of calculating real dense fields. Extensive experiments reveal that DeMatch outperforms state-of-the-art methods in multiple tasks and shows promising low computational usage and piecewise smoothness property. The code and trained models are publicly available at https://github.com/SuhZhang/DeMatch. Zizhuo Li, Yuan Gao 0015, Jiayi Ma 0001 |
CVPR | 2 |
| 2024 | U-Match: Exploring Hierarchy-Aware Local Context for Two-View Correspondence LearningabstractRejecting outlier correspondences is one of the critical steps for successful feature-based two-view geometry estimation, and contingent heavily upon local context exploration. Recent advances focus on devising elaborate local context extractors whereas typically adopting explicit neighborhood relationship modeling at a specific scale, which is intrinsically flawed and inflexible, because 1) severe outliers often populated in putative correspondences and 2) the uncertainty in the distribution of inliers and outliers make the network incapable of capturing adequate and reliable local context from such neighborhoods, therefore resulting in the failure of pose estimation. This prospective study proposes a novel network called U-Match that has the flexibility to enable implicit local context awareness at multiple levels, naturally circumventing the aforementioned issues that plague most existing studies. Specifically, to aggregate multi-level local context implicitly, a hierarchy-aware graph representation module is designed to flexibly encode and decode hierarchical features. Moreover, considering that global context always works collaboratively with local context, an orthogonal local-and-global information fusion module is presented to integrate complementary local and global context in a redundancy-free manner, thus yielding compact feature representations to facilitate correspondence learning. Thorough experimentation across relative pose estimation, homography estimation, visual localization, and point cloud registration affirms U-Match's remarkable capabilities. Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | A more reliable local-global-guided network for correspondence pruning
Chengli Peng, Zizhuo Li, Qiwen Jin |
Pattern Recognit. Lett. | 4 |
| 2024 | MC-Net: Integrating Multi-Level Geometric Context for Two-View Correspondence LearningabstractIn two-view correspondence learning, prevalent multi-layer perceptron (MLP)-based methods struggle with context capturing. To remedy this issue, recent advances innovatively stack convolutional neural network (CNN)-based Resblocks sequentially, showing an inherent proficiency in local context extraction. Yet, such non-issue-specific designs inherit the drawback of CNN’s difficulty in aggregating global context, leading to performance bottlenecks. To address this problem, this prospective study further explores the potential of the CNN-based framework and proposes MC-Net, a top-performing network that integrates both local and global context elegantly and seamlessly. Specifically, considering that sparse motion vectors and a dense motion field can be converted into each other through interpolation and sampling, we first transform unordered matches into image-structured data by estimating the dense motion field implicitly. Then, we design a hierarchical rectifying module to rectify the error of each ordered motion vector with CNN at multiple levels, enabling MC-Net to perceive global context from coarse-level features and local context from fine-level features simultaneously, which facilitates to tackle the discontinuities of the motion field in case of large scene disparity. Finally, we reconstruct comprehensive context-embedded features from rectified motion fields at all levels. Also, instead of using the residuals between rectified and pre-rectified motion vectors at the same layer to reject outliers as in previous studies, which seriously affects the inlier prediction accuracy, we rethink this operation meticulously and modify it to the difference between motion vectors obtained from each layer’s reconstruction and ones from the first layer before transformation, ensuring purer residuals and enhancing the matching performance without extra computational burden. Extensive experiments show that MC-Net outperforms state-of-the-arts on multiple domains and datasets. Zizhuo Li, Chunbao Su, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | U-Match: Two-view Correspondence Learning with Hierarchy-aware Local Context AggregationabstractLocal context capturing has become the core factor for achieving leading performance in two-view correspondence learning. Recent advances have devised various local context extractors whereas typically adopting explicit neighborhood relation modeling that is restricted and inflexible. To address this issue, we introduce U-Match, an attentional graph neural network that has the flexibility to enable implicit local context awareness at multiple levels. Specifically, a hierarchy-aware graph representation (HAGR) module is designed and fleshed out by local context pooling and unpooling operations. The former encodes local context by adaptively sampling a set of nodes to form a coarse-grained graph, while the latter decodes local context by recovering the coarsened graph back to its original size. Moreover, an orthogonal fusion module is proposed for the collaborative use of HAGR module, which integrates complementary local and global information into compact feature representations without redundancy. Extensive experiments on different visual tasks prove that our method significantly surpasses the state-of-the-arts. In particular, U-Match attains an AUC at 5 degree threshold of 60.53% on the challenging YFCC100M dataset without RANSAC, outperforming the strongest prior model by 8.61 absolute percentage points. Our code is publicly available at https://github.com/ZizhuoLi/U-Match. Zizhuo Li, Jiayi Ma 0001 |
IJCAI | 1 |
| 2023 | Density-Guided Incremental Dominant Instance Exploration for Two-View Geometric Model FittingabstractExisting two-view multi-model fitting methods typically follow a two-step manner, i.e., model generation and selection, without considering their interaction. Therefore, in the first step, these methods have to generate a considerable number of instances in order to cover all desired ones, which not only offers no guarantees, but also introduces unnecessary expensive calculations. To address this challenge, this study presents a new algorithm, termed as D2Fitting, that incrementally explores dominant instances. Particularly, rather than viewing model generation and selection as two disjoint parts, D2Fitting fully considers their interaction, and thus performs these two subroutines alternatively under a simple yet effective optimization framework. This design can avoid generating too many redundant instances, thus reducing computational overhead and allowing the proposed D2Fitting being real-time. Meanwhile, we further design a novel density-guided sampler to sample high-quality minimal subsets during the model generation process, so as to fully exploit the spatial distribution of the input data. Also, to mitigate the influence of noise on the subsets sampled by the proposed sampler, a global-residual optimization strategy is investigated for the minimal subset refinement. With all the ingredients mentioned above, the proposed D2Fitting can accurately estimate the number and parameters of geometric models and efficiently segment the input data simultaneously. Extensive experiments on several public datasets demonstrate the significant superiority of D2Fitting over several state-of-the-arts. Zizhuo Li, Jiayi Ma 0001, Guobao Xiao |
IEEE Trans. Image Process. | 1 |
| 2023 | Loop Closure Detection With Bidirectional Manifold Representation ConsensusabstractLoop closure detection (LCD) is an indispensable module in simultaneous localization and mapping. It is responsible to recognize pre-visited areas during the navigation of a robot, providing auxiliary information to revise pose estimation. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. Specifically, to avoid unnecessary memory costs, consecutive images are segmented into sequences as per the similarity of their global features. Then the sequence descriptor is incrementally inserted into hierarchical navigable small world for the construction of reference database, from which the most similar image for the query one is searched parallelly. To further identify whether the candidate pair is geometry-consistent, a feature matching method termed as bidirectional manifold representation consensus (BMRC) is proposed. It constructs local neighborhood structures of feature points via manifold representation, and formulates the matching problem into an optimization model, enabling linearithmic time complexity via a closed-form solution. Meanwhile, an accelerated version of it is introduced (BMRC*), which performs about 63% faster than BMRC in an image pair with 352 initial correspondences. Extensive experiments on nine publicly available datasets demonstrate that BMRC and BMRC* perform well in feature matching and the proposed pipeline has remarkable performance in the LCD task. Kaining Zhang, Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Guided neighborhood affine subspace embedding for feature matching
Zizhuo Li, Yong Ma 0001, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001 |
Pattern Recognit. | 1 |
| 2021 | Appearance-based Loop Closure Detection via Bidirectional Manifold Representation ConsensusabstractLoop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. When a query image occurs, we select candidates on-line according to the similarity of global semantic features in the first stage, and then conduct robust geometric confirmation to verify true loop-closing pairs in the second stage. To this end, a robust feature matching algorithm, termed as bidirectional manifold representation consensus (BMRC), is proposed. In particular, we utilize manifold representation to construct local neighborhood structures of feature points and formulate the matching problem into an optimization model, enabling linearithmic time complexity via a closed-form solution. Furthermore, we propose a dynamic place partition strategy based on BMRC to segment image streams with similar content into a place, which can mine more valid candidate frames, improving the recall rate of the whole system. Extensive experiments on several publicly available datasets reveal that BMRC has a good performance in the general feature matching task and the proposed pipeline outperforms the current state-of-the-art approaches in the LCD task. Kaining Zhang, Zizhuo Li, Jiayi Ma 0001 |
ICRA | 2 |