VLDB 2026 Research / reviewers in the wild / expert
Mengkai Li
dblp:338/3707
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
diffusion bridge |
0.9 | 1 | 2025 | Stabilizing Holistic Semantics in Diffusion Bridge for Image Inpainting · IJCAI 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Stabilizing Holistic Semantics in Diffusion Bridge for Image Inpainting · IJCAI 2025 |
Machine learning › Generative modeling › diffusion model › image restoration
image inpainting |
0.9 | 1 | 2025 | Stabilizing Holistic Semantics in Diffusion Bridge for Image Inpainting · IJCAI 2025 |
Image and video processing › image restoration
image inpainting |
0.9 | 1 | 2025 | Stabilizing Holistic Semantics in Diffusion Bridge for Image Inpainting · IJCAI 2025 |
Methods — techniques the papers use, named apart from their topics
structure restorer · 1.7semantic fusion schedule · 1.7posterior sampling · 1.7diffusion bridge · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Sequential Context Modeling for High-Fidelity Image InpaintingabstractImage inpainting aims to restore missing regions by leveraging surrounding spatial context, where nearby pixels provide crucial structural cues and distant regions offer complementary semantic guidance. To jointly model these complementary dependencies, this paper proposes Hierarchical Sequential Context Modeling (HSCM), a novel inpainting framework that employs state-space models for multi-scale autoregressive sequence modeling. Unlike existing single-scale SSM-based approaches, HSCM explicitly separates pixel-level and semanticlevel modeling into two complementary branches. The Local Perception Unit preserves fine-grained textures, and the Global Compensation Unit propagates high-level semantics across patches to enhance overall coherence. The asynchronous hierarchical design first reconstructs local textures and then performs semantic compensation, achieving notable performance gains with minimal computational overhead. Leveraging its four-directional architecture, HSCM maintains linear computational growth with spatial resolution and effectively establishes a comprehensive global receptive field. Furthermore, a Cross-Gated Feedforward Network is proposed to alleviate patch boundary artifacts and enhance inter-channel feature consistency. Built upon a multi-scale encoder–decoder architecture, HSCM delivers state-of-the-art inpainting quality and robust generalization across diverse benchmarks, including CelebA-HQ, FFHQ, Paris Street View, and Places2. Zexuan Sun, Jinjia Peng, Mengkai Li, Huibing Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Stabilizing Holistic Semantics in Diffusion Bridge for Image InpaintingabstractImage inpainting aims to restore the original image from a damaged version. Recently, a special type of diffusion bridge model has achieved promising performance by directly mapping the degradation process and restoring corrupted images through the corresponding reverse process. However, due to the lack of explicit semantic priors during the denoising process, the inpainted results typically exhibit inferior context-stability and semantic consistency. To this end, this paper proposes a novel Global Structure-Guided Diffusion Bridge framework (GSGDiff), which incorporates an additional structure restorer to stabilize the generation of holistic semantics. Specifically, to acquire richer semantic structure priors, this paper proposes a posterior sampling approach that captures semantically global and consistent structures at each timestep, efficiently integrating them into the texture generation through the corresponding guidance module. Additionally, considering the characteristics of diffusion models with low denoising levels at larger timesteps, this paper proposes a semantic fusion schedule to avoid noise interference by reducing the weight of ineffective guided semantics in the early stages. By applying the proposed posterior sampling to the texture denoising process, GSGDiff can achieve more stable and superior inpainting results over competitive baselines. Experiments on Places2, Paris Street View and CelebA-HQ datasets validate the efficacy of the proposed method. Jinjia Peng, Mengkai Li, Huibing Wang |
IJCAI | 2 |
| 2025 | Omni Contextual Aggregation Networks for High-Fidelity Image InpaintingabstractImage inpainting aims to restore a realistic image from a damaged or incomplete version. Although Transformer-based methods have achieved impressive results by modeling long-range dependencies, the inherent quadratic complexity of canonical self-attention has typically led to these approaches adopting uni-dimensional modeling, which limits the model’s ability to capture complex relationships from both spatial and channel dimensions. To this end, this paper exploits a novel attention paradigm termed Dynamic Omni-Attention Mechanism (DOAM) for simultaneously modeling pixel-interaction from both spatial and channel dimensions, and implements the information interaction across the omni-axis (i.e., spatial and channel) with linear computational complexity. In addition, to handle large-scale degradation, this paper proposes a Multi-band Feature Enhancement (MFE) module to enhance feature representation in downsampling, thus unlocking the potential of subsequent attentional interactions. Moreover, motivated by recent advances in image restoration, this paper incorporates a domain-related prior representation from CNN-based Network to modulate the features during proposed attention mechanism and feed-forward networks. Integrating the above designs into an encoder-decoder architecture, the proposed Omni Contextual Aggregation Networks (OCANet) achieve superior performance at lower parameters and time costs than the competitive baselines. Extensive experiments on CelebA-HQ, Paris Street View, FFHQ and Dunhuang datasets validate the efficacy of the proposed method. Jinjia Peng, Mengkai Li, Huibing Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | A review of vision-based road detection technology for unmanned vehiclesabstractWith the development of unmanned vehicle technology, unmanned vehicles have played a huge role in logistics transportation, emergency rescue and disaster relief, etc., so the research on unmanned vehicles is becoming more and more important. Road detection is an important part of environmental perception and an important factor in the realization of assisted driving and unmanned driving technology. High-precision road detection technology can provide important environmental information for efficient planning and reasonable decision-making of unmanned vehicles. Firstly, the technical framework of road detection is given, and the road detection process is introduced in detail. Then, the vision-based road detection algorithm is introduced. Finally, some related data sets in the field of road detection are collected, which provides new ideas and methods for road detection researchers. Chaoyang Liu, Qi Liu 0020, Fan Yang 0098, Mengkai Li |
IV | 6 |
| 2022 | Lidar-only 3D SLAM System Comparative StudyabstractSimultaneous localization and mapping (SLAM) is an attractive and hot research topic in computer vision, robotics, and artificial intelligence. Autonomous vehicles driving in unknown environments try to perceive and map the surrounding environment while recognizing their location and trajectory. In this paper, five state-of-the-art open-source 3D lidar-only SLAM algorithms are reviewed: LOAM, LeGO-LOAM, F-LOAM, BALM, and MULLS. We briefly introduce the characteristics of these algorithms. Finally, the experimental comparison is carried out to compare the absolute pose error (APE), efficiency, and operation memory occupation of each algorithm. Wenhu Ren, Mengkai Li, Qi Liu 0020 |
ICARCV | 3 |