EDBT 2026 Demo / reviewers in the wild / expert
Mingqi Gao 0003
dblp:191/2698-3
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-8688-8228ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ContA-HOI: Towards Physically Plausible Human-Object Interaction Generation via Contact-Aware ModelingabstractIn human-object interaction (HOI), physical contact between the body and objects is a primary determinant of realism and plausibility. Prior HOI methods typically encode relations via global joint-to-centroid or joint-to-boundary. Such strategies neglect contact anchors that are essential for defining joint-to-contact relations-where and how HOI occurs, thereby implicitly reducing the problem to nearest-distance optimization. Without explicit contact anchors and joint-to-contact dynamics, previous models drift toward artifacts: human-object penetration or unnatural object floating. We argue that modeling contact relationships by contact anchors is important for generating realistic HOIs, as it directly captures where and how humans physically interact with objects rather than merely minimizing spatial proximity. To address these limitations, we propose Contact-Aware HOI (ContA-HOI), a progressive framework that decomposes HOI generation into three synergistic stages: discovering where contact occurs, modeling how contact evolves, and guiding generation with contact constraints. First, a Contact Affordance Predictor (CAP) addresses the “where” by predicting precise object-surface contact anchors from text, human pose, and object geometry. Second, these anchors seed a Contact Relation Field (CRF) that captures “how” by modeling spatiotemporal dynamics of joint-to-contact relations throughout the interaction. Finally, a Contact Dynamics Model (CDM) learns a prior CRF evolution pattern and guides motion diffusion sampling by aligning the generated motion's CRF with this learned prior. On the FullBodyManipulation dataset, ContA-HOI yields more realistic and physically plausible HOIs, improving foot sliding and contact percentage over recent baselines. Zhe Li 0008, Mingqi Gao 0003, Feng Zheng 0001 |
3DV | 3 |
| 2026 | SCORP: Scene-Consistent Object Refinement via Proxy Generation and TuningabstractViewpoint missing of objects is common in scene reconstruction, as camera paths typically prioritize capturing the overall scene structure rather than individual objects. This makes it highly challenging to achieve high-fidelity object-level modeling while maintaining accurate scene-level representation. Addressing this issue is critical for advancing downstream tasks requiring high-fidelity object reconstruction. In this paper, we introduce Scene-Consistent Object Refinement via Proxy Generation and Tuning (SCORP), a novel 3D enhancement framework that leverages 3D generative priors to recover fine-grained object geometry and appearance under missing views. Starting with proxy generation by substituting degraded objects using a 3D generation model, SCORP then progressively refines geometry and texture by aligning each proxy to its degraded counter-part in 7-DoF pose, followed by correcting spatial and appearance inconsistencies through registration-constrained enhancement. This two-stage proxy tuning ensures the high-fidelity geometry and appearance of the original object in unseen views while maintaining consistency in spatial positioning, observed geometry, and appearance. Across challenging benchmarks, SCORP achieves consistent gains over recent state-of-the-art baselines on both novel view synthesis and geometry completion tasks. SCORP is available at https://github.com/PolySummit/SCORP. Ziling Liu, Zitong Huang, Mingqi Gao 0003, Feng Zheng 0001 |
WACV | 4 |
| 2025 | Unlocking the Potential of Diffusion Priors in Blind Face RestorationabstractAlthough diffusion prior is rising as a powerful solution for blind face restoration (BFR), the inherent gap between the vanilla diffusion model and BFR settings hinders its seamless adaptation. The gap mainly stems from the discrepancy between 1) high-quality (HQ) and low-quality (LQ) images and 2) synthesized and real-world images. The vanilla diffusion model is trained on images with no or less degradations, whereas BFR handles moderately to severely degraded images. Additionally, LQ images used for training are synthesized by a naive degradation model with limited degradation patterns, which fails to simulate complex and unknown degradations in real-world scenarios. In this work, we use a unified network FLIPNET that switches between two modes to resolve specific gaps. In Restoration mode, the model gradually integrates BFR-oriented features and face embeddings from LQ images to achieve authentic and faithful face restoration. In Degradation mode, the model synthesizes real-world like degraded images based on the knowledge learned from real-world degradation datasets. Extensive evaluations on benchmark datasets show that our model 1) outperforms previous diffusion prior based BFR methods in terms of authenticity and fidelity, and 2) outperforms the naive degradation model in modeling the real-world degradations. Yunqi Miao, Zhiyu Qu, Mingqi Gao 0003, Changrui Chen, Jifei Song, Jungong Han, Jiankang Deng |
ICCV | 3 |
| 2025 | ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-Term Tracking
Haofeng Liu, Mingqi Gao 0003, Xuxiao Luo, Ziyue Wang 0005, Guanyi Qin, Yueming Jin |
MICCAI (10) | 2 |
| 2025 | OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference OptimizationabstractAutomatic indoor layout generation has attracted increasing attention due to its potential in interior design, virtual environment construction, and embodied AI. Existing methods fall into two categories: prompt-driven approaches that leverage proprietary LLM services (e.g., GPT APIs), and learning-based methods trained on layout data upon diffusion-based models. Prompt-driven methods often suffer from spatial inconsistency and high computational costs, while learning-based methods are typically constrained by coarse relational graphs and limited datasets, restricting their generalization to diverse room categories. In this paper, we revisit LLM-based indoor layout generation and present 3D-SynthPlace, a large-scale dataset that combines synthetic layouts generated via a `GPT synthesize, Human inspect' pipeline, upgraded from the 3D-Front dataset. 3D-SynthPlace contains nearly 17,000 scenes, covering four common room types—bedroom, living room, kitchen, and bathroom—enriched with diverse objects and high-level spatial annotations. We further introduce OptiScene, a strong open-source LLM optimized for indoor layout generation, fine-tuned based on our 3D-SynthPlace dataset through our two-stage training. For the warum-up stage I, we adopt supervised fine-tuning (SFT), which is taught to first generate high-level spatial descriptions then conditionally predict concrete object placements. For the reinforcing stage II, to better align the generated layouts with human design preferences, we apply multi-turn direct preference optimization (DPO), which significantly improving layout quality and generation success rates. Extensive experiments demonstrate that OptiScene outperforms traditional prompt-driven and learning-based baselines. Moreover, OptiScene shows promising potential in interactive tasks such as scene editing and robot navigation, highlighting its applicability beyond static layout generation. Tongsheng Ding, Junru Lu, Mingqi Gao 0003, Victor Sanchez, Feng Zheng 0001 |
NeurIPS | 5 |
| 2025 | Few-Shot Referring Video Single- and Multi-Object Segmentation Via Cross-Modal Affinity with Instance Sequence Matching
Heng Liu 0002, Mingqi Gao 0003, Xiantong Zhen, Feng Zheng 0001, Yang Wang 0023 |
Int. J. Comput. Vis. | 3 |
| 2024 | Place Anything into Any Video
Ziling Liu, Mingqi Gao 0003, Feng Zheng 0001 |
IJCAI | 3 |
| 2024 | Unveiling the Power of Visible-Thermal Video Object SegmentationabstractDespite recent progress, Video Object Segmentation (VOS) remains challenging in complex situations such as low light and dark scenes. In this paper, we tackle the visibility limitations by introducing thermal information as auxillary for VOS. Specifically, we generate a hybrid benchmark dataset for Visible-Thermal VOS, named VisT300, which contains 300 challenging videos with visible light and thermal frames and corresponding object mask annotations. Besides, a Visible-Thermal integration Network, named as VTiNet, is proposed to use both cross-modal and cross-frame propagation for accurate video object segmentation. It is advantageous in two aspects: 1) effective cross-modal feature fusion and propagation for strong expressions on visible, thermal, and fused modalities; 2) effective modality-sensitive memory bank enables preserving the most valuable historical contexts in each modality. Extensive experiments demonstrate our VTiNet outperforms the state-of-the-art VOS works by a large margin (over 5% than RGB SotAs in Mean J&F). Our preliminary research clearly recovers that importing complementary modalities can effectively increase the strength of models to achieve robust segmentation in challenging scenarios. Data and code are released at https://github.com/yjybuaa/vtinet, and we hope this work will promote the progress of visible-thermal VOS. Mingqi Gao 0003, Runmin Cong, Chengjie Wang 0001, Feng Zheng 0001, Ales Leonardis |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Weakly-Supervised RGBD Video Object SegmentationabstractDepth information opens up new opportunities for video object segmentation (VOS) to be more accurate and robust in complex scenes. However, the RGBD VOS task is largely unexplored due to the expensive collection of RGBD data and time-consuming annotation of segmentation. In this work, we first introduce a new benchmark for RGBD VOS, named DepthVOS, which contains 350 videos (over 55k frames in total) annotated with masks and bounding boxes. We futher propose a novel, strong baseline model - Fused Color-Depth Network (FusedCDNet), which can be trained solely under the supervision of bounding boxes, while being used to generate masks with a bounding box guideline only in the first frame. Thereby, the model possesses three major advantages: a weakly-supervised training strategy to overcome the high-cost annotation, a cross-modal fusion module to handle complex scenes, and weakly-supervised inference to promote ease of use. Extensive experiments demonstrate that our proposed method performs on par with top fully-supervised algorithms. We will open-source our project on https://github.com/yjybuaa/depthvos/ to facilitate the development of RGBD VOS. Mingqi Gao 0003, Feng Zheng 0001, Xiantong Zhen, Rongrong Ji, Ling Shao 0001, Ales Leonardis |
IEEE Trans. Image Process. | 2 |
| 2023 | Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited SamplesabstractReferring video object segmentation (RVOS), as a supervised learning task, relies on sufficient annotated data for a given scene. However, in more realistic scenarios, only minimal annotations are available for a new scene, which poses significant challenges to existing RVOS methods. With this in mind, we propose a simple yet effective model with a newly designed cross-modal affinity (CMA) module based on a Transformer architecture. The CMA module builds multimodal affinity with a few samples, thus quickly learning new semantic information, and enabling the model to adapt to different scenarios. Since the proposed method targets limited samples for new scenes, we generalize the problem as - few-shot referring video object segmentation (FS-RVOS). To foster research in this direction, we build up a new FS-RVOS benchmark based on currently available datasets. The benchmark covers a wide range and includes multiple situations, which can maximally simulate real-world scenarios. Extensive experiments show that our model adapts well to different scenarios with only a few samples, reaching state-of-the-art performance on the benchmark. On Mini-Ref-YouTube-VOS, our model achieves an average performance of 53.1 ${\mathcal{J}}$ and 54.8 ${\mathcal{F}}$, which are 10% better than the baselines. Furthermore, we show impressive results of 77.7 ${\mathcal{J}}$ and 74.8 ${\mathcal{F}}$ on Mini-Ref-SAIL-VOS, which are significantly better than the baselines. Code is publicly available at https://github.com/hengliusky/Few_shot_RVOS. Mingqi Gao 0003, Heng Liu 0002, Xiantong Zhen, Feng Zheng 0001 |
ICCV | 2 |
| 2023 | Filter pruning with uniqueness mechanism in the frequency domain for efficient neural networks
Mingqi Gao 0003, Qiang Ni, Jungong Han |
Neurocomputing | 2 |
| 2023 | Video Object Segmentation using Point-based Memory NetworkabstractRecent years have witnessed the prevalence of memory-based methods for Semi-supervised Video Object Segmentation (SVOS) which utilise past frames efficiently for label propagation. When conducting feature matching, fine-grained multi-scale feature matching has typically been performed using all query points, which inevitably results in redundant computations and thus makes the fusion of multi-scale results ineffective. In this paper, we develop a new Point-based Memory Network, termed as PMNet, to perform fine-grained feature matching on hard samples only, assuming that easy samples can already obtain satisfactory matching results without the need for complicated multi-scale feature matching. Our approach first generates an uncertainty map from the initial decoding outputs. Next, the fine-grained features at uncertain locations are sampled to match the memory features on the same scale. Finally, the matching results are further decoded to provide a refined output. The point-based scheme works with the coarsest feature matching in a complementary and efficient manner. Furthermore, we propose an approach to adaptively perform global or regional matching based on the motion history of memory points, making our method more robust against ambiguous backgrounds. Experimental results on several benchmark datasets demonstrate the superiority of our proposed method over state-of-the-art methods. Mingqi Gao 0003, Jungong Han, Feng Zheng 0001, James Jian Qiao Yu, Giovanni Montana |
Pattern Recognit. | 1 |
| 2023 | Decoupling Multimodal Transformers for Referring Video Object SegmentationabstractReferring Video Object Segmentation (RVOS) aims to segment the text-depicted object from video sequences. With excellent capabilities in long-range modelling and information interaction, transformers have been increasingly applied in existing RVOS architectures. To better leverage multimodal data, most efforts focus on the interaction between visual and textual features. However, they ignore the syntactic structures of the text during the interaction, where all textual components are intertwined, resulting in ambiguous vision-language alignment. In this paper, we improve the multimodal interaction by DECOUPLING the interweave. Specifically, we train a lightweight subject perceptron, which extracts the subject part from the input text. Then, the subject and text features are fed into two parallel branches to interact with visual features. This enables us to perform subject-aware and context-aware interactions, respectively, thus encouraging more explicit and discriminative feature embedding and alignment. Moreover, we find the decoupled architecture also facilitates incorporating the vision-language pre-trained alignment into RVOS, further improving the segmentation performance. Experimental results on all RVOS benchmark datasets demonstrate the superiority of our proposed method over the state-of-the-arts. The code of our method is available at:https://github.com/gaomingqi/dmformer. Mingqi Gao 0003, Jungong Han, Ke Lu 0002, Feng Zheng 0001, Giovanni Montana |
IEEE Trans. Circuits Syst. Video Technol. | 1 |