VLDB 2026 Research / reviewers in the wild / expert
Zuheng Ming
dblp:46/9230
· DBLP profile ↗
20ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-1094-3112ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NVC-GS: Monocular Dynamic Scene Reconstruction via Normal-Regularized and Multi-View-Consistent 3D Gaussian SplattingabstractHigh-quality, real-time dynamic scene reconstruction and rendering are essential for immersive applications. While techniques like 3D Gaussian Splatting (3DGS) succeed in static scenes, dynamic monocular scenarios still suffer from deformation, surface noise, and inconsistent view-dependent effects due to insufficient geometric constraints and inadequate multi-view consistency enforcement for monocular input. To address these challenges, we present Normal-Regularized and Multi-View-Consistent Gaussian Splatting (NVC-GS), a novel approach combining geometry-aware normal regularization with diffusion-based multi-view consistency. Our method explicitly preserves geometric consistency in dynamic objects through tailored normal constraints, while leveraging diffusion-driven latent space regularization to ensure cross-view rendering consistency, particularly for complex materials such as reflective surfaces. Experimental results demonstrate that our approach effectively improves geometric accuracy and visual quality in dynamic scenes while maintaining real-time capabilities, outperforming existing methods in terms of deformation handling, surface noise reduction, and rendering of reflective materials. Huiwen Xue, Kaixing Zhao, Tingcheng Li, Zuheng Ming |
3DV | 5 |
| 2026 | Tri-HGNet: A feature-driven dynamic hypergraph framework for medical image segmentation
Xiaoyan Kui, Lingxiao Liu, Qinsong Li, Haonan Yan, Weixin Si, Zuheng Ming, Beiji Zou 0001 |
Neurocomputing | 6 |
| 2025 | Distance-Aware and Knowledge-Driven Vision Mamba U-Net for Radiotherapy Dose PredictionabstractDose planning is essential in radiotherapy for cancer patients, yet current practice relies on iterative manual optimization, underscoring the need for automated prediction. Existing deep learning approaches remain limited because they often ignore the 3D spatial relationships between tumors and surrounding organs at risk (OARs), and clinical priors on safe dose thresholds. To overcome these limitations, we propose DKVMU-Net, a distance-aware and knowledge-driven Vision Mamba U-Net for automated dose prediction. Our framework incorporates Vision Mamba blocks to capture global, long-range dependencies from CT scans and OAR signed distance field (SDF) maps, which naturally encode spatial information. Additionally, we introduce a deformable dynamic feature enhancement module (DDFEM) for texture refinement, followed by a linear crossattention fusion module to improve cross-modality integration. A customized loss function is also designed to incorporate prior knowledge of OAR dose constraints, ensuring optimal target coverage and OAR protection. To alleviate the scarcity of doseplanning datasets, we collect an in-house radiotherapy lung cancer dataset (RLCD), consisting of CT volumes, OAR masks, and corresponding SDF maps from 116 patients. We evaluate our DKVMU-Net on both the in-house dataset and public available OpenKBP dataset. Compared with the sate-of-the-art method, our approach achieves an 11.6 % improvement in dose score (1.641 vs. 1.857) and 26.3 % in DVH score (6.481 vs. 8.799) on RLCD, and a 7.8 % improvement in dose score (2.421 vs. 2.626) and 13.9 % in DVH score (1.057 vs. 1.227) on OpenKBP. These results demonstrate the robustness and effectiveness of our approach. Yangyang Shi, Xiaoyan Kui, Yucong Zhang, Shihao Zou, Zuheng Ming, Weixin Si, Azeddine Beghdadi, Beiji Zou 0001 |
BIBM | 5 |
| 2025 | From Global to Local: Mamba-Based Hierarchical Registration for Respiratory Lung DeformationabstractDeformable image registration is essential in medical applications, as accurately estimating organ displacements across respiratory phases enables precise radiation dose planning in dynamic environments, mitigates damage to organs at risk (OARs), and thus improves patients' health-related quality of life. Although current learning-based methods have achieved impressive performance in small deformation registration, challenges remain due to their limited ability to capture large deformations occurring during respiration. To address this issue, we propose a novel Mamba-based hierarchical registration framework that effectively extracts both global and local features for accurate deformation prediction. Specifically, given a pair of source and target 3DCT volumes, we incorporate a foundation model pretrained on medical image registration tasks to enhance alignment accuracy. We further propose a directional-deformable Mamba scheme to facilitate global context extraction and local motion awareness. The directional Mamba component scans input features from multiple orientations to achieve broad contextual perception, while the deformable Mamba module employs adaptive directional scanning strategies to capture dynamic local variations. To overcome the scarcity of annotated respiratory data, we also collect a new respiratory lung cancer dataset comprising 100 annotated phases from 20 patients. Experimental results on our in-house dataset demonstrate that our method outperforms state-of-the-art approaches, achieving a 1.3 % improvement in overall Dice accuracy and a 1.6 dB increase in PSNR, underscoring its strong potential for clinical deployment. Code and test data are available at: https://github.com/yangyangshi806/Mamba_based_Registration. Yangyang Shi, Yucong Zhang, Beiji Zou 0001, Xiaoyan Kui, Zexin Ji, Zuheng Ming, Azeddine Beghdadi, Weixin Si |
BIBM | 6 |
| 2025 | GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and ClassificationabstractVisual document understanding (VDU) has rapidly advanced with the development of powerful multi-modal language models. However, these models typically require extensive document pre-training data to learn intermediate representations and often suffer a significant performance drop in real-world online industrial settings. A primary issue is their heavy reliance on OCR engines to extract local positional information within document pages, which limits the models' ability to capture global information and hinders their generalizability, flexibility, and robustness. In this paper, we introduce GlobalDoc, a cross modal transformer-based architecture pre-trained in a self supervised manner using three novel pretext objective tasks. GlobalDoc improves the learning of richer semantic concepts by unifying language and visual representations, resulting in more transferable models. For proper evaluation, we also propose two novel document-level downstream VDU tasks, Few-Shot Document Image Classification (DIC) and Content-based Document Image Retrieval (DIR), designed to simulate industrial scenarios more closely. Extensive experimentation has been conducted to demonstrate GlobalDoc's effectiveness in practical settings. Souhail Bakkali, Sanket Biswas, Zuheng Ming, Mickaël Coustaty, Marçal Rusiñol, Oriol Ramos Terrades, Josep Lladós 0001 |
WACV | 3 |
| 2025 | PK-Net: A prior knowledge-driven dual-path network for enhanced glaucoma screening
Xiaoyan Kui, Zeru Hai, Beiji Zou 0001, Yang Li 0111, Wei Liang 0005, Zuheng Ming, Liming Chen 0002 |
Knowl. Based Syst. | 6 |
| 2024 | Multimodal Transformer Using Cross-Channel Attention For Object Detection In Remote Sensing ImagesabstractObject detection in Remote Sensing Images (RSI) is a critical task for numerous applications in Earth Observation (EO). Differing from object detection in natural images, object detection in remote sensing images faces challenges of scarcity of annotated data and the presence of small objects represented by only a few pixels. Multi-modal fusion has been determined to enhance the accuracy by fusing data from multiple modalities such as RGB, infrared (IR), lidar, and synthetic aperture radar (SAR). To this end, the fusion of representations at the mid or late stage, produced by parallel subnetworks, is dominant, with the disadvantages of increasing computational complexity in the order of the number of modalities and the creation of additional engineering obstacles. Using the cross-attention mechanism, we propose a novel multi-modal fusion strategy for mapping relationships between different channels at the early stage, enabling the construction of a coherent input by aligning the different modalities. By addressing fusion in the early stage, as opposed to mid or late-stage methods, our method achieves competitive and even superior performance compared to existing techniques. Additionally, we enhance the SWIN transformer by integrating convolution layers into the feed-forward of non-shifting blocks. This augmentation strengthens the model’s capacity to merge separated windows through local attention, thereby improving small object detection. Extensive experiments prove the effectiveness of the proposed multimodal fusion module and the architecture, demonstrating their applicability to object detection in multimodal aerial imagery. Our code is available at here. Bissmella Bahaduri, Zuheng Ming, Fangchen Feng, Anissa Zergaïnoh-Mokraoui |
ICIP | 2 |
| 2024 | Identifying fraudulent identity documents by analyzing imprinted guilloche patterns
Musab Al-Ghadi, Tanmoy Mondal, Zuheng Ming, Petra Gomez-Krämer, Mickaël Coustaty, Nicolas Sidere, Jean-Christophe Burie |
Multim. Tools Appl. | 3 |
| 2023 | Guilloche Detection for ID Authentication: A Dataset and BaselinesabstractIn cases of digital enrolment via mobile and online services, identity documents (IDs) verification is critical to efficiently detect forgery and therefore build user trust in the digital world. In this paper, we propose a copy-move public dataset, called FMIDV (forged mobile ID video dataset) containing forged IDs with respect to guilloche patterns. Also, we propose two fraud detection models on guilloche patterns of IDs, which are based on contrastive and adversarial learning. In the sequel, each proposed model manages to read the entire ID and to recognize the guilloche pattern to check its similarity to the pattern of an authentic ID. The objective of the similarity check is to validate its authenticity or its rejection. Experiments are conducted on MIDV and FMIDV datasets to analyze and identify the most proper parameters to achieve higher authentication performance. The code and the dataset are available at https://github.com/malghadi/CheckID. Musab Al-Ghadi, Zuheng Ming, Petra Gomez-Krämer, Jean-Christophe Burie, Mickaël Coustaty, Nicolas Sidere |
MMSP | 2 |
| 2023 | VLCDoC: Vision-Language contrastive pre-training model for cross-Modal document classification
Souhail Bakkali, Zuheng Ming, Mickaël Coustaty, Marçal Rusiñol, Oriol Ramos Terrades |
Pattern Recognit. | 2 |
| 2022 | Vitranspad: Video Transformer Using Convolution And Self-Attention For Face Presentation Attack DetectionabstractFace Presentation Attack Detection (PAD) is an important measure to prevent spoof attacks for face biometric systems. Many works based on Convolution Neural Networks (CNNs) for face PAD formulate the problem as an image-level binary classification task without considering the context. Alternatively, Vision Transformers (ViT) using self-attention to attend the context of an image become the mainstreams in face PAD. Inspired by ViT, we propose a Video-based Transformer for face PAD (ViTransPAD) with short/long-range spatio-temporal attention which can not only focus on local details with short-range attention within a frame but also capture long-range dependencies over frames. Instead of using coarse image patches with single-scale as in ViT, we pro-pose the Multi-scale Multi-Head Self-Attention (MsMHSA) module to accommodate multi-scale patch partitions of Q, K, V feature maps to different heads on a single transformer in a coarse-to-fine manner, which enables to learn a fine-grained representation to perform pixel-level discrimination for face PAD. Due to lack inductive biases of convolutions in pure transformers, we also introduce convolutions to our ViTransPAD to integrate the desirable properties of CNNs. The extensive experiments show the effectiveness of our proposed ViTransPAD with a preferable accuracy-computation balance, which can serve as a new backbone for face PAD. Zuheng Ming, Zitong Yu, Musab Al-Ghadi, Muriel Visani, Muhammad Muzzamil Luqman, Jean-Christophe Burie |
ICIP | 1 |
| 2022 | Exploring multi-tasking learning in document attribute classification
Tanmoy Mondal, Zuheng Ming |
Pattern Recognit. Lett. | 3 |
| 2021 | EAML: ensemble self-attention-based mutual learning network for document image classification
Souhail Bakkali, Zuheng Ming, Mickaël Coustaty, Marçal Rusiñol |
Int. J. Document Anal. Recognit. | 2 |
| 2021 | Cross-modal photo-caricature face recognition based on dynamic multi-task learning
Zuheng Ming, Jean-Christophe Burie, Muhammad Muzzamil Luqman |
Int. J. Document Anal. Recognit. | 1 |
| 2020 | Cross-Modal Deep Networks For Document Image ClassificationabstractAs a fundamental step of document related tasks, document classification has been widely adopted to various document image processing applications. Unlike the general image classification problem in the computer vision field, text document images contain both the visual cues and the corresponding text within the image. However, how to bridge these two different modalities and leverage textual and visual features to classify text document images remains challenging. In this paper, we present a cross-modal deep network that enables to capture both the textual content and the visual information included in document images. Thanks to the efficient jointly learning of text and image features, the proposed cross-modal approach shows its superiority to the state-of-the-art single-modal methods. In this paper, we propose to use NASNet-Large and Bert to extract image and text features respectively. Experimental results demonstrate that the proposed cross-modal approach achieves new state-of-the-art results for text document image classification on the benchmark Tobacco-3482 dataset, outperforming the current state-of-the-art method by 3.91% of classification accuracy. Souhail Bakkali, Zuheng Ming, Mickaël Coustaty, Marçal Rusiñol |
ICIP | 2 |
| 2019 | Classification of Hyperspectral and Lidar with Deep Rotation ForestabstractIn this work, a novel deep rotation forest is proposed to fuse hyperspectral (HS) and LiDAR. First, we extract the spatial and elevation information of two datasets by using morphological filters. Then, each feature source is applied to superpixel segmentation and then are treated as the input of deep rotation forest. In the deep rotation forest, the spatial relationships are fully considered, and the output probability of each layer is used as the input of the next layer. Experimental results demonstrate that the excellent performance of the proposed method. Junshi Xia, Zuheng Ming |
ICASSP | 2 |
| 2018 | FaceLiveNet: End-to-End Networks Combining Face Verification with Interactive Facial Expression-Based Liveness DetectionabstractThe effectiveness of the state-of-the-art face verifi-cation/recognition algorithms and the convenience of face recognition greatly boost the face-related biometric authentication applications. However, existing face verification architectures seldom integrate any liveness detection or keep such stage isolated from face verification as if it was irrelevant. This may potentially result in the system being exposed to spoof attacks between the two stages. This work introduces FaceLiveNet, a holistic end-to-end deep networks which can perform face verification and liveness detection simultaneously. An interactive scheme for facial expression recognition is proposed to perform liveness detection, providing better generalization capacity and higher security level. The proposed framework is low-cost as it relies on commodity hardware instead of costly sensors, and lightweight with much fewer parameters comparing to the other popular deep networks such as VGG16 and FaceNet. Experimental results on the benchmarks LFW, YTF, CK+, OuluCASIA, SFEW, FER2013 demonstrate that the proposed FaceLiveNet can achieve state-of-art performance or better for both face verification and facial expression recognition. We also introduce a new protocol to evaluate the global performance for face authentication with the fusion of face verification and interactive facial expression-based liveness detection. Zuheng Ming, Joseph Chazalon, Muhammad Muzzamil Luqman, Muriel Visani, Jean-Christophe Burie |
ICPR | 1 |
| 2018 | Multiple Sources Data Fusion Via Deep ForestabstractIn this paper, we propose to fuse multiple sources remotely sensed datasets, such as hyperspectral (HS) and Light Detection and Ranging (LiDAR)-derived digital surface model (DSM) using a novel deep learning method. Morphological openings and closings with partial reconstruction are taken into account to model spatial and elevation information for both sources. Then, the stacked features directly input to a deep learning classifier, namely Deep Forest (DF). In particular, Deep Forest can be viewed as the cascade or the ensembles of Rotation Forests (RoF) and Random Forests (RF). We applied the proposed method to the datasets obtained from Tama forest, Japan. Experimental results demonstrate that Deep Forest can achieve better classification results than other approaches. Compared to deep neural networks, deep forest pays little effort in parameter tuning and has a significant reduction in computational complexity. Junshi Xia, Zuheng Ming, Akira Iwasaki |
IGARSS | 2 |
| 2015 | Synthetic Evidential Study as Augmented Collective Thought Process - Preliminary Report
Toyoaki Nishida, Masakazu Abe, Takashi Ookaki, Divesh Lala, Sutasinee Thovutikul, Hengjie Song, Yasser Mohammad, Christian Nitschke, Yoshimasa Ohmoto, Atsushi Nakazawa, Takaaki Shochi, Jean-Luc Rouas, Aurélie Bugeau, Fabien Lotte, Zuheng Ming, Geoffrey Letournel, Marine Guerry, Dominique Fourer |
ACIIDS (1) | 15 |
| 2010 | Estimation of speech lip features from discrete cosinus transformabstractInternational audience Zuheng Ming, Denis Beautemps, Gang Feng 0002, Sébastien Schmerber |
INTERSPEECH | 1 |