EDBT 2026 Demo / reviewers in the wild / expert
Zimeng Li 0001
dblp:168/2865-1
· DBLP profile ↗
15ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-2798-3134ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech SynthesisabstractHigh-fidelity character voice synthesis is a cornerstone of immersive multimedia applications, particularly for interacting with anime avatars and digital humans. However, existing systems struggle to maintain consistent persona traits across diverse emotional contexts. To bridge this gap, we present ATRIE, a unified framework utilizing a Persona-Prosody Dual-Track (P2-DT) architecture. Our system disentangles generation into a static Timbre Track (via Scalar Quantization) and a dynamic Prosody Track (via Hierarchical Flow-Matching), distilled from a 14B LLM teacher. This design enables robust identity preservation (Zero-Shot Speaker Verification EER: 0.04) and rich emotional expression. Evaluated on our extended AnimeTTS-Bench (50 characters), ATRIE achieves state-of-the-art performance in both generation and cross-modal retrieval (mAP: 0.75), establishing a new paradigm for persona-driven multimedia content creation. The code is available at Github. Aoduo Li, Hongjian Xu, Shengmin Li, Sihao Qin, Zimeng Li 0001, Chi-Man Pun, Xuhang Chen 0002 |
ICMR | 6 |
| 2025 | FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI TranslationabstractUltra-high-field 7T MRI offers enhanced spatial resolution and tissue contrast that enable the detection of subtle pathological changes in neurological disorders. However, the limited availability of 7T scanners restricts widespread clinical adoption due to substantial infrastructure costs and technical demands. Computational approaches for synthesizing 7T-quality images from accessible 3T acquisitions present a viable solution to this accessibility challenge. Existing CNN approaches suffer from limited spatial coverage, while Transformer models demand excessive computational overhead. RWKV architectures offer an efficient alternative for global feature modeling in medical image synthesis, combining linear computational complexity with strong long-range dependency capture. Building on this foundation, we propose Frequency Spatial-RWKV (FS-RWKV), an RWKV-based framework for 3T-to-7T MRI translation. To better address the challenges of anatomical detail preservation and global tissue contrast recovery, FS- RWKV incorporates two key modules: (1) Frequency-Spatial Omnidirectional Shift (FSO-Shift), which performs discrete wavelet decomposition followed by omnidirectional spatial shifting on the low-frequency branch to enhance global contextual representation while preserving high-frequency anatomical details; and (2) Structural Fidelity Enhancement Block (SFEB), a module that adaptively reinforces anatomical structure through frequency-aware feature fusion. Comprehensive experiments on UNC and BNU datasets demonstrate that FS- RWKV consistently outperforms existing CNN-, Transformer-, GAN-, and RWKV-based baselines across both T1 wand T2w modalities, achieving superior anatomical fidelity and perceptual quality. Yingtie Lei, Zimeng Li 0001, Yupeng Liu 0003, Xuhang Chen 0002, Chi-Man Pun |
BIBM | 2 |
| 2025 | DTEA: Dynamic Topology Weaving and Instability-Driven Entropic Attenuation for Medical Image SegmentationabstractIn medical image segmentation, skip connections are used to merge global context and reduce the semantic gap between encoder and decoder. Current methods often struggle with limited structural representation and insufficient contextual modeling, affecting generalization in complex clinical scenarios. We propose the DTEA model, featuring a new skip connection framework with the Semantic Topology Reconfiguration (STR) and Entropic Perturbation Gating (EPG) modules. STR reorganizes multi-scale semantic features into a dynamic hypergraph to better model cross-resolution anatomical dependencies, enhancing structural and semantic representation. EPG assesses channel stability after perturbation and filters high-entropy channels to emphasize clinically important regions and improve spatial attention. Extensive experiments on three benchmark datasets show our framework achieves superior segmentation accuracy and better generalization across various clinical settings. The code is available at https://github.com/LWX-Research/DTEA. Quanjun Li, Zimeng Li 0001, Chi-Man Pun, Yupeng Liu 0003, Xuhang Chen 0002 |
BIBM | 5 |
| 2025 | HAFT: Hierarchical Attentional Fusion Transformer for Adaptive Feature Fusion in Medical Image SegmentationabstractMedical image segmentation remains fundamentally challenged by the complexity of anatomical structures and severe class imbalance. Existing methods often fall short in multi-scale feature integration and long-tailed distribution modeling, limiting their ability to simultaneously capture fine-grained local structures and holistic contextual semantics. To overcome these issues, we propose HAFT, a hierarchical attentional fusion transformer for adaptive feature fusion in medical image segmentation. We further introduce a Bayesian Adaptive Loss (BAL), which incorporates Bayesian uncertainty modeling to effectively alleviate the long-tailed distribution prevalent in medical datasets. Extensive experiments on multiple public benchmarks demonstrate that our method consistently outperforms existing approaches, particularly in segmenting intricate anatomical regions and rare pathological lesions. The code is available at https://github.com/QuincyQAQ/HAFT. Quanjun Li, Fuchen Zheng, Junhua Zhou, Changwei Gong, Zimeng Li 0001, Yihua Shao, Xuhang Chen 0002 |
BIBM | 7 |
| 2025 | Elevating Medical Image Security: A Cryptographic Framework Integrating Hyperchaotic Map and GRUabstractChaotic systems play a key role in modern image encryption due to their sensitivity to initial conditions, ergodicity, and complex dynamics. However, many existing chaos-based encryption methods suffer from vulnerabilities, such as inade-quate permutation and diffusion, and suboptimal pseudorandom properties. To address these issues, this paper presents the Knot-like Unique Novel-Scan Image Encryption (Kun-IE). The framework comprises two main components: The 2D Sin–Cos Pi Hyperchaotic Map (2D-SCPHM), which provides a broader chaotic range and superior pseudorandom sequence generation, and the Knot-like Unique Novel-Scan Algorithm (Kun-SCAN), a novel permutation mechanism that markedly reduces pixel correlations and enhances resistance to statistical attacks. Kun-IE is flexible and supports encryption for images of any size. Experimental results and security analyses demonstrate its robustness against various cryptanalytic attacks, making it a strong solution for secure image communication. The code is available at this link. Quanjun Li, Junhua Zhou, Yihang Dong, Mengqian Wang, Zimeng Li 0001, Changwei Gong, Xuhang Chen 0002 |
BIBM | 8 |
| 2025 | PC-UNet: An Enforcing Poisson Statistics U-Net for Positron Emission Tomography Denoising
Jingchao Wang 0002, Liangsi Lu, Mingxuan Huang, Ruixin He, Yifeng Xie, Hanqian Liu, Minzhe Guo, Yangyang Liang, Zimeng Li 0001, Xuhang Chen 0002 |
BIBM | 11 |
| 2025 | EEMS: Edge-Prompt Enhanced Medical Image Segmentation Based on Learnable Gating MechanismabstractMedical image segmentation is vital for diagnosis, treatment planning, and disease monitoring but is challenged by complex factors like ambiguous edges and background noise. We introduce EEMS, a new model for segmentation, combining an Edge-Aware Enhancement Unit (EAEU) and a Multi-scale Prompt Generation Unit (MSPGU). EAEU enhances edge perception via multi-frequency feature extraction, accurately defining boundaries. MSPGU integrates high-level semantic and low-level spatial features using a prompt-guided approach, ensuring precise target localization. The Dual-Source Adaptive Gated Fusion Unit (DAGFU) merges edge features from EAEU with semantic features from MSPGU, enhancing segmentation accuracy and robustness. Tests on datasets like ISIC2018 confirm EEMS's superior performance and reliability as a clinical tool. Quanjun Li, Zimeng Li 0001, Hongbin Ye, Yupeng Liu 0003, Haolun Li 0001, Xuhang Chen 0002 |
BIBM | 4 |
| 2025 | TDADL-IE: A Deep Learning-Driven Cryptographic Architecture for Medical Image SecurityabstractThe rise of digital medical imaging, like MRI and CT, demands strong encryption to protect patient data in telemedicine and cloud storage. Chaotic systems are popular for image encryption due to their sensitivity and unique characteristics, but existing methods often lack sufficient security. This paper presents the Three-dimensional Diffusion Algorithm and Deep Learning Image Encryption system (TDADL-IE), built on three key elements. First, we propose an enhanced chaotic generator using an LSTM network with a 1D-Sine Quadratic Chaotic Map (1D-SQCM) for better pseudorandom sequence generation. Next, a new three-dimensional diffusion algorithm (TDA) is applied to encrypt permuted images. TDADL-IE is versatile for images of any size. Experiments confirm its effectiveness against various security threats. The code is available at https://github.com/QuincyQAQ/TDADL-IE. Junhua Zhou, Quanjun Li, Yihua Shao, Yihang Dong, Mengqian Wang, Zimeng Li 0001, Changwei Gong, Xuhang Chen 0002 |
BIBM | 8 |
| 2025 | SWAN: A Synergistic Wavelet Attention Network for Enhanced Underwater Image Enhancement
Junhua Zhou, Quanjun Li, Yihang Dong, Zimeng Li 0001, Zexiao Liang, Xuhang Chen 0002 |
CGI (3) | 4 |
| 2025 | Underwater Image Restoration via Polymorphic Large Kernel CNNsabstractUnderwater Image Restoration (UIR) remains a challenging task in computer vision due to the complex degradation of images in underwater environments. While recent approaches have leveraged various deep learning techniques, including Transformers and complex, parameter-heavy models to achieve significant improvements in restoration effects, we demonstrate that pure CNN architectures with lightweight parameters can achieve comparable results. In this paper, we introduce UIR-PolyKernel, a novel method for underwater image restoration that leverages Polymorphic Large Kernel CNNs. Our approach uniquely combines large kernel convolutions of diverse sizes and shapes to effectively capture long-range dependencies within underwater imagery. Additionally, we introduce a Hybrid Domain Attention module that integrates frequency and spatial domain attention mechanisms to enhance feature importance. By leveraging the frequency domain, we can capture hidden features that may not be perceptible to humans but are crucial for identifying patterns in both underwater and on-air images. This approach enhances the generalization and robustness of our UIR model. Extensive experiments on benchmark datasets demonstrate that UIR-PolyKernel achieves state-of-the-art performance in underwater image restoration tasks, both quantitatively and qualitatively. Our results show that well-designed pure CNN architectures can effectively compete with more complex models, offering a balance between performance and computational efficiency. This work provides new insights into the potential of CNN-based approaches for challenging image restoration tasks in underwater environments. The code is available at https://github.com/CXH-Research/UIR-PolyKernel. Xiaojiao Guo, Yihang Dong, Xuhang Chen 0002, Weiwen Chen, Zimeng Li 0001, Fuchen Zheng, Chi-Man Pun |
ICASSP | 5 |
| 2025 | SFormer: SNR-Guided Transformer for Underwater Image Enhancement from the Frequency Domain
Yingtie Lei, Zimeng Li 0001, Chi-Man Pun, Xuhang Chen 0002 |
PRICAI (5) | 4 |
| 2025 | High-Fidelity Document Stain Removal via A Large-Scale Real-World Dataset and A Memory-Augmented TransformerabstractDocument images are often degraded by various stains, significantly impacting their readability and hindering downstream applications such as document digitization and analysis. The absence of a comprehensive stained document dataset has limited the effectiveness of existing document enhancement methods in removing stains while preserving fine-grained details. To address this challenge, we construct StainDoc, the first large-scale, high-resolution (2145 x 2245) dataset specifically designed for document stain removal. StainDoc comprises over 5,000 pairs of stained and clean document images across multiple scenes. This dataset encompasses a diverse range of stain types, severities, and document backgrounds, facilitating robust training and evaluation of document stain removal algorithms. Furthermore, we propose StainRestorer, a Transformer-based document stain removal approach. StainRestorer employs a memory-augmented Transformer architecture that captures hierarchical stain representations at part, instance, and semantic levels via the DocMemory module. The Stain Removal Transformer (SRTransformer) leverages these feature representations through a dual attention mechanism: an enhanced spatial attention with an expanded receptive field, and a channel attention captures channel-wise feature importance. This combination enables precise stain removal while preserving document content integrity. Extensive experiments demonstrate StainRestorer's superior performance over state-of-the-art methods on the Stain-Doc dataset and its variants StainDoc.Mark and Stain-Doc.Seal, establishing a new benchmark for document stain removal. Our work highlights the potential of memory-augmented Transformers for this task and contributes a valuable dataset to advance future research. Mingxian Li, Yingtie Lei, Xiaofeng Zhang 0006, Yihang Dong, Zimeng Li 0001, Xuhang Chen 0002 |
WACV | 7 |
| 2025 | An asymmetric calibrated transformer network for underwater image restoration
Xiaojiao Guo, Shenghong Luo, Yihang Dong, Zexiao Liang, Zimeng Li 0001, Xuhang Chen 0002 |
Vis. Comput. | 5 |
| 2024 | Test-Time Intensity Consistency Adaptation for Shadow Detection
Leyi Zhu, Weihuang Liu, Zimeng Li 0001, Xuhang Chen 0002, Chi-Man Pun |
ICONIP (7) | 4 |
| 2024 | Encoding Enhanced Complex CNN for Accurate and Highly Accelerated MRIabstractMagnetic resonance imaging (MRI) using hyperpolarized noble gases provides a way to visualize the structure and function of human lung, but the long imaging time limits its broad research and clinical applications. Deep learning has demonstrated great potential for accelerating MRI by reconstructing images from undersampled data. However, most existing deep convolutional neural networks (CNN) directly apply square convolution to k-space data without considering the inherent properties of k-space sampling, limiting k-space learning efficiency and image reconstruction quality. In this work, we propose an encoding enhanced (EN2) complex CNN for highly undersampled pulmonary MRI reconstruction. EN2 complex CNN employs convolution along either the frequency or phase-encoding direction, resembling the mechanisms of k-space sampling, to maximize the utilization of the encoding correlation and integrity within a row or column of k-space. We also employ complex convolution to learn rich representations from the complex k-space data. In addition, we develop a feature-strengthened modularized unit to further boost the reconstruction performance. Experiments demonstrate that our approach can accurately reconstruct hyperpolarized 129Xe and 1H lung MRI from 6-fold undersampled k-space data and provide lung function measurements with minimal biases compared with fully sampled images. These results demonstrate the effectiveness of the proposed algorithmic components and indicate that the proposed approach could be used for accelerated pulmonary MRI in research and clinical lung disease patient care. Zimeng Li 0001, Xiuchao Zhao, Caohui Duan, Qiuchen Rao, Junshuai Xie, Fumin Guo, Chaohui Ye, Xin Zhou 0004 |
IEEE Trans. Medical Imaging | 1 |