EDBT 2026 Demo / reviewers in the wild / expert
Chi-Man Pun
dblp:p/ChiManPun · also Chi Man Pun
· DBLP profile ↗
28ranked-venue papers in the field
1as first author
12since 2021 · last 2026
0000-0003-1788-3746ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 17 (1 first)Database Systems & Data Management · 4Information Retrieval & Web Search · 4Data Mining & Knowledge Discovery · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoupling Vocal and Rhythmic Conditioning for Music-Driven Singing Avatar AnimationabstractSynthetic media generation is a burgeoning field in multimedia research. While audio-driven avatar animation has garnered significant attention in digital entertainment, yet music-driven singing avatar animation remains relatively underexplored due to its unique challenges. Distinct from speech, singing animation necessitates the simultaneous modeling of lip articulation governed by singing vocal, and global facial dynamics synchronized with musical rhythm. Existing methods typically rely on 3D intermediate representations, which impose geometric constraints and often degrade visual details. Furthermore, some approaches that simply concatenate vocal and BGM features fail to capture the distinct roles of these signals in driving specific facial regions. To address these limitations, we propose MusicAvatar, a diffusion-based framework that directly synthesizes 2D singing avatars without relying on 3D priors. Moreover, we design a dual-stream music attention module that decouples the roles of singing voice and BGM. Specifically, one cross-attention stream extracts vocal cues from the singing track to drive lip movements, while a parallel stream captures rhythmic patterns from the BGM to modulate facial motion. This parallel yet synergistic design ensures that precise lip movement and rhythmic facial motion are modeled explicitly without interference. Extensive experiments demonstrate that MusicAvatar generates highly natural, expressive, and rhythmically synchronized singing avatars, outperforming state-of-the-art approaches. Yiguo Jiang, Xiaodong Cun, Chen-Bin Feng, Jian Sun 0038, Chi-Man Pun |
ICMR | 5 |
| 2026 | ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech SynthesisabstractHigh-fidelity character voice synthesis is a cornerstone of immersive multimedia applications, particularly for interacting with anime avatars and digital humans. However, existing systems struggle to maintain consistent persona traits across diverse emotional contexts. To bridge this gap, we present ATRIE, a unified framework utilizing a Persona-Prosody Dual-Track (P2-DT) architecture. Our system disentangles generation into a static Timbre Track (via Scalar Quantization) and a dynamic Prosody Track (via Hierarchical Flow-Matching), distilled from a 14B LLM teacher. This design enables robust identity preservation (Zero-Shot Speaker Verification EER: 0.04) and rich emotional expression. Evaluated on our extended AnimeTTS-Bench (50 characters), ATRIE achieves state-of-the-art performance in both generation and cross-modal retrieval (mAP: 0.75), establishing a new paradigm for persona-driven multimedia content creation. The code is available at Github. Aoduo Li, Hongjian Xu, Shengmin Li, Sihao Qin, Zimeng Li 0001, Chi-Man Pun, Xuhang Chen 0002 |
ICMR | 7 |
| 2025 | Evidential Prototype Learning for Semi-supervised Medical Image SegmentationabstractAlthough current semi-supervised medical segmentation methods can achieve decent performance, they are still affected by the uncertainty in unlabeled data and model predictions, and there is currently a lack of effective strategies that can explore the uncertain aspects of both simultaneously. To address the aforementioned issues, we propose Evidential Prototype Learning (EPL), which utilizes an extended probabilistic framework to effectively fuse voxel-level evidential predictions from different classifiers and achieves prototype fusion utilization of labeled and unlabeled data under a generalized evidential framework, leveraging voxel-level dual uncertainty masking. The uncertainty measure not only enables the model to self-correct predictions but also improves the guided learning process with pseudo-labels and is able to feed back into the construction of hidden features. The method proposed in this paper has been experimented on LA, Pancreas-CT and TBAD datasets, achieving the state-of-the-art performance in three different labeled ratios, which strongly demonstrates the effectiveness of our strategy. The source code will be made publicly available. Yuanpeng He, Lijian Li 0003, Tianxiang Zhan, Chi-Man Pun, Wenpin Jiao, Zhi Jin 0001 |
KDD (2) | 4 |
| 2024 | UIE-UnFold: Deep Unfolding Network with Color Priors and Vision Transformer for Underwater Image EnhancementabstractUnderwater image enhancement (UIE) plays a crucial role in various marine applications, but it remains challenging due to the complex underwater environment. Current learning-based approaches frequently lack explicit incorporation of prior knowledge about the physical processes involved in underwater image formation, resulting in limited optimization despite their impressive enhancement results. This paper proposes a novel deep unfolding network (DUN) for UIE that integrates color priors and inter-stage feature transformation to improve enhancement performance. The proposed DUN model combines the iterative optimization and reliability of model-based methods with the flexibility and representational power of deep learning, offering a more explainable and stable solution compared to existing learning-based UIE approaches. The proposed model consists of three key components: a Color Prior Guidance Block (CPGB) that establishes a mapping between color channels of degraded and original images, a Nonlinear Activation Gradient Descent Module (NAGDM) that simulates the underwater image degradation process, and an Inter Stage Feature Transformer (ISF-Former) that facilitates feature exchange between different network stages. By explicitly incorporating color priors and modeling the physical characteristics of underwater image for-mation, the proposed DUN model achieves more accurate and reliable enhancement results. Extensive experiments on multiple underwater image datasets demonstrate the superiority of the proposed model over state-of-the-art methods in both quantitative and qualitative evaluations. The proposed DUN-based approach offers a promising solution for UIE, enabling more accurate and reliable scientific analysis in marine research. The code is available at https://github.com/CXH-Research/UIE-UnFold. Yingtie Lei, Yihang Dong, Changwei Gong, Ziyang Zhou 0001, Chi-Man Pun |
DSAA | 6 |
| 2024 | Multi-task subspace clustering
Guo Zhong, Chi-Man Pun |
Inf. Sci. | 2 |
| 2023 | Supervised Discriminative Discrete Hashing for Cross-Modal Retrieval
Chi-Man Pun |
ADMA (2) | 2 |
| 2023 | Data Representation by Joint Hypergraph Embedding and Sparse Coding (Extended Abstract)abstractMatrix factorization (MF), a popular unsupervised learning technique for data representation, has been widely applied in data mining and machine learning. According to different application scenarios, one can impose different constraints on the factorization to find the desired basis, which captures high-level semantics for the given data, and learns the compact representation corresponding to the basis. We note that almost all previous work on MF in data mining has ignored to find such a basis, which can carry high-order semantics in the data. In this work, we propose a novel MF framework called Joint Hypergraph Embedding and Sparse Coding, in which the obtained basis captures high-order semantic information in data. Experimental results on data clustering demonstrate that the proposed method consistently outperforms the other state-of-the-art matrix factorization methods. Guo Zhong, Chi-Man Pun |
ICDE | 2 |
| 2023 | Deep self-learning based dynamic secret key generation for novel secure and efficient hashing algorithm
Fasee Ullah, Chi-Man Pun |
Inf. Sci. | 2 |
| 2022 | Learning ordinal constraint binary codes for fast similarity search
Zheng Zhang 0006, Chi-Man Pun |
Inf. Process. Manag. | 2 |
| 2022 | Data Representation by Joint Hypergraph Embedding and Sparse CodingabstractMatrix factorization (MF), a popular unsupervised learning technique for data representation, has been widely applied in data mining and machine learning. According to different application scenarios, one can impose different constraints on the factorization to find the desired basis, which captures high-level semantics for the given data, and learns the compact representation corresponding to the basis. We note that almost all previous work on MF in data mining has ignored to find such a basis, which can carry high-order semantics in the data. In this article, we propose a novel MF framework called Joint Hypergraph Embedding and Sparse Coding (JHESC), in which the obtained basis captures high-order semantic information in data. Specifically, we first propose a new hypergraph learning model to obtain a more discriminative basis by hypergraph-based Laplacian Eigenmap, then sparse coding is conducted on the learned basis such that the new representation has stronger identification capability. In addition, we extend the proposed method to the reproducing kernel Hilbert space for dealing with nonlinear data more effectively. Extensive experimental results on data clustering demonstrate that the proposed method consistently outperforms the other state-of-the-art matrix factorization methods. Guo Zhong, Chi-Man Pun |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Latent Low-rank Graph Learning for Multimodal ClusteringabstractMultimodal clustering has become a fundamental and important problem in the data mining community since the development of multimedia technology over the last two decades has led to a tremendous increase in unlabeled multimodal data. Although a panoply of multimodal subspace clustering methods shows promising performance via fusing information from different views of multimodal data, most of them consist of two sequential steps, i.e., learning a consensus affinity matrix from the original data and then feeding the resulting affinity matrix into the framework of spectral clustering. However, this leads to the suboptimal clustering performance due to the following limitations: 1) the two steps of learning the affinity matrix and clustering are carried out independently; 2) the affinity matrix may be unreliable; 3) the post-processing requirement, such as K-means. To address these issues, we propose a novel multimodal subspace clustering method via adaptively learning a similarity graph on a latent low-rank representation space. In particular, the number of connected components of the learned graph is precisely equal to the number of clusters, i.e., the optimal solution of the associated problem directly reveals the clustering structure of data. Extensive evaluations on several benchmark multimodal datasets demonstrate that the proposed approach outperforms state-of-the-art methods. Guo Zhong, Chi-Man Pun |
ICDE | 2 |
| 2021 | Improving adversarial attacks on deep neural networks via constricted gradient-based perturbations
Yatie Xiao, Chi-Man Pun |
Inf. Sci. | 2 |
| 2020 | A Unified Framework for Multi-view Spectral ClusteringabstractIn the era of big data, multi-view clustering has drawn considerable attention in machine learning and data mining communities due to the existence of a large number of unlabeled multi-view data in reality. Traditional spectral graph theoretic methods have recently been extended to multi-view clustering and shown outstanding performance. However, most of them still consist of two separate stages: learning a fixed common real matrix (i.e., continuous labels) of all the views from original data, and then applying K-means to the resulting common label matrix to obtain the final clustering results. To address these, we design a unified multi-view spectral clustering scheme to learn the discrete cluster indicator matrix in one stage. Specifically, the proposed framework directly obtain clustering results without performing K-means clustering. Experimental results on several famous benchmark datasets verify the effectiveness and superiority of the proposed method compared to the state-of-the-arts. Guo Zhong, Chi-Man Pun |
ICDE | 2 |
| 2020 | Exposing splicing forgery in realistic scenes using deep fusion network
Bo Liu 0047, Chi-Man Pun |
Inf. Sci. | 2 |
| 2020 | Adversarial example generation with adaptive gradient search for single and ensemble deep neural network
Yatie Xiao, Chi-Man Pun, Bo Liu 0047 |
Inf. Sci. | 2 |
| 2020 | Two-pass hashing feature representation and searching method for copy-move forgery detection
Chi-Man Pun |
Inf. Sci. | 2 |
| 2020 | Nonnegative self-representation with a fixed rank constraint for subspace clustering
Guo Zhong, Chi-Man Pun |
Inf. Sci. | 2 |
| 2020 | Dense moment feature index and best match algorithms for video copy-move forgery detection
Chi-Man Pun |
Inf. Sci. | 2 |
| 2019 | Adaptive multi-scale deep neural networks with perceptual loss for panchromatic and multispectral images classification
Cheng Shi 0002, Chi-Man Pun |
Inf. Sci. | 2 |
| 2018 | Reversible data-hiding in encrypted images by redundant space transfer
Chi-Man Pun |
Inf. Sci. | 2 |
| 2018 | A two-stage localization for copy-move forgery detection
Chi-Man Pun, Jim-Lee Chung |
Inf. Sci. | 1 |
| 2017 | Fast reflective offset-guided searching method for copy-move forgery detection
Xiuli Bi, Chi-Man Pun |
Inf. Sci. | 2 |
| 2017 | 3D multi-resolution wavelet convolutional neural networks for hyperspectral image classification
Cheng Shi 0002, Chi-Man Pun |
Inf. Sci. | 2 |
| 2016 | Multi-Level Dense Descriptor and Hierarchical Feature Matching for Copy-Move Forgery Detection
Xiuli Bi, Chi-Man Pun, Xiaochen Yuan |
Inf. Sci. | 2 |
| 2016 | An efficient image segmentation method based on a hybrid particle swarm algorithm with learning strategy
Hao Gao 0005, Chi-Man Pun, Sam Kwong |
Inf. Sci. | 2 |
| 2015 | 2D Sine Logistic modulation map for image encryption
Zhongyun Hua, Yicong Zhou, Chi-Man Pun, C. L. Philip Chen |
Inf. Sci. | 3 |
| 2015 | Robust Mel-Frequency Cepstral coefficients feature detection and dual-tree complex wavelet transform for digital audio watermarking
Xiaochen Yuan, Chi-Man Pun, C. L. Philip Chen |
Inf. Sci. | 2 |
| 2003 | Rotation and Scale Invariant Wavelet Feature for Content-based Texture Image RetrievalabstractAbstract This article introduces an effective rotation and scale invariant log‐polar wavelet texture feature for image retrieval. The proposed feature is an attempt to enhance the existing content‐based image retrieval systems that largely present difficulty in coping with images with changes in orientations and scales. The underlying feature extraction process involves a log‐polar transform followed by an adaptive row shift invariant wavelet packet transform. The log‐polar transform converts a given image into a rotation and scale invariant but row‐shifted image, which is then further processed through an adaptive row‐shift invariant wavelet packet transform operation to generate adaptively selected subbands of rotation and scale invariant wavelet coefficients, based on an information cost function. An energy signature is computed for each subband of these wavelet coefficients. To reduce feature dimensionality, only the most dominant log‐polar wavelet energy signatures are selected for the feature vector for image retrieval. The overall feature extraction process is quite efficient and involves only O(n · log n) complexity. Experimental results show that this rotation and scale invariant wavelet feature is quite effective for image retrieval and outperforms the traditional wavelet packet signatures. Moon-Chuen Lee, Chi-Man Pun |
J. Assoc. Inf. Sci. Technol. | 2 |