EDBT 2026 Demo / reviewers in the wild / expert
Chuankun Li
dblp:192/1909
· DBLP profile ↗
22ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task-Aware Information Decoupling for Multimodal ClusteringabstractMultimodal clustering (MMC) overcomes the limitations of unimodal methods by integrating information from multiple sources, but the complexity of heterogeneous information coupling hinders effective feature extraction. Critically, existing MMC paradigms primarily focus on capturing consensus through coarse-grained cross-modal alignment. However, such task-agnostic strategies overlook the differences in the utility of feature information across varying task environments. In the absence of task-centric guidance, models often struggle to effectively distinguish task-relevant critical information from task-irrelevant redundant noise during the disentanglement process, leading to information confusion in the representation space. To address this challenge, we propose a deep disentangled multimodal clustering method guided by information theory, named DRLMMC, which employs a tripartite information optimization mechanism to achieve deep disentanglement of cross-modal representations. 1) We design modality-specific encoders to construct nonlinear mapping spaces, transforming the reconstruction mechanism of autoencoders into an information-theoretic mutual information (MI) constraint problem, preserving the unique features of different modalities; 2) To establish cross-modal semantic associations, it constructs a cross-modal shared information extraction module, and, based on an information-theoretic framework, designs an optimization objective function to progressively align multimodal feature subspaces through MI maximization and contrastive learning, capturing task-relevant invariant features across modalities; 3) A unique information dynamic perception module is proposed, which employs a conditional MI projection network combined with learning distribution regularization to adaptively extract and enhance modality-specific task-relevant unique information. Experimental results demonstrate that DRLMMC outperforms existing state-of-the-art methods on multimodal benchmark datasets, exhibiting excellent generalization ability. Notably, it achieves precise disentanglement of cross-omics features in multi-omics analysis, offering a novel methodological approach for handling complex biomedical data. Zixiao Jin, Chang Tang, Chuankun Li, Yuanyuan Liu 0004, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Large-scale Multi-view Tensor Clustering with Implicit Linear KernelsabstractMulti-view clustering is a long-standing hot topic in machine learning communities, due to its capability of integrating data information from multiple sources and modalities. By utilizing tensor Singular Value Decomposition (t-SVD) technique with the tensor rotation trick, recent advances have achieved remarkable improvements on clustering performance. However, we find this is attributed to the inadvertent use of sequential information of sorted data samples, i.e. inadvertent label use, which violates the unsupervised learning setting. On the other hand, existing large-scale approaches are mostly developed on the basis of matrix factorization or anchor techniques, thereby fail to consider the similarities among all data samples, preventing from further performance improvement. To address the above issues, we first analyze the tensor rotation trick and recommend to remove it from tensor clustering. On its basis, a novel large-scale multi-view tensor clustering method is developed by incorporating the pair-wise similarities with implicit linear kernel function. To solve the resultant optimization problem, we design an efficient algorithm of linear complexity. Moreover, extensive experiments are conducted and corresponding results well support the aforementioned finding and validate the effectiveness and efficiency of the proposed method. Jiyuan Liu 0003, Xinwang Liu 0002, Chuankun Li, Xinhang Wan, Yi Zhang 0104, Weixuan Liang, Qian Qu, Renxiang Guan, Ke Liang 0006 |
CVPR | 3 |
| 2025 | MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural RepresentationabstractThis paper presents MetricGrids, a novel grid-based neural representation that combines elementary metric grids in various metric spaces to approximate complex nonlinear signals. While grid-based representations are widely adopted for their efficiency and scalability, the existing feature grids with linear indexing for continuous-space points can only provide degenerate linear latent space representations, and such representations cannot be adequately compensated to represent complex nonlinear signals by the following compact decoder. To address this problem while keeping the simplicity of a regular grid structure, our approach builds upon the standard grid-based paradigm by constructing multiple elementary metric grids as high-order terms to approximate complex nonlinearities, following the Taylor expansion principle. Furthermore, we enhance model compactness with hash encoding based on different sparsities of the grids to prevent detrimental hash collisions, and a high-order extrapolation decoder to reduce explicit grid storage requirements. experimental results on both 2D and 3D reconstructions demonstrate the superior fitting and rendering accuracy of the proposed method across diverse signal types, validating its robustness and generalizability. Code is available at https://github.com/wangshu31/MetricGrids. Yanbo Gao, Shuai Li 0005, Chong Lv, Chuankun Li, Hui Yuan 0001, Jinglin Zhang 0001 |
CVPR | 6 |
| 2025 | DACMF-DTI: Dual attention embedded cross-modality fusion for drug-target interaction prediction
Chuankun Li, Minhui Wang, Chang Tang |
Knowl. Based Syst. | 3 |
| 2025 | Unsupervised Feature Enrichment and Fidelity Preservation Learning Framework for Skeleton-Based Action RecognitionabstractUnsupervised skeleton-based action recognition has achieved remarkable progress recently. Existing unsupervised learning methods suffer from severe overfitting problem, and thus small networks are used, significantly reducing the representation capability. To address this problem, the overfitting mechanism behind the unsupervised learning for skeleton-based action recognition is first investigated. It is observed that skeleton is already a relatively high-level and low-dimension feature, but not in the same manifold as the features for action recognition. Simply applying the existing unsupervised learning method tends to produce features that discriminate the different samples rather than action classes, resulting in the overfitting problem. To address this problem, this paper proposes an Unsupervised spatial-temporal Feature Enrichment and Fidelity Preservation (U-FEFP) learning framework to generate rich distributed features that contain all the information of a skeleton sample. A spatial-temporal feature transformation subnetwork is developed using channel-wise topology refinement graph convolutional block and graph convolutional gated recurrent unit block as the basic feature extraction network. The unsupervised Bootstrap Your Own Latent-based learning is utilized to generate rich distributed features, and the unsupervised pretext task-based learning is employed to preserve the information contained in the skeleton. The two unsupervised learning ways are collaborated as U-FEFP to produce robust and discriminative representations. Experimental results on four widely used benchmarks, namely NTU-RGB+D-60, PKU-MMD, NTU-RGB+D-120 and AAV-Human dataset, demonstrate that the proposed U-FEFP obtains the best result compared with the state-of-the-art unsupervised learning methods. Chuankun Li, Shuai Li 0005, Yanbo Gao, Xingyu Gao 0001, Ping Chen 0004, Wanqing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Spectral Discrepancy and Cross-Modal Semantic Consistency Learning for Object Detection in Hyperspectral ImagesabstractHyperspectral images with high spectral resolution provide new insights into recognizing subtle differences in similar substances. However, object detection in hyperspectral images faces significant challenges in intra- and inter-class similarity due to the spatial differences in hyperspectral inter-bands and unavoidable interferences, e.g., sensor noises and illumination. To alleviate the hyperspectral inter-bands inconsistencies and redundancy, we propose a novel network termedSpectralDiscrepancy andCross-Modal semantic consistency learning (SDCM), which facilitates the extraction of consistent information across a wide range of hyperspectral bands while utilizing the spectral dimension to pinpoint regions of interest. Specifically, we leverage a semantic consistency learning (SCL) module that utilizes inter-band contextual cues to diminish the heterogeneity of information among bands, yielding highly coherent spectral dimension representations. On the other hand, we incorporate a spectral gated generator (SGG) into the framework that filters out the redundant data inherent in hyperspectral information based on the importance of the bands. Then, we design the spectral discrepancy aware (SDA) module to enrich the semantic representation of high-level information by extracting pixel-level spectral features. Extensive experiments on two hyperspectral datasets demonstrate that our proposed method achieves state-of-the-art performance when compared with other ones. Xiao He 0010, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, Zhimin Gao, Chuankun Li, Shaohua Qiu, Jiangfeng Xu |
IEEE Trans. Multim. | 6 |
| 2025 | LiftFormer: Lifting and Frame Theory Based Monocular Depth Estimation Using Depth and Edge Oriented Subspace RepresentationabstractMonocular depth estimation (MDE) has attracted increasing interest in the past few years, owing to its important role in 3D vision. MDE is the estimation of a depth map from a monocular image/video to represent the 3D structure of a scene, which is a highly ill-posed problem. To solve this problem, in this paper, we propose a LiftFormer based on lifting theory topology, for constructing an intermediate subspace that bridges the image color features and depth values, and a subspace that enhances the depth prediction around edges. MDE is formulated by transforming the depth value prediction problem into depth-oriented geometric representation (DGR) subspace feature representation, thus bridging the learning from color values to geometric depth values. A DGR subspace is constructed based on frame theory by using linearly dependent vectors in accordance with depth bins to provide a redundant and robust representation. The image spatial features are transformed into the DGR subspace, where these features correspond directly to the depth values. Moreover, considering that edges usually present sharp changes in a depth map and tend to be erroneously predicted, an edge-aware representation (ER) subspace is constructed, where depth features are transformed and further used to enhance the local features around edges. The experimental results demonstrate that our LiftFormer achieves state-of-the-art performance on widely used datasets, and an ablation study validates the effectiveness of both proposed lifting modules in our LiftFormer. Shuai Li 0005, Huibin Bai, Yanbo Gao, Chong Lv, Hui Yuan 0001, Chuankun Li, Wei Hua 0002, Tian Xie 0011 |
IEEE Trans. Multim. | 6 |
| 2024 | Heterogeneous Graph Guided Contrastive Learning for Spatially Resolved Transcriptomics DataabstractSpatial transcriptomics provides revolutionary insights into cellular interactions and disease development mechanisms by combining high-throughput gene sequencing and spatially resolved imaging technologies to analyze genes naturally associated with spatially variable tissue genes. However, existing methods typically map aggregated multi-view features into a unified representation, ignoring the heterogeneity and view independence of genes and spatial information. To this end, we construct a heterogeneous Graph guided Contrastive Learning (stGCL) for aggregating spatial transcriptomics data. The method is guided by the inherent heterogeneity of cellular molecules by dynamically coordinating triple-level node attributes through comparative learning loss distributed across view domains, thus maintaining view independence during the aggregation process. In addition, we introduce a cross-view hierarchical feature alignment module employing a parallel approach to decouple spatial and genetic views on molecular structures while aggregating multi-view features according to information theory, thereby enhancing the integrity of inter- and intra-views. Rigorous experiments demonstrate that stGCL outperforms existing methods in various tasks and related downstream applications. Xiao He 0010, Chang Tang, Xinwang Liu 0002, Chuankun Li, Shan An, Zhenglai Li |
ACM Multimedia | 4 |
| 2024 | Frequency-Domain Transformation-Based Dynamic Gesture Recognition with Skeleton
Chuankun Li, Shuai Li 0005, Wanqing Li 0001, Danyan Xie |
PRCV (3) | 2 |
| 2024 | Static graph convolution with learned temporal and channel-wise graph topology generation for skeleton-based action recognition
Chuankun Li, Shuai Li 0005, Yanbo Gao, Lijuan Zhou 0002, Wanqing Li 0001 |
Comput. Vis. Image Underst. | 1 |
| 2024 | DFN: A deep fusion network for flexible single and multi-modal action recognitionabstractMulti-modal action recognition methods can be generally classified into two categories: (1) fusing multi-modal features with simple concatenation or fusing the classification scores of individual modalities without considering the interaction among the multi-modalities; (2) using one of the modalities as privileged information in training to boost the recognition on the other modalities in inference. The former approach usually is not able to deal with the cases where one of the modalities is missing. In the latter, the trained classifier does not work on the privileged modality. To address these shortcomings, this paper presents a novel end-to-end trainable deep fusion network (DFN) that is able to improve the performance not only in the cases where all modalities are available and also in the cases where there is a missing modality. The DFN is simple yet effective with the capability of retrieving an estimation of one modality by using another modality through a Multilayer Perceptron (MLP). In order to better preserve structure information, the DFN first maps the individual modality features to a high dimensional Kronecker-product space and subsequently learns a low-dimensional discriminative space for classification. The effectiveness of the proposed DFN has been verified on three benchmark datasets: the large NTU RGB+D, UTD-MHAD, and SYSU-3D datasets and it has achieved state-of-the-art results. Chuankun Li, Yonghong Hou, Wanqing Li 0001, Zewei Ding, Pichao Wang |
Expert Syst. Appl. | 1 |
| 2024 | Color and Geometric Contrastive Learning Based Intra-Frame Supervision for Self-Supervised Monocular Depth EstimationabstractIn recent years, self-supervised monocular depth estimation has become popular due to its advantage in estimating the depth without the need of groundtruth depth labels. Instead, it takes an inter-frame supervision using depth based view synthesis to reconstruct temporal adjacent frames to indirectly supervise the generated depth. However, such supervision weakens the depth estimation at temporal incoherent regions containing small changes among consecutive frames. To overcome the above problem, we propose a color and geometric contrastive learning based intra-frame supervision framework to enhance self-supervised monocular depth estimation. Color-contrastive learning is proposed to guide the network to learn color invariant features considering color information is irrelevant to depth data. To improve the local details of the learned feature, a pixel-level contrastive learning is further used to optimize the learning. In view that the depth estimation, as a pixel-level task, is sensitive to the geometric transformation, geometric-contrastive learning is developed using an inverse geometric transformation to learn features that are equivariant to the geometric data augmentation. A local plane guidance layer (LPG) with contrastive learning is further used to decompose the geometric information and enhance the geometric contrastive learning. Experiments demonstrate that the proposed method achieves the best result compared to the state-of-the-art methods in all tested quality metrics, with the largest improvement of 22.8% over baseline Monodepth2 and 3.2% over Monovit, in terms of SqRel reduction. Yanbo Gao, Xianye Wu, Shuai Li 0005, Chuankun Li |
IEEE Signal Process. Lett. | 5 |
| 2024 | Aligned Intra Prediction and Hyper Scale Decoder Under Multistage Context Model for JPEG AIabstractLearning-based image compression has raised increasing interests in the last few years. Currently, Joint Photographic Experts Group (JPEG) is working on the standardization of learning-based image compression as JPEG AI. It adopts a deep neural network based encoder-decoder architecture with hyperprior based probability formulation for entropy coding. JPEG AI currently contains two coding profiles, including the Base Operating Point (BaseOP) and High Operating Point (HighOP). Among the various techniques developed in JPEG AI, Multistage Context Model (MCM) was adopted as the context model to perform intra prediction in HighOP. It transforms the spatially progressive context prediction into sub-image feature prediction among channels via feature down-shuffling. However, in this prediction process, sub-image features are not spatially aligned to each other, and directly using the neighboring sub-image features cannot provide accurate prediction. Moreover, the distributions of residual features generated by MCM are also not consistent with that of the hyper scale decoder, which is used to construct the probability model in the entropy coding of residual features, leading to suboptimal residual coding. To address the above problems, we propose an Aligned Intra Prediction (AIP) and Aligned Hyper Scale Decoder (AHSD) under Multistage Context Model for JPEG AI coding. AIP aligns the reference sub-image features to the to-be-predicted feature in MCM with an offset prediction network and deformable convolution. AHSD further generates hyper scale features with matched distributions to the residual features, in order to enhance the probability formulation in its entropy coding. Experimental results demonstrate that the proposed method improves the coding performance by 1.3% in terms of BD-rate saving over the JPEG AI reference software and the effectiveness of each module is verified in ablation study. Shuai Li 0005, Yanbo Gao, Chuankun Li, Hui Yuan 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Geometric Warping Error Aware Spatial-Temporal Enhancement for DIBR Oriented View Synthesis
Rui Peng 0009, Shuai Li 0005, Yanbo Gao, Chuankun Li |
IEEE Signal Process. Lett. | 5 |
| 2023 | Motion saliency based hierarchical attention network for action recognition
Zihui Guo, Yonghong Hou, Renyi Xiao, Chuankun Li, Wanqing Li 0001 |
Multim. Tools Appl. | 4 |
| 2023 | Improved Shift Graph Convolutional Network for Action Recognition With SkeletonabstractShift graph convolutional network (Shift-GCN) achieves remarkable performance for skeleton based action recognition with lower computational complexity than other GCN based methods. However, the current Shift-GCN, with one spatial shift, a static mask and a local temporal convolution, cannot fully explore the spatial-temporal features among skeleton joints of different frames. In order to address these problems, an improved shift graph convolutional network (Ishift-GCN) is proposed in this letter. The Ishift-GCN consists of two parts including a bidirectional spatial shift graph convolution with a dynamic mask, and a multi-scale temporal shift graph convolution. The bidirectional spatial shift graph convolution exploits more spatial information among joints, and the dynamic mask with stronger generalization ability can learn different correlations among features of different joints for different actions. The multi-scale temporal shift graph convolution captures more temporal information by complementing the shifted features with multi-scale convolution. Furthermore, knowledge distillation is used to reduce computational complexity. Compared with Shift-GCN, the proposed Ishift-GCN achieves better results with less computation complexity on two widely used benchmarks, namely the NTU-RGB+D and UAV-Human dataset. Chuankun Li, Shuai Li 0005, Yanbo Gao, Wanqing Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2020 | ConvNets-based action recognition from skeleton motion maps
Yanfang Chen, Chuankun Li, Yonghong Hou, Wanqing Li 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Self-Attention Guided Deep Features for Action RecognitionabstractSkeleton based human action recognition is an important task in computer vision. However, it is very challenging due to the complex spatio-temporal variations of skeleton joints. In this work, we propose an end-to-end trainable network consisting of a Deep Convolutional Model (DCM) and a Self-Attention Model (SAM) for human action recognition from skeleton data. Specifically, skeleton sequences are encoded into color images and fed into DCM to extract deep features. In the SAM, handcrafted features representing the motion degree of joints are extracted and the attention weights are learned by a simple yet effective linear mapping. The effectiveness of proposed method has been verified on NTU RGB+D, SYSU-3D and UTD-MHAD datasets and achieved state-of-the-art results. Renyi Xiao, Yonghong Hou, Zihui Guo, Chuankun Li, Pichao Wang, Wanqing Li 0001 |
ICME | 4 |
| 2019 | Learning attentive dynamic maps (ADMs) for Understanding Human Actions
Chuankun Li, Yonghong Hou, Wanqing Li 0001, Pichao Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Multiview-Based 3-D Action Recognition Using Deep NetworksabstractIn multiview learning, views may be obtained from multiple sources or extracted from a single source as different features. In this paper, effective multiple views from skeleton sequences are proposed to learn the discriminative features using multiple networks for three-dimensional human action recognition. Specifically, three views are constructed in the spatial domain and fed to a stack of long short-term memory networks to exploit temporal information and three views are constructed using the improved joint trajectory maps and fed to three convolutional neural networks to exploit spatial information. Multiply fusion is used to combine the recognition scores of all views. The proposed method has been verified and achieved the state-of-the-art results on the widely used UTD-MHAD, MSRC-12 Kinect Gesture, and NTU red, green, blue (RGB)+D datasets. Chuankun Li, Yonghong Hou, Pichao Wang, Wanqing Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2018 | Action recognition based on joint trajectory maps with convolutional neural networks
Pichao Wang, Wanqing Li 0001, Chuankun Li, Yonghong Hou |
Knowl. Based Syst. | 3 |
| 2017 | Joint Distance Maps Based Action Recognition With Convolutional Neural NetworksabstractMotivated by the promising performance achieved by deep learning, an effective yet simple method is proposed to encode the spatio-temporal information of skeleton sequences into color texture images, referred to as joint distance maps (JDMs), and convolutional neural networks are employed to exploit the discriminative features from the JDMs for human action and interaction recognition. The pair-wise distances between joints over a sequence of single or multiple person skeletons are encoded into color variations to capture temporal information. The efficacy of the proposed method has been verified by the state-of-the-art results on the large RGB+D Dataset and small UTD-MHAD Dataset in both single-view and cross-view settings. Chuankun Li, Yonghong Hou, Pichao Wang, Wanqing Li 0001 |
IEEE Signal Process. Lett. | 1 |