Kangkang Lu 0002

dblp:229/6641-2 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-0983-758XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021
YearPublicationVenuePosition
2026 From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification Framework
abstract
The impressive performance of large language models (LLMs) also brings inherent toxicity risks, prompting the need for effective detoxification to support responsible deployment. Prevailing methods generally follow an inflexible model-specific fashion, addressing only individual models or model families. Moreover, overlooking the underlying toxic risks involved in the input prefix can lead to toxic accumulation during autoregressive generation. Existing methods rely on external strong attribute interventions to address this issue, which further exacerbates contextual semantic inconsistencies and makes it difficult to balance toxicity efficacy and generation quality. To address these concerns, we propose a novel Model-Agnostic Adaptive Detoxification (MAAD) framework. To address accumulating toxicity, we present prefix heuristics that serve as contextual signals, guiding the base LLM toward safer generation. Along this line, we construct an antidote dataset to support a lightweight model, Detoxifier, which steers the base LLM to make in-scope and reliable detoxifying distribution adjustments while preserving fluency and contextual understanding. Designed as an easy-to-deploy module, Detoxifier requires a small amount of data and can be seamlessly applied to various base LLMs with one-off training. Since over-purifying often reduces diversity, we also propose a dynamic truncation method called CW-cutoff sampling to trade off language model quality and diversity. Extensive experiments demonstrate that MAAD strikes a better balance between detoxification effectiveness and generation quality, while also maintaining model utility.
Yuhu Shang, Xiang Cheng 0003, Yimeng Ren 0001, Huijia Wu, Xuexiong Luo, Kangkang Lu 0002, Jian Zhao 0018, Zhaofeng He 0001
AAAI6
2026 Addressing graph heterogeneity and heterophily from a spectral perspective
Kangkang Lu 0002, Yanhua Yu, Ruopei Guo, Zhiyong Huang 0010, Yunshan Ma 0002, Meiyu Liang, Xiting Qin, Yimeng Ren 0001, Tat-Seng Chua
Neurocomputing1
2025 R2DQG: A Quality Meets Diversity Framework for Question Generation over Knowledge Bases
abstract
The task of Knowledge-Based Question Generation (KBQG) involves generating natural language questions from structured knowledge sources, posing unique challenges in balancing linguistic diversity and semantic relevance. Existing models often focus on maximizing surface-level similarity to ground-truth questions, neglecting the need for diverse syntactic forms and leading to semantic drift during generation. To overcome these challenges, we propose Refine-Reinforced Diverse Question Generation (R2DQG), a two-phase framework leveraging a generation-then-refinement paradigm. The Generator first constructs a diverse set of expressive templates using dependency parse tree similarity, capturing a wide range of syntactic patterns and styles. These templates guide the creation of question drafts, ensuring both diversity and semantic relevance. In the second phase, a Corrector module refines the drafts to mitigate semantic drift and enhance overall coherence and quality. Experiments on public datasets show that R2DQG outperforms state-of-the-art models in generating diverse, contextually accurate questions. Moreover, synthetic datasets generated by R2DQG enhance downstream QA performance, underscoring the practical utility of our approach.
Yimeng Ren 0001, Yanhua Yu, Lizi Liao, Yuhu Shang, Kangkang Lu 0002, Mingliang Yan
IJCAI5
2025 Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Retrieval
abstract
In recent years, pre-trained multimodal large models have attracted widespread attention due to their outstanding performance in various multimodal applications. Nonetheless, the extensive computational resources and vast datasets required for their training present significant hurdles for deployment in environments with limited computational resources. Many existing methods attempt to compress pre-trained multimodal large models through knowledge distillation, typically focusing on a single optimization objective. While such methods successfully reduce model parameters, they often incur significant performance degradation. Moreover, single-scale optimization fails to ensure comprehensive learning of the teacher model's knowledge across different aspects. In this work, we propose, for the first time, a dynamic self-adaptive multiscale distillation (DSMD) from pre-trained multi-modal large model for efficient cross-modal retrieval method, considering multiple scales from the perspectives of fine granularity, global structure, and hard negative sample mining. Furthermore, we design a dynamic loss balancer, eliminating the need to manually tune objective weights during distillation. This dynamic mechanism ensures that all objectives are optimized in a balanced and adaptive manner throughout the training process. Experiments demonstrate that our multiscale distillation framework achieves significant performance improvements over traditional single-scale distillation methods. Additionally, our proposed dynamic balancer effectively stabilizes the distillation process, ensuring consistent optimization across objectives. The distilled student model achieves 90% of the teacher model's performance while using only 10% of its parameters. Notably, our model also achieves state-of-the-art performance on cross-modal retrieval tasks, outperforming existing approaches. Codes are available at https://github.com/chrisx599/DSMD.
Zhengyang Liang, Meiyu Liang, Yawen Li 0001, Wu Liu 0005, Yingxia Shao, Kangkang Lu 0002
ACM Multimedia7
2025 DeepMolTex: Deep Alignment of Molecular Graphs with Large Language Models via Mixture of Modality Experts
abstract
Recent advances in Molecular Graph-Language Models (MGLMs) have demonstrated promising capabilities in molecular understanding tasks. However, existing approaches face critical limitations: (1) shallow alignment methods which employ identical processing modules for both modalities, resulting in compromised expressiveness and catastrophic forgetting of pre-trained language capabilities; and (2) over-reliance on high-level molecular representations that inadequately capture fine-grained structural information essential for comprehensive molecular understanding. To address these challenges, we present DeepMolTex, a novel framework for Deep fusion of Molecular structure and Textual representations across multiple scales. Our approach introduces a Mixture of Modality Experts (MoME) architecture that facilitates deep alignment between molecular graph features and large language models while preserving language capabilities, and a multi-scale graph projector that extracts and aligns molecular features at atom, motif, and molecule levels. Experimental results demonstrate that DeepMolTex significantly outperforms existing methods on fundamental molecular understanding tasks, including molecule description generation and IUPAC name prediction, while effectively preserving the language capabilities of the pre-trained LLM.
Mingliang Yan, Yanhua Yu, Ruochi Zhang, Zhiyuan Liu 0010, Ruicheng Zhang, Yimeng Ren 0001, Kangkang Lu 0002, Zhiyong Huang 0010, Feng Luo 0004
ACM Multimedia7
2025 Asymmetric Pre-aligned Anchor Contrastive Enhanced Diffusion Hashing Model for Incomplete Multimodal Retrieval
abstract
Multimodal hashing stands as an efficient approach for multimodal retrieval, yet it frequently grapples with the challenge of misaligned representation spaces across different modalities. This misalignment can degrade the consistency and discrimination of multimodal representations, complicating the learning of effective representations for image and text pairs. Particularly, the task becomes arduous when the system must handle incomplete data while ensuring accurate and relevant retrieval outcomes. To address these challenges, we propose the Asymmetric Pre-aligned Anchor Contrastive Enhanced Diffusion Hashing Model (AADH) for Incomplete Multimodal Retrieval. Our model is specifically tailored to robustly manage multimodal incomplete data scenarios. Initially, we develop an Asymmetric Pre-alignment Strategy that utilizes asymmetric contrastive learning to preliminarily align the semantic disparities between various modalities. Subsequently, we propose an innovative Anchor Contrastive Reinforcement Diffusion Hashing Model, which integrates image and text modalities to varying extents during the reverse diffusion process. It constructs an anchor space that not only facilitates the learning of incomplete multimodal hashing representations through anchor contrastive learning but also leverages inter-modal and intra-modal contrastive learning to enhance the representations. Moreover, we effectively bridge the modality gap between different modal hash codes by employing the anchor space to constrain the representations of different modal hashes. By adjusting the initial noise of the diffusion model, we indirectly expand the data volume, which in turn bolsters the model's robustness. Our extensive experimental results across multiple datasets demonstrate that the proposed AADH model achieves state-of-the-art (SOTA) results.
Meiyu Liang, Juncheng Zheng, Kangkang Lu 0002, Yawen Li 0001, Junping Du 0001, Zhe Xue, Wu Liu 0005
ACM Multimedia5
2025 HiGraph-LLM: Hierarchical Graph Encoding and Integration with Large Language Models
Yanhua Yu, Xidian Wang, Kangkang Lu 0002, Tu Ao, Mingliang Yan, Liang Pang 0001, Pinghui Wang, Tat-Seng Chua
PRICAI4
2024 Improving Expressive Power of Spectral Graph Neural Networks with Eigenvalue Correction
abstract
In recent years, spectral graph neural networks, characterized by polynomial filters, have garnered increasing attention and have achieved remarkable performance in tasks such as node classification. These models typically assume that eigenvalues for the normalized Laplacian matrix are distinct from each other, thus expecting a polynomial filter to have a high fitting ability. However, this paper empirically observes that normalized Laplacian matrices frequently possess repeated eigenvalues. Moreover, we theoretically establish that the number of distinguishable eigenvalues plays a pivotal role in determining the expressive power of spectral graph neural networks. In light of this observation, we propose an eigenvalue correction strategy that can free polynomial filters from the constraints of repeated eigenvalue inputs. Concretely, the proposed eigenvalue correction strategy enhances the uniform distribution of eigenvalues, thus mitigating repeated eigenvalues, and improving the fitting capacity and expressive power of polynomial filters. Extensive experimental results on both synthetic and real-world datasets demonstrate the superiority of our method.
Kangkang Lu 0002, Yanhua Yu, Hao Fei 0001, Zixuan Yang 0001, Zirui Guo, Meiyu Liang, Mengran Yin, Tat-Seng Chua
AAAI1
2024 Hop-based Heterogeneous Graph Transformer
abstract
The Graph Transformer (GT) has shown significant ability in processing graph-structured data, addressing limitations in graph neural networks, such as over-smoothing and over-squashing. However, the implementation of GT in real-world heterogeneous graphs (HGs) with complex topology continues to present numerous challenges. Firstly, a challenge arises in designing a tokenizer that is compatible with heterogeneity. Secondly, the complexity of the transformer hampers the acquisition of high-order neighbor information in HGs. In this paper, we propose a novel Hop-based Heterogeneous Graph Transformer (H2Gormer) framework, paving a promising path for HGs to benefit from the capabilities of Transformers. We propose a Heterogeneous Hop-based Token Generation module to obtain high-order information in a flexible way. Specifically, to enrich the fine-grained heterogeneous semantics of each token, we propose a tailored multi-relational encoder to encode the hop-based neighbors. In this way, the resulting token embeddings are input to the Hop-based Transformer to obtain node representations, which are then combined with position embeddings to obtain the final encoding. Extensive experiments on four datasets are conducted to demonstrate the effectiveness of H2Gormer.
Zixuan Yang 0001, Xiao Wang 0017, Yanhua Yu, Kangkang Lu 0002, Zirui Guo, Xiting Qin, Yunshan Ma 0002, Tat-Seng Chua
ECAI5
2024 Unsupervised Multimodal Graph Contrastive Semantic Anchor Space Dynamic Knowledge Distillation Network for Cross-Media Hash Retrieval
abstract
Cross-media hash retrieval are efficient and effective techniques for retrieval on multi-media database. The success of the Multimodal Large Models (MLM) provides a valuable direction to enhance the accuracy of multimodal hash retrieval, which achieves decent retrieval accuracy with finetuning the pretrained multimodal large models, but their massive model parameters significantly reduce retrieval efficiency. Knowledge Distillation (KD) methods enable small models to learn from the knowledge of larger models, achieving a reduction in model parameter count while ensuring a certain level of accuracy. However, current KD methods face challenges when applied in the multimodal domain, as it requires preserving the multimodal semantic information while minimizing accuracy degradation. To address these challenges, we propose a novel unsupervised multimodal graph contrastive semantic anchor space dynamic knowledge distillation network for cross-media hash retrieval (GASKN). Firstly, to obtain a multimodal semantic anchor space, we construct a large multimodal fusion teacher model using the BEiT-3 model as the backbone. This teacher model is capable of encoding data from different modalities, such as images and text, using the same multimodal encoder to acquire multimodal hash codes that contain rich information from both modalities simultaneously. Secondly, to ensure efficient retrieval capabilities for the student model, we utilize the ALBERT text encoding model and the BiFormer image encoding model as the compact student model's backbones. This allows us to build a lightweight student model with only a twentieth of the parameter count of the teacher model. We propose a dynamic knowledge distillation technique to transfer the multimodal semantic anchor space knowledge embedded in the multimodal large teacher model to the lightweight student model as much as possible. Thirdly, to further distill the structural knowledge of the semantic anchor space from the teacher model to the student model, we propose a graph attention contrastive learning mechanism, which enables structural semantic space learning, thereby mining implicit fine-grained cross-media semantic information. By evaluating our method using three widely-used datasets, we demonstrate that GASKN is able to significantly outperform existing state-of-the-art hashing algorithms.
Meiyu Liang, Mengran Yin, Kangkang Lu 0002, Junping Du 0001, Zhe Xue
ICDE4
2024 Adversary and Attention Guided Knowledge Graph Reasoning Based on Reinforcement Learning
Yanhua Yu, Xiuxiu Cai, Ang Ma, Yimeng Ren 0001, Shuai Zhen, Kangkang Lu 0002, Zhiyong Huang 0010, Tat-Seng Chua
KSEM (5)7
2024 EE-LCE: An Event Extraction Framework Based on LLM-Generated CoT Explanation
Yanhua Yu, Yunshan Ma 0002, Kangkang Lu 0002, Zhiyong Huang 0010, Tat-Seng Chua
KSEM (1)5
2024 Information-Controllable Graph Contrastive Learning for Recommendation
abstract
In the evolving landscape of recommender systems, Graph Contrastive Learning (GCL) has become a prominent method for enhancing recommendation performance by alleviating the issue of data sparsity. However, existing GCL-based recommendations often overlook the control of shared information between the contrastive views. In this paper, we initially analyze and experimentally demonstrate these methods often lead to the issue of augmented representation collapse, where the representations between views become excessively similar, diminishing their distinctiveness. To address this issue, we propose the Information-Controllable Graph Contrastive Learning (IGCL) framework, a novel approach that focuses on optimizing the shared information between views to include as much relevant information for the recommendation task as possible while maintaining an appropriate level. In particular, we design the Collaborative Signals Enhanced Augmentation module to infuse the augmented representation with rich, task-relevant collaborative signals. Furthermore, the Information-Controllable Contrastive Learning module is designed to direct control over the magnitude of shared information between the contrastive views to avoid over-similarity. Extensive experiments on three public datasets demonstrate the effectiveness of IGCL, showcasing significant improvements in performance and the capability to alleviate augmented representation collapse.
Zirui Guo, Yanhua Yu, Kangkang Lu 0002, Zixuan Yang 0001, Liang Pang 0001, Tat-Seng Chua
RecSys4
2024 Structures Aware Fine-Grained Contrastive Adversarial Hashing for Cross-Media Retrieval
abstract
Deep cross-media hashing provides an efficient semantic representation learning solution for large-scale cross-media retrieval. The existing methods only consider the inter-media or intra-media semantic association learning, ignore the guiding of semantic structure information, and have weak reasoning ability for implicit fine-grained semantic associations. To tackle this problem, we propose a novel structures aware fine-grained contrastive adversarial hashing method for cross-media retrieval. A novel cross-media contrastive adversarial hash network is constructed for the first time, which integrates the cross-media and intra-media contrastive learning and multi-modal adversarial learning, aiming at maximizing the semantic association between different modalities, and improving the semantic discrimination and consistency of cross-media unified hash representation, thereby the inter-media and intra-media semantic preserving ability can be well enhanced; A fine-grained cross-media semantic feature learning method based on fine-grained semantic reasoning with transformers is proposed, which captures fine-grained salient features of different modalities for semantic association learning, and enhances the reasoning ability of fine-grained implicit semantic association; A semantic label graph convolutional network guided cross-media semantic association learning strategy is proposed, which makes full use of semantic structure information to enhance the learning ability of implicit cross-media semantic associations. Extensive experiments on several large-scale cross-media benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods.
Meiyu Liang, Yawen Li 0001, Xiaowen Cao 0003, Zhe Xue, Ang Li 0015, Kangkang Lu 0002
IEEE Trans. Knowl. Data Eng.7
2023 Deep Unsupervised Momentum Contrastive Hashing for Cross-modal Retrieval
abstract
Unsupervised cross-modal hashing (UCMH) methods often start from the similarity of sample features and design a reconstruction loss to achieve similarity preservation. However, these methods suffer from inaccurate similarity problems, be-cause different feature representations may share similar semantic information. In this paper, we propose Deep Unsupervised Momentum Contrastive Hashing (DUMCH). Specifically, we introduce momentum contrastive learning for unsupervised cross-modal hashing, which allows us to flexibly define a robust loss by comparing positive and negative samples. Moreover, in order to achieve similarity retention of hash codes in Hamming space and fully utilize the potential of contrastive learning in Hamming space, we remove the L2 normalization corresponding to cosine similarity and design a novel normalization method called hash normalization, which has been proved to greatly improve the model performance. We conducted extensive experiments on three datasets, and the experimental results demonstrate the superiority of DUMCH.
Kangkang Lu 0002, Yanhua Yu, Meiyu Liang, Xiaowen Cao 0003, Zehua Zhao, Mengran Yin, Zhe Xue
ICME1
2022 Semantic Structure Enhanced Contrastive Adversarial Hash Network for Cross-media Representation Learning
abstract
Deep cross-media hashing technology provides an efficient cross-media representation learning solution for cross-media search. However, the existing methods do not consider both fine-grained semantic features and semantic structures to mine implicit cross-media semantic associations, which leads to weaker semantic discrimination and consistency for cross-media representation. To tackle this problem, we propose a novel semantic structure enhanced contrastive adversarial hash network for cross-media representation learning (SCAHN). Firstly, in order to capture more fine-grained cross-media semantic associations, a fine-grained cross-media attention feature learning network is constructed, thus the learned saliency features of different modalities are more conducive to cross-media semantic alignment and fusion. Secondly, for further improving learning ability of implicit cross-media semantic associations, a semantic label association graph is constructed, and the graph convolutional network is utilized to mine the implicit semantic structures, thus guiding learning of discriminative features of different modalities. Thirdly, a cross-media and intra-media contrastive adversarial representation learning mechanism is proposed to further enhance the semantic discriminativeness of different modal representations, and a dual-way adversarial learning strategy is developed to maximize cross-media semantic associations, so as to obtain cross-media unified representations with stronger discriminativeness and semantic consistency preserving power. Extensive experiments on several cross-media benchmark datasets demonstrate that the proposed SCAHN outperforms the state-of-the-art methods.
Meiyu Liang, Junping Du 0001, Xiaowen Cao 0003, Kangkang Lu 0002, Zhe Xue
ACM Multimedia5