EDBT 2026 Demo / reviewers in the wild / expert
Chen Chen 0128
dblp:65/4423-128
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-0579-2353ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Color-Shape Disentangled Representation Learning with channel augmentation in interactive image retrieval
Chen Chen 0128, Bin Song 0001 |
Neurocomputing | 1 |
| 2025 | Frequency-domain keyframe interpolation denoising for text and image-guided video editing acceleration with diffusion models
Chenhao Pang, Chen Chen 0128, Bin Song 0001 |
Neurocomputing | 2 |
| 2025 | Frequency-Based Comprehensive Prompt Learning for Vision-Language ModelsabstractThis paper targets to learn multiple comprehensive text prompts that can describe the visual concepts from coarse to fine, thereby endowing pre-trained VLMs with better transfer ability to various downstream tasks. We focus on exploring this idea on transformer-based VLMs since this kind of architecture achieves more compelling performances than CNN-based ones. Unfortunately, unlike CNNs, the transformer-based visual encoder of pre-trained VLMs cannot naturally provide discriminative and representative local visual information. To solve this problem, we propose Frequency-based Comprehensive Prompt Learning (FCPrompt) to excavate representative local visual information from the redundant output features of the visual encoder. FCPrompt transforms these features into frequency domain via Discrete Cosine Transform (DCT). Taking the advantages of energy concentration and information orthogonality of DCT, we can obtain compact, informative and disentangled local visual information by leveraging specific frequency components of the transformed frequency features. To better fit with transformer architectures, FCPrompt further adopts and optimizes different text prompts to respectively align with the global and frequency-based local visual information via a dual-branch framework. Finally, the learned text prompts can thus describe the entire visual concepts from coarse to fine comprehensively. Extensive experiments indicate that FCPrompt achieves the state-of-the-art performances on various benchmarks. Liangchen Liu 0001, Nannan Wang 0001, Chen Chen 0128, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | HSMH: A Hierarchical Sequence Multi-Hop Reasoning Model With Reinforcement LearningabstractThe incompleteness of knowledge graphs (KGs) negatively impacts the performance of KGs in downstream applications (e.g., recommendation systems and information retrieval). This phenomenon has brought an increasing rise in research related to knowledge graph reasoning. Recently, emerged reinforcement learning (RL)-based multi-hop reasoning methods can infer missing information through multi-hop reasoning according to the existing information in KGs, which has better reasoning performance and interpretability. However, these methods always use relation-entity pairs that have been pre-cropped as the action space of agents for path reasoning, which leads to two problems: 1) insufficient learning and reasoning ability of reasoning models and 2) the hard convergence of the training process of agents. To address these problems, we propose aHierarchicalSequenceMultiHop (HSMH) reasoning framework, which consists of the interactive search reasoning model, local-global knowledge fusion mechanism, and action optimization mechanism. We use interactive search reasoning models to select relations and entities independently, thus fully mining the semantic information of relations and entities and improving the learning and reasoning ability of reasoning models. In the HSMH framework, we design the local-global knowledge fusion and action optimization mechanisms for path reasoning, which can enhance agents' state information and action space. Specifically, the local-global knowledge fusion mechanism is designed to acquire the local knowledge of entities and neighboring relations and the global knowledge about KG structure. This local-global knowledge can improve the learning ability of reasoning models. In addition, the action optimization mechanism can combine the filtered action space and the additional action space for efficient path reasoning for agents. Experimental results on five benchmark datasets show that our proposed HSMH framework comprehensively outperforms the state-of-the-art multi-hop reasoning model. Dan Wang 0002, Bo Li 0034, Bin Song 0001, Chen Chen 0128, F. Richard Yu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Inter-Intra Modal Representation Augmentation With Trimodal Collaborative Disentanglement Network for Multimodal Sentiment AnalysisabstractRecently, Multimodal Sentiment Analysis (MSA) is a challenging research area given its complex nature, and humans express emotional cues across various modalities such as language, facial expressions, and speech. Representation and fusion of features are the most crucial tasks in multimodal sentiment analysis research. However, in the current research, most methods ignore the importance of eliminating potential irrelevant features in the original features of each modality and cross-modal common feature. Moreover, the features extracted from all the modalities contain cluttered background noise and different occlusions noise, which negatively affects feature alignment. Different from these methods, we propose a novel Trimodal Collaborative Disentanglement Network (TCDN) to solve these problems in this paper. This work can obtain effective sentiment results on two aspects: i) Trimodal collaborative uses L1-norm to eliminate irrelevant features and unify the characteristics of the three modals (inter-modal). ii) Disentanglement network introduces an adversary noise by combining the original features of various single modalities and the common representation, alleviating the background noises within each modality (intramodal). This inter-intra modal feature augmentation method is the first work to obtain the common representation by implementing data augmentation as far as we know. Extensive experiments are completed on two benchmark datasets, including MOSI and MOSEI, demonstrating the superiority of the TCDN model over the state-of-the-art methods. Chen Chen 0128, Hansheng Hong, Jie Guo 0008, Bin Song 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | A VAE-Based User Preference Learning and Transfer Framework for Cross-Domain RecommendationabstractThe core idea of cross-domain recommendation is to alleviate the problem of data scarcity. Previous methods have made brilliant successes. However, many of them mainly focus on learning an ideal mapping function across-domains, ignoring the user preferences within a specific domain, which leads to suboptimal results. In this paper, we propose a Cross-Domain Recommendation Variational AutoEncoder framework (CDRVAE), a novel extension of a variational autoencoder on cross-domain recommendations for user behaviour distribution modeling. It applies a new hybrid architecture of VAE as the backbone and simultaneously constructs two information flows, within-domain and cross-domain modeling. For the former, an asymmetric codec structure is designed to reconstruct preference distribution from domain-specific latent factors. To relieve the posterior collapse dilemma, a combined prior is employed to increase the distribution complexity. The equivalent transition by a transformation matrix and the unobserved interaction generation by cross-domain reconstruction contribute to the latter. We combine all the above components for the more accurate and reliable user features. Extensive experiments are conducted on three public benchmark datasets to validate the effectiveness of the proposed CDRVAE. Experimental results demonstrate that CDRVAE is consistently superior to other state-of-the-art alternative baseline models. Tong Zhang 0015, Chen Chen 0128, Dan Wang 0002, Jie Guo 0008, Bin Song 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Trust-Aware Multi-Task Knowledge Graph for RecommendationabstractData sparsity and cold start problems are common in recommender systems. Adding some side information, such as knowledge graph and users' trust relationship, is an effective method to alleviate these problems. However, few work jointly explore the fine-grained implicit relationships between the external heterogeneous graphs to enhance the recommendation accuracy. To address this issue, in this paper, we propose a new method named Trust-aware Multi-task Knowledge Graph (TMKG), which uses multi-task learning to integrate two kinds of side information of trust graph and knowledge graph in an end-to-end manner. Firstly, we mine the intra-graph and inter-graph high-order connections through the node propagation and aggregation, and optimize the embedding of nodes through the implicit relationships obtained. Furthermore, through the shared cross unit, the connection relationships between each layer is mined, and the high-order interaction of nodes of different layers is obtained. We conduct extensive experiments on real-world datasets and prove that our model has the superior performance compared with the state-of-the-art models. Jie Guo 0008, Bin Song 0001, Chen Chen 0128, Jianglong Chang, F. Richard Yu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Inter-Intra Modal Representation Augmentation With DCT-Transformer Adversarial Network for Image-Text MatchingabstractImage-text matching has become a challenging task in the multimedia analysis field. Many advanced methods have been used to explore local and global cross-modal correspondence in matching. However, most methods ignore the importance of eliminating potential irrelevant features in the original features of each modality and cross-modal common feature. Moreover, the features extracted from regions in images and words in sentences contain cluttered background noise and different occlusion noise, which negatively affects alignment. Different from these methods, we propose a novel DCT-Transformer Adversarial Network (DTAN) for image-text matching in this paper. This work can obtain an effective metric based on two aspects: i) DCT-Transformer uses DCT (Discrete Cosine Transform) method based on a transformer mechanism to extract multi-domain common representations and eliminate irrelevant features from different modalities (inter-modal). Among them, DCT divides multi-modal content into chunks of different frequencies and quantifies them. ii) The adversarial network introduces an adversary idea by combining the original features of various single modalities and the multi-domain common representation, alleviating the background noise within each modality (intra-modal). The proposed adversarial feature augmentation method can easily obtain the common representation that is only useful for alignment. Extensive experiments are completed on the benchmark datasets Flickr30K and MS-COCO, demonstrating the superiority of the DTAN model over the state-of-the-art methods. Chen Chen 0128, Dan Wang 0002, Bin Song 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Dual Attention Transfer in Session-based Recommendation with Multi-dimensional IntegrationabstractSession-based recommendation (SBR) is widely used in e-commerce to predict the anonymous user's next click action according to a short sequence. Many previous studies have shown the potential advantages of applying Graph Neural Networks (GNN) to SBR tasks. However, the existing SBR models using GNN to solve user preference problems are only based on one single dataset to obtain one recommendation model during training. While the single dataset has the problems including the excessive sparse data source and the long-distance relationship of items. Therefore, introducing the dual transfer, which can enrich the data source, to SBR is absolutely necessary. To this end, a new method is proposed in this paper, which is called dual attention transfer based on multi-dimensional integration (DAT-MDI): (i) DAT uses a potential mapping method based on a slot attention mechanism to extract the user's representation information in different sessions between multiple domains. (ii) MDI combines the graph neural network for the graphs (session graph and global graph) and the gate recurrent unit (GRU) for the sequence to learn the item representation in each session. Then the multi-level session representation are combined by a soft-attention mechanism. We do a variety of experiments on four benchmark datasets which have shown that the superiority of the DAT-MDI model over the state-of-the-art methods. Chen Chen 0128, Jie Guo 0008, Bin Song 0001 |
SIGIR | 1 |