EDBT 2026 Demo / reviewers in the wild / expert
Keyan Jin
dblp:330/1689
· DBLP profile ↗
21ranked-venue papers
8as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distribution-aware Re-representations for Multi-Scenario RecommendationsabstractModern applications provided personalized recommendations across diverse scenarios, including the homepage, local pages, and live streams on platforms like TikTok. These scenarios exhibit varying user behavior patterns, resulting in heterogeneous yet interrelated distributions. Existing Multi-Scenario Recommendation (MSR) methods usually use parameter-sharing networks for shared features and scenario-specific networks for unique features. However, these methods fail to handle different distribution across scenarios, resulting in representation entanglement and localization, which hinder effective knowledge transfer and compromise performance. In this paper, we propose a Distribution-aware Re-representations (DAR) method for MSR. Its core idea is to construct distribution-aware prototype spaces and learn disentangled re-representations around global prototypes. Specifically, DAR employs a Multi-gate Mixture of Experts (MMoE) to obtain scenario-shared representations, and uses independent networks to learn scenario-specific representations. These representations are then projected into scenario-shared and scenario-specific prototype spaces, producing scenario-shared re-representations (capturing global information) and scenario-specific re-representations (focusing on distributional differences). During this process, DAR utilizes Unbalanced Optimal Transport (UOT) to compute the transport relationships between representations and global prototypes, taking these as pseudo-labels for re-representation learning. Moreover, to prevent prototype entanglement, a matrix orthogonalization constraint ensures independence among global prototypes. The effectiveness of DAR is demonstrated through extensive offline experiments conducted on four datasets, as well as online A/B tests on a video platform. Xiaoyu Kang, Keyan Jin, Jiechao Gao |
SIGIR | 5 |
| 2026 | HiCORE: Enhancing consumer intent prediction via hybrid attention and contrastive learning in sequential recommendationabstractAccurately predicting consumer intent from sequential interactions is essential for enhancing recommendation accuracy and driving effective marketing strategies. In this paper, we introduce HiCORE, a Transformer-based sequential recommendation model integrating Layer-wise Hybrideak Attention, Rotary Positional Embeddings (RoPE), and a Contrastive Self-Supervised Learning (SSL) objective. HiCORE effectively captures both global consumer preferences and local contextual interests, explicitly encodes relative positional information, and ensures robust sequence-level representations, addressing common limitations such as data sparsity and noise. Extensive experiments on three real-world datasets demonstrate that HiCORE significantly outperforms state-of-the-art methods in recommendation accuracy metrics. Moreover, HiCORE’s enhanced capability in consumer intent prediction provides valuable insights for consumer segmentation and targeted marketing, bridging advanced computational methods and practical business applications. Keyan Jin, Francisco Javier Blanco-Encomienda |
Expert Syst. Appl. | 1 |
| 2026 | Reasoning or not? A comprehensive evaluation of reasoning LLMs for dialogue summarization
Keyan Jin, Yapeng Wang 0001, Leonel Santos, Xu Yang 0010, Sio Kei Im, Hugo Gonçalo Oliveira |
Expert Syst. Appl. | 1 |
| 2026 | Densifying Knowledge Hypergraphs with LLM-based review distillation for conversational consumer decision supportabstract• Review-driven structural induction alleviates sparsity in conversational recommendation. • LLM-distilled review semantics are grounded as entities to densify hypergraph structures. • Item-centric hypergraph learning supports joint recommendation and response generation. • Two-stage training stabilizes structure learning under sparse conversational supervision. Keyan Jin, Francisco Javier Blanco-Encomienda |
Inf. Process. Manag. | 1 |
| 2025 | FDS-Net: A Frequency-Decomposed Synergistic Network for Multivariate Time Series PredictionabstractTime series prediction is widely used in traffic planning, weather forecasting, and energy management. However, real-world time series often exhibit complex dynamic characteristics, such as multi-scale patterns and dynamic dependencies, making predictions challenging. Existing methods can model these patterns but struggle with component interactions and timestamp utilization. To address these issues, we propose FDS-Net (Frequency-Decomposed Synergistic Network), which enhances interactions and timestamp utilization through three core modules: Frequency-Decomposed and Reconstructed Block, which decomposes time series into components for explicit modeling and facilitates interactions through signal reconstruction; Multi-Scale Temporal Semantic Extraction Block, which uses multi-scale convolutions to extract diverse temporal features; and Temporal Feature Interaction Block which integrates timestamp features with reconstructed signals and uses attention mechanisms to capture dynamic dependencies. Additionally, we introduce an adaptive dynamic weighted loss function to optimize training. Experimental results show FDS-Net’s superior performance in long-term and short-term forecasting, providing new insights for time series prediction. Zhijiang Wang, Jinzhe Liang, Keyan Jin |
ECAI | 4 |
| 2025 | SSCM: Self-Supervised Critical Model for Reducing Hallucinations in Chinese Financial Text GenerationabstractLarge Language Models (LLMs) show strong performance in natural language processing tasks, but their application in the financial domain is limited. Current methods rely on large datasets and manual prompt engineering, resulting in high data demands, long inference times, and frequent hallucinations. To address these limitations, we propose a novel self-supervised prompt optimization framework tailored for the financial domain. Our approach involves training a critical model that evaluates and ranks generated outputs using both good and bad answers generated from various revised prompts. Experiments on a large Chinese financial corpus show that our framework significantly improves performance on tasks such as summarization and event-based question answering, as evidenced by higher scores on both automated metrics like ROUGE, BLEU, and BERTScore, and also through human evaluations. These results validate the effectiveness of our method in reducing hallucinations and improving the quality of financial text generation. Keyan Jin, Yapeng Wang 0001, Leonel Santos, Xu Yang 0010, Sio Kei Im |
ICASSP | 1 |
| 2025 | LDGNet: LLMs Debate-Guided Network for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection aims to uncover the sarcasm emotions expressed through various modalities such as text and image. Previous work has made enlightening exploration in detecting sarcastic sentiments with given domains. However, there remains a gap in utilizing deeper contextual information to capture elusive sarcastic clues, hidden in open-world knowledge such as history, politics, and common sense of life that has not been touched by previous models. To address this gap, a natural idea is to simulate the process of a debate, involving debaters with different viewpoints and judges to collaboratively drive the judgment of emotional expressions. Benefiting from the development of large multimodal language models, and building upon previous advancements, we propose a novel framework called LLMs Debate-Guided Network (LDGNet) for Multimodal Sarcasm Detection. LDGNet effectively leverages large language model debates to uncover subtle emotional information and uses an innovative Judge Network for more realible and accurate sentiment judgments. Extensive experiments on in-domain and out-of-distribution (OOD) datasets have validated the superiority of our proposed method. Hengyang Zhou, Jinwu Yan, Rongman Hong, Wenbo Zuo, Keyan Jin |
ICASSP | 6 |
| 2025 | SM-CBNet: A Speech-Based Parkinson's Disease Diagnosis Model with SMOTE-ENN and CNN + BiLSTM Integration
Weichao Pan, Ruida Liu, Zhen Tian 0002, Keyan Jin |
ICIC (26) | 5 |
| 2025 | Multi-Object Grounding via Hierarchical Contrastive Siamese TransformersabstractMulti-object grounding in 3D scenes involves localizing multiple objects based on natural language input. While previous work has primarily focused on single-object grounding, real-world scenarios often demand the localization of several objects. To tackle this challenge, we propose Hierarchical Contrastive Siamese Transformers (H-COST), which employs a Hierarchical Processing strategy to progressively refine object localization, enhancing the understanding of complex language instructions. Additionally, we introduce a Contrastive Siamese Transformer framework, where two networks with the identical structure are used: one auxiliary network processes robust object relations from ground-truth labels to guide and enhance the second network, the reference network, which operates on segmented point-cloud data. This contrastive mechanism strengthens the model’s semantic understanding and significantly enhances its ability to process complex point-cloud data. Our approach outperforms previous state-of-the-art methods by 9.5% on challenging multi-object grounding benchmarks. Chengyi Du, Keyan Jin |
IJCNN | 2 |
| 2025 | DRMoE: Discourse-Aware Rhetorical Structure for Meeting Summarization with Mixture of ExpertsabstractMeeting summarization is a challenging task that requires condensing multi-party conversations while preserving their discourse coherence and key information. Despite advancements in large language models (LLMs), current approaches often lack the ability to explicitly model discourse structures, leading to potential information loss and incoherent summaries. To address these limitations, we propose DRMoE, a novel discourse-aware rhetorical structure framework for meeting summarization. DRMoE leverages probabilistic Rhetorical Structure Theory (RST) parsing to model hierarchical discourse relations and integrates a Mixture of Experts (MoE) mechanism for dynamic expert allocation, enabling the model to process diverse rhetorical functions effectively. Our framework incorporates discourse relations into the attention mechanism of LLaMA-3, guiding content selection and enhancing coherence. Evaluated on the QMSum dataset, DRMoE achieves state-of-the-art results across Product, Academic, and Committee domains, outperforming baselines such as BART and PEGASUS in both automatic and human evaluations. Additionally, a case study illustrates the model’s ability to generate domain-specific summaries closer to human references, while ablation studies confirm the critical contributions of RST and MoE components. These findings highlight DRMoE’s effectiveness in producing informative, coherent, and contextually faithful summaries, paving the way for advancements in meeting summarization systems. Keyan Jin |
IJCNN | 1 |
| 2025 | DABART: Dynamic Semantic Optimization Framework for Dialogue Summarization via Adaptive Topic Analysis and Semantic BridgingabstractWith the increasing prevalence of online communication and automated services, dialogue summarization technology plays a vital role in meeting minutes, customer service, and online Q&A scenarios. However, existing methods often suffer from insufficient flexibility in topic segmentation, low efficiency in semantic information transfer, and limited role adaptability. To address these challenges, we propose DABART, a dynamic semantic optimization framework. The framework employs a dynamic semantic topic segmentation mechanism to adaptively segment topics based on the distribution characteristics of sentence embeddings within dialogues, effectively identifying key information while overcoming the limitations of fixed-parameter methods in complex dialogue scenarios. Additionally, a dynamic semantic bridging module integrates semantic and positional information, further enhancing the coherence and consistency of dialogue summarization. Experimental results demonstrate that DABART achieves superior performance on widely-used benchmarks such as SAMSum and CSDS. Notably, it surpasses state-of-the-art open-source models on the CSDS dataset in ROUGE and BERTScore metrics, while achieving more balanced and accurate role-oriented summarization. Extensive experimental analyses further validate the robustness and applicability of the DABART across diverse dialogue scenarios. Keyan Jin, Yapeng Wang 0001, Leonel Santos, Xu Yang 0010, Sio Kei Im |
IJCNN | 1 |
| 2025 | Exploring Disentangled Appearance-Motion Contexts for Temporal Activity LocalizationabstractTemporal Activity Localization (TAL) is crucial and fundamental for multimedia understanding. Although many works have made great efforts and achieved significant progress on this task, most of them directly utilize the mixed visual features extracted by the 3D backbone network to match with the complicated query semantic, thus failing to capture the subtly distinct visual features associated with the interested entities or events for better activity modeling. To overcome this challenge, in this paper, we present a novel Disentangled Appearance-Motion Learning (DAML) framework that is able to learn the disentangled representations and capture finer levels of granularity across different modalities, such as nouns-related visual appearance or verbs-related visual motion for more interpretable cross-modal alignment. Specifically, without introducing any large feature extraction model, we disentangle the mixed video feature extracted by 3D backbone into separate appearance and motion contexts with the help of vector quantization. In this way, we can achieve more fine-grained correspondence between the visual appearance and textual nouns, visual motion and textual verbs for better modeling the object entities, events of the target activity. Extensive experiments on three challenging datasets (Charades-STA, TACoS and ActivityNet) show the effectiveness of DAML. Huashuo Lei, Xiaowen Cai 0001, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Jixiang Yu, Keyan Jin |
IJCNN | 8 |
| 2025 | MonoAttack: A Strong Attack Framework with Depth-Migration and Attribute-Tampering for Monocular 3D Object DetectionabstractAlthough many efforts have been made into attacks on deep neural networks (DNNs) in recent years, no research explores the vulnerability of monocular 3D object detection (M3D) models. This M3D task is fundamental but essential in safety-critical 3D applications, potentially bringing hazards to autonomous driving. In this paper, we thoroughly investigate the sensitivity of current M3D models to adversarial noise and propose a novel M3D adversarial attack method called MonoAttack. The key insight of our method is exploring both depth-migration and attribute-tampering for generating M3D adversarial samples. Specifically, in addition to the general misleading of the detection model, we deceive the M3D model by changing the potential object depth into its opposite position. We also guide the M3D model to mis-recognize the class attribute of its detected object for generating low-confidence bounding boxes. Moreover, we further disentangle the depth knowledge from the geometric and semantic perspectives to auxiliary correlate the detection and attribute information for jointly generating the latent perturbation. In this manner, our attack framework is strong and can effectively attack M3D models with trivial perturbations. Experimental results on the KITTI dataset demonstrate that our attack achieves high adversarial ability against current monocular 3D detection models. Xiayue Zhang, Huashuo Lei, Daizong Liu, Xiaoye Qu, Runwei Guan, Keyan Jin |
IJCNN | 7 |
| 2025 | Manipulating the Bounding Box: Multimodal Controlled Backdoor Attacks on 3D Visual Grounding Modelsabstract3D visual grounding models, pivotal in interpreting and aligning objects within 3D spaces with textual descriptions, have become integral to the advancement of the multimedia community. As these models are widely used in daily life as real-world applications, they become more susceptible to be attacked. Backdoor attacks are designed to corrupt a model in such a way that it responds with adversary-wanted outputs when specific trigger patterns are introduced, while responding normally to clean inputs. Unlike traditional backdoor attack methods that focus on attacking simple classification models, attacking 3D visual grounding models presents unique challenges due to their multi-modal inputs and the nature of their output, which is the localization box of objects described by the text within the 3D scene. This necessitates distinct attack strategies and trigger designs, adding complexity to executing successful attacks. To this end, in this paper, we present a novel multimodal controlled backdoor attack aimed at manipulating the positioning and size of bounding boxes in the challenging multi-modal 3D visual grounding models. Specifically, we design triggers for both point cloud and textual modalities, along with specialized placement strategies for each, to enhance the stealth and precision of the attack. Furthermore, we develop optimization strategies to enhance the efficacy of the point cloud trigger. Experimental results across various standard models confirm the effectiveness of our backdoor attack method, with negligible impact on performance in clean datasets. Xiayue Zhang, Huashuo Lei, Daizong Liu, Xiaoye Qu, Runwei Guan, Keyan Jin |
IJCNN | 7 |
| 2025 | NILMixer: A Novel Multi-Seq2Seq Model for Load Disaggregation in Long-Term WindowsabstractNeural network models have markedly enhanced the field of Non-Intrusive Load Monitoring (NILM) compared to traditional approaches. However, two significant challenges persist. First, many studies overlook the design of model input paradigms and the role of feature interactions, leading to models with increasingly complex architectures and larger parameter counts, but with diminishing returns in performance. Second, existing models struggle with long-window input scenarios for load disaggregation, failing to balance the influences of local and global features effectively. This typically leads to a high rate of false positives and missed detections in identifying device energy usage events.To address the challenges associated with long-window disaggregation, this paper introduces a model named NILMixer. Constructed solely from linear and convolutional layers, NILMixer achieves an optimal balance between cost and performance through a layered approach to multi-scale feature interaction. Experimental results indicate that the proposed model not only outperforms several state-of-the-art methods but also effectively handles load disaggregation tasks across a spectrum from short to long window durations. Additionally, interpretability and ablation studies provide robust evidence and insights into the rational design of the model and the foundational role of its multi-scale features. Houyi Zhu, Zonghui Wang, Youlong Zhang, Keyan Jin |
IJCNN | 6 |
| 2025 | Prototype-Guided Representation Projection for Multi-Domain Multi-Task RecommendationabstractMulti-domain and multi-task learning enhance the efficiency and performance of industrial recommendation systems by integrating information from different domains/tasks to model user interests uniformly. However, existing methods suffer from the problem of representation entanglement, which limits the effective handling of commonality and specificity among various domains/tasks. In this paper, we propose a Prototype-guided Representation Projection (PRP) model to address this issue, which explores a novel direction of applying prototype learning to deal with complex domain/task relationships in the recommendation field. To identify inter-domain/task commonality, PRP initially uses a shared Mixture of Experts (MoE) architecture to learn representations for each sample, projecting them into a common prototype space across all domains/tasks. For domain/task specificity, specific feature extraction experts are employed, and sample representations are projected to the corresponding prototype spaces, constrained by an orthogonal loss to ensure the independence of those spaces. Moreover, PRP utilizes Optimal Transport (OT) to guide the correct representation projection within the prototype spaces, employing the linear combination of prototypes as the new sample representation. We conduct offline experiments on two open-source datasets and deploy our approach in an online system for A/B testing. Extensive experimental results consistently demonstrate that our approach outperforms existing methods. Binrui Wu, Haochen Sui, Jiaye Lin, Jiechao Gao, Keyan Jin |
ACM Multimedia | 6 |
| 2025 | HiSum: Hierarchical Topic-Driven Approach for Role-Oriented Dialogue SummarisationabstractABSTRACT As the volume of information on online communication platforms continues to grow, the task of dialogue summarisation becomes increasingly critical for understanding and extracting key information from diverse conversations. Traditional approaches often struggle to cope with the dynamic nature of dialogues, such as managing perspectives from multiple speakers and seamlessly transitioning between different topics. We propose a novel hierarchical topic‐driven approach to generate role‐oriented dialogue summarisation (HiSum) to address these challenges. First, we utilise VarGMM clustering technology for in‐depth topic segmentation, which enables the model to capture the key topics in a dialogue. Second, we employ a LayerAttn hierarchical attention mechanism to dynamically adjust the focus of dialogue content based on participants' importance and the topics' relevance. Experimental results on three public dialogue summarisation data sets (CSDS, MC and SAMSUM) demonstrate that our method significantly outperforms most existing strong baseline methods across various evaluation metrics and surpasses the current state‐of‐the‐art methods in certain metrics. Detailed analysis demonstrates that HiSum can perform more precise topic segmentation and effectively identify critical information. Our code is publicly available at: https://github.com/kjin0119/HiSum . Keyan Jin, Yapeng Wang 0001, Xu Yang 0010, Sio Kei Im |
Expert Syst. J. Knowl. Eng. | 1 |
| 2025 | LLMCL-GEC: Advancing grammatical error correction with LLM-driven curriculum learningabstractWhile large-scale language models (LLMs) have demonstrated remarkable capabilities in specific natural language processing (NLP) tasks, they may still lack proficiency compared to specialized models in certain domains, such as grammatical error correction (GEC). Drawing inspiration from the concept of curriculum learning, we have delved into refining LLMs into proficient GEC experts by devising effective curriculum learning (CL) strategies. In this paper, we introduce a novel approach, termed LLM-based curriculum learning, which capitalizes on the robust semantic comprehension and discriminative prowess inherent in LLMs to gauge the complexity of GEC training data. Unlike traditional curriculum learning techniques, our method closely mirrors human expert-designed curriculums. Leveraging the proposed LLM-based CL method, we sequentially select varying levels of curriculums ranging from easy to hard, and iteratively train and refine using the pretrianed T5 and LLaMA series models. Through rigorous testing and analysis across diverse benchmark assessments in English GEC, including the CoNLL14 test, BEA19 test, and BEA19 development sets, our approach showcases a significant performance boost over baseline models and conventional curriculum learning methodologies. Specifically, our method achieves a new state-of-the-art (SOTA) result on the CoNLL14 test set, with an F 0 . 5 score of 69.6. Additionally, on the BEA19 test set and BEA19 development set, our approach outperforms conventional curriculum learning methodologies by 1.0 and 0.3 F 0 . 5 points, respectively. Derek F. Wong, Keyan Jin, Lusheng Zhang, Qiang Zhang 0055, Tianjiao Li 0001, Jinlong Hou, Lidia S. Chao |
Expert Syst. Appl. | 4 |
| 2024 | Cross-Domain Transfer in Residual Networks for Clinical Image PartitioningabstractThe accurate recognition and comprehensive understanding of medical images depicting human tissue represent a central focus in computer vision research. Many tasks within medical imaging rely on deep neural networks, particularly those with U-shaped architectures and skip connections. The advancement of computer vision technologies demands the application of convolutional neural networks (CNNs). Despite progress, two major challenges remain in medical image processing: (1) developing a model framework with low computational complexity that allows for efficient inference without compromising accuracy, and (2) designing a model with strong generalization capabilities across various datasets derived from patients with differing pathologies, thereby mitigating domain shift challenges. In response to the first issue, we propose a novel unsupervised domain adaptation method utilizing Interoperable Batch Normalization (IBN) to integrate multiple channels within deep neural networks, enhancing adversarial domain adaptation. Our experimental evaluation on the Hubmap and Synapse multiorgan segmentation datasets reveals that the proposed RRUNet model achieves superior performance compared to existing methods, setting a new standard in the domain. Mingshuo Wang, Keyan Jin, Wenzhuo Bao, Zhenghan Chen |
BIBM | 2 |
| 2023 | Impacts of Word of Mouth (WOM) on E-Business Online PricingabstractIn response to the rising power of electronic word of mouth (eWOM), marketers are gradually placing importance on the effects of their marketing strategies. The main objective of the study is to use both secondary as well as primary data to investigate the relationship and interaction between eWOM and online pricing. This study took BizRate UK, a price comparison website, as the database to investigate 100 mobile phone and tablet products and conducted interviews with online shoppers and managers from online retailers. After applying the test from quantitative methodologies, the result indicated the positive and negative reviews from online shoppers had an effect on pricing performance. Two types of reviews can be sorted out in a specific way as well. Finally, because the subject covers both quantitative (pricing) and qualitative (eWOM) aspects, future studies that are interested in this area should still apply the two methodologies for more accurate investigation. Keyan Jin |
J. Glob. Inf. Manag. | 1 |
| 2022 | Financial Risk Early Warning Model of Listed Companies Under Rough Set Theory Using BPNNabstractIn order to reduce the risk of enterprise management, the financial risk early warning methods of listed companies are mainly studied. The financial risk characteristics of listed companies are analysed. With the help of rough set theory, the financial risk indicators are selected, and the financial risk early warning index system is established. The financial risk early warning model is constructed by using back propagation neural network (BPNN) algorithm based on deep learning. Finally, the accuracy and feasibility of the constructed neural network model are verified. The results show that rough set theory can be used to screen financial risk indicators and select important indicators, which can simplify the data and reduce the complexity of calculation. BPNN can calculate the simplified data and identify and evaluate the financial risk. Empirical analysis shows that the proposed method can shorten the training time of the model to a certain extent, and improve the accuracy of financial risk prediction. Chengai Li, Keyan Jin, Ziqi Zhong, Kunzhi Tang |
J. Glob. Inf. Manag. | 2 |