VLDB 2026 Research / reviewers in the wild / expert
Kehui Song
dblp:197/1051
· DBLP profile ↗
35ranked-venue papers
3as first author
35since 2021 · last 2026
0000-0001-9694-6359ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Security and privacy · 8 · 8 since 2021Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information ExtractionabstractMultimodal information extraction (MIE) constitutes a set of essential tasks aimed at extracting structural information from Web texts with integrating images, to facilitate the structural construction of Web-based semantic knowledge. To address the expanding category set including newly emerging entity types or relations on websites, prior research proposed the zero-shot MIE (ZS-MIE) task which aims to extract unseen structural knowledge with textual and visual modalities. However, the ZS-MIE models are limited to recognizing the samples that fall within the unseen category set, and they struggle to deal with real-world scenarios that encompass both seen and unseen categories. The shortcomings of existing methods can be ascribed to two main aspects. On one hand, these methods construct representations of samples and categories within Euclidean space, failing to capture the hierarchical semantic relationships between the two modalities within a sample and their corresponding category prototypes. On the other hand, there is a notable gap in the distribution of semantic similarity between seen and unseen category sets, which impacts the generative capability of the ZS-MIE models. To overcome the above disadvantages, we delve into the generalized zero-shot MIE (GZS-MIE) task and propose the hyperbolic multimodal generative representation learning framework (HMGRL). The variational information bottleneck and autoencoder networks are reconstructed with hyperbolic space for modeling the multi-level hierarchical semantic correlations among samples and prototypes. Furthermore, the proposed model is trained with the unseen samples generated by the decoder, and we introduce the semantic similarity distribution alignment loss to enhance the model's generalization performance. Experimental evaluations on two benchmark datasets underscore the superiority of HMGRL compared to existing baseline methods. Baohang Zhou, Kehui Song, Rize Jin, Yu Zhao 0043, Xuhui Sui, Xinying Qian, Xingyue Guo, Ying Zhang 0015 |
WWW | 2 |
| 2026 | ByteDance: Let bytes perform brilliantly in multi-view encrypted traffic classification
Yuwei Xu 0001, Zhiyuan Liang, Xiaotian Fang, Kehui Song, Qiao Xiang, Guang Cheng 0001 |
Comput. Networks | 4 |
| 2026 | PacketPatch: Practical generation and deployment of adversarial packets for byte-feature-based encrypted traffic classification
Yuwei Xu 0001, Yunpeng Bai, Kehui Song, Jie Cao 0009, Qiao Xiang, Guang Cheng 0001 |
Comput. Secur. | 5 |
| 2026 | A duet of perception and reasoning: CLIP and LLM brainstorming for scene text recognition
Zeguang Jia, Kehui Song, Zhilan Wang, Rize Jin |
Neurocomputing | 3 |
| 2025 | PacketMorph: Generation of Recoverable Adversarial Packets Against Encrypted Traffic Classification via Class-Wise Universal Perturbation
Yuwei Xu 0001, Yunpeng Bai, Jie Cao 0009, Kehui Song, Guang Cheng 0001 |
ICA3PP (4) | 5 |
| 2025 | Mind Map-guided Meta-prompting for ADHD Intervention with Large Language Models
Xuguang Qiu, Kehui Song, Rize Jin |
ICONIP (2) | 3 |
| 2025 | MaCSE: Multi-Agent Ranking Distillation for Contrastive Learning of Sentence EmbeddingsabstractSentence embedding models are typically trained using the Contrastive Learning (CL) method, which works by pulling similar semantics closer and pushing dissimilar ones away. Recent studies have shown that utilizing a multi-teacher ranking distillation approach, which assigns fine-grained rankings to sentences, enables the generation of smoother sentence similarity representations and results in higher-quality sentence embeddings. However, the effectiveness of distillation may be limited by the capacity of the student model. A simple student model with fewer parameters may struggle to approximate a highly complex teacher model, potentially leading to overfitting on certain datasets or specific aspects of the task. To address this, we propose MaCSE, a multi-agent ranking distillation framework that dynamically selects and optimizes teacher model contributions across training stages. MaCSE employs a Centralized Training with Decentralized Execution (CTDE) paradigm, enabling collaborative agent interactions to adaptively adjust teacher fusion weights based on training dynamics. Experimental results on Semantic Textual Similarity and transfer tasks demonstrate that MaCSE outperforms most existing baselines and even rivals methods using large language models for sentence representation. Our implementation is available at GitHub1. Zekai Zhi, Zhilan Wang, Rize Jin, Kehui Song, Da-Jung Cho |
IJCNN | 4 |
| 2025 | From Chain to Loop: Improving Reasoning Capability in Small Language Models via Loop-of-Thought
Mingxin Ji, Kehui Song, Rize Jin, Xuguang Qiu |
NLPCC (1) | 2 |
| 2025 | Introducing high correlation and high quality instances for few-shot entity linking
Xuhui Sui, Ying Zhang 0015, Kehui Song, Baohang Zhou, Xiaojie Yuan |
Neural Networks | 3 |
| 2025 | ZS-MNET: A zero-shot learning based approach to multimodal named entity typing
Baohang Zhou, Ying Zhang 0015, Kehui Song, Xuhui Sui, Yu Zhao 0043, Xiaojie Yuan |
Neural Networks | 3 |
| 2024 | Bring Invariant to Variant: A Contrastive Prompt-based Framework for Temporal Knowledge Graph ForecastingabstractTemporal knowledge graph forecasting aims to reason over known facts to complete the missing links in the future. Existing methods are highly dependent on the structures of temporal knowledge graphs and commonly utilize recurrent or graph neural networks for forecasting. However, entities that are infrequently observed or have not been seen recently face challenges in learning effective knowledge representations due to insufficient structural contexts. To address the above disadvantages, in this paper, we propose a Contrastive Prompt-based framework with Entity background information for TKG forecasting, which we named CoPET. Specifically, to bring the time-invariant entity background information to time-variant structural information, we employ a dual encoder architecture consisting of a candidate encoder and a query encoder. A contrastive learning framework is used to encourage the query representation to be closer to the candidate representation. We further propose three kinds of trainable time-variant prompts aimed at capturing temporal structural information. Experiments on two datasets demonstrate that our method is effective and stays competitive in inference with limited structural information. Our code is available at https://github.com/qianxinying/CoPET. Ying Zhang 0015, Xinying Qian, Yu Zhao 0043, Baohang Zhou, Kehui Song, Xiaojie Yuan |
LREC/COLING | 5 |
| 2024 | MCIL: Multimodal Counterfactual Instance Learning for Low-resource Entity-based Multimodal Information ExtractionabstractMultimodal information extraction (MIE) is a challenging task which aims to extract the structural information in free text coupled with the image for constructing the multimodal knowledge graph. The entity-based MIE tasks are based on the entity information to complete the specific tasks. However, the existing methods only investigated the entity-based MIE tasks under supervised learning with adequate labeled data. In the real-world scenario, collecting enough data and annotating the entity-based samples are time-consuming, and impractical. Therefore, we propose to investigate the entity-based MIE tasks under the low-resource settings. The conventional models are prone to overfitting on limited labeled data, which can result in poor performance. This is because the models tend to learn the bias existing in the limited samples, which can lead them to model the spurious correlations between multimodal features and task labels. To provide a more comprehensive understanding of the bias inherent in multimodal features of MIE samples, we decompose the features into image, entity, and context factors. Furthermore, we investigate the causal relationships between these factors and model performance, leveraging the structural causal model to delve into the correlations between the input features and output labels. Based on this, we propose the multimodal counterfactual instance learning framework to generate the counterfactual instances by the interventions on the limited observational samples. In the framework, we analyze the causal effect of the counterfactual instances and exploit it as a supervisory signal to maximize the effect for reducing the bias and improving the generalization of the model. Empirically, we evaluate the proposed method on the two public MIE benchmark datasets and the experimental results verify the effectiveness of it. Baohang Zhou, Ying Zhang 0015, Kehui Song, Hongru Wang 0003, Yu Zhao 0043, Xuhui Sui, Xiaojie Yuan |
LREC/COLING | 3 |
| 2024 | TimeR⁴ : Time-aware Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question AnsweringabstractTemporal Knowledge Graph Question Answering (TKGQA) aims to answer temporal questions using knowledge in Temporal Knowledge Graphs (TKGs).Previous works employ pre-trained TKG embeddings or graph neural networks to incorporate the knowledge of TKGs.However, these methods fail to fully understand the complex semantic information of time constraints.In contrast, Large Language Models (LLMs) have shown exceptional performance in knowledge graph reasoning, unifying both semantic understanding and structural reasoning.To further enhance LLMs' temporal reasoning ability, this paper aims to integrate temporal knowledge from TKGs into LLMs through a Time-aware Retrieve-Rewrite-Retrieve-Rerank framework, which we named TimeR 4 .Specifically, to reduce temporal hallucination in LLMs, we propose a retrieve-rewrite module to rewrite questions using background knowledge stored in the TKGs, thereby acquiring explicit time constraints.Then, we implement a retrievererank module aimed at retrieving semantically and temporally relevant facts from the TKGs and reranking according to the temporal constraints.To achieve this, we fine-tune a retriever using the contrastive time-aware learning framework.Our approach achieves great improvements, with relative gains of 47.8% and 22.5% on two datasets, underscoring its effectiveness in boosting the temporal reasoning abilities of LLMs.Our code is available at https://github.com/qianxinying/TimeR4 . Xinying Qian, Ying Zhang 0015, Yu Zhao 0043, Baohang Zhou, Xuhui Sui, Kehui Song |
EMNLP | 7 |
| 2024 | NuanceTracker: A Website Fingerprinting Attack against Tor Hidden Services through Burst patternsabstractHidden services (HS) allow users to experience anonymity, but they also provide shelter for criminal activities. The widespread attention towards deanonymizing HS has brought website fingerprinting attack (WFA) into the spotlight, which is considered highly promising. However, most HS websites are designed simply and have high similarity in resource structures, making it difficult to represent the HS access traffic well, and existing work often directly applies traffic representation methods in the field of web research, resulting in poor effects of the model. Besides, features of HS access traffic are closely related to the resource access sequence of websites. Current studies build models based on convolutional neural network (CNN), ignoring the global correlation of HS access traffic parts. To address the short-comings, we have proposed an efficient WFA to deanonymize HS, and named it NuanceTracker. The contribution of our work lies in three points. Firstly, a burst-based HS fingerprint generation algorithm is proposed to describe the sequence of HS access traffic. Secondly, we propose NuanceTracker, which is designed by introducing multi-scale global attention (MGA) into a basic CNN model for global information extraction. Finally, comparison experiments are conducted in closed-world and open-world scenarios. Our NuanceTracker has proven to outperform three state-of-the-art WFA methods. Yuwei Xu 0001, Yujie Hou, Kehui Song, Guang Cheng 0001 |
ISCC | 5 |
| 2024 | FullView: Using Bidirectional Group Sequences to Achieve Accurate Encrypted Traffic Classification
Yuwei Xu 0001, Zhiyuan Liang, Zhengxin Xu, Kehui Song, Qiao Xiang, Guang Cheng 0001 |
SecureComm (2) | 4 |
| 2024 | Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph CompletionabstractA large number of studies have emerged for Multimodal Knowledge Graph Completion (MKGC) to predict the missing links in MKGs. However, fewer studies have been proposed to study the inductive MKGC (IMKGC) involving emerging entities unseen during training. Existing inductive approaches focus on learning textual entity representations, which neglect rich semantic information in visual modality. Moreover, they focus on aggregating structural neighbors from existing KGs, which of emerging entities are usually limited. However, the semantic neighbors are decoupled from the topology linkage and usually imply the true target entity. In this paper, we propose the IMKGC task and a semantic neighbor retrieval-enhanced IMKGC framework CMR, where the contrast brings the helpful semantic neighbors close, and then the memorize supports semantic neighbor retrieval to enhance inference. Specifically, we first propose a unified cross-modal contrastive learning to simultaneously capture the textual-visual and textual-textual correlations of query-entity pairs in a unified representation space. The contrastive learning increases the similarity of positive query-entity pairs, therefore making the representations of helpful semantic neighbors close. Then, we explicitly memorize the knowledge representations to support the semantic neighbor retrieval. At test time, we retrieve the nearest semantic neighbors and interpolate them to the query-entity similarity distribution to augment the final prediction. Extensive experiments validate the effectiveness of CMR on three inductive MKGC datasets. Codes are available at https://github.com/OreOZhao/CMR. Yu Zhao 0043, Ying Zhang 0015, Baohang Zhou, Xinying Qian, Kehui Song, Xiangrui Cai |
SIGIR | 5 |
| 2024 | M-ETC: Improving Multi-Task Encrypted Traffic Classification by Reducing Inter-Task InterferenceabstractWith the rapid evolution of deep learning (DL), its integration in encrypted traffic classification (ETC) can automatically extract key features from raw traffic data, enhancing classification performance. So far, researchers have proposed many DL-based models for ETC. However, the complexity and dynamism of network applications lead to the diversification of ETC tasks. Current models, mostly tailored for single tasks, overlook real-world multi-tasking needs of network devices. Deploying task-specific complex models concurrently on resource-limited devices is impractical. In response to the increasing number of tasks, researchers have introduced multi-task learning frameworks for ETC, demonstrating its potential as a promising technical approach. However, current research overlooks the interference between tasks, resulting in flawed models when it comes to sharing parameters, setting learning rates, and determining loss values. Aiming at these deficiencies, we propose $\mathcal{M}$-ETC, a multi-task ETC method reducing inter-task interference. The innovation of $\mathcal{M}$-ETC lies in two aspects. Firstly, we design a hierarchical multi-task learning model (HMLM) to provide effective features for each task and prevent the impact of invalid features. Secondly, we propose a learning rate balancing strategy (LRB) for modules and a dynamic weight average strategy (DWA) for tasks’ loss values. During model training, LRB prevents overfitting and underfitting of tasks, while DWA prevents bias towards tasks with large loss values. To validate $\mathcal{M}$-ETC, we carry out comparative experiments using four encrypted traffic datasets. The experimental results show that the classification performance of $\mathcal{M}$-ETC on multiple tasks exceeds those of five state-of-the-art methods. Yuwei Xu 0001, Xiaotian Fang, Zhengxin Xu, Kehui Song, Yali Yuan, Guang Cheng 0001 |
TrustCom | 4 |
| 2024 | A High-Accuracy Unknown Traffic Identification Method Based on Multi-View Contrastive LearningabstractThe technology for AI-based encrypted traffic classification (ETC) is advancing rapidly. However, many current studies are conducted in closed network environments where traffic is classified into pre-determined classes. In the actual network environment, new traffic is constantly emerging, making anomaly detection of unknown traffic a pressing issue. The current research attempts different approaches for unknown traffic identification (UTI) from the perspective of model construction, including UTI based on n-classification, UTI based on multiple 2-classifiction, and UTI baed on (n+1)-classification. The above three approaches mentioned are prone to misclassification and have low accuracy due to issues with threshold setting, poor generalization of binary classifiers, and low credibility of generated samples. Researchers have used contrastive learning (CL) for UTI because of its strengths in feature representation. However, there are two shortcomings in the current studies. First, the feature representation of a single view cannot fully capture different classes of traffic features. Second, the existing CL schemes distinguish whether they belong to the same category by constructing pairs of positive and negative samples, but they are still unable to distinguish between known classes and unknown classes in the feature space, resulting in low accuracy. To address the above shortcomings, we propose UTI-MCL, an unknown traffic identification method using multi-view contrastive learning. Firstly, in terms of feature expression, we extract the packet length sequence and the packet byte sequence respectively to learn a more comprehensive feature representation. Secondly, in terms of model construction, we introduce an anchor and compare the distance with both positive and negative samples to help the model better separate different classes in the feature space, enhancing feature distinction. Furthermore, the distance between samples is optimized through the adaptive weights triplet loss function to balance samples from different classes. A series of experiments have proved the effectiveness of UTI-MCL. Even with unknown traffic accounting for 60%, the Fβ-Score can still exceed 94%. Yuwei Xu 0001, Zizhi Zhu, Chufan Zhang, Kehui Song, Guang Cheng 0001 |
TrustCom | 4 |
| 2024 | WCDGA: BERT-Based and Character-Transforming Adversarial DGA with High Anti-Detection AbilityabstractDomain Generation Algorithms (DGAs) are essential for creating numerous domain names automatically, commonly used to make malicious domains more stealthy and persistent online. To counteract DGAs, deep learning-based detection methods have been proposed, significantly reducing the effectiveness of traditional DGAs. However, due to inherent vulnerabilities in deep learning models, these detection methods are susceptible to adversarial attacks. Existing adversarial DGAs focus on character-based detection methods but overlook word-based structures, leading to weak performance against advanced word-based detection methods like graph neural networks. In this paper, we propose an innovative adversarial method and name it Word-Character DGA (WCDGA). The main idea is to create domain names by combining high-frequency words with common prefixes and suffixes found in reputable domains. This process utilizes a bidirectional encoder (BERT) and implements character transformations based on edit distance. We evaluate WCDGA against established character-based DGA detection methods (LSTM.MI, MIT, NYU) and the latest word-based method DGGCN. The results demonstrate that WCDGA out-performs existing adversarial DGAs in its evasion capabilities, successfully bypassing multiple detection methods simultaneously. Index Terms—Domain generation algorithm, Anti-detection, Dictionary generation, Character transformation Zhujie Guan, Mengmeng Tian, Yuwei Xu 0001, Kehui Song, Guang Cheng 0001 |
TrustCom | 4 |
| 2024 | GateKeeper: An UltraLite malicious traffic identification method with dual-aspect optimization strategies on IoT gateways
Jie Cao 0009, Yuwei Xu 0001, Enze Yu, Qiao Xiang, Kehui Song, Liang He 0002, Guang Cheng 0001 |
Comput. Networks | 5 |
| 2024 | An intelligent garment for online fetal well-being monitoring
Kehui Song, Xianyi Zeng, Julien de Jonckheere, Ludovic Koehl, Xiaojie Yuan |
Expert Syst. Appl. | 1 |
| 2024 | Multi-level feature interaction for open knowledge base canonicalization
Xuhui Sui, Ying Zhang 0015, Kehui Song, Baohang Zhou, Xiaojie Yuan |
Knowl. Based Syst. | 3 |
| 2023 | BioFEG: Generate Latent Features for Biomedical Entity LinkingabstractBiomedical entity linking is an essential task in biomedical text processing, which aims to map entity mentions in biomedical text to standard terms in a given knowledge base.However, this task is challenging due to the rarity of many biomedical entities in real-world scenarios, which leads to a lack of annotated data for them.Limited by understanding these unseen entities, traditional biomedical entity linking models suffer from multiple types of linking errors.In this paper, we propose a novel latent feature generation framework BioFEG to address these challenges.Specifically, our BioFEG leverages domain knowledge to train a generative adversarial network, which generates latent semantic features of corresponding mentions for unseen entities.Utilizing these features, we fine-tune our entity encoder to capture fine-grained coherence information of unseen entities and better understand them.This allows models to make linking decisions more accurately, particularly for ambiguous mentions involving rare entities.Extensive experiments on the two benchmark datasets demonstrate the superiority of our proposed method. Xuhui Sui, Ying Zhang 0015, Xiangrui Cai, Kehui Song, Baohang Zhou, Xiaojie Yuan, Wensheng Zhang 0002 |
EMNLP | 4 |
| 2023 | Zoomer: A Website Fingerprinting Attack Against Tor Hidden Services
Yuwei Xu 0001, Kehui Song, Yali Yuan |
ICICS | 4 |
| 2023 | Cerberus: Efficient OSPS Traffic Identification through Multi-Task LearningabstractThe privacy protection capabilities of open source proxy software (OSPS) while browsing the Internet have sparked great interest from both industry and academia, bringing forth pressing security concerns. Currently, using artificial intelligence for traffic identification is the most promising direction. Due to the wide variety and rich configuration of OSPS, it is not feasible to train models for all tasks and deploy them on the same network device. It is a novel idea to improve efficiency by leveraging multi-task learning. However, the related studies still have three shortcomings. First, improving the performance of the main task through auxiliary tasks does not apply to equally important OSPS identification tasks. Second, the model’s ability to characterize traffic is weak, resulting in performance gaps between different tasks. Finally, the influence of task difficulty on convergence speed is ignored, which is easy to cause overfitting and underfitting. Aiming at the shortcomings, we propose Cerberus, an OSPS traffic identification scheme based on multi-task learning. The main contributions of our work can be summarized in three aspects. Firstly, a high-quality dataset is constructed through traffic collection, and three OSPS traffic identification tasks are defined on it. Secondly, an identification model is designed by optimizing the ability to characterize traffic and balancing the convergence speed of multiple tasks. Finally, Cerberus is verified through comparative experiments. Its classification performance is better than both single-task and multi-task solutions. Besides, Cerberus runs fast and consumes few resources, making it suitable for deployment on network devices. Yuwei Xu 0001, Xiaotian Fang, Jie Cao 0009, Rou Yu, Kehui Song, Guang Cheng 0001 |
TrustCom | 5 |
| 2023 | SharpEye: Identify mKCP Camouflage Traffic through Feature OptimizationabstractAs a new self-developed protocol of V2Ray, mKCP disguises users’ network access as communication of four network applications by forging application layer headers to evade traffic-based detection. The emergence of mKCP has received widespread attention. Whether mKCP can provide secure network access that protects user privacy is the focus. Traditional methods cannot identify mKCP camouflage traffic, but machine learning (ML)-based traffic identification is considered a promising direction. Unlike the previous network traffic classification, mKCP camouflage traffic identification introduces new challenges. First, existing work has neither published any dataset containing mKCP camouflage traffic nor designed specific traffic features. Second, no researchers have optimized the identification scheme for deployment on network devices. Aiming at the shortcomings, we propose SharpEye, an ML-based mKCP camouflage traffic identification scheme. The novelty of our work lies in three points. Firstly, a complete dataset containing mKCP camouflage traffic is constructed through long-term traffic collection. Secondly, by analyzing the communication patterns of mKCP traffic, a feature set mFS is designed to improve identification accuracy. Finally, a two-stage feature selection method mGBFS is proposed to improve the operation efficiency. The experimental results show that mFS can enhance the performance of classifiers in identifying mKCP camouflage traffic, and mGBFS reduces the running time and overhead while ensuring high accuracy. Therefore, SharpEye achieves accurate and efficient mKCP camouflage traffic identification. Yuwei Xu 0001, Zizhi Zhu, Yunpeng Bai, Lilanyi Wu, Kehui Song, Guang Cheng 0001 |
TrustCom | 5 |
| 2023 | FastTraffic: A lightweight method for encrypted traffic fast classification
Yuwei Xu 0001, Jie Cao 0009, Kehui Song, Qiao Xiang, Guang Cheng 0001 |
Comput. Networks | 3 |
| 2023 | Corrigendum to "FastTraffic: A lightweight method for encrypted traffic fast classification" [Computer Networks, Volume 235, November 2023, 109965]
Yuwei Xu 0001, Jie Cao 0009, Kehui Song, Qiao Xiang, Guang Cheng 0001 |
Comput. Networks | 3 |
| 2023 | Location Recommendation Based on Mobility Graph With Individual and Group InfluencesabstractWith the rapid development of mobile technology, it is very convenient to share people’s current locations by checking-in on Location-Based Social Networks (LBSNs). Using users’ check-in histories to study mobility preferences and recommend new locations is a typical application to LBSNs. Most existing models explore reasonable representations for users and locations. However, a lack of behavioral mobility modeling would hamper a better understanding of users’ mobility patterns. This paper proposes a location recommendation model to serve the personalized LBSNs application, called Spatio-temporal Individual mobility graph encoding network with Group Mobility Assistance (SIGMA). We design a spatio-temporal interaction enhanced graph neural network to encode the mobility graphs to represent individual mobility behaviors. Furthermore, we provide a novel stacked scoring approach to generate the recommendation score by combining the stacked individual mobility graphs with the group influences. We conduct extensive experiments on two real-world LBSNs data, Foursquare and Gowalla. The result demonstrates SIGMA outperforms ten state-of-the-art models and further confirms that both the individual and the group mobility behaviors play essential roles in the practical scenario of location recommendation. Xuan Pan, Xiangrui Cai, Kehui Song, Thar Baker, G. Thippa Reddy, Xiaojie Yuan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Improving Zero-Shot Entity Linking Candidate Generation with Ultra-Fine Entity Type InformationabstractEntity linking, which aims at aligning ambiguous entity mentions to their referent entities in a knowledge base, plays a key role in multiple natural language processing tasks. Recently, zero-shot entity linking task has become a research hotspot, which links mentions to unseen entities to challenge the generalization ability. For this task, the training set and test set are from different domains, and thus entity linking models tend to be overfitting due to the tendency of memorizing the properties of entities that appear frequently in the training set. We argue that general ultra-fine-grained type information can help the linking models to learn contextual commonality and improve their generalization ability to tackle the overfitting problem. However, in the zero-shot entity linking setting, any type information is not available and entities are only identified by textual descriptions. Thus, we first extract the ultra-fine entity type information from the entity textual descriptions. Then, we propose a hierarchical multi-task model to improve the high-level zero-shot entity linking candidate generation task by utilizing the entity typing task as an auxiliary low-level task, which introduces extracted ultra-fine type information into the candidate generation task. Experimental results demonstrate the effectiveness of utilizing the ultra-fine entity type information and our proposed method achieves state-of-the-art performance. Xuhui Sui, Ying Zhang 0015, Kehui Song, Baohang Zhou, Xiaojie Yuan |
COLING | 3 |
| 2022 | A Span-based Multimodal Variational Autoencoder for Semi-supervised Multimodal Named Entity RecognitionabstractMultimodal named entity recognition (MNER) on social media is a challenging task which aims to extract named entities in free text and incorporate images to classify them into userdefined types.The existing semi-supervised named entity recognition methods focus on the text modal and are utilized to reduce labeling costs in traditional NER.However, the previous methods are not efficient for semisupervised MNER.Because the MNER task is defined to combine the text information with image one and needs to consider the mismatch between the posted text and image.To fuse the text and image features for MNER effectively under semi-supervised setting, we propose a novel span-based multimodal variational autoencoder (SMVAE) model for semisupervised MNER.The proposed method exploits modal-specific VAEs to model text and image latent features, and utilizes product-ofexperts to acquire multimodal features.In our approach, the implicit relations between labels and multimodal features are modeled by multimodal VAE.Thus, the useful information of unlabeled data can be exploited in our method under semi-supervised setting.Experimental results on two benchmark datasets demonstrate that our approach not only outperforms baselines under supervised setting, but also improves MNER performance with less labeled data than existing semi-supervised methods. Baohang Zhou, Ying Zhang 0015, Kehui Song, Wenya Guo, Xiaojie Yuan |
EMNLP | 3 |
| 2022 | A Multi-Task Learning Framework for Chinese Medical Procedure Entity NormalizationabstractMedical entity normalization is a fundamental task in medical natural language processing and clinical applications. The task aims to map medical mentions to standard entities in a given knowledge base. In this paper, we focus on Chinese medical procedure entity normalization. This task brings an extra multi-implication challenge that a mention may link to multiple standard entities. To perform the task, we propose a novel deep neural multi-task learning framework to jointly model implication number prediction and entity normalization. Our model utilizes the multi-head attention mechanism to provide mutual benefits between the two tasks. Experimental results show that our method achieves comparable performance compared with the baseline methods. Xuhui Sui, Kehui Song, Baohang Zhou, Ying Zhang 0015, Xiaojie Yuan |
ICASSP | 2 |
| 2021 | Multimodal Topic Detection in Social Networks with Graph Fusion
Kehui Song, Xiangrui Cai, Yierxiati Tuergong, Ling Yuan, Ying Zhang 0015 |
WISA | 2 |
| 2021 | A Decision Support System for Heart Failure Risk Prediction Based on Weighted Naive Bayes
Kehui Song, Samson Shenglong Yu, Haiwei Zhang 0001, Ying Zhang 0015, Xiangrui Cai, Xiaojie Yuan |
DASFAA (3) | 1 |
| 2021 | An interpretable knowledge-based decision support system and its applications in pregnancy diagnosis
Kehui Song, Xianyi Zeng, Ying Zhang 0015, Julien de Jonckheere, Xiaojie Yuan, Ludovic Koehl |
Knowl. Based Syst. | 1 |