EDBT 2026 Demo / reviewers in the wild / expert
Haoran Li 0003
dblp:50/10038-3
· DBLP profile ↗
28ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0003-1656-1278ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 6 first-author · 21 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language ModelsabstractYiyang Gu, Junwei Yang, Junyu Luo, Ye Yuan, Bin Feng, Yingce Xia, Shufang Xie, Kaili Liu, Bohan Wu, Qi Shi, Haoran Li, Beier Xiao, Zhiping Xiao, Xiao Luo, Weizhi Zhang, Philip S. Yu, Zequn Liu, Ming Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yiyang Gu, Junyu Luo 0002, Ye Yuan 0016, Yingce Xia, Shufang Xie 0003, Kaili Liu, Bohan Wu, Haoran Li 0003, Beier Xiao, Zhiping Xiao 0001, Xiao Luo 0001, Weizhi Zhang 0001, Philip S. Yu, Zequn Liu, Ming Zhang 0004 |
ACL (1) | 11 |
| 2026 | Into the Gray Zone: Domain Contexts Can Blur LLM Safety BoundariesabstractKi Sen Hung, Xi Yang, Chang Liu, Haoran Li, Kejiang Chen, Changxuan Fan, Tsun On Kwok, Weiming Zhang, Xiaomeng Li, Yangqiu Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ki Sen Hung, Chang Liu 0089, Haoran Li 0003, Kejiang Chen, Changxuan Fan, Tsun On Kwok, Weiming Zhang 0001, Xiaomeng Li 0001, Yangqiu Song |
ACL (1) | 4 |
| 2026 | ContextLens: Modeling Imperfect Privacy and Safety Context for Legal ComplianceabstractHaoran Li, Yulin Chen, Huihao Jing, Wenbin Hu, Tsz Ho Li, Chanhou Lou, Hong Ting Tsang, Sirui Han, Yangqiu Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Haoran Li 0003, Huihao Jing, Wenbin Hu 0001, Tsz Ho Li, Chanhou Lou, Hong Ting Tsang, Sirui Han, Yangqiu Song |
ACL (1) | 1 |
| 2026 | Activation-Guided Local Editing for Jailbreaking AttacksabstractJiecong Wang, Haoran Li, Hao Peng, Ziqian Zeng, Zihao Wang, Haohua Du, Zhengtao Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiecong Wang, Haoran Li 0003, Hao Peng 0001, Ziqian Zeng, Zihao Wang 0001, Haohua Du, Zhengtao Yu 0001 |
ACL (1) | 2 |
| 2025 | Simulate and Eliminate: Revoke Backdoors for Generative Large Language ModelsabstractWith rapid advances, generative large language models (LLMs) dominate various Natural Language Processing (NLP) tasks from understanding to reasoning. Yet, language models' inherent vulnerabilities may be exacerbated due to increased accessibility and unrestricted model training on massive data. A malicious adversary may publish poisoned data online and conduct backdoor attacks on the victim LLMs pre-trained on the poisoned data. Backdoored LLMs behave innocuously for normal queries and generate harmful responses when the backdoor trigger is activated. Despite significant efforts paid to LLMs' safety issues, LLMs are still struggling against backdoor attacks. As Anthropic recently revealed, existing safety training strategies, including supervised fine-tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), fail to revoke the backdoors once the LLM is backdoored during the pre-training stage. In this paper, we present Simulate and Eliminate (SANDE) to erase the undesired backdoored mappings for generative LLMs. We initially propose Overwrite Supervised Fine-tuning (OSFT) for effective backdoor removal when the trigger is known. Then, to handle scenarios where trigger patterns are unknown, we integrate OSFT into our two-stage framework, SANDE. Unlike other works that assume access to cleanly trained models, our safety-enhanced LLMs are able to revoke backdoors without any reference. Consequently, our safety-enhanced LLMs no longer produce targeted responses when the backdoor triggers are activated. We conduct comprehensive experiments to show that our proposed SANDE is effective against backdoor attacks while bringing minimal harm to LLMs' powerful capability. Haoran Li 0003, Chunkit Chan, Heshan Liu, Yangqiu Song |
AAAI | 1 |
| 2025 | PrivaCI-Bench: Evaluating Privacy with Contextual Integrity and Legal ComplianceabstractHaoran Li, Wenbin Hu, Huihao Jing, Yulin Chen, Qi Hu, Sirui Han, Tianshu Chu, Peizhao Hu, Yangqiu Song. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haoran Li 0003, Wenbin Hu 0001, Huihao Jing, Sirui Han, Peizhao Hu, Yangqiu Song |
ACL (1) | 1 |
| 2025 | Defense Against Prompt Injection Attack by Leveraging Attack TechniquesabstractWith the advancement of technology, large language models (LLMs) have achieved remarkable performance across various natural language processing (NLP) tasks, powering LLMintegrated applications like Microsoft Copilot.However, as LLMs continue to evolve, new vulnerabilities, especially prompt injection attacks arise.These attacks trick LLMs into deviating from the original input instructions and executing the attacker's instructions injected in data content, such as retrieved results.Recent attack methods leverage LLMs' instruction-following abilities and their inabilities to distinguish instructions injected in the data content, and achieve a high attack success rate (ASR).When comparing the attack and defense methods, we interestingly find that they share similar design goals, of inducing the model to ignore unwanted instructions and instead to execute wanted instructions.Therefore, we raise an intuitive question: Could these attack techniques be utilized for defensive purposes?In this paper, we invert the intention of prompt injection methods to develop novel defense methods based on previous trainingfree attack methods, by repeating the attack process but with the original input instruction rather than the injected instruction.Our comprehensive experiments demonstrate that our defense techniques outperform existing defense approaches, achieving state-of-the-art results. 1 Haoran Li 0003, Dekai Wu, Yangqiu Song, Bryan Hooi |
ACL (1) | 2 |
| 2025 | Can Indirect Prompt Injection Attacks Be Detected and Removed?abstractPrompt injection attacks manipulate large language models (LLMs) by misleading them to deviate from the original input instructions and execute maliciously injected instructions, because of their instruction-following capabilities and inability to distinguish between the original input instructions and maliciously injected instructions.To defend against such attacks, recent studies have developed various detection mechanisms.If we restrict ourselves specifically to works which perform detection rather than direct defense, most of them focus on direct prompt injection attacks, while there are few works for the indirect scenario, where injected instructions are indirectly from external tools, such as a search engine.Moreover, current works mainly investigate injection detection methods and pay less attention to the post-processing method that aims to mitigate the injection after detection.In this paper, we investigate the feasibility of detecting and removing indirect prompt injection attacks, and we construct a benchmark dataset for evaluation.For detection, we assess the performance of existing LLMs and open-source detection models, and we further train detection models using our crafted training datasets.For removal, we evaluate two intuitive methods: (1) the segmentation removal method, which segments the injected document and removes parts containing injected instructions, and (2) the extraction removal method, which trains an extraction model to identify and remove injected instructions.1 Haoran Li 0003, Yuan Sui 0001, Yue Liu 0008, Yangqiu Song, Bryan Hooi |
ACL (1) | 2 |
| 2025 | TopicAttack: An Indirect Prompt Injection Attack via Topic TransitionabstractLarge language models (LLMs) have shown remarkable performance across a range of NLP tasks.However, their strong instructionfollowing capabilities and inability to distinguish instructions from data content make them vulnerable to indirect prompt injection attacks.In such attacks, instructions with malicious purposes are injected into external data sources, such as web documents.When LLMs retrieve this injected data through tools, such as a search engine and execute the injected instructions, they provide misled responses.Recent attack methods have demonstrated potential, but their abrupt instruction injection often undermines their effectiveness.Motivated by the limitations of existing attack methods, we propose Topi-cAttack, which prompts the LLM to generate a fabricated conversational transition prompt that gradually shifts the topic toward the injected instruction, making the injection smoother and enhancing the plausibility and success of the attack.Through comprehensive experiments, TopicAttack achieves state-of-the-art performance, with an attack success rate (ASR) over 90% in most cases, even when various defense methods are applied.We further analyze its effectiveness by examining attention scores.We find that a higher injected-to-original attention ratio leads to a greater success probability, and our method achieves a much higher ratio than the baseline methods. 1 Haoran Li 0003, Yuexin Li, Yue Liu 0008, Yangqiu Song, Bryan Hooi |
EMNLP | 2 |
| 2025 | Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement LearningabstractWenbin Hu, Haoran Li, Huihao Jing, Qi Hu, Ziqian Zeng, Sirui Han, Xu Heli, Tianshu Chu, Peizhao Hu, Yangqiu Song. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Wenbin Hu 0001, Haoran Li 0003, Huihao Jing, Ziqian Zeng, Sirui Han, Heli Xu, Peizhao Hu, Yangqiu Song |
EMNLP | 2 |
| 2025 | MCIP: Protecting MCP Safety via Model Contextual Integrity ProtocolabstractAs Model Context Protocol (MCP) introduces an easy-to-use ecosystem for users and developers, it also brings underexplored safety risks.Its decentralized architecture, which separates clients and servers, poses unique challenges for systematic safety analysis.This paper proposes a novel framework to enhance MCP safety.Guided by the MAESTRO framework, we first analyze the missing safety mechanisms in MCP, and based on this analysis, we propose the Model Contextual Integrity Protocol (MCIP), a refined version of MCP that addresses these gaps.Next, we develop a fine-grained taxonomy that captures a diverse range of unsafe behaviors observed in MCP scenarios.Building on this taxonomy, we develop benchmark and training data that support the evaluation and improvement of LLMs' capabilities in identifying safety risks within MCP interactions.Leveraging the proposed benchmark and training data, we conduct extensive experiments on state-of-the-art LLMs.The results highlight LLMs' vulnerabilities in MCP interactions and demonstrate that our approach substantially improves their safety performance.1 Huihao Jing, Haoran Li 0003, Wenbin Hu 0001, Heli Xu, Peizhao Hu, Yangqiu Song |
EMNLP | 2 |
| 2025 | Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity TheoryabstractHaoran Li, Wei Fan, Yulin Chen, Cheng Jiayang, Tianshu Chu, Xuebing Zhou, Peizhao Hu, Yangqiu Song. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Haoran Li 0003, Wei Fan 0001, Cheng Jiayang, Xuebing Zhou, Peizhao Hu, Yangqiu Song |
NAACL (Long Papers) | 1 |
| 2024 | PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language ModelsabstractHaoran Li, Dadi Guo, Donghao Li, Wei Fan, Qi Hu, Xin Liu, Chunkit Chan, Duanyi Yao, Yuan Yao, Yangqiu Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Haoran Li 0003, Dadi Guo, Wei Fan 0001, Xin Liu 0039, Chunkit Chan, Duanyi Yao, Yuan Yao 0001, Yangqiu Song |
ACL (1) | 1 |
| 2024 | Table-Filling via Mean Teacher for Cross-domain Aspect Sentiment Triplet ExtractionabstractCross-domain Aspect Sentiment Triplet Extraction (ASTE) aims to extract fine-grained sentiment elements from target domain sentences by leveraging the knowledge acquired from the source domain. Due to the absence of labeled data in the target domain, recent studies tend to rely on pre-trained language models to generate large amounts of synthetic data for training purposes. However, these approaches entail additional computational costs associated with the generation process. Different from them, we discover a striking resemblance between table-filling methods in ASTE and two-stage Object Detection (OD) in computer vision, which inspires us to revisit the cross-domain ASTE task and approach it from an OD standpoint. This allows the model to benefit from the OD extraction paradigm and region-level alignment. Building upon this premise, we propose a novel method named Table-Filling via Mean Teacher (TFMT). Specifically, the table-filling methods encode the sentence into a 2D table to detect word relations, while TFMT treats the table as a feature map and utilizes a region consistency to enhance the quality of those generated pseudo labels. Additionally, considering the existence of the domain gap, a cross-domain consistency based on Maximum Mean Discrepancy is designed to alleviate domain shift problems. Our method achieves state-of-the-art performance with minimal parameters and computational costs, making it a strong baseline for cross-domain ASTE. Lei Jiang 0003, Qian Li 0033, Haoran Li 0003, Li Sun 0008, Yanxian Bi, Hao Peng 0001 |
CIKM | 4 |
| 2024 | Adaptive Differentially Private Structural Entropy Minimization for Unsupervised Social Event DetectionabstractSocial event detection refers to extracting relevant message clusters from social media data streams to represent specific events in the real world. Social event detection is important in numerous areas, such as opinion analysis, social safety, and decision-making. Most current methods are supervised and require access to large amounts of data. These methods need prior knowledge of the events and carry a high risk of leaking sensitive information in the messages, making them less applicable in open-world settings. Therefore, conducting unsupervised detection while fully utilizing the rich information in the messages and protecting data privacy remains a significant challenge. To this end, we propose a novel social event detection framework, ADP-SEMEvent, an unsupervised social event detection method that prioritizes privacy. Specifically, ADP-SEMEvent is divided into two stages, i.e., the construction stage of the private message graph and the clustering stage of the private message graph. In the first stage, an adaptive differential privacy approach is used to construct a private message graph. In this process, our method can adaptively apply differential privacy based on the events occurring each day in an open environment to maximize the use of the privacy budget. In the second stage, to address the reduction in data utility caused by noise, a novel 2-dimensional structural entropy minimization algorithm based on optimal subgraphs is used to detect events in the message graph. The highlight of this process is unsupervised and does not compromise differential privacy. Extensive experiments on two public datasets demonstrate that ADP-SEMEvent can achieve detection performance comparable to state-of-the-art methods while maintaining reasonable privacy budget parameters. Zhiwei Yang 0009, Yuecen Wei, Haoran Li 0003, Qian Li 0033, Lei Jiang 0003, Li Sun 0008, Chunming Hu, Hao Peng 0001 |
CIKM | 3 |
| 2024 | Audience Persona Knowledge-Aligned Prompt Tuning Method for Online DebateabstractDebate is the process of exchanging viewpoints or convincing others on a particular issue. Recent research has provided empirical evidence that the persuasiveness of an argument is determined not only by language usage but also by communicator characteristics. Researchers have paid much attention to aspects of languages, such as linguistic features and discourse structures, but combining argument persuasiveness and impact with the social personae of the audience has not been explored due to the difficulty and complexity. We have observed the impressive simulation and personification capability of ChatGPT, indicating a giant pre-trained language model may function as an individual to provide personae and exert unique influences based on diverse background knowledge. Therefore, we propose a persona knowledge-aligned framework for argument quality assessment tasks from the audience side. This is the first work that leverages the emergence of ChatGPT and injects such audience personae knowledge into smaller language models via prompt tuning. The performance of our pipeline demonstrates significant and consistent improvement compared to competitive architectures. Chunkit Chan, Cheng Jiayang, Xin Liu 0039, Yauwai Yim, Zheye Deng, Haoran Li 0003, Yangqiu Song, Ginny Y. Wong, Simon See |
ECAI | 7 |
| 2024 | LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingabstractJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, Wenpeng Yin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jiangshu Du, Yibo Wang 0001, Wenting Zhao 0006, Zhongfen Deng, Shuaiqi Liu 0002, Renze Lou, Henry Peng Zou, Pranav Venkit, Mukund Srinath, Ranran Haoran Zhang, Tao Li 0039, Fei Wang 0060, Qin Liu 0010, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang 0003, Raj Sanjay Shah, Ruohao Guo, Haoran Li 0003, Kangda Wei, Zihao Wang 0001, Lu Cheng 0001, Surangika Ranathunga, Fei Liu 0004, Ruihong Huang, Eduardo Blanco 0002, Yixin Cao 0002, Rui Zhang 0037, Philip S. Yu, Wenpeng Yin 0001 |
EMNLP | 27 |
| 2024 | GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity TheoryabstractPrivacy issues arise prominently during the inappropriate transmission of information between entities.Existing research primarily studies privacy by exploring various privacy attacks, defenses, and evaluations within narrowly predefined patterns, while neglecting that privacy is not an isolated, context-free concept limited to traditionally sensitive data (e.g., social security numbers), but intertwined with intricate social contexts that complicate the identification and analysis of potential privacy violations.The advent of Large Language Models (LLMs) offers unprecedented opportunities for incorporating the nuanced scenarios outlined in privacy laws to tackle these complex privacy issues.However, the scarcity of open-source relevant case studies restricts the efficiency of LLMs in aligning with specific legal statutes.To address this challenge, we introduce a novel framework, GOLDCOIN 1 , designed to efficiently ground LLMs in privacy laws for judicial assessing privacy violations.Our framework leverages the theory of contextual integrity as a bridge, creating numerous synthetic scenarios grounded in relevant privacy statutes (e.g., HIPAA), to assist LLMs in comprehending the complex contexts for identifying privacy risks in the real world.Extensive experimental results demonstrate that GOLD-COIN markedly enhances LLMs' capabilities in recognizing privacy risks across real court cases, surpassing the baselines on different judicial tasks. Wei Fan 0001, Haoran Li 0003, Zheye Deng, Weiqi Wang 0001, Yangqiu Song |
EMNLP | 2 |
| 2024 | Privacy-Preserved Neural Graph DatabasesabstractIn the era of large language models (LLMs), efficient and accurate data retrieval has become increasingly crucial for the use of domain-specific or private data in the retrieval augmented generation (RAG). Neural graph databases (NGDBs) have emerged as a powerful paradigm that combines the strengths of graph databases (GDBs) and neural networks to enable efficient storage, retrieval, and analysis of graph-structured data which can be adaptively trained with LLMs. The usage of neural embedding storage and Complex neural logical Query Answering (CQA) provides NGDBs with generalization ability. When the graph is incomplete, by extracting latent patterns and representations, neural graph databases can fill gaps in the graph structure, revealing hidden relationships and enabling accurate query answering. Nevertheless, this capability comes with inherent trade-offs, as it introduces additional privacy risks to the domain-specific or private databases. Malicious attackers can infer more sensitive information in the database using well-designed queries such as from the answer sets of where Turing Award winners born before 1950 and after 1940 lived, the living places of Turing Award winner Hinton are probably exposed, although the living places may have been deleted in the training stage due to the privacy concerns. In this work, we propose a privacy-preserved neural graph database (P-NGDB) framework to alleviate the risks of privacy leakage in NGDBs. We introduce adversarial training techniques in the training stage to enforce the NGDBs to generate indistinguishable answers when queried with private information, enhancing the difficulty of inferring sensitive information through combinations of multiple innocuous queries. Extensive experimental results on three datasets show that our framework can effectively protect private information in the graph database while delivering high-quality public answers responses to queries. The code is available at https://github.com/HKUST-KnowComp/PrivateNGDB. Haoran Li 0003, Jiaxin Bai, Zihao Wang 0001, Yangqiu Song |
KDD | 2 |
| 2022 | You Don't Know My Favorite Color: Preventing Dialogue Representations from Revealing Speakers' Private PersonasabstractSocial chatbots, also known as chit-chat chatbots, evolve rapidly with large pretrained language models.Despite the huge progress, privacy concerns have arisen recently: training data of large language models can be extracted via model inversion attacks.On the other hand, the datasets used for training chatbots contain many private conversations between two individuals.In this work, we further investigate the privacy leakage of the hidden states of chatbots trained by language modeling which has not been well studied yet.We show that speakers' personas can be inferred through a simple neural network with high accuracy.To this end, we propose effective defense objectives to protect persona leakage from hidden states.We conduct extensive experiments to demonstrate that our proposed defense objectives can greatly reduce the attack accuracy from 37.6% to 0.5%.Meanwhile, the proposed objectives preserve language models' powerful generation ability. Haoran Li 0003, Yangqiu Song, Lixin Fan |
NAACL-HLT | 1 |
| 2021 | Differentially Private Federated Knowledge Graphs EmbeddingabstractKnowledge graph embedding plays an important role in knowledge representation, reasoning, and data mining applications. However, for multiple cross-domain knowledge graphs, state-of-the-art embedding models cannot make full use of the data from different knowledge domains while preserving the privacy of exchanged data. In addition, the centralized embedding model may not scale to the extensive real-world knowledge graphs. Therefore, we propose a novel decentralized scalable learning framework, Federated Knowledge Graphs Embedding (FKGE), where embeddings from different knowledge graphs can be learnt in an asynchronous and peer-to-peer manner while being privacy-preserving. FKGE exploits adversarial generation between pairs of knowledge graphs to translate identical entities and relations of different domains into near embedding spaces. In order to protect the privacy of the training data, FKGE further implements a privacy-preserving neural network structure to guarantee no raw data leakage. We conduct extensive experiments to evaluate FKGE on 11 knowledge graphs, demonstrating a significant and consistent improvement in model quality with at most 17.85% and 7.90% increases in performance on triple classification and link prediction tasks. Hao Peng 0001, Haoran Li 0003, Yangqiu Song, Vincent Wenchen Zheng, Jianxin Li 0002 |
CIKM | 2 |
| 2020 | Self-supervised Dance Video Synthesis Conditioned on MusicabstractWe present a self-supervised approach with pose perceptual loss for automatic dance video generation. Our method can produce a realistic dance video that conforms to the beats and rhymes of given music. To achieve this, we firstly generate a human skeleton sequence from music and then apply the learned pose-to-appearance mapping to generate the final video. In the stage of generating skeleton sequences, we utilize two discriminators to capture different aspects of the sequence and propose a novel pose perceptual loss to produce natural dances. Besides, we also provide a new cross-modal evaluation metric to evaluate the dance quality, which is able to estimate the similarity between two modalities (music and dance). Finally, our experimental qualitative and quantitative results demonstrate that our dance video synthesis approach produces realistic and diverse results. Our source code and data are available at https://github.com/xrenaa/Music-Dance-Video-Synthesis. Xuanchi Ren, Haoran Li 0003, Zijian Huang 0002, Qifeng Chen 0001 |
ACM Multimedia | 2 |
| 2019 | Learning Hierarchical Representations of Electronic Health Records for Clinical Outcome Prediction
Luchen Liu, Haoran Li 0003, Zhiting Hu, Haoran Shi 0001, Zichang Wang, Jian Tang 0005, Ming Zhang 0004 |
AMIA | 2 |
| 2019 | Predictive Multi-level Patient Representations from Electronic Health RecordsabstractThe advent of the Internet era has led to an explosive growth in the Electronic Health Records (EHR) in the past decades. The EHR data can be regarded as a collection of clinical events, including laboratory results, medication records, physiological indicators, etc, which can be used for clinical outcome prediction tasks to support constructions of intelligent health systems. Learning patient representation from these clinical events for the clinical outcome prediction is an important but challenging step. Most related studies transform EHR data of a patient into a sequence of clinical events in temporal order and then use sequential models to learn patient representations for outcome prediction. However, clinical event sequence contains thousands of event types and temporal dependencies. We further make an observation that clinical events occurring in a short period are not constrained by any temporal order but events in a long term are influenced by temporal dependencies. The multi-scale temporal property makes it difficult for traditional sequential models to capture the short-term co-occurrence and the long-term temporal dependencies in clinical event sequences. In response to the above challenges, this paper proposes a Multilevel Representation Model (MRM). MRM first uses a sparse attention mechanism to model the short-term co-occurrence, then uses interval-based event pooling to remove redundant information and reduce sequence length and finally predicts clinical outcomes through Long Short-Term Memory (LSTM). Experiments on real-world datasets indicate that our proposed model largely improves the performance of clinical outcome prediction tasks using EHR data. Zichang Wang, Haoran Li 0003, Luchen Liu, Haoxian Wu, Ming Zhang 0004 |
BIBM | 2 |
| 2018 | Unsupervised meta-path selection for text similarity measure based on heterogeneous information networks
Chenguang Wang 0001, Yangqiu Song, Haoran Li 0003, Ming Zhang 0004, Jiawei Han 0001 |
Data Min. Knowl. Discov. | 3 |
| 2017 | Distant Meta-Path Similarities for Text-Based Heterogeneous Information NetworksabstractMeasuring network similarity is a fundamental data mining problem. The mainstream similarity measures mainly leverage the structural information regarding to the entities in the network without considering the network semantics. In the real world, the heterogeneous information networks (HINs) with rich semantics are ubiquitous. However, the existing network similarity doesn't generalize well in HINs because they fail to capture the HIN semantics. The meta-path has been proposed and demonstrated as a right way to represent semantics in HINs. Therefore, original meta-path based similarities (e.g., PathSim and KnowSim) have been successful in computing the entity proximity in HINs. The intuition is that the more instances of meta-path(s) between entities, the more similar the entities are. Thus the original meta-path similarity only applies to computing the proximity of two neighborhood (connected) entities. In this paper, we propose the distant meta-path similarity that is able to capture HIN semantics between two distant (isolated) entities to provide more meaningful entity proximity. The main idea is that even there is no shared neighborhood entities of (i.e., no meta-path instances connecting) the two entities, but if the more similar neighborhood entities of the entities are, the more similar the two entities should be. We then find out the optimum distant meta-path similarity by exploring the similarity hypothesis space based on different theoretical foundations. We show the state-of-the-art similarity performance of distant meta-path similarity on two text-based HINs and make the datasets public available. Chenguang Wang 0001, Yangqiu Song, Haoran Li 0003, Yizhou Sun, Ming Zhang 0004, Jiawei Han 0001 |
CIKM | 3 |
| 2016 | Text Classification with Heterogeneous Information Network KernelsabstractText classification is an important problem with many applications. Traditional approaches represent text as a bag-of-words and build classifiers based on this representation. Rather than words, entity phrases, the relations between the entities, as well as the types of the entities and relations carry much more information to represent the texts. This paper presents a novel text as network classification framework, which introduces 1) a structured and typed heterogeneous information networks (HINs) representation of texts, and 2) a meta-path based approach to link texts. We show that with the new representation and links of texts, the structured and typed information of entities and relations can be incorporated into kernels. Particularly, we develop both simple linear kernel and indefinite kernel based on meta-paths in the HIN representation of texts, where we call them HIN-kernels. Using Freebase, a well-known world knowledge base, to construct HIN for texts, our experiments on two benchmark datasets show that the indefinite HIN kernel based on weighted meta-paths outperforms the state-of-the-art methods and other HIN-kernels. Chenguang Wang 0001, Yangqiu Song, Haoran Li 0003, Ming Zhang 0004, Jiawei Han 0001 |
AAAI | 3 |
| 2015 | KnowSim: A Document Similarity Measure on Structured Heterogeneous Information NetworksabstractAs a fundamental task, document similarity measure has broad impact to document-based classification, clustering and ranking. Traditional approaches represent documents as bag-of-words and compute document similarities using measures like cosine, Jaccard, and dice. However, entity phrases rather than single words in documents can be critical for evaluating document relatedness. Moreover, types of entities and links between entities/words are also informative. We propose a method to represent a document as a typed heterogeneous information network (HIN), where the entities and relations are annotated with types. Multiple documents can be linked by the words and entities in the HIN. Consequently, we convert the document similarity problem to a graph distance problem. Intuitively, there could be multiple paths between a pair of documents. We propose to use the meta-path defined in HIN to compute distance between documents. Instead of burdening user to define meaningful meta-paths, an automatic method is proposed to rank the meta-paths. Given the meta-paths associated with ranking scores, an HIN-based similarity measure, KnowSim, is proposed to compute document similarities. Using Freebase, a well-known world knowledge base, to conduct semantic parsing and construct HIN for documents, our experiments on 20Newsgroups and RCV1 datasets show that KnowSim generates impressive high-quality document clustering. Chenguang Wang 0001, Yangqiu Song, Haoran Li 0003, Ming Zhang 0004, Jiawei Han 0001 |
ICDM | 3 |