Tengfei Ma 0001

dblp:94/9023-1 · DBLP profile ↗
← Back
60ranked-venue papers
10as first author
34since 2021 · last 2025
0000-0002-1086-529XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 9 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
abstract
Large Language Models (LLMs) have demonstrated remarkable proficiency in generating code. However, the misuse of LLM-generated (synthetic) code has raised concerns in both educational and industrial contexts, underscoring the urgent need for synthetic code detectors. Existing methods for detecting synthetic content are primarily designed for general text and struggle with code due to the unique grammatical structure of programming languages and the presence of numerous ``low-entropy'' tokens. Building on this, our work proposes a novel zero-shot synthetic code detector based on the similarity between the original code and its LLM-rewritten variants. Our method is based on the observation that differences between LLM-rewritten and original code tend to be smaller when the original code is synthetic. We utilize self-supervised contrastive learning to train a code similarity model and evaluate our approach on two synthetic code detection benchmarks. Our results demonstrate a significant improvement over existing SOTA synthetic content detectors, delivering notable gains in both performance and robustness on the APPS and MBPP benchmarks.
Yangkai Du, Tengfei Ma 0001, Lingfei Wu 0001, Xuhong Zhang 0002, Shouling Ji, Wenhai Wang
AAAI3
2025 Influencer: Empowering Everyday Users in Creating Promotional Posts via AI-infused Exploration and Customization
abstract
Figure 1: A design novice uses Infuencer to ideate and make promotional posts to promote their homemade juice.Infuencer has the following core features: (A) The user can input a topic via a text block and explores the related images and captions in three dimensions.(B) Context-aware exploration is supported which updates the image and caption recommendation by dragging a brand/product image or message to the initial image and caption recommendation.(C) Various materials (i.e., image and text) can be fexibly fused to make a new image or caption.(D) Infuencer allows the user to not only easily create harmonious promotional posts but also quickly obtain multiple post alternatives.Steps in (A), (B), and (C) can be fexibly combined or skipped; as soon as the user fnds satisfed image and/or caption, they can go to (D) for post generation.
Xuye Liu, Annie Sun, Pengcheng An, Tengfei Ma 0001, Jian Zhao 0010
CHI4
2025 Shedding Light on Time Series Classification using Interpretability Gated Networks
abstract
In time-series classification, interpretable models can bring additional insights but be outperformed by deep models since human-understandable features have limited expressivity and flexibility. In this work, we present InterpGN, a framework that integrates an interpretable model and a deep neural network. Within this framework, we introduce a novel gating function design based on the confidence of the interpretable expert, preserving interpretability for samples where interpretable features are significant while also identifying samples that require additional expertise. For the interpretable expert, we incorporate shapelets to effectively model shape-level features for time-series data. We introduce a variant of Shapelet Transforms to build logical predicates using shapelets. Our proposed model achieves comparable performance with state-of-the-art deep learning models while additionally providing interpretable classifiers for various benchmark datasets. We further show that our models improve on quantitative shapelet quality and interpretability metrics over existing shapelet-learning formulations. Finally, we show that our models can integrate additional advanced architectures and be applied to real-world tasks beyond standard benchmarks such as the MIMIC-III and time series extrinsic regression datasets.
Yunshi Wen, Tengfei Ma 0001, Ronny Luss, Debarun Bhattacharjya, Achille Fokoue, A. Agung Julius
ICLR2
2025 Understanding and Tackling Over-Dilution in Graph Neural Networks
abstract
Message Passing Neural Networks (MPNNs) hold a key position in machine learning on graphs, but they struggle with unintended behaviors, such as over-smoothing and over-squashing, due to irregular data structures. The observation and formulation of these limitations have become foundational in constructing more informative graph representations. In this paper, we delve into the limitations of MPNNs, focusing on aspects that have previously been overlooked. Our observations reveal that even within a single layer, the information specific to an individual node can become significantly diluted. To delve into this phenomenon in depth, we present the concept of Over-dilution and formulate it with two dilution factors: intra-node dilution for attribute-level and inter-node dilution for node-level representations. We also introduce a transformer-based solution that alleviates over-dilution and complements existing node embedding methods like MPNNs. Our findings provide new insights and contribute to the development of informative representations. The implementation and supplementary materials are publicly available at https://github.com/LeeJunHyun/NATR.
Jun-Hyun Lee, Veronika Thost, Bumsoo Kim 0005, Jaewoo Kang, Tengfei Ma 0001
KDD (2)5
2025 MACEDON : Supporting Programmers with Real-Time Multi-Dimensional Code Evaluation and Optimization
Xuye Liu, Yuzhe You, Xinrong Qiu, Tengfei Ma 0001, Jian Zhao 0010
UIST4
2025 Enhancing Graph Representation Learning with Localized Topological Features
abstract
Representation learning on graphs is a fundamental problem that can be crucial in various tasks. Graph neural networks, the dominant approach for graph representation learning, are limited in their representation power. Therefore, it can be beneficial to explicitly extract and incorporate high-order topological and geometric information into these models. In this paper, we propose a principled approach to extract the rich connectivity information of graphs based on the theory of persistent homology. Our method utilizes the topological features to enhance the representation learning of graph neural networks and achieve state-of-the-art performance on various node classification and link prediction benchmarks. We also explore the option of end-to-end learning of the topological features, i.e., treating topological computation as a differentiable operator during learning. Our theoretical analysis and empirical study provide insights and potential guidelines for employing topological features in graph learning tasks.
Zuoyu Yan, Qi Zhao 0007, Ze Ye, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Yusu Wang 0001, Chao Chen 0012
J. Mach. Learn. Res.4
2025 Wasserstein Graph Neural Networks for Graphs With Missing Attributes
abstract
Missing node attributes pose a common problem in real-world graphs, impacting the performance of graph neural networks' representation learning. Existing GNNs often struggle to effectively leverage incomplete attribute information, as they are not specifically designed for graphs with missing attributes. To address this issue, we propose a novel node representation learning framework called Wasserstein Graph Neural Network (WGNN). Our approach aims to maximize the utility of limited observed attribute information and account for uncertainty caused by missing values. We achieve this by representing nodes as low-dimensional distributions obtained through attribute matrix decomposition. Additionally, we enhance representation expressiveness by introducing a unique message-passing schema that aggregates distributional information from neighboring nodes in the Wasserstein space. We evaluate the performance of WGNN in node classification tasks using both synthetic and real-world datasets under two missing-attribute scenarios. Moreover, we demonstrate the applicability of WGNN in recovering missing values and tackling matrix completion problems, specifically in graphs involving users and items. Experimental results on both tasks convincingly demonstrate the superiority of our proposed method.
Tengfei Ma 0001, Yangqiu Song, Yang Wang 0020
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone Detection
abstract
Code Clone Detection, which aims to retrieve functionally similar programs from large code bases, has been attracting increasing attention. Modern software often involves a diverse range of programming languages. However, current code clone detection methods are generally limited to only a few popular programming languages due to insufficient annotated data as well as their own model design constraints. To address these issues, we present AdaCCD, a novel cross-lingual adaptation method that can detect cloned codes in a new language without annotations in that language. AdaCCD leverages language-agnostic code representations from pre-trained programming language models and propose an Adaptively Refined Contrastive Learning framework to transfer knowledge from resource-rich languages to resource-poor languages. We evaluate the cross-lingual adaptation results of AdaCCD by constructing a multilingual code clone detection benchmark consisting of 5 programming languages. AdaCCD achieves significant improvements over other baselines, and achieve comparable performance to supervised fine-tuning.
Yangkai Du, Tengfei Ma 0001, Lingfei Wu 0001, Xuhong Zhang 0002, Shouling Ji
AAAI2
2024 TrojVLM: Backdoor Attack Against Vision Language Models
Weimin Lyu, Lu Pang 0006, Tengfei Ma 0001, Haibin Ling, Chao Chen 0012
ECCV (65)3
2024 FAC²E: Better Understanding Large Language Model Capabilities by Dissociating Language and Cognition
abstract
Large language models (LLMs) are primarily evaluated by overall performance on various text understanding and generation tasks.However, such a paradigm fails to comprehensively differentiate the fine-grained language and cognitive skills, rendering the lack of sufficient interpretation to LLMs' capabilities.In this paper, we present FAC 2 E, a framework for Fine-grAined and Cognition-grounded LLMs' Capability Evaluation.Specifically, we formulate LLMs' evaluation in a multi-dimensional and explainable manner by dissociating the language-related capabilities and the cognitionrelated ones.Besides, through extracting the intermediate reasoning from LLMs, we further break down the process of applying a specific capability into three sub-steps: recalling relevant knowledge, utilizing knowledge, and solving problems.Finally, FAC 2 E evaluates each sub-step of each fine-grained capability, providing a two-faceted diagnosis for LLMs.Utilizing FAC 2 E, we identify a common shortfall in knowledge utilization among models and propose a straightforward, knowledge-enhanced method to mitigate this issue.Our results not only showcase promising performance enhancements but also highlight a direction for future LLM advancements.
Xiaoqiang Wang 0007, Lingfei Wu 0001, Tengfei Ma 0001, Bang Liu 0003
EMNLP3
2024 What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
abstract
Graph Transformers, which incorporate self-attention and positional encoding, have recently emerged as a powerful architecture for various graph learning tasks. Despite their impressive performance, the complex non-convex interactions across layers and the recursive graph structure have made it challenging to establish a theoretical foundation for learning and generalization. This study introduces the first theoretical investigation of a shallow Graph Transformer for semi-supervised node classification, comprising a self-attention layer with relative positional encoding and a two-layer perception. Focusing on a graph data model with discriminative nodes that determine node labels and non-discriminative nodes that are class-irrelevant, we characterize the sample complexity required to achieve a desirable generalization error by training with stochastic gradient descent (SGD). This paper provides the quantitative characterization of the sample complexity and number of iterations for convergence dependent on the fraction of discriminative nodes, the dominant patterns, and the initial model errors. Furthermore, we demonstrate that self-attention and positional encoding enhance generalization by making the attention map sparse and promoting the core neighborhood during training, which explains the superior feature representation of Graph Transformers. Our theoretical results are supported by empirical experiments on synthetic and real-world benchmarks.
Hongkang Li, Meng Wang 0003, Tengfei Ma 0001, Sijia Liu 0001, Zaixi Zhang
ICML3
2024 Abstracted Shapes as Tokens - A Generalizable and Interpretable Model for Time-series Classification
abstract
In time-series analysis, many recent works seek to provide a unified view and representation for time-series across multiple domains, leading to the development of foundation models for time-series data. Despite diverse modeling techniques, existing models are black boxes and fail to provide insights and explanations about their representations. In this paper, we present VQShape, a pre-trained, generalizable, and interpretable model for time-series representation learning and classification. By introducing a novel representation for time-series data, we forge a connection between the latent space of VQShape and shape-level features. Using vector quantization, we show that time-series from different domains can be described using a unified set of low-dimensional codes, where each code can be represented as an abstracted shape in the time domain. On classification tasks, we show that the representations of VQShape can be utilized to build interpretable classifiers, achieving comparable performance to specialist models. Additionally, in zero-shot learning, VQShape and its codebook can generalize to previously unseen datasets and domains that are not included in the pre-training process. The code and pre-trained weights are available at https://github.com/YunshiWen/VQShape.
Yunshi Wen, Tengfei Ma 0001, Lily Weng, Lam M. Nguyen, A. Agung Julius
NeurIPS2
2023 Slide4N: Creating Presentation Slides from Computational Notebooks with Human-AI Collaboration
abstract
Data scientists often have to use other presentation tools (e.g., Microsoft PowerPoint) to create slides to communicate their analysis obtained using computational notebooks. Much tedious and repetitive work is needed to transfer the routines of notebooks (e.g., code, plots) to the presentable contents on slides (e.g., bullet points, figures). We propose a human-AI collaborative approach and operationalize it within Slide4N, an interactive AI assistant for data scientists to create slides from computational notebooks. Slide4N leverages advanced natural language processing techniques to distill key information from user-selected notebook cells and then renders them in appropriate slide layouts. The tool also provides intuitive interactions that allow further refinement and customization of the generated slides. We evaluated Slide4N with a two-part user study, where participants appreciated this human-AI collaborative approach compared to fully-manual or fully-automatic methods. The results also indicate the usefulness and effectiveness of Slide4N in slide creation tasks from notebooks.
Fengjie Wang, Xuye Liu, Oujing Liu, Ali Neshati, Tengfei Ma 0001, Min Zhu 0005, Jian Zhao 0010
CHI5
2023 Knowledge Graph Compression Enhances Diverse Commonsense Generation
abstract
Generating commonsense explanations requires reasoning about commonsense knowledge beyond what is explicitly mentioned in the context.Existing models use commonsense knowledge graphs such as ConceptNet to extract a subgraph of relevant knowledge pertaining to concepts in the input.However, due to the large coverage and, consequently, vast scale of ConceptNet, the extracted subgraphs may contain loosely related, redundant and irrelevant information, which can introduce noise into the model.We propose to address this by applying a differentiable graph compression algorithm that focuses on more salient and relevant knowledge for the task.The compressed subgraphs yield considerably more diverse outputs when incorporated into models for the tasks of generating commonsense and abductive explanations.Moreover, our model achieves better quality-diversity tradeoff than a large language model with 100 times the number of parameters.Our generic approach can be applied to additional NLP tasks that can benefit from incorporating external knowledge.1
Eunjeong Hwang, Veronika Thost, Vered Shwartz, Tengfei Ma 0001
EMNLP4
2023 CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code
abstract
Automatically generating function summaries for binaries is an extremely valuable but challenging task, since it involves translating the execution behavior and semantics of the low-level language (assembly code) into human-readable natural language.However, most current works on understanding assembly code are oriented towards generating function names, which involve numerous abbreviations that make them still confusing.To bridge this gap, we focus on generating complete summaries for binary functions, especially for stripped binary (no symbol table and debug information in reality).To fully exploit the semantics of assembly code, we present a control flow graph and pseudo code guided binary code summarization framework called CP-BCS.CP-BCS utilizes a bidirectional instruction-level control flow graph and pseudo code that incorporates expert knowledge to learn the comprehensive binary function execution behavior and logic semantics.We evaluate CP-BCS on 3 different binary optimization levels (O1, O2, and O3) for 3 different computer architectures (X86, X64, and ARM).The evaluation results demonstrate CP-BCS is superior and significantly improves the efficiency of reverse engineering. * Corresponding author.with limited high-level information, making it difficult to read and understand, as shown in Figure 1.Even an experienced reverse engineer needs to spend a significant amount of time determining the functionality of an assembly code snippet.
Lingfei Wu 0001, Tengfei Ma 0001, Xuhong Zhang 0002, Yangkai Du, Peiyu Liu 0003, Shouling Ji, Wenhai Wang
EMNLP3
2023 Weighted Clock Logic Point Process
Ruixuan Yan, Yunshi Wen, Debarun Bhattacharjya, Ronny Luss, Tengfei Ma 0001, Achille Fokoue, A. Agung Julius
ICLR5
2023 IGB: Addressing The Gaps In Labeling, Features, Heterogeneity, and Size of Public Graph Datasets for Deep Learning Research
abstract
Graph neural networks (GNNs) have shown high potential for a variety of real-world, challenging applications, but one of the major obstacles in GNN research is the lack of large-scale flexible datasets. Most existing public datasets for GNNs are relatively small, which limits the ability of GNNs to generalize to unseen data. The few existing large-scale graph datasets provide very limited labeled data. This makes it difficult to determine if the GNN model's low accuracy for unseen data is inherently due to insufficient training data or if the model failed to generalize. Additionally, datasets used to train GNNs need to offer flexibility to enable a thorough study of the impact of various factors while training GNN models.
Arpandeep Khatua, Vikram S. Mailthody, Bhagyashree Taleka, Tengfei Ma 0001, Xiang Song 0003, Wen-Mei W. Hwu
KDD4
2023 SyncTREE: Fast Timing Analysis for Integrated Circuit Design through a Physics-informed Tree-based Graph Neural Network
abstract
Nowadays integrated circuits (ICs) are underpinning all major information technology innovations including the current trends of artificial intelligence (AI). Modern IC designs often involve analyses of complex phenomena (such as timing, noise, and power etc.) for tens of billions of electronic components, like resistance (R), capacitance (C), transistors and gates, interconnected in various complex structures. Those analyses often need to strike a balance between accuracy and speed as those analyses need to be carried out many times throughout the entire IC design cycles. With the advancement of AI, researchers also start to explore news ways in leveraging AI to improve those analyses. This paper focuses on one of the most important analyses, timing analysis for interconnects. Since IC interconnects can be represented as an RC-tree, a specialized graph as tree, we design a novel tree-based graph neural network, SyncTREE, to speed up the timing analysis by incorporating both the structural and physical properties of electronic circuits. Our major innovations include (1) a two-pass message-passing (bottom-up and top-down) for graph embedding, (2) a tree contrastive loss to guide learning, and (3) a closed formular-based approach to conduct fast timing. Our experiments show that, compared to conventional GNN models, SyncTREE achieves the best timing prediction in terms of both delays and slews, all in reference to the industry golden numerical analyses results on real IC design data.
Jiajie Li 0002, Florian Klemme, Gi-Joon Nam, Tengfei Ma 0001, Hussam Amrouch, Jinjun Xiong
NeurIPS5
2023 Federated learning of models pre-trained on different features with consensus graphs
abstract
Learning an effective global model on private and decentralized datasets has become an increasingly important challenge of machine learning when applied in practice. Existing distributed learning paradigms, such as Federated Learning, enable this via model aggregation which enforces a strong form of modeling homogeneity and synchronicity across clients. This is however not suitable to many practical scenarios. For example, in distributed sensing, heterogeneous sensors reading data from different views of the same phenomenon would need to use different models for different data modalities. Local learning therefore happens in isolation but inference requires merging the local models to achieve consensus. To enable consensus among local models, we propose a feature fusion approach that extracts local representations from local models and incorporates them into a global representation that improves the prediction performance. Achieving this requires addressing two non-trivial problems. First, we need to learn an alignment between similar feature components which are arbitrarily arranged across clients to enable representation aggregation. Second, we need to learn a consensus graph that captures the high-order interactions between local feature spaces and how to combine them to achieve a better prediction. This paper presents solutions to these problems and demonstrates them in real-world applications on time series data such as power grids and traffic networks.
Tengfei Ma 0001, Trong Nghia Hoang, Jie Chen 0007
UAI1
2023 Multilevel Graph Matching Networks for Deep Graph Similarity Learning
abstract
While the celebrated graph neural networks (GNNs) yield effective representations for individual nodes of a graph, there has been relatively less success in extending to the task of graph similarity learning. Recent work on graph similarity learning has considered either global-level graph-graph interactions or low-level node-node interactions, however, ignoring the rich cross-level interactions (e.g., between each node of one graph and the other whole graph). In this article, we propose a multilevel graph matching network (MGMN) framework for computing the graph similarity between any pair of graph-structured objects in an end-to-end fashion. In particular, the proposed MGMN consists of a node-graph matching network (NGMN) for effectively learning cross-level interactions between each node of one graph and the other whole graph, and a siamese GNN to learn global-level interactions between two input graphs. Furthermore, to compensate for the lack of standard benchmark datasets, we have created and collected a set of datasets for both the graph-graph classification and graph-graph regression tasks with different sizes in order to evaluate the effectiveness and robustness of our models. Comprehensive experiments demonstrate that MGMN consistently outperforms state-of-the-art baseline models on both the graph-graph classification and graph-graph regression tasks. Compared with previous work, multilevel graph matching network (MGMN) also exhibits stronger robustness as the sizes of the two input graphs increase.
Xiang Ling 0001, Lingfei Wu 0001, Saizhuo Wang, Tengfei Ma 0001, Fangli Xu, Alex X. Liu, Chunming Wu 0001, Shouling Ji
IEEE Trans. Neural Networks Learn. Syst.4
2023 GNNLens: A Visual Analytics Approach for Prediction Error Diagnosis of Graph Neural Networks
abstract
Graph Neural Networks (GNNs) aim to extend deep learning techniques to graph data and have achieved significant progress in graph analysis tasks (e.g., node classification) in recent years. However, similar to other deep neural networks like Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), GNNs behave like a black box with their details hidden from model developers and users. It is therefore difficult to diagnose possible errors of GNNs. Despite many visual analytics studies being done on CNNs and RNNs, little research has addressed the challenges for GNNs. This paper fills the research gap with an interactive visual analysis tool, GNNLens, to assist model developers and users in understanding and analyzing GNNs. Specifically, Parallel Sets View and Projection View enable users to quickly identify and validate error patterns in the set of wrong predictions; Graph View and Feature Matrix View offer a detailed analysis of individual nodes to assist users in forming hypotheses about the error patterns. Since GNNs jointly model the graph structure and the node features, we reveal the relative influences of the two types of information by comparing the predictions of three models: GNN, Multi-Layer Perceptron (MLP), and GNN Without Using Features (GNNWUF). Two case studies and interviews with domain experts demonstrate the effectiveness of GNNLens in facilitating the understanding of GNN models and their errors.
Zhihua Jin, Yong Wang 0021, Qianwen Wang 0001, Yao Ming, Tengfei Ma 0001, Huamin Qu
IEEE Trans. Vis. Comput. Graph.5
2022 Neuro-symbolic Models for Interpretable Time Series Classification using Temporal Logic Description
abstract
Most existing Time series classification (TSC) models lack interpretability and are difficult to inspect. Interpretable machine learning models can aid in discovering patterns in data as well as give easy-to-understand insights to domain specialists. In this study, we present Neuro-Symbolic Time Series Classification (NSTSC), a neuro-symbolic model that leverages signal temporal logic (STL) and neural network (NN) to accomplish TSC tasks using multi-view data representation and expresses the model as a human-readable, interpretable formula. In NSTSC, each neuron is linked to a symbolic expression, i.e., an STL (sub)formula. The output of NSTSC is thus interpretable as an STL formula akin to natural language, describing temporal and logical relations hidden in the data. We propose an NSTSC-based classifier that adopts a decision-tree approach to learn formula structures and accomplish a multiclass TSC task. The proposed smooth activation functions enable the model to be learned in an end-to-end fashion. We test NSTSC on a real-world wound healing dataset from mice and benchmark datasets from the UCR time-series repository, demonstrating that NSTSC achieves comparable performance with the state-of-the-art models. Furthermore, NSTSC can generate interpretable formulas that match domain knowledge.
Ruixuan Yan, Tengfei Ma 0001, Achille Fokoue, Maria Chang 0001, A. Agung Julius
ICDM2
2022 Cycle Representation Learning for Inductive Relation Prediction
abstract
In recent years, algebraic topology and its modern development, the theory of persistent homology, has shown great potential in graph representation learning. In this paper, based on the mathematics of algebraic topology, we propose a novel solution for inductive relation prediction, an important learning task for knowledge graph completion. To predict the relation between two entities, one can use the existence of rules, namely a sequence of relations. Previous works view rules as paths and primarily focus on the searching of paths between entities. The space of rules is huge, and one has to sacrifice either efficiency or accuracy. In this paper, we consider rules as cycles and show that the space of cycles has a unique structure based on the mathematics of algebraic topology. By exploring the linear structure of the cycle space, we can improve the searching efficiency of rules. We propose to collect cycle bases that span the space of cycles. We build a novel GNN framework on the collected cycles to learn the representations of cycles, and to predict the existence/non-existence of a relation. Our method achieves state-of-the-art performance on benchmarks.
Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Chao Chen 0012
ICML2
2022 Improving Inductive Link Prediction Using Hyper-Relational Facts (Extended Abstract)
abstract
For many years, link prediction on knowledge. graphs has been a purely transductive task, not allowing for reasoning on unseen entities. Recently, increasing efforts are put into exploring semi- and fully inductive scenarios, enabling inference over unseen and emerging entities. Still, all these approaches only consider triple-based KGs, whereas their richer counterparts, hyper-relational KGs (e.g., Wikidata), have not yet been properly studied. In this work, we classify different inductive settings and study the benefits of employing hyper-relational KGs on a wide range of semi- and fully inductive link prediction tasks powered by recent advancements in graph neural networks. Our experiments on a novel set of benchmarks show that qualifiers over typed edges can lead to performance improvements of 6% of absolute gains (for the Hits@10 metric) compared to triple-only baselines. Our code is available at https://github.com/mali-git/hyper_relational_ilp.
Mehdi Ali, Max Berrendorf, Michael Galkin, Veronika Thost, Tengfei Ma 0001, Volker Tresp, Jens Lehmann 0001
IJCAI5
2022 MalGraph: Hierarchical Graph Neural Networks for Robust Windows Malware Detection
abstract
With the ever-increasing malware threats, malware detection plays an indispensable role in protecting information systems. Although tremendous research efforts have been made, there are still two key challenges hindering them from being applied to accurately and robustly detect malwares. Firstly, most of them represent executables with shallow features, but ignore their semantic and structural information. Secondly, they are primarily based on representations that can be easily modified by attackers and thus cannot provide robustness against adversarial attacks. To tackle the challenges, we present MalGraph, which first represents executables with hierarchical graphs and then uses an end-to-end learning framework based on graph neural networks for malware detection. In particular, a hierarchical graph consists of a function call graph that captures the interaction semantics among different functions at the inter-function level and corresponding control-flow graphs for learning the structural semantics of each function at the intra-function level. We argue the abstraction and hierarchy nature of hierarchical graphs makes them not only easy to capture rich structural information of executables, but also be immune to adversarial attacks. Evaluations show that MalGraph not only outperforms state-of-the-art malware detection, but also exhibits stronger robustness against adversarial attacks by a large margin.
Xiang Ling 0001, Lingfei Wu 0001, Zhenqing Qu, Jiangyu Zhang, Tengfei Ma 0001, Bin Wang 0062, Chunming Wu 0001, Shouling Ji
INFOCOM7
2022 A Study of the Attention Abnormality in Trojaned BERTs
abstract
Trojan attacks raise serious security concerns.In this paper, we investigate the underlying mechanism of Trojaned BERT models.We observe the attention focus drifting behavior of Trojaned models, i.e., when encountering an poisoned input, the trigger token hijacks the attention focus regardless of the context.We provide a thorough qualitative and quantitative analysis of this phenomenon, revealing insights into the Trojan mechanism.Based on the observation, we propose an attention-based Trojan detector to distinguish Trojaned models from clean ones.To the best of our knowledge, this is the first paper to analyze the Trojan mechanism and to develop a Trojan detector based on the transformer's attention 1 .
Weimin Lyu, Songzhu Zheng, Tengfei Ma 0001, Chao Chen 0012
NAACL-HLT3
2022 Neural Approximation of Graph Topological Features
abstract
Topological features based on persistent homology capture high-order structural information so as to augment graph neural network methods. However, computing extended persistent homology summaries remains slow for large and dense graphs and can be a serious bottleneck for the learning pipeline. Inspired by recent success in neural algorithmic reasoning, we propose a novel graph neural network to estimate extended persistence diagrams (EPDs) on graphs efficiently. Our model is built on algorithmic insights, and benefits from better supervision and closer alignment with the EPD computation algorithm. We validate our method with convincing empirical results on approximating EPDs and downstream graph representation learning tasks. Our method is also efficient; on large and dense graphs, we accelerate the computation by nearly 100 times.
Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Yusu Wang 0001, Chao Chen 0012
NeurIPS2
2022 Exploiting Heterogeneous Graph Neural Networks with Latent Worker/Task Correlation Information for Label Aggregation in Crowdsourcing
abstract
Crowdsourcing has attracted much attention for its convenience to collect labels from non-expert workers instead of experts. However, due to the high level of noise from the non-experts, a label aggregation model that infers the true label from noisy crowdsourced labels is required. In this article, we propose a novel framework based on graph neural networks for aggregating crowd labels. We construct a heterogeneous graph between workers and tasks and derive a new graph neural network to learn the representations of nodes and the true labels. Besides, we exploit the unknown latent interaction between the same type of nodes (workers or tasks) by adding a homogeneous attention layer in the graph neural networks. Experimental results on 13 real-world datasets show superior performance over state-of-the-art models.
Hanlu Wu, Tengfei Ma 0001, Lingfei Wu 0001, Fangli Xu, Shouling Ji
ACM Trans. Knowl. Discov. Data2
2022 CHEER: Rich Model Helps Poor Model via Knowledge Infusion
abstract
There is a growing interest in applying deep learning (DL) to healthcare, driven by the availability of data with multiple feature channels inrich-dataenvironments (e.g., intensive care units). However, in many other practical situations, we can only access data with much fewer feature channels in apoor-dataenvironments (e.g., at home), which often results in predictive models with poor performance. How can we boost the performance of models learned from suchpoor-dataenvironment by leveraging knowledge extracted from existing models trained usingrich datain a related environment? To address this question, we develop a knowledge infusion framework namedCHEERthat can succinctly summarize suchrich modelinto transferable representations, which can be incorporated into thepoor modelto improve its performance. The infused model is analyzed theoretically and evaluated empirically on several datasets. Our empirical results showed thatCHEERoutperformed baselines by 5.60 to 46.80 percent in terms of the macro-F1 score on multiple physiological datasets.
Cao Xiao, Trong Nghia Hoang, Shenda Hong, Tengfei Ma 0001, Jimeng Sun 0001
IEEE Trans. Knowl. Data Eng.4
2021 Unsupervised Learning of Graph Hierarchical Abstractions with Differentiable Coarsening and Optimal Transport
abstract
Hierarchical abstractions are a methodology for solving large-scale graph problems in various disciplines. Coarsening is one such approach: it generates a pyramid of graphs whereby the one in the next level is a structural summary of the prior one. With a long history in scientific computing, many coarsening strategies were developed based on mathematically driven heuristics. Recently, resurgent interests exist in deep learning to design hierarchical methods learnable through differentiable parameterization. These approaches are paired with downstream tasks for supervised learning. In practice, however, supervised signals (e.g., labels) are scarce and are often laborious to obtain. In this work, we propose an unsupervised approach, coined OTCoarsening, with the use of optimal transport. Both the coarsening matrix and the transport cost matrix are parameterized, so that an optimal coarsening strategy can be learned and tailored for a given set of graphs. We demonstrate that the proposed approach produces meaningful coarse graphs and yields competitive performance compared with supervised methods for graph classification and regression.
Tengfei Ma 0001, Jie Chen 0007
AAAI1
2021 Timeline Summarization based on Event Graph Compression via Time-Aware Optimal Transport
abstract
Timeline Summarization identifies major events from a news collection and describes them following temporal order, with key dates tagged.Previous methods generally generate summaries separately for each date after they determine the key dates of events.These methods overlook the events' intra-structures (arguments) and inter-structures (event-event connections).Following a different route, we propose to represent the news articles as an event-graph, thus the summarization task becomes compressing the whole graph to its salient sub-graph.The key hypothesis is that the events connected through shared arguments and temporal order depict the skeleton of a timeline, containing events that are semantically related, structurally salient, and temporally coherent in the global event graph.A time-aware optimal transport distance is then introduced for learning the compression model in an unsupervised manner.We show that our approach significantly improves the state of the art on three real-world datasets, including two public standard benchmarks and our newly collected Timeline 100 dataset. 1
Manling Li, Tengfei Ma 0001, Mo Yu, Lingfei Wu 0001, Tian Gao 0007, Heng Ji 0001, Kathy McKeown
EMNLP (1)2
2021 Link Prediction with Persistent Homology: An Interactive View
abstract
Link prediction is an important learning task for graph-structured data. In this paper, we propose a novel topological approach to characterize interactions between two nodes. Our topological feature, based on the extended persistent homology, encodes rich structural information regarding the multi-hop paths connecting nodes. Based on this feature, we propose a graph neural network method that outperforms state-of-the-arts on different benchmarks. As another contribution, we propose a novel algorithm to more efficiently compute the extended persistence diagrams for graphs. This algorithm can be generally applied to accelerate many other topological methods for graph learning tasks.
Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Chao Chen 0012
ICML2
2021 Improving Inductive Link Prediction Using Hyper-relational Facts
Mehdi Ali, Max Berrendorf, Michael Galkin, Veronika Thost, Tengfei Ma 0001, Volker Tresp, Jens Lehmann 0001
ISWC5
2021 Deep Graph Matching and Searching for Semantic Code Retrieval
abstract
Code retrieval is to find the code snippet from a large corpus of source code repositories that highly matches the query of natural language description. Recent work mainly uses natural language processing techniques to process both query texts (i.e., human natural language) and code snippets (i.e., machine programming language), however, neglecting the deep structured features of query texts and source codes, both of which contain rich semantic information. In this article, we propose an end-to-end deep graph matching and searching (DGMS) model based on graph neural networks for the task of semantic code retrieval. To this end, we first represent both natural language query texts and programming language code snippets with the unified graph-structured data, and then use the proposed graph matching and searching model to retrieve the best matching code snippet. In particular, DGMS not only captures more structural information for individual query texts or code snippets, but also learns the fine-grained similarity between them by cross-attention based semantic matching operations. We evaluate the proposed DGMS model on two public code retrieval datasets with two representative programming languages (i.e., Java and Python). Experiment results demonstrate that DGMS significantly outperforms state-of-the-art baseline models by a large margin on both datasets. Moreover, our extensive ablation studies systematically investigate and illustrate the impact of each part of DGMS.
Xiang Ling 0001, Lingfei Wu 0001, Saizhuo Wang, Tengfei Ma 0001, Fangli Xu, Alex X. Liu, Chunming Wu 0001, Shouling Ji
ACM Trans. Knowl. Discov. Data5
2020 Online Planner Selection with Graph Neural Networks and Adaptive Scheduling
abstract
Automated planning is one of the foundational areas of AI. Since no single planner can work well for all tasks and domains, portfolio-based techniques have become increasingly popular in recent years. In particular, deep learning emerges as a promising methodology for online planner selection. Owing to the recent development of structural graph representations of planning tasks, we propose a graph neural network (GNN) approach to selecting candidate planners. GNNs are advantageous over a straightforward alternative, the convolutional neural networks, in that they are invariant to node permutations and that they incorporate node labels for better inference.Additionally, for cost-optimal planning, we propose a two-stage adaptive scheduling method to further improve the likelihood that a given task is solved in time. The scheduler may switch at halftime to a different planner, conditioned on the observed performance of the first one. Experimental results validate the effectiveness of the proposed method against strong baselines, both deep learning and non-deep learning based.The code is available at https://github.com/matenure/GNN_planner.
Tengfei Ma 0001, Patrick Ferber, Siyu Huo, Jie Chen 0007, Michael Katz 0001
AAAI1
2020 EvolveGCN: Evolving Graph Convolutional Networks for Dynamic Graphs
abstract
Graph representation learning resurges as a trending research subject owing to the widespread use of deep learning for Euclidean data, which inspire various creative designs of neural networks in the non-Euclidean domain, particularly graphs. With the success of these graph neural networks (GNN) in the static setting, we approach further practical scenarios where the graph dynamically evolves. Existing approaches typically resort to node embeddings and use a recurrent neural network (RNN, broadly speaking) to regulate the embeddings and learn the temporal dynamics. These methods require the knowledge of a node in the full time span (including both training and testing) and are less applicable to the frequent change of the node set. In some extreme scenarios, the node sets at different time steps may completely differ. To resolve this challenge, we propose EvolveGCN, which adapts the graph convolutional network (GCN) model along the temporal dimension without resorting to node embeddings. The proposed approach captures the dynamism of the graph sequence through using an RNN to evolve the GCN parameters. Two architectures are considered for the parameter evolution. We evaluate the proposed approach on tasks including link prediction, edge classification, and node classification. The experimental results indicate a generally higher performance of EvolveGCN compared with related approaches. The code is available at https://github.com/IBM/EvolveGCN.
Aldo Pareja, Giacomo Domeniconi, Jie Chen 0007, Tengfei Ma 0001, Toyotaro Suzumura, Hiroki Kanezashi, Tim Kaler, Tao B. Schardl, Charles E. Leiserson
AAAI4
2020 Supervised Topic Compositional Neural Language Model for Clinical Narrative Understanding
abstract
Clinical narratives that describe complex medical events are often accompanied by meta-information such as a patient’s demographics, diagnoses and medications. This structured information implicitly relates to the logical and semantic structure of the entire narrative, and thus affects vocabulary choices for the narrative composition. To leverage this meta-information, we propose a supervised topic compositional neural language model, called MeTRNN, that integrates the strength of supervised topic modeling in capturing global semantics with the capacity of contextual recurrent neural networks (RNN) in modeling local word dependencies. MeTRNN generates interpretable topics from global meta-information and uses them to facilitate contextual RNNs in modeling local dependencies of text. For efficient training of MeTRNN, we develop an autoencoding variational Bayes inference method. We evaluate MeTRNN on the word prediction tasks using public text datasets. MeTRNN consistently outperforms all baselines across all datasets in perplexity ranging from 5% to 40%. Our case studies on real world electronic health records (EHR) data show that MeTRNN can learn and benefit from meaningful topics.
Xiao Qin 0003, Cao Xiao, Tengfei Ma 0001, Tabassum Kakar, Susmitha Wunnava, Xiangnan Kong, Elke A. Rundensteiner, Fei Wang 0001
IEEE BigData3
2020 Unsupervised Reference-Free Summary Quality Evaluation via Contrastive Learning
abstract
Evaluation of a document summarization system has been a critical factor to impact the success of the summarization task. Previous approaches, such as ROUGE, mainly consider the informativeness of the assessed summary and require human-generated references for each test summary. In this work, we propose to evaluate the summary qualities without reference summaries by unsupervised contrastive learning. Specifically, we design a new metric which covers both linguistic qualities and semantic informativeness based on BERT. To learn the metric, for each summary, we construct different types of negative samples with respect to different aspects of the summary qualities, and train our model with a ranking loss. Experiments on Newsroom and CNN/Daily Mail demonstrate that our new evaluation method outperforms other metrics even without reference summaries. Furthermore, we show that our method is general and transferable across datasets.
Hanlu Wu, Tengfei Ma 0001, Lingfei Wu 0001, Tariro Manyumwa, Shouling Ji
EMNLP (1)2
2020 Curvature Graph Network
Ze Ye, Kin Sum Liu, Tengfei Ma 0001, Jie Gao 0001, Chao Chen 0012
ICLR3
2020 Deep Graph Learning: Foundations, Advances and Applications
abstract
Many real data come in the form of non-grid objects, i.e. graphs, from social networks to molecules. Adaptation of deep learning from grid-alike data (e.g. images) to graphs has recently received unprecedented attention from both machine learning and data mining communities, leading to a new cross-domain field---Deep Graph Learning (DGL). Instead of painstaking feature engineering, DGL aims to learn informative representations of graphs in an end-to-end manner. It has exhibited remarkable success in various tasks, such as node/graph classification, link prediction, etc.
Yu Rong 0001, Tingyang Xu, Junzhou Huang, Wenbing Huang 0001, Hong Cheng 0001, Yao Ma 0001, Yiqi Wang 0001, Tyler Derr, Lingfei Wu 0001, Tengfei Ma 0001
KDD10
2020 DeepDrawing: A Deep Learning Approach to Graph Drawing
abstract
Node-link diagrams are widely used to facilitate network explorations. However, when using a graph drawing technique to visualize networks, users often need to tune different algorithm-specific parameters iteratively by comparing the corresponding drawing results in order to achieve a desired visual effect. This trial and error process is often tedious and time-consuming, especially for non-expert users. Inspired by the powerful data modelling and prediction capabilities of deep learning techniques, we explore the possibility of applying deep learning techniques to graph drawing. Specifically, we propose using a graph-LSTM-based approach to directly map network structures to graph drawings. Given a set of layout examples as the training dataset, we train the proposed graph-LSTM-based model to capture their layout characteristics. Then, the trained model is used to generate graph drawings in a similar style for new networks. We evaluated the proposed approach on two special types of layouts (i.e., grid layouts and star layouts) and two general types of layouts (i.e., ForceAtlas2 and PivotMDS) in both qualitative and quantitative ways. The results provide support for the effectiveness of our approach. We also conducted a time cost assessment on the drawings of small graphs with 20 to 50 nodes. We further report the lessons we learned and discuss the limitations and future work.
Yong Wang 0021, Zhihua Jin, Qianwen Wang 0001, Weiwei Cui 0001, Tengfei Ma 0001, Huamin Qu
IEEE Trans. Vis. Comput. Graph.5
2019 GAMENet: Graph Augmented MEmory Networks for Recommending Medication Combination
abstract
Recent progress in deep learning is revolutionizing the healthcare domain including providing solutions to medication recommendations, especially recommending medication combination for patients with complex health conditions. Existing approaches either do not customize based on patient health history, or ignore existing knowledge on drug-drug interactions (DDI) that might lead to adverse outcomes. To fill this gap, we propose the Graph Augmented Memory Networks (GAMENet), which integrates the drug-drug interactions knowledge graph by a memory module implemented as a graph convolutional networks, and models longitudinal patient records as the query. It is trained end-to-end to provide safe and personalized recommendation of medication combination. We demonstrate the effectiveness and safety of GAMENet by comparing with several state-of-the-art methods on real EHR data. GAMENet outperformed all baselines in all effectiveness measures, and also achieved 3.60% DDI rate reduction from existing EHR data.
Junyuan Shang, Cao Xiao, Tengfei Ma 0001, Hongyan Li 0002, Jimeng Sun 0001
AAAI3
2019 Reaching Data Confidentiality and Model Accountability on the CalTrain
abstract
Distributed collaborative learning (DCL) paradigms enable building joint machine learning models from distrusted multi-party participants. Data confidentiality is guaranteed by retaining private training data on each participant's local infrastructure. However, this approach makes today's DCL design fundamentally vulnerable to data poisoning and backdoor attacks. It limits DCL's model accountability, which is key to backtracking problematic training data instances and their responsible contributors. In this paper, we introduce CALTRAIN, a centralized collaborative learning system that simultaneously achieves data confidentiality and model accountability. CALTRAIN enforces isolated computation via secure enclaves on centrally aggregated training data to guarantee data confidentiality. To support building accountable learning models, we securely maintain the links between training instances and their contributors. Our evaluation shows that the models generated by CALTRAIN can achieve the same prediction accuracy when compared to the models trained in non-protected environments. We also demonstrate that when malicious training participants tend to implant backdoors during model training, CALTRAIN can accurately and precisely discover the poisoned or mislabeled training data that lead to the runtime mispredictions.
Zhongshu Gu, Hani Jamjoom, Dong Su, Heqing Huang 0001, Jialong Zhang 0001, Tengfei Ma 0001, Dimitrios E. Pendarakis, Ian M. Molloy
DSN6
2019 Pre-Training BERT on Domain Resources for Short Answer Grading
abstract
Chul Sung, Tejas Dhamecha, Swarnadeep Saha, Tengfei Ma, Vinay Reddy, Rishi Arora. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chul Sung, Tejas I. Dhamecha, Swarnadeep Saha, Tengfei Ma 0001, Vinay Reddy, Rishi Arora
EMNLP/IJCNLP (1)4
2019 RDPD: Rich Data Helps Poor Data via Imitation
abstract
In many situations, we need to build and deploy separate models in related environments with different data qualities. For example, an environment with strong observation equipments (e.g., intensive care units) often provides high-quality multi-modal data, which are acquired from multiple sensory devices and have rich-feature representations. On the other hand, an environment with poor observation equipment (e.g., at home) only provides low-quality, uni-modal data with poor-feature representations. To deploy a competitive model in a poor-data environment without requiring direct access to multi-modal data acquired from a rich-data environment, this paper develops and presents a knowledge distillation (KD) method (RDPD) to enhance a predictive model trained on poor data using knowledge distilled from a high-complexity model trained on rich, private data. We evaluated RDPD on three real-world datasets and shown that its distilled model consistently outperformed all baselines across all datasets, especially achieving the greatest performance improvement over a model trained only on low-quality data by 24.56% on PR-AUC and 12.21% on ROC-AUC, and over that of a state-of-the-art KD model by 5.91% on PR-AUC and 4.44% on ROC-AUC.
Shenda Hong, Cao Xiao, Trong Nghia Hoang, Tengfei Ma 0001, Hongyan Li 0002, Jimeng Sun 0001
IJCAI4
2019 MINA: Multilevel Knowledge-Guided Attention for Modeling Electrocardiography Signals
abstract
Electrocardiography (ECG) signals are commonly used to diagnose various cardiac abnormalities. Recently, deep learning models showed initial success on modeling ECG data, however they are mostly black-box, thus lack interpretability needed for clinical usage. In this work, we propose MultIlevel kNowledge-guided Attention networks (MINA) that predict heart diseases from ECG signals with intuitive explanation aligned with medical knowledge. By extracting multilevel (beat-, rhythm- and frequency-level) domain knowledge features separately, MINA combines the medical knowledge and ECG data via a multilevel attention model, making the learned models highly interpretable. Our experiments showed MINA achieved PR-AUC 0.9436 (outperforming the best baseline by 5.51%) in real world ECG dataset. Finally, MINA also demonstrated robust performance and strong interpretability against signal distortion and noise contamination.
Shenda Hong, Cao Xiao, Tengfei Ma 0001, Hongyan Li 0002, Jimeng Sun 0001
IJCAI3
2019 Pre-training of Graph Augmented Transformers for Medication Recommendation
abstract
Medication recommendation is an important healthcare application. It is commonly formulated as a temporal prediction task. Hence, most existing works only utilize longitudinal electronic health records (EHRs) from a small number of patients with multiple visits ignoring a large number of patients with a single visit (selection bias). Moreover, important hierarchical knowledge such as diagnosis hierarchy is not leveraged in the representation learning process. Despite the success of deep learning techniques in computational phenotyping, most previous approaches have two limitations: task-oriented representation and ignoring hierarchies of medical codes. To address these challenges, we propose G-BERT, a new model to combine the power of Graph Neural Networks (GNNs) and BERT (Bidirectional Encoder Representations from Transformers) for medical code representation and medication recommendation. We use GNNs to represent the internal hierarchical structures of medical codes. Then we integrate the GNN representation into a transformer-based visit encoder and pre-train it on EHR data from patients only with a single visit. The pre-trained visit encoder and representation are then fine-tuned for downstream predictive tasks on longitudinal EHRs from patients with multiple visits. G-BERT is the first to bring the language model pre-training schema into the healthcare domain and it achieved state-of-the-art performance on the medication recommendation task.
Junyuan Shang, Tengfei Ma 0001, Cao Xiao, Jimeng Sun 0001
IJCAI2
2018 Preliminary Evaluations of a Dialogue-Based Digital Tutor
Matthew Ventura, Maria Chang 0001, Peter W. Foltz, Nirmal Mukhi, Jessica Yarbro, Anne Pier Salverda, John T. Behrens, Jae-wook Ahn, Tengfei Ma 0001, Tejas I. Dhamecha, Smit Marvaniya, Patrick Watson, Cassius D'Helon, Ravi Tejwani, Shazia Afzal
AIED (2)9
2018 FastGCN: Fast Learning with Graph Convolutional Networks via Importance Sampling
Jie Chen 0007, Tengfei Ma 0001, Cao Xiao
ICLR (Poster)2
2018 Drug Similarity Integration Through Attentive Multi-view Graph Auto-Encoders
abstract
Drug similarity has been studied to support downstream clinical tasks such as inferring novel properties of drugs (e.g. side effects, indications, interactions) from known properties. The growing availability of new types of drug features brings the opportunity of learning a more comprehensive and accurate drug similarity that represents the full spectrum of underlying drug relations. However, it is challenging to integrate these heterogeneous, noisy, nonlinear-related information to learn accurate similarity measures especially when labels are scarce. Moreover, there is a trade-off between accuracy and interpretability. In this paper, we propose to learn accurate and interpretable similarity measures from multiple types of drug features. In particular, we model the integration using multi-view graph auto-encoders, and add attentive mechanism to determine the weights for each view with respect to corresponding tasks and features for better interpretability. Our model has flexible design for both semi-supervised and unsupervised settings. Experimental results demonstrated significant predictive accuracy improvement. Case studies also showed better model capacity (e.g. embed node features) and interpretability.
Tengfei Ma 0001, Cao Xiao, Fei Wang 0001
IJCAI1
2018 Constrained Generation of Semantically Valid Graphs via Regularizing Variational Autoencoders
abstract
Deep generative models have achieved remarkable success in various data domains, including images, time series, and natural languages. There remain, however, substantial challenges for combinatorial structures, including graphs. One of the key challenges lies in the difficulty of ensuring semantic validity in context. For example, in molecular graphs, the number of bonding-electron pairs must not exceed the valence of an atom; whereas in protein interaction networks, two proteins may be connected only when they belong to the same or correlated gene ontology terms. These constraints are not easy to be incorporated into a generative model. In this work, we propose a regularization framework for variational autoencoders as a step toward semantic validity. We focus on the matrix representation of graphs and formulate penalty terms that regularize the output distribution of the decoder to encourage the satisfaction of validity constraints. Experimental results confirm a much higher likelihood of sampling valid graphs in our approach, compared with others reported in the literature.
Tengfei Ma 0001, Jie Chen 0007, Cao Xiao
NeurIPS1
2018 Health-ATM: A Deep Architecture for Multifaceted Patient Health Record Representation and Risk Prediction
abstract
Leveraging massive electronic health records (EHR) brings tremendous promises to advance clinical and precision medicine informatics research. However, it is very challenging to directly work with multifaceted patient information encoded in their EHR data. Deriving effective representations of patient EHRs is a crucial step to bridge raw EHR information and the endpoint analytical tasks, such as risk prediction or disease subtyping. In this paper, we propose Health-ATM, a novel and integrated deep architecture to uncover patients' comprehensive health information from their noisy, longitudinal, heterogeneous and irregular EHR data. Health-ATM extracts comprehensive multifaceted patient information patterns with attentive and time-aware modulars (ATM) and a hybrid network structure composed of both Recurrent Neural Network (RNN) and Convolutional Neural Network (CNN). The learned features are finally fed into a prediction layer to conduct the risk prediction task. We evaluated the Health-ATM on both artificial and real world EHR corpus and demonstrated its promising utility and efficacy on representation learning and disease onset predictions.
Tengfei Ma 0001, Cao Xiao, Fei Wang 0001
SDM1
2017 Wizard's Apprentice: Cognitive Suggestion Support for Wizard-of-Oz Question Answering
Jae-wook Ahn, Patrick Watson, Maria Chang 0001, Sharad Sundararajan, Tengfei Ma 0001, Nirmal Mukhi, Srijith Prabhu
AIED5
2017 Multilingual Training of Crosslingual Word Embeddings
abstract
Long Duong, Hiroshi Kanayama, Tengfei Ma, Steven Bird, Trevor Cohn. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Long Duong, Hiroshi Kanayama, Tengfei Ma 0001, Steven Bird, Trevor Cohn
EACL (1)3
2017 Inverted Bilingual Topic Models for Lexicon Extraction from Non-parallel Data
abstract
Topic models have been successfully applied in lexicon extraction. However, most previous methods are limited to document-aligned data. In this paper, we try to address two challenges of applying topic models to lexicon extraction in non-parallel data: 1) hard to model the word relationship and 2) noisy seed dictionary. To solve these two challenges, we propose two new bilingual topic models to better capture the semantic information of each word while discriminating the multiple translations in a noisy seed dictionary. We extend the scope of topic models by inverting the roles of "word" and "document". In addition, to solve the problem of noise in seed dictionary, we incorporate the probability of translation selection in our models. Moreover, we also propose an effective measure to evaluate the similarity of words in different languages and select the optimal translation pairs. Experimental results using real world data demonstrate the utility and efficacy of the proposed models.
Tengfei Ma 0001, Tetsuya Nasukawa
IJCAI1
2016 Learning Crosslingual Word Embeddings without Bilingual Corpora
abstract
Crosslingual word embeddings represent lexical items from different languages in the same vector space, enabling transfer of NLP tools. However, previous attempts had expensive resource requirements, difficulty incorporating monolingual data or were unable to handle polysemy. We address these drawbacks in our method which takes advantage of a high coverage dictionary in an EM style training algorithm over monolingual corpora in two languages. Our model achieves state-of-the-art performance on bilingual lexicon induction task exceeding models using large bilingual corpora, and competitive results on the monolingual word similarity and cross-lingual document classification task.
Long Duong, Hiroshi Kanayama, Tengfei Ma 0001, Steven Bird, Trevor Cohn
EMNLP3
2015 The Hybrid Nested/Hierarchical Dirichlet Process and its Application to Topic Modeling with Word Differentiation
abstract
The hierarchical Dirichlet process (HDP) is a powerful nonparametric Bayesian approach to modeling groups of data which allows the mixture components in each group to be shared. However, in many cases the groups themselves are also in latent groups (categories) which may impact the modeling a lot. In order to utilize the unknown category information of grouped data, we present the hybrid nested/ hierarchical Dirichlet process (hNHDP), a prior that blends the desirable aspects of both the HDP and the nested Dirichlet Process (NDP). Specifically, we introduce a clustering structure for the groups. The prior distribution for each cluster is a realization of a Dirichlet process. Moreover, the set of cluster-specific distributions can share part of atoms between groups, and the shared atoms and specific atoms are generated separately. We apply the hNHDP to document modeling and bring in a mechanism to identify discriminative words and topics. We derive an efficient Markov chain Monte Carlo scheme for posterior inference and present experiments on document modeling.
Tengfei Ma 0001, Issei Sato, Hiroshi Nakagawa
AAAI1
2013 Automatically Determining a Proper Length for Multi-Document Summarization: A Bayesian Nonparametric Approach
abstract
Document summarization is an important task in the area of natural language processing, which aims to extract the most important information from a single document or a cluster of documents.In various summarization tasks, the summary length is manually defined.However, how to find the proper summary length is quite a problem; and keeping all summaries restricted to the same length is not always a good choice.It is obviously improper to generate summaries with the same length for two clusters of documents which contain quite different quantity of information.In this paper, we propose a Bayesian nonparametric model for multidocument summarization in order to automatically determine the proper lengths of summaries.Assuming that an original document can be reconstructed from its summary, we describe the "reconstruction" by a Bayesian framework which selects sentences to form a good summary.Experimental results on DUC2004 data sets and some expanded data demonstrate the good quality of our summaries and the rationality of the length determination.
Tengfei Ma 0001, Hiroshi Nakagawa
EMNLP1
2011 Named Entity Recognition in Chinese News Comments on the Web
Xiaojun Wan 0001, Liang Zong, Xiaojiang Huang, Tengfei Ma 0001, Houping Jia, Yuqian Wu, Jianguo Xiao
IJCNLP4
2010 Multi-document Summarization Using Minimum Distortion
abstract
Document summarization plays an important role in the area of natural language processing and text mining. This paper proposes several novel information-theoretic models for multi-document summarization. They consider document summarization as a transmission system and assume that the best summary should have the minimum distortion. By defining a proper distortion measure and a new representation method, the combination of the last two models (the linear representation model and the facility location model) gains good experimental results on the DUC2002 and DUC2004 datasets. Moreover, we also indicate that the model has high interpretability and extensibility.
Tengfei Ma 0001, Xiaojun Wan 0001
ICDM1