VLDB 2026 Research / reviewers in the wild / expert
Bayu Distiawan Trisedya
dblp:115/5506 · also Bayu Distiawan
· DBLP profile ↗
13ranked-venue papers
7as first author
7since 2021 · last 2024
0000-0002-1672-9483ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hierarchical Shared Encoder With Task-Specific Transformer Layer Selection for Emotion-Cause Pair ExtractionabstractEmotion Cause Pair Extraction (ECPE) aims to extract emotions and their causes from a document. Powerful emotion and cause extraction abilities have proven essential in achieving accurate ECPE. However, most existing methods employ shared feature learning of emotion extraction and cause extraction, which can harm the abilities of both tasks as they focus on different information (i.e., task-specific features). Moreover, shared feature learning of the two tasks also leads to the label imbalance problem. To address these issues, this paper proposes a multi-task learning framework named Hierarchical Shared Encoder with Task-specific Transformer Layer Selection (HSE-TTLS). The model achieves ECPE via two subtasks: Emotion Extraction (EE) and Emotion Cause Extraction (ECE). The design of two subtasks for ECPE corresponds to the fact that cause clauses are emotion-dependent and significantly alleviates the label imbalance problem. To effectively extract task-specific features for EE and ECE, we employ BERT as the token-level encoder and select task-specific optimal layers for the two subtasks. Focal loss is used as the objective function for EE to further alleviate the label imbalance problem. Extensive experiments on benchmark ECPE corpus demonstrate the effectiveness of HSE-TTLS, which outperforms state-of-the-art baseline methods by at least 1.56% on the F1 score. Xinxin Su, Zhen Huang 0006, Yixin Su 0001, Bayu Distiawan Trisedya, Yong Dou |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | AutoAlign: Fully Automatic and Effective Knowledge Graph Alignment Enabled by Large Language ModelsabstractThe task of entity alignment between knowledge graphs (KGs) aims to identify every pair of entities from two different KGs that represent the same entity. Many machine learning-based methods have been proposed for this task. However, to our best knowledge, existing methods all requiremanually craftedseed alignments, which are expensive to obtain. In this paper, we propose the first fully automatic alignment method named AutoAlign, which does not require any manually crafted seed alignments. Specifically, for predicate embeddings, AutoAlign constructs a predicate-proximity-graph with the help of large language models to automatically capture the similarity between predicates across two KGs. For entity embeddings, AutoAlign first computes the entity embeddings of each KG independently using TransE, and then shifts the two KGs' entity embeddings into the same vector space by computing the similarity between entities based on their attributes. Thus, both predicate alignment and entity alignment can be done without manually crafted seed alignments. AutoAlign is not only fully automatic, but also highly effective. Experiments using real-world KGs show that AutoAlign improves the performance of entity alignment significantly compared to state-of-the-art methods. Our source code is available at ruizhang-ai/AutoAlign. Rui Zhang 0003, Yixin Su 0001, Bayu Distiawan Trisedya, Xiaoyan Zhao 0005, Min Yang 0007, Hong Cheng 0001, Jianzhong Qi 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | i-Align: an interpretable knowledge graph alignment modelabstractAbstract Knowledge graphs (KGs) are becoming essential resources for many downstream applications. However, their incompleteness may limit their potential. Thus, continuous curation is needed to mitigate this problem. One of the strategies to address this problem is KG alignment, i.e., forming a more complete KG by merging two or more KGs. This paper proposes i-Align, an interpretable KG alignment model. Unlike the existing KG alignment models, i-Align provides an explanation for each alignment prediction while maintaining high alignment performance. Experts can use the explanation to check the correctness of the alignment prediction. Thus, the high quality of a KG can be maintained during the curation process (e.g., the merging process of two KGs). To this end, a novel Transformer-based Graph Encoder (Trans-GE) is proposed as a key component of i-Align for aggregating information from entities’ neighbors (structures). Trans-GE uses Edge-gated Attention that combines the adjacency matrix and the self-attention matrix to learn a gating mechanism to control the information aggregation from the neighboring entities. It also uses historical embeddings, allowing Trans-GE to be trained over mini-batches, or smaller sub-graphs, to address the scalability issue when encoding a large KG. Another component of i-Align is a Transformer encoder for aggregating entities’ attributes. This way, i-Align can generate explanations in the form of a set of the most influential attributes/neighbors based on attention weights. Extensive experiments are conducted to show the power of i-Align. The experiments include several aspects, such as the model’s effectiveness for aligning KGs, the quality of the generated explanations, and its practicality for aligning large KGs. The results show the effectiveness of i-Align in these aspects. Bayu Distiawan Trisedya, Flora D. Salim, Jeffrey Chan, Damiano Spina, Falk Scholer, Mark Sanderson |
Data Min. Knowl. Discov. | 1 |
| 2023 | TransCP: A Transformer Pointer Network for Generic Entity Description Generation With Explicit Content-PlanningabstractWe study neural data-to-text generation to generate a sentence to describe a target entity based on its attributes. Specifically, we address two problems of the encoder-decoder framework for data-to-text generation: i) how to encode a non-linear input (e.g., a set of attributes); and ii) how to order the attributes in the generated description. Existing studies focus on the encoding problem but do not address the ordering problem, i.e., they learn the content-planning implicitly. The other approaches focus on two-stage models but overlook the encoding problem. To address the two problems at once, we propose a model namedTransCPto explicitly learn content-planning and integrate them into a description generation model in an end-to-end fashion. We propose a novel Transformer-based Pointer Network withgated residual attentionandimportance maskingto learn a content-plan. To integrate the content-plan with a description generator, we propose a tracking mechanism to trace the extent to which the content-plan is exposed in the previous decoding time-step. This helps the description generator select the attributes to be mentioned in proper order. Experimental results show that our model consistently outperforms state-of-the-art baselines by up to 2% and 3% in terms of BLEU score on two real-world datasets. Bayu Distiawan Trisedya, Jianzhong Qi 0001, Hai-Tao Zheng 0002, Flora D. Salim, Rui Zhang 0003 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Learning Region Similarities via Graph-Based Deep Metric LearningabstractRegion similarity learning plays an essential role in applications such as business site selection, region recommendation, and urban planning. Earlier studies mainly represent regions as bags of points of interest (POIs) for region similarity comparisons, which cannot fully exploit the spatial features of the regions. Recently, researchers propose to use deep neural networks to exploit spatial features such as POI geo-coordinates and categories, which have produced more accurate and robust region similarity learning results. However, many useful features such as the height and size of a POI, and the distance and relative importance between the POIs, are still overlooked in these methods. To take advantage of such features, we propose to represent regions as graphs, where nodes are POIs with rich features such as height, size, and hexagonal coordinates, while edges are the relationships between POIs formulated by their road network distances. To capture POIs’ importance, we weigh them by their height and size. Since there is limited availability of ground-truth region similarity data, we propose a contrastive learning-based multi-relational graph neural network (C-MPGCN) for region similarity learning based on the graph representations. To generate data for model training, we propose a soft graph edit distance (SGED) based algorithm to generate triples of similar and dissimilar graphs of a given graph (representing a given region) based on the POI weights. Experimental results show that C-MPGCN outperforms the state-of-the-art methods for region similarity learning consistently with an improvement of at least 8.6% and 9.4% in terms of MRR and HR@1, respectively. Jianzhong Qi 0001, Bayu Distiawan Trisedya, Yixin Su 0001, Rui Zhang 0003, Hongguang Ren |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | GCP: Graph Encoder With Content-Planning for Sentence Generation From Knowledge BasesabstractA knowledge base is a large repository of facts usually represented as triples, each consisting of a subject, a predicate, and an object. The triples together form a graph, i.e., a knowledge graph. The triple representation in a knowledge graph offers a simple interface for applications to access the facts. However, this representation is not in a natural language form, which is difficult for humans to understand. We address this problem by proposing a system to translate a set of triples (i.e., a graph) into natural sentences. We take an encoder-decoder based approach. Specifically, we propose a Graph encoder with Content-Planning capability (GCP) to encode an input graph. GCP not only works as an encoder but also serves as a content-planner by using an entity-order aware topological traversal to encode a graph. This way, GCP can capture the relationships between entities in a knowledge graph as well as providing information regarding the proper entity order for the decoder. Hence, the decoder can generate sentences with a proper entity mention ordering. Experimental results show that GCP achieves improvements over state-of-the-art models by up to 3.6%, 4.1%, and 3.8% in three common metrics BLEU, METEOR, and TER, respectively. The code is available at (https://github.com/ruizhang-ai/GCP/). Bayu Distiawan Trisedya, Jianzhong Qi 0001, Wei Wang 0011, Rui Zhang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | A benchmark and comprehensive survey on knowledge graph entity alignment via representation learning
Rui Zhang 0003, Bayu Distiawan Trisedya, Yong Jiang 0001, Jianzhong Qi 0001 |
VLDB J. | 2 |
| 2020 | Sentence Generation for Entity Description with Content-Plan AttentionabstractWe study neural data-to-text generation. Specifically, we consider a target entity that is associated with a set of attributes. We aim to generate a sentence to describe the target entity. Previous studies use encoder-decoder frameworks where the encoder treats the input as a linear sequence and uses LSTM to encode the sequence. However, linearizing a set of attributes may not yield the proper order of the attributes, and hence leads the encoder to produce an improper context to generate a description. To handle disordered input, recent studies propose two-stage neural models that use pointer networks to generate a content-plan (i.e., content-planner) and use the content-plan as input for an encoder-decoder model (i.e., text generator). However, in two-stage models, the content-planner may yield an incomplete content-plan, due to missing one or more salient attributes in the generated content-plan. This will in turn cause the text generator to generate an incomplete description. To address these problems, we propose a novel attention model that exploits content-plan to highlight salient attributes in a proper order. The challenge of integrating a content-plan in the attention model of an encoder-decoder framework is to align the content-plan and the generated description. We handle this problem by devising a coverage mechanism to track the extent to which the content-plan is exposed in the previous decoding time-step, and hence it helps our proposed attention model select the attributes to be mentioned in the description in a proper order. Experimental results show that our model outperforms state-of-the-art baselines by up to 3% and 5% in terms of BLEU score on two real-world datasets, respectively. Bayu Distiawan Trisedya, Jianzhong Qi 0001, Rui Zhang 0003 |
AAAI | 1 |
| 2019 | Entity Alignment between Knowledge Graphs Using Attribute EmbeddingsabstractThe task of entity alignment between knowledge graphs aims to find entities in two knowledge graphs that represent the same real-world entity. Recently, embedding-based models are proposed for this task. Such models are built on top of a knowledge graph embedding model that learns entity embeddings to capture the semantic similarity between entities in the same knowledge graph. We propose to learn embeddings that can capture the similarity between entities in different knowledge graphs. Our proposed model helps align entities from different knowledge graphs, and hence enables the integration of multiple knowledge graphs. Our model exploits large numbers of attribute triples existing in the knowledge graphs and generates attribute character embeddings. The attribute character embedding shifts the entity embeddings from two knowledge graphs into the same space by computing the similarity between entities based on their attributes. We use a transitivity rule to further enrich the number of attributes of an entity to enhance the attribute character embedding. Experiments using real-world knowledge bases show that our proposed model achieves consistent improvements over the baseline models by over 50% in terms of hits@1 on the entity alignment task. Bayu Distiawan Trisedya, Jianzhong Qi 0001, Rui Zhang 0003 |
AAAI | 1 |
| 2019 | Neural Relation Extraction for Knowledge Base EnrichmentabstractWe study relation extraction for knowledge base (KB) enrichment.Specifically, we aim to extract entities and their relationships from sentences in the form of triples and map the elements of the extracted triples to an existing KB in an end-to-end manner.Previous studies focus on the extraction itself and rely on Named Entity Disambiguation (NED) to map triples into the KB space.This way, NED errors may cause extraction errors that affect the overall precision and recall.To address this problem, we propose an end-to-end relation extraction model for KB enrichment based on a neural encoder-decoder model.We collect high-quality training data by distant supervision with co-reference resolution and paraphrase detection.We propose an n-gram based attention model that captures multi-word entity names in a sentence.Our model employs jointly learned word and entity embeddings to support named entity disambiguation.Finally, our model uses a modified beam search and a triple classifier to help generate high-quality triples.Our model outperforms state-of-theart baselines by 15.51% and 8.38% in terms of F1 score on two real-world datasets. Bayu Distiawan Trisedya, Gerhard Weikum, Jianzhong Qi 0001, Rui Zhang 0003 |
ACL (1) | 1 |
| 2018 | GTR-LSTM: A Triple Encoder for Sentence Generation from RDF DataabstractA knowledge base is a large repository of facts that are mainly represented as RDF triples, each of which consists of a subject, a predicate (relationship), and an object.The RDF triple representation offers a simple interface for applications to access the facts.However, this representation is not in a natural language form, which is difficult for humans to understand.We address this problem by proposing a system to translate a set of RDF triples into natural sentences based on an encoder-decoder framework.To preserve as much information from RDF triples as possible, we propose a novel graph-based triple encoder.The proposed encoder encodes not only the elements of the triples but also the relationships both within a triple and between the triples.Experimental results show that the proposed encoder achieves a consistent improvement over the baseline models by up to 17.6%, 6.0%, and 16.4% in three common metrics BLEU, METEOR, and TER, respectively. Bayu Distiawan Trisedya, Jianzhong Qi 0001, Rui Zhang 0003, Wei Wang 0011 |
ACL (1) | 1 |
| 2014 | Automatically Building a Corpus for Sentiment Analysis on Indonesian Tweets
Alfan Farizki Wicaksono, Clara Vania, Bayu Distiawan Trisedya, Mirna Adriani |
PACLIC | 3 |
| 2010 | Developing an Online Indonesian Corpora Repository
Ruli Manurung, Bayu Distiawan Trisedya, Desmond Darma Putra |
PACLIC | 2 |