Chenyi Hu

dblp:23/6432 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
5since 2021 · last 2022
0000-0003-1982-7537ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2022 Towards Explainable Summary of Crowdsourced Reviews Through Text Mining
Aaron Moody, Chenyi Hu, Huixin Zhan, Makenzie Spurling, Victor S. Sheng
IPMU (1)2
2022 Anomaly Detection in Crowdsourced Work with Interval-Valued Labels
Makenzie Spurling, Chenyi Hu, Huixin Zhan, Victor S. Sheng
IPMU (1)2
2021 K2-GNN: Multiple Users' Comments Integration with Probabilistic K-Hop Knowledge Graph Neural Networks
abstract
Integrating multiple comments into a concise statement for any online products or web services requires a non-trivial understanding of the input. Recently, graph neural networks (GNN) has been successfully applied to learn from highly-structured graph representations to mitigate the relationship between entities, such as co-references. However, current inter-sentence relation extraction cannot leverage discrete reasoning chains over multiple comments. To address this issue, in this paper, we propose a probabilistic $K$-hop knowledge graph (KKG) to extend existing knowledge graphs with inferred relations via discrete intra-sentence and inter-sentence reasoning chains. KKG associates each inferred relation with a confidence value through Bayesian inference. We further answer how a knowledge graph with inferred relations can help the multiple comments integration through integrating KKG with GNN ($\text{K}^2$-GNN). Our extensive experimental results show that our $\text{K}^2$-GNN outperforms all baseline graph models on multiple comments integration.
Huixin Zhan, Kun Zhang 0012, Chenyi Hu, Victor S. Sheng
ACML3
2021 HGATs: hierarchical graph attention networks for multiple comments integration
abstract
For decades, research in natural language processing (NLP) has focused on summarization. Sequence-to-sequence models for abstractive summarization have been studied extensively, yet generated summaries commonly suffer from fabricated content, and are often found to be near-extractive. We argue that, to address these issues, summarizers need to acquire the co-references that form multiple types of relations over input sentences, e.g., 1-to-N, N-to-1, and N-to-N relations, since the structured knowledge for text usually appears on these relations. By allowing the decoder to pay different attention to the input sentences for the same entity at different generation states, the structured graph representations generate more informative summaries. In this paper, we propose a hierarchical graph attention networks (HGATs) for abstractive summarization with a topic-sensitive PageRank augmented graph. Specifically, we utilize dual decoders, a sequential sentence decoder, and a graph-structured decoder (which are built hierarchically) to maintain the global context and local characteristics of entities, complementing each other. We further design a greedy heuristic to extract salient users' comments while avoiding redundancy to drive a model to better capture entity interactions. Our experimental results show that our models produce significantly higher ROUGE scores than variants without graph-based attention on both SSECIF and CNN/Daily Mail (CNN/DM) datasets.
Huixin Zhan, Kun Zhang 0012, Chenyi Hu, Victor S. Sheng
ASONAM3
2021 Multi-objective Privacy-preserving Text Representation Learning
abstract
Private information can either take the form of key phrases that are explicitly contained in the text or be implicit. For example, demographic information about the author of a text can be predicted with above-chance accuracy from linguistic cues in the text itself. Letting alone its explicitness, some of the private information correlates with the output labels and therefore can be learned by a neural network. In such a case, there is a tradeoff between the utility of the representation (measured by the accuracy of the classification network) and its privacy. This problem is inherently a multi-objective problem because these two objectives may conflict, necessitating a trade-off. Thus, we explicitly cast this problem as multi-objective optimization (MOO) with the overall objective of finding a Pareto stationary solution. We, therefore, propose a multiple-gradient descent algorithm (MGDA) that enables the efficient application of the Frank-Wolfe algorithm [10] using the line search. Experimental results on sentiment analysis and part-of-speech (POS) tagging show that MGDA produces higher-performing models than most recent proxy objective approaches, and performs as well as single objective baselines.
Huixin Zhan, Kun Zhang 0012, Chenyi Hu, Victor S. Sheng
CIKM3
2020 On Statistics, Probability, and Entropy of Interval-Valued Datasets
Chenyi Hu, Zhihui H. Hu
IPMU (3)1
2020 A Computational Study on the Entropy of Interval-Valued Datasets from the Stock Market
Chenyi Hu, Zhihui H. Hu
IPMU (3)1
2015 An Interval-Radial Algorithm for Hierarchical Clustering Analysis
abstract
Hierarchical clustering analysis (HCA) produces a structure that is more informative than an unstructured set of clusters. However, the advantage comes at the cost of lower efficiency. In analyzing large dataset with HCA, it is important to improve its efficiency. Motivated by the fact that small quantitative differences may not necessarily reflect changes of qualitative property, we report an interval-radial algorithm for HCA. By grouping data points within a neighborhood, the interval-radial algorithm is O(N^2) for both agglomerative and divisive approaches under an easy to satisfy weak condition. The algorithm can adaptively adjust radius during its execution. Furthermore, the algorithm provides flexibility to users for them to select initial radius and step size such that to produce customized output automatically. We report the algorithm, its analysis, and results of computational experiments on several benchmark datasets. Examples and illustrative dendrograms are included.
Christopher Rhodes, James Lemon, Chenyi Hu
ICMLA3
2012 Interval-Valued Centroids in K-Means Algorithms
abstract
The K-Means algorithms are fundamental in machine learning and data mining. In this study, we investigate interval-valued rather than commonly used point-valued centroids in the K-Means algorithm. Using a proposed interval peak method to select initial interval centroids, we have obtained overall quality improvement of clusters on a set of test problems in the Fundamental Clustering Problem Suite (FCPS).
Benjamine Nordin, Chenyi Hu, Bernard Chen 0001, Victor S. Sheng
ICMLA (1)2
2011 Efficient Calculation of Structural Similarity Threshold for the SCAN Network Clustering Algorithm
abstract
Community detection algorithms play an important role in discovering knowledge in networks. The Structural Clustering Algorithm for Network (SCAN) is a community detection algorithm which is capable of detecting hubs and outliers, in addition to cluster members. The term hub means node with the ability of collecting and delivering information among clusters while outlier is considered as a noise in the data. Currently, researchers use exhaustive search to determine the structural similarity threshold value (ε) in the SCAN. This paper reports a new approach of using interval ε value to narrow the searching domain for proper ε value for the SCAN. The approach first adopts computational results produced by the Fast Modularity and the Walktrap algorithms to bind the number of clusters of a network and then determine the interval for ε value. For each of our test datasets, the interval prediction reliably finds the true number of clusters. More importantly, the proposed prediction method helps users to eliminate an average of 67.7% of inappropriate ε values used to generate clusters.
Vincent Yip, Sinan Kockara, Chenyi Hu
BIBM3
2008 Studying interval valued matrix games with fuzzy logic
W. Dwayne Collins, Chenyi Hu
Soft Comput.2
2004 Generating and Applying Rules for Interval Valued Fuzzy Observations
André de Korvin, Chenyi Hu, Ping Chen 0001
IDEAL2
2003 Icon-based Visualization of Large High-Dimensional Datasets
abstract
High dimensional data visualization is critical to data analysts since it gives a direct view of original data. We present a method to visualize large amount of high dimensional data. We divide dimensions of data into several groups. Then, we use one icon to represent each group, and associate visual properties of each icon with dimensions in each group. A high dimensional data record will be represented by multiple different types of icons located in the same position. Furthermore, we use summary icons to display local details of viewer's interests and the whole data set at meantime. We show its effectiveness and efficiency through a case study on a real large data set.
Ping Chen 0001, Chenyi Hu, Wei Ding 0003, Heloise Lynn, Yves Simon
ICDM2
2002 Association analysis with interval valued fuzzy sets and body of evidence
abstract
Association analysis is proven to be one of the most useful techniques in data mining to analyze large datasets. We apply association analysis to fuzzy datasets and show how it works in decision making processes. We define three fundamental concepts: belief, compatibility, and plausibility for association analysis in terms of fuzzy logic. We further extend our study on association analysis with interval valued fuzzy sets and body of evidence. Applications with sample data are discussed.
Ping Chen 0001, André de Korvin, Chenyi Hu
FUZZ-IEEE3
1994 Parallel All-Row Preconditioned Interval Linear Solver for Nonlinear Equations on Multiprocessors
Qi Gan, Qing Yang 0001, Chenyi Hu
Parallel Comput.3
1994 Algorithm 737; INTLIB: a portable Fortran 77 interval standard-function library
abstract
INTLIB is meant to be a readily available, portable, exhaustively documented interval arithmetic library, written in standard Fortran 77. Its underlying philosophy is to provide a standard for interval operations to aid in efficiently transporting programs involving interval arithmetic. The model is the BLAS package, for basic linear algebra operations. The library is composed of elementary interval arithmetic routines, standard function routines for interval data and values, and utility routines. The library can be used with INTBIS (Algorithm 681), and a Fortran 90 module to use the library to define an interval data type is available from the first author.
R. Baker Kearfott, Milind Dawande, Kaisheng Du, Chenyi Hu
ACM Trans. Math. Softw.4