EDBT 2026 Demo / reviewers in the wild / expert
Xiaohua Hu 0001
dblp:h/XiaohuaHu · also Xiaohua Tony Hu
· DBLP profile ↗
96ranked-venue papers in the field
13as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 27 (5 first)Information Retrieval & Web Search · 24 (3 first)Big Data, Cloud & Distributed Data Systems · 24Database Systems & Data Management · 11 (2 first)Other / Interdisciplinary · 7 (3 first)Business Process & Enterprise Data · 2Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Secure Retrieval-Augmented Generation Against Poisoning Attacks
Zirui Cheng, Jikai Sun, Anjun Gao, Yueyang Quan, Zhuqing Liu, Xiaohua Hu 0001, Minghong Fang |
IEEE Big Data | 6 |
| 2025 | HyperGAT: Hypergraph Attention Network for Stock Movement Prediction
Jianliang Gao, Shilin Xie, Shujin Wang, Xiaohua Hu 0001 |
IEEE Big Data | 4 |
| 2023 | Investigating Data Reusability in Density Functional Theory StudiesabstractOver the last decade, there has been a significant increase in supporting reproducible computational research (RCR) [1]. The global adoption of the FAIR principles [2] stands as a key indicator of this trend. Specifically, federal and global research funding agencies have increasingly mandated scientific data and related products, such as code and algorithms, be made Findable, Accessible, Interoperable, and Reusable (FAIR) [2]. Rob Fleur, Addy Ireland, Xintong Zhao, Scott McClellan, Eric Paltoo, Channyung Lee, Xiaohua Hu 0001, Elif Ertekin, Jane Greenberg |
IEEE Big Data | 9 |
| 2023 | When LLM Meets Material Science: An Investigation on MOF Synthesis LabelingabstractRecent developments in Large Language Models (LLMs) have advanced the natural language processing (NLP) studies to a new era [1], [2], [4]–[6]. In generic domains, LLMs have become a key component in wide variety of state-of-the-art NLP tasks. In addition, prompt learning enables LLMs-based models to reach robust performance with much smaller training data. Xintong Zhao, Kyle Langlois, Jacob Furst 0002, Scott McClellan, Rob Fleur, Xiaohua Hu 0001, Fernando J. Uribe-Romo, Diego A. Gómez-Gualdrón, Jane Greenberg |
IEEE Big Data | 7 |
| 2023 | The First Workshop on Personalized Generative AI @ CIKM 2023: Personalization Meets Large Language ModelsabstractThe First Workshop on Personalized Generative AI1 aims to be a cornerstone event fostering innovation and collaboration in the dynamic field of personalized AI. Leveraging the potent capabilities of Large Language Models (LLMs) to enhance user experiences with tailored responses and recommendations, the workshop is designed to address a range of pressing challenges including knowledge gap bridging, hallucination mitigation, and efficiency optimization in handling extensive user profiles. As a nexus for academics and industry professionals, the event promises rich discussions on a plethora of topics such as the development and fine-tuning of foundational models, strategies for multi-modal personalization, and the imperative ethical and privacy considerations in LLM deployment. Through a curated series of keynote speeches, insightful panel discussions, and hands-on sessions, the workshop aspires to be a catalyst in the development of more precise, contextually relevant, and user-centric AI systems. It aims to foster a landscape where generative AI systems are not only responsive but also anticipatory of individual user needs, marking a significant stride in personalized experiences. Zheng Chen 0010, Ziyan Jiang, Fan Yang 0155, Zhankui He, Yupeng Hou, Eunah Cho, Julian J. McAuley, Aram Galstyan, Xiaohua Hu 0001, Jie Yang 0028 |
CIKM | 9 |
| 2022 | Exploring Pre-Trained Language Models to Build Knowledge Graph for Metal-Organic Frameworks (MOFs)abstractBuilding a knowledge graph is a time-consuming and costly process which often applies complex natural language processing (NLP) methods for extracting knowledge graph triples from text corpora. Pre-trained large Language Models (PLM) have emerged as a crucial type of approach that provides readily available knowledge for a range of AI applications. However, it is unclear whether it is feasible to construct domain-specific knowledge graphs from PLMs. Motivated by the capacity of knowledge graphs to accelerate data-driven materials discovery, we explored a set of state-of-the-art pre-trained general-purpose and domain-specific language models to extract knowledge triples for metal-organic frameworks (MOFs). We created a knowledge graph benchmark with 7 relations for 1248 published MOF synonyms. Our experimental results showed that domain-specific PLMs consistently outperformed the general-purpose PLMs for predicting MOF related triples. The overall benchmarking results, however, show that using the present PLMs to create domain-specific knowledge graphs is still far from being practical, motivating the need to develop more capable and knowledgeable pre-trained language models for particular applications in materials science. Jane Greenberg, Xiaohua Hu 0001, Alexander Kalinowski, Xintong Zhao, Scott McClellan, Fernando J. Uribe-Romo, Kyle Langlois, Jacob Furst 0002, Diego A. Gómez-Gualdrón, Fernando Fajardo-Rojas, Katherine Ardila, Semion Saikin, Corey A. Harper, Ron Daniel Jr. 0001 |
IEEE Big Data | 3 |
| 2021 | Towards Explainable Visual Emotion UnderstandingabstractEmotion understanding from an image is an important computer vision research topic, but the reasoning behind the justification has been largely unexplored. We propose to utilize language, and more specifically, explanation of the judgement to increase the explainability for this task. We collect a dataset with image, emotion tag and explanation in natural language, and conduct analysis on the dataset to gain insights on the ambiguity of the human emotion perceptual. We examine baseline methods to predict emotion from image, explanation and general image description, and unifying both modalities. Our experiments shed lights on effects from different modalities and we also identify opportunities for future visual emotion categorization research based on the analysis. We release our dataset to advance future research. Wanying Ding, Xiaohua Hu 0001 |
IEEE BigData | 4 |
| 2021 | Fine-Tuning BERT Model for Materials Named Entity RecognitionabstractScientific literature presents a wellspring of cutting-edge knowledge for materials science, including valuable data (e.g., numerical data from experiment results, material properties and structure). These data are critical for accelerating materials discovery by data-driven machine learning (ML) methods. The challenge is, it is impossible for humans to manually extract and retain this knowledge due to the extensive and growing volume of publications.To this end, we explore a fine-tuned BERT model for extracting knowledge. Our preliminary results show that our fine-tuned Bert model reaches an f-score of 85% for the materials named entity recognition task. The paper covers background, related work, methodology including tuning parameters, and our overall performance evaluation. Our discussion offers insights into our results, and points to directions for next steps. Xintong Zhao, Jane Greenberg, Xiaohua Hu 0001 |
IEEE BigData | 4 |
| 2021 | Knowledge Graph-Empowered Materials DiscoveryabstractIn this position paper, we describe research on knowledge graph-empowered materials science prediction and discovery. The research consists of several key components including ontology mapping, materials data annotation, and information extraction from unstructured scholarly articles. We argue that although big data generated by simulations and experiments have motivated and accelerated the data-driven science, the distribution and heterogeneity of materials science-related big data hinders major advancements in the field. Knowledge graphs, as semantic hubs, integrate disparate data and provide a feasible solution to addressing this challenge. We design a knowledge-graph based approach for data discovery, extraction, and integration in materials science. Xintong Zhao, Jane Greenberg, Scott McClellan, Yong-Jie Hu, Steven Lopez, Semion Saikin, Xiaohua Hu 0001 |
IEEE BigData | 7 |
| 2021 | An effective framework for semistructured document classification via hierarchical attention modelabstractRecent years have witnessed the rapidly growing of the amount of semistructured documents in real-world applications. Due to the huge size of the real-world data, how to manage semistructured documents effectively is a big challenge for researchers. As a fundamental task in natural language processing field, document classification is a feasible way to handle the large-scale semistructured documents. However, existing methods fail to explicitly take advantage of the hierarchical semantics in semistructured documents. It's known that the contained semantics is beneficial for understanding the semistructured documents. Considering the hierarchical structure of a given semistructured document, we propose a semistructured document classification framework which explicitly utilizes the semantic hierarchical attention mechanism. More specifically, the hierarchical attention mechanism and graph neural network are employed to model semistructured documents, by which the multilevel semantic relationships and grammatical information are considered. Moreover, we propose an adaptive class cost learning method to treat the issue of data imbalance. Comprehensive experiments are conducted on two real-world data sets, and the results demonstrate that our framework performs better than selected baselines for semistructured document classification. Weizhong Zhao, Dandan Fang, Jinyong Zhang, Xiaowei Xu 0001, Xingpeng Jiang, Xiaohua Hu 0001, Tingting He 0003 |
Int. J. Intell. Syst. | 7 |
| 2020 | Hierarchy construction and classification of heterogeneous information networks based on RSDAEf
Jinli Zhang, Zongli Jiang, Yongping Du, Tong Li 0001, Xiaohua Hu 0001 |
Data Knowl. Eng. | 6 |
| 2019 | Hierarchical-Document-Structure-Aware Attention with Adaptive Cost Sensitive Learning for Biomedical Document ClassificationabstractBiomedical document classification is a fundamental task in biomedical field. Existing methods do not make full use of the hierarchically semantic structures in biomedical documents which can be utilized to improve the performance of biomedical document classification. In this paper, according to the hierarchical structures in given biomedical documents, we propose two models for biomedical document classification, which are based on the semantically hierarchical attention mechanism. Specifically, we utilize a hierarchical attention mechanism to model biomedical documents, taking into account simultaneously multiple-level semantic relationships in documents. In addition, an adaptive cost sensitive learning method is proposed to address the data imbalance issue. Extensive experiments on two real-world datasets demonstrate the effectiveness of the proposed methods. Dandan Fang, Jinyong Zhang, Weizhong Zhao, Xiaowei Xu 0001, Xingpeng Jiang, Xiaohua Hu 0001, Tingting He 0003 |
IEEE BigData | 6 |
| 2019 | End-to-End Joint Opinion Role Labeling with BERTabstractOpinion mining has raised growing interest both in industry and academia in the past decade. Opinion role labeling (ORL) is a task to extract opinion holder and target from natural language to answer the question “who express what”. Recent years, neural network based methods with additional lexical and syntactic features have achieved state-of-the-art performances in similar tasks. Moreover, Bidirectional Encoder Representations from Transformers (BERT) has shown impressive performances among a variety of natural language processing (NLP) tasks. To investigate BERT based end-to-end model in ORL, we propose models using BERT, Bidirectional Long short-term Memory (BiLSTM) and Conditional Random Field (CRF) to jointly extract opinion roles (e.g., opinion holder and target). Experimental results show that our models achieve remarkable scores without using extra syntactic and/or semantic features. To our best knowledge, we are among the pioneers to successfully integrate BERT in this manner. Our work contributes to the improvement of state-of-the-art aspect-level opinion mining methods and providing strong baselines for future work. Jinli Zhang, Xiaohua Hu 0001 |
IEEE BigData | 3 |
| 2018 | Correlated Anomaly Detection from Large Streaming DataabstractCorrelated anomaly detection (CAD) from streaming data is a type of group anomaly detection and an essential task in useful real-time data mining applications like botnet detection, financial event detection, industrial process monitor, etc. The primary approach for this type of detection in previous researches is based on principal score (PS) of divided batches or sliding windows by computing top eigenvalues of the correlation matrix, e.g. the Lanczos algorithm. However, this paper brings up the phenomenon of principal score degeneration for large data set, and then mathematically and practically prove current PS-based methods are likely to fail for CAD on large-scale streaming data even if the number of correlated anomalies grows with the data size at a reasonable rate; in reality, anomalies tend to be the minority of the data, and this issue can be more serious. We propose a framework with two novel randomized algorithms rPS and gPS for better detection of correlated anomalies from large streaming data of various correlation strength. The experiment shows high and balanced recall and estimated accuracy of our framework for anomaly detection from a large server log data set and a U.S. stock daily price data set in comparison to direct principal score evaluation and some other recent group anomaly detection algorithms. Moreover, our techniques significantly improve the computation efficiency and scalability for principal score calculation. Zheng Chen 0010, Xinli Yu 0002, Yuan Ling, Xiaohua Hu 0001, Erjia Yan |
IEEE BigData | 6 |
| 2018 | Comparative Study of CNN and LSTM based Attention Neural Networks for Aspect-Level Opinion MiningabstractAspect-level opinion mining aims to find and aggregate opinions on opinion targets. Previous work has demonstrated that precise modeling of opinion targets within the surrounding context can improve performances. However, how to effectively and efficiently learn hidden word semantics and better represent targets and the context still needs to be further studied. In this paper, we propose and compare two interactive attention neural networks for aspect-level opinion mining, one employs two bi-directional Long-Short-Term-Memory (BLSTM) and the other employs two Convolutional Neural Networks (CNN). Both frameworks learn opinion targets and the context respectively, followed by an attention mechanism that integrates hidden states learned from both the targets and context. We compare our model with state-of-the-art baselines on two SemEval 2014 datasets1. Experiment results show that our models obtain competitive performances against the baselines on both datasets. Our work contributes to the improvement of state-of-the-art aspect-level opinion mining methods and offers a new approach to support human decision-making process based on opinion mining results. The quantitative and qualitative comparisons in our work aim to give basic guidance for neural network selection in similar tasks. Zheng Chen 0010, Jianliang Gao, Xiaohua Hu 0001 |
IEEE BigData | 4 |
| 2017 | Fast botnet detection from streaming logs using online lanczos methodabstractBotnet, a group of coordinated bots, is becoming the main platform of malicious Internet activities like DDOS, click fraud, web scraping, spam/rumor distribution, etc. This paper focuses on design and experiment of a new approach for botnet detection from streaming web server logs, motivated by its wide applicability, real-time protection capability, ease of use and better security of sensitive data. Our algorithm is inspired by a Principal Component Analysis (PCA) to capture correlation in data, and we are first to recognize and adapt Lanczos method to improve the time complexity of PCA-based botnet detection from cubic to sub-cubic, which enables us to more accurately and sensitively detect botnets with sliding time windows rather than fixed time windows. We contribute a generalized online correlation matrix update formula, and a new termination condition for Lanczos iteration for our purpose based on error bound and non-decreasing eigenvalues of symmetric matrices. On our dataset of an ecommerce website logs, experiments show the time cost of Lanczos method with different time windows are consistently only 20% to 25% of PCA. Zheng Chen 0010, Xinli Yu 0002, Cui Lin, Jianliang Gao, Xiaohua Hu 0001, Wei-Shih Yang, Erjia Yan |
IEEE BigData | 8 |
| 2017 | Visualization of non-metric relationships by adaptive learning multiple maps t-SNE regularizationabstractKnown as phenotypic overlapping, some disease-rel ated symptoms share a common pathologi cal and physiological mechanism. Researchers attempt to visualize the phenotypic relationships between different human diseases from the perspective of machine learning, but traditional visualization methods may be subject to fundamental limitations of metric spaces. Multiple maps t-SNE regularization method, a probabilistic method for visualizing data points in multiple low-dimensional spaces has been proposed to address the limitation. However, the convergence speed is low when apply on the scale dataset. We use the RMSProp with Nesterov momentum method to learn the objective loss function. This method normalize the gradients by applying an exponential moving average of gradient magnitude for each iteration parameter and use Nesterov momentum to counterweigh too high velocities by “peeking ahead” actual objective values in the candidate search direction. This method convergent faster than the original method of convergence speed. Experiments results on several dataset shows that the proposed method outperforms the several version of mm-tSNE with or without regularization, as measured by the neighborhood preservation ratio and error rate. This suggests the modified mm-tSNE regularization can be applied directly in other domain including social, biological and microbiomic datasets. Xianjun Shen, Xianchao Zhu, Xingpeng Jiang, Tingting He 0003, Xiaohua Hu 0001 |
IEEE BigData | 6 |
| 2017 | Large-scale joint topic, sentiment & user preference analysis for online reviewsabstractThis paper presents a non-trivial reconstruction of a previous joint topic-sentiment-preference review model TSPRA with stick-breaking representation under the framework of variational inference (VI) and stochastic variational inference (SVI). TSPRA is a Gibbs Sampling based model that solves topics, word sentiments and user preferences altogether and has been shown to achieve good performance, but for large dataset it can only learn from a relatively small sample. We develop the variational models vTSPRA and svTSPRA to improve the time use, and our new approach is capable of processing millions of reviews. We rebuild the generative process, improve the rating regression, solve and present the coordinate-ascent updates of variational parameters, and show the time complexity of each iteration is theoretically linear to the corpus size, and the experiments on Amazon datasets show it converges faster than TSPRA and attains better results given the same amount of time. In addition, we tune svTSPRA into an online algorithm ovTSPRA that can monitor oscillations of sentiment and preference overtime. Some interesting fluctuations are captured and possible explanations are provided. The results give strong visual evidence that user preference is better treated as an independent factor from sentiment. Xinli Yu 0002, Zheng Chen 0010, Wei-Shih Yang, Xiaohua Hu 0001, Erjia Yan, Guangrong Li |
IEEE BigData | 4 |
| 2017 | Heterogeneous knowledge transfer via domain regularization for improving cross-domain collaborative filteringabstractCross-Domain Collaborative Filtering(CDCF) methods transfer knowledge from auxiliary domains (e.g., books) to improve recommendation in a target domain (e.g. movies). Most CDCF methods exploit homogeneous user feedback, e.g. numeric ratings, from auxiliary domains as the knowledge source. However, in a typical recommender system, the usage data is usually heterogeneous and therefore is potential to better improve recommendation in other domains. In this paper, we propose a novel and generic CDCF solution called Heterogeneous Knowledge Transfer via Domain Regularization (HKT-DR). Our solution is able to mine high quality knowledge from heterogeneous knowledge sources, i.e. both explicit and implicit feedbacks from multiple auxiliary domains, by building a fused user similarity network, and to incorporate the knowledge by imposing domain regularization to constrain matrix factorization objective function. Extensive experiments on real world datasets show that the proposed HKT-DR model outperforms the state-of-the-art CDCF solutions. Yizhou Zang, Xiaohua Hu 0001 |
IEEE BigData | 2 |
| 2017 | A map-based visual analysis method for patterns discovery of mobile learning in education with big dataabstractBig data in education relate closely to a wealth of activities from the teachers, students and parents, as well as a substantial resources of knowledge that can be represented in hierarchical structures. The activities have characteristic of geolocation which can be projected onto a map, while the resources of knowledge can also be converted into the map. A map-based management and visual analysis method will largely benefit the users and the researchers from taking advantages of the big data in education. In this paper, we propose a novel map based method to manage and analyze the big data of mobile learning in education. With this method, the activities of users scattered among the space are reorganized on a geographic map with location changes in time series, and the resources are geo-tagged with the information from the developers or adopters, which are converted to a map style according to their hierarchical structures even when the users' information are unavailable. We first present the basic framework to organize the data by a map-based technology, and then a platform is proposed to perform the visual analysis. The method is adapt to the construction of massive online learning system and the mobile learning system of Central China Normal University to serve a national wide big data cloud learning program. Dongbo Zhou, Sannyuya Liu, Xiaohua Hu 0001 |
IEEE BigData | 5 |
| 2017 | Community-Based Network Alignment for Large Attributed NetworkabstractNetwork alignment is becoming an active topic in network data analysis. Despite extensive research, we realize that efficient use of topological and attribute information for large attributed network alignment has not been sufficiently addressed in previous studies. In this paper, based on Stochastic Block Model (SBM) and Dirichlet-multinomial, we propose "divide-and-conquer" models CAlign that jointly consider network alignment, community discovery and community alignment in one framework for large networks with node attributes, in an effort to reduce both the computation time and memory usage while achieving better or competitive performance. It is provable that the algorithms derived from our model have sub-quadratic time complexity and linear space complexity on a network with small densification power, which is true for most real-world networks. Experiments show CAlign is superior to two recent state-of-art models in terms of accuracy, time and memory on large networks, and CAlign is capable of handling millions of nodes on a modern desktop machine. Zheng Chen 0010, Xinli Yu 0002, Jianliang Gao, Xiaohua Hu 0001, Wei-Shih Yang |
CIKM | 5 |
| 2017 | LKT-FM: A Novel Rating Pattern Transfer Model for Improving Non-overlapping Cross-Domain Collaborative Filtering
Yizhou Zang, Xiaohua Hu 0001 |
ECML/PKDD (2) | 2 |
| 2017 | Counter Deanonymization Query: H-index Based k-Anonymization Privacy Protection for Social NetworksabstractIn this paper, we propose a novel k-anonymization scheme to counter deanonymization queries on social networks. With this scheme, all entities are protected by k-anonymization, which means the attackers cannot re-identify a target with confidence higher than 1/k. The proposed scheme minimizes the modification on original networks, and accordingly maximizes the utility preservation of published data while achieving k-anonymization privacy protection. Extensive experiments on real data sets demonstrate the effectiveness of the proposed scheme, where the efficacy of the k-anonymized networks is verified with the distributions of pagerank, betweenness, and their Kolmogorov-Smirnov (K-S) test. Jianliang Gao, Zheng Chen 0010, Weimao Ke, Wanying Ding, Xiaohua Hu 0001 |
SIGIR | 6 |
| 2017 | Which used product is more sellable? A time-aware approach
Mengwen Liu, Wanying Ding, Dae Hoon Park, Yi Fang 0008, Rui Yan 0001, Xiaohua Hu 0001 |
Inf. Retr. J. | 6 |
| 2017 | Product review summarization through question retrieval and diversification
Mengwen Liu, Yi Fang 0008, Alexander G. Choulos, Dae Hoon Park, Xiaohua Hu 0001 |
Inf. Retr. J. | 5 |
| 2016 | Semi-supervised Dirichlet-Hawkes process with applications of topic detection and tracking in TwitterabstractUnderstanding ongoing topics and their evolutions in social media is of great importance. Although topic analysis is not a novel research question, social media environment has presented new challenges. First, with insufficient co-occurrence information, short text have undermined many word co-occurrence oriented topic models' applicability. Second, real time message streams make traditional discretized topic tracking methods hard to function. Third, topics' evolution mechanisms are of great importance in social media context, but many studies have ignored them. Forth, topics have more complicated correlation among each other. Considering the existing problems, this paper has proposed a Semi-Supervised Dirichlet-Hawkes Process (SDHP) to deal with topic detection and tracking from social media. The main contributions of this paper are reflected in: (1) SDHP can handle short text problem efficiently; (2) SDHP can track topics from continuous message stream; (3) SDHP can reveal topics' underlying evolution patterns; and (4) SDHP can capture topics' correlations We have evaluated SDHP's ability in both topic detection and tracking in 8 real datasets from Twitter, and the algorithm's performances are very promising. Wanying Ding, Chaomei Chen, Xiaohua Hu 0001 |
IEEE BigData | 4 |
| 2016 | Parallel top-k subgraph query in massive graphs: Computing from the perspective of single vertexabstractIn the real world, many problems on massive graphs can be mapped to an underlying critical problem of discovering top-k subgraphs. For massive graphs, subgraph queries may have enormous number of matches, and so it is inefficient to compute all matches when only top-k matches are desired. Meanwhile, parallel algorithm is urgent for the scalability of massive graph computing. In this paper, we address the challenges of top-k subgraph query in massive graph. Firstly, we present a new graph matching notion: “approximate graph simulation”. With approximate graph simulation, top-k subgraph query can be customized by appointing a weighted query graph, which provides good flexibility for different application scenarios. Secondly, we propose a parallel top-k subgraph query algorithm at the level of vertex. With such algorithm, each vertex in massive graph obtains its matching state separately without requiring global graph information. In the algorithm, we also design a filter mechanism to speed up the the computation and a aggregation mechanism to obtain top-k vertices for query focus. Using real-life datasets, we experimentally verify that our approach of parallel top-k subgraph query are efficient. Jianliang Gao, Weimao Ke, Jianxin Wang 0001, Xiaohua Hu 0001 |
IEEE BigData | 6 |
| 2016 | Pairwise topic model and its application to topic transition and evolutionabstractNowadays the explosion of Web information has led to the boom of massive web documents such as news webpages, online literature, etc. The latent topics behind the documents spread by self-evolution and mutual transition. Understanding how topics in documents evolve and transit is an important and challenging problem. Topic model is a set of powerful toolkits to model documents generation to find their underlying topics, usually at the unigram level, making it difficult to model the relationship between terms and their underlying topics. In this paper, we propose a pairwise topic modeling method to incorporate a pairwise relationship into topic modeling methods. We manage to discover latent topics as well as topic transitions at the same time in a natural way. We show that the pairwise topic model can facilitate discovering of individual topics as well as topic evolution. The results indicate our proposed method leads to a significant performance improvement over the traditional topic modeling methods, such as Latent Dirichlet Allocation (LDA) in terms of language perplexity. Besides, we conduct a series of empirical studies to show the topic words and topic transitions discovered. From the case studies, we show that with the help of PTM methods, people are able to explicitly understand how topics evolve and transit between each other. Xiaoli Song, Yan Rui, Xiaohua Hu 0001 |
IEEE BigData | 3 |
| 2016 | Semantic pattern mining for text miningabstractPattern mining is a fundamental topic in data mining area. Many pattern mining techniques, such as closed and maximal pattern mining have been proposed for different applications. However, when calculating the frequency of a pattern, the existing techniques treat each word equally. For example, although the word `pie' in `I love eating pie.' is quite different from `pie' in `american pie', `pie' in `american pie' will still be added up to the counts of `pie' when calculating its frequency. Therefore, this paper aims to overcome the drawback to find the valid patterns tailored to text mining. We will approach pattern mining from a different perspective and introduce a novel problem of frequent semantic pattern mining. We then propose an algorithm to solve this problem via suffix array sorting. The algorithm can be implemented to run in linear time. Compared with traditional pattern representations, our results show the semantic patterns extracted are more than 13% compact. Also, classifier built on these features is no less or more powerful. Xiaoli Song, Xiaohua Hu 0001 |
IEEE BigData | 3 |
| 2016 | Retrieving Non-Redundant Questions to Summarize a Product ReviewabstractProduct reviews have become an important resource for customers before they make purchase decisions. However, the abundance of reviews makes it difficult for customers to digest them and make informed choices. In our study, we aim to help customers who want to quickly capture the main idea of a lengthy product review before they read the details. In contrast with existing work on review analysis and document summarization, we aim to retrieve a set of real-world user questions to summarize a review. In this way, users would know what questions a given review can address and they may further read the review only if they have similar questions about the product. Specifically, we design a two-stage approach which consists of question retrieval and question diversification. We first propose probabilistic retrieval models to locate candidate questions that are relevant to a review. We then design a set function to re-rank the questions with the goal of rewarding diversity in the final question set. The set function satisfies submodularity and monotonicity, which results in an efficient greedy algorithm of submodular optimization. Evaluation on product reviews from two categories shows that the proposed approach is effective for discovering meaningful questions that are representative for individual reviews. Mengwen Liu, Yi Fang 0008, Dae Hoon Park, Xiaohua Hu 0001, Zhengtao Yu 0001 |
SIGIR | 4 |
| 2016 | Socialized Language Model Smoothing via Bi-directional Influence Propagation on Social NetworksabstractIn recent years, online social networks are among the most popular websites with high PV (Page View) all over the world, as they have renewed the way for information discovery and distribution. Millions of users have registered on these websites and hence generate formidable amount of user-generated contents every day. The social networks become "giants", likely eligible to carry on any research tasks. However, we have pointed out that these giants still suffer from their "Achilles Heel", i.e., extreme sparsity. Compared with the extremely large data over the whole collection, individual posting documents such as microblogs seem to be too sparse to make a difference under various research scenarios, while actually these postings are different. In this paper we propose to tackle the Achilles Heel of social networks by smoothing the language model via influence propagation. To further our previously proposed work to tackle the sparsity issue, we extend the socialized language model smoothing with bi-directional influence learned from propagation. Intuitively, it is insufficient not to distinguish the influence propagated between information source and target without directions. Hence, we formulate a bi-directional socialized factor graph model, which utilizes both the textual correlations between document pairs and the socialized augmentation networks behind the documents, such as user relationships and social interactions. These factors are modeled as attributes and dependencies among documents and their corresponding users, and then are distinguished on the direction level. We propose an effective learning algorithm to learn the proposed factor graph model with directions. Finally we propagate term counts to smooth documents based on the estimated influence. We run experiments on two instinctive datasets of Twitter and Weibo. The results validate the effectiveness of the proposed model. By incorporating direction information into the socialized language model smoothing, our approach obtains improvement over several alternative methods on both intrinsic and extrinsic evaluations measured in terms of perplexity, nDCG and MAP measurements. Rui Yan 0001, Cheng-Te Li, Hsun-Ping Hsieh, Po Hu 0001, Xiaohua Hu 0001, Tingting He 0003 |
WWW | 5 |
| 2016 | Cross-lingual sentiment classification with stacked autoencoders
Guangyou Zhou, Tingting He 0003, Xiaohua Hu 0001 |
Knowl. Inf. Syst. | 4 |
| 2015 | Video Popularity Prediction by Sentiment Propagation via Implicit NetworkabstractVideo popularity prediction plays a foundational role in many aspects of life, such as recommendation systems and investment consulting. Because of its technological and economic importance, this problem has been extensively studied for years. However, four constraints have limited most related works' usability. First, most feature oriented models are inadequate in the social media environment, because many videos are published with no specific content features, such as a strong cast or a famous script. Second, many studies assume that there is a linear correlation existing between view counts from early and later days, but this is not the case in every scenario. Third, numerous works just take view counts into consideration, but discount associated sentiments. Nevertheless, it is the public opinions that directly drive a video's final success/failure. Also, many related approaches rely on a network topology, but such topologies are unavailable in many situations. Here, we propose a Dual Sentimental Hawkes Process (DSHP) to cope with all the problems above. DSHP's innovations are reflected in three ways: (1) it breaks the "Linear Correlation" assumption, and implements Hawkes Process; (2) it reveals deeper factors that affect a video's popularity; and (3) it is topology free. We evaluate DSHP on four types of videos: Movies, TV Episodes, Music Videos, and Online News, and compare its performance against 6 widely used models, including Translation Model, Multiple Linear Regression, KNN Regression, ARMA, Reinforced Poisson Process, and Univariate Hawkes Process. Our model outperforms all of the others, which indicates a promising application prospect. Wanying Ding, Lifan Guo, Xiaohua Hu 0001, Rui Yan 0001, Tingting He 0003 |
CIKM | 4 |
| 2015 | Learning Focused Hierarchical Topic Models with Semi-Supervision in Microblogs
Anton Slutsky, Xiaohua Hu 0001 |
PAKDD (2) | 2 |
| 2015 | Tackling the Achilles Heel of Social Networks: Influence Propagation based Language Model SmoothingabstractOnline social networks nowadays enjoy their worldwide prosperity, as they have revolutionized the way for people to discover, to share, and to distribute information. With millions of registered users and the proliferation of user-generated contents, the social networks become "giants", likely eligible to carry on any research tasks. However, the giants do have their Achilles Heel: extreme data sparsity. Compared with the massive data over the whole collection, individual posting documents, (e.g., a microblog less than 140 characters), seem to be too sparse to make a difference under various research scenarios, while actually they are different. In this paper we propose to tackle the Achilles Heel of social networks by smoothing the language model via influence propagation. We formulate a socialized factor graph model, which utilizes both the textual correlations between document pairs and the socialized augmentation networks behind the documents, such as user relationships and social interactions. These factors are modeled as attributes and dependencies among documents and their corresponding users. An efficient algorithm is designed to learn the proposed factor graph model. Finally we propagate term counts to smooth documents based on the estimated influence. Experimental results on Twitter and Weibo datasets validate the effectiveness of the proposed model. By leveraging the smoothed language model with social factors, our approach obtains significant improvement over several alternative methods on both intrinsic and extrinsic evaluations measured in terms of perplexity, nDCG and MAP results. Rui Yan 0001, Ian En-Hsu Yen, Cheng-Te Li, Xiaohua Hu 0001 |
WWW | 5 |
| 2014 | Pairwise Topic Model via relation extractionabstractTopic modeling is a powerful tool to model documents to find their underlying topics. However, the unstructured nature of the raw text makes it hard to model the semantic relationship between the text units, which may be the words, phrases or sentences, and thus even harder to model their corresponding underlying topics. In our work, we try to examine the pairwise relationship of the underlying topics through relation extraction. We first extract the entity pairs within one relation tuple out of the raw text. Then, we model the relationship between the entity pairs by adding the dependencies between entities and their corresponding topics. We propose six different versions of Pairwise Topic Model (PTM) to simultaneously discover the latent topics and their pairwise relationship. The experiment on four data sets (AP news articles, DUC 2004 task2, Clinical Notes and Neuroscience Papers) shows the PTM models are better-structured language model than the traditional topic model Latent Dirichlet Allocation (LDA). Also, empirical results show that the proposed Pairwise Topic Models (PTMs) can explicitly explain how two topics are related. Xiaoli Song, Yuan Ling, Mengwen Liu, Xiaohua Hu 0001 |
IEEE BigData | 5 |
| 2014 | Identifying top Chinese network buzzwords from social media big data set based on time-distribution featuresabstractBuzzwords are the main embodiment of Internet culture, which play an important role in public opinion analysis, social focus tracking and language evolution study. At present, questionnaire has been wildly used as a standard method to obtain network buzzwords, which is subjective and costly. In this paper, we will propose a novel algorithm relying on the time-distribution feature of words and a KL-divergence measure to estimate words' popularity so as to figure out buzzwords in a specific period. The time-distribution feature simply states the fact that buzzwords' usage has a sharp increase during a very short period, which is then modeled formally with the KL-divergence measure. Compared with traditional method involving much workforce, the automatic algorithm presented here is clearly more efficient. Moreover, buzzwords identified in this manner will not be affected by individual's subjective opinions, so they can reflect the language usage in practice better. When applying the algorithm to a social media big data set, our experimental results show that the proposed approach can accurately identify buzzwords in a certain period, which is highly coincident with results tagged manually. Yongli Tang, Tingting He 0003, Xiaohua Hu 0001 |
IEEE BigData | 4 |
| 2014 | A Two-level Approach for Subtitle Alignment
Hao Ding 0006, Xiaohua Hu 0001, Yong Liu 0013 |
ECIR | 3 |
| 2014 | Hash-Based Stream LDA: Topic Modeling in Social Streams
Anton Slutsky, Xiaohua Hu 0001 |
PAKDD (1) | 2 |
| 2014 | Document Clustering with an Augmented Nonnegative Matrix Factorization Model
Zunyan Xiong, Yizhou Zang, Xingpeng Jiang, Xiaohua Hu 0001 |
PAKDD (2) | 4 |
| 2014 | AOBA: Recognizing Object Behavior in Pervasive Urban ManagementabstractAccurately recognizing the object's behavior from the uncertain sensor data is a key issue of Internet of Things application. For example, in urban management monitoring system, it is necessary to have an autonomous analyzing module that can online monitor object's behavior based on environmental monitoring information in order to prevent an emergent situation in advance. In this work, we present an approximate object's behavior analysis method, called AOBA, which can recognize behavioral patterns of the hybrid objects which include patrolman, watering cart, street lamp etc. In intelligent urban management. AOBA consists of two phases: filtering phase and recognizing phase. In the filtering phase, a -approximate pre-matching algorithm based on q-grams distance is introduced to select possible pattern rapidly, which can discard huge amount insignificant or dirty data; in the recognizing phase, aiming to the temporal and the spatial characteristics of sensor data, an improved bit-parallel string matching algorithm is proposed to recognize the k-approximate multiple patterns over event sequences selected by the filtering phase. Experiments on real urban monitoring data and synthetic data show that the proposed method can efficiently discriminate object's behavior. Compared with the existing method, the proposed method provides a fault-tolerant approximate pattern recognition solution. Xiaohua Hu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Tree Labeled LDA: A Hierarchical model for web summariesabstractWe study the applications of hierarchical topic models to represent the content of website summaries. We concentrate on the DMOZ collection of Web extracts and propose a novel Tree Labeled LDA (tLLDA) algorithm to infer topic models using its manually compiled ontology. The algorithm takes advantage of the ontology structure and infers topic models by jointly modeling word and ontology node assignments for documents. We evaluate the performance of our topic modeling approach against that of four state-of-the-art algorithms (Labeled LDA, Hierarchically Labeled LDA, Hierarchically Supervised LDA and Supervised LDA) and show improvement in terms of perplexity and accuracy. Our evaluation shows that topic models produced by tLLDA outperform other algorithms in terms of perplexity for all test sets and all but one test case in terms of accuracy. Anton Slutsky, Xiaohua Hu 0001 |
IEEE BigData | 2 |
| 2013 | A Novel Hybrid HDP-LDA Model for Sentiment AnalysisabstractSentiment analysis studies the public opinions towards an entity, and it is an important research area in data mining. Recently, a lot of sentiment analysis models have been proposed, including supervised and unsupervised approaches. However, the role of supervised models has been undermined by the phenomenon of big data, and the unsupervised ones are drawing more and more attention. But, most current unsupervised methods are based on Latent Dirichlet Allocation (LDA), and they need to specify the number of aspects in advance, making them subjective. In addition, these methods treat factual words and opinioned words the same, and assume that one sentence contains only one aspect, all of which make the existing unsupervised methods unsatisfactory. To solve these problems, this paper proposes a novel hybrid Hierarchical Dirichlet Process-Latent Dirichlet Allocation (HDP-LDA) model. This model can automatically determine the number of aspects, distinguish factual words from opinioned words, and further effectively extracts the aspect specific sentiment words. Experiment result shows that our model can clearly capture the aspects people mentioned and the specific sentiment words they use in each aspect, improving the performance of sentiment analysis efficiently. At last, we compared our model with the influential topic models, namely, JST, AUSM and Maxine-LDA, on the online restaurant review, and found our model performs very well. Wanying Ding, Xiaoli Song, Lifan Guo, Zunyan Xiong, Xiaohua Hu 0001 |
Web Intelligence | 5 |
| 2012 | Learning to discover complex mappings from web forms to ontologiesabstractIn order to realize the Semantic Web, various structures on the Web including Web forms need to be annotated with and mapped to domain ontologies. We present a machine learning-based automatic approach for discovering complex mappings from Web forms to ontologies. A complex mapping associates a set of semantically related elements on a form to a set of semantically related elements in an ontology. Existing schema mapping solutions mainly rely on integrity constraints to infer complex schema mappings. However, it is difficult to extract rich integrity constraints from forms. We show how machine learning techniques can be used to automatically discover complex mappings between Web forms and ontologies. The challenge is how to capture and learn the complicated knowledge encoded in existing complex mappings. We develop an initial solution that takes a naive Bayesian approach. We evaluated the performance of the solution on various domains. Our experimental results show that the solution returns the expected mappings as the top-1 results usually among several hundreds candidate mappings for more than 80% of the test cases. Furthermore, the expected mappings are always returned as the top-k results with k<4. The experiments have demonstrated that the approach is effective and has the potential to save significant human efforts. Xiaohua Hu 0001, Il-Yeol Song |
CIKM | 2 |
| 2012 | Modeling semantic relations between visual attributes and object categories via dirichlet forest priorabstractIn this paper, we deal with two research issues: the automation of visual attribute identification and semantic relation learning between visual attributes and object categories. The contribution is two-fold, firstly, we provide uniform framework to reliably extract both categorical attributes and depictive attributes. Secondly, we incorporate the obtained semantic associations between visual attributes and object categories into a text-based topic model and extract descriptive latent topics from external textual knowledge sources. Specifically, we show that in mining natural language descriptions from external knowledge sources, the relation between semantic visual attributes and object categories can be encoded as Must-Links and Cannot-Links, which can be represented by Dirichlet-Forest prior. To alleviate the workload of manual supervision and labeling in image categorization process, we introduce a semi-supervised training framework using soft-margin semi-supervised SVM classifier. We also show that the large-scale image categorization results can be significantly improved by combining automatically acquired visual attributes. Experimental results show that the proposed model achieves better ability in describing object-related attributes and makes the inferred latent topics more descriptive. Xin Chen 0041, Xiaohua Hu 0001, Zhongna Zhou, Tingting He 0003, E. K. Park |
CIKM | 2 |
| 2012 | Incorporating word correlation into tag-topic model for semantic knowledge acquisitionabstractThis paper presents a tag-topic model with Dirichlet Forest prior (TTM-DF) for semantic knowledge acquisition from blog. The TTM-DF model extends the tag-topic model (TTM) by replacing the Dirichlet prior with the Dirichlet Forest prior over the topic-word multinomial. The correlation between words are calculated to generate a set of Must-Links and Cannot-Links, then the structures of Dirichlet trees are obtained though encoding the constraints of Must-Links and Cannot-Links. Words under the same subtrees are expected to be more correlated than words under different subtrees. We conduct experiments on a synthetic and a blog dataset. Both of the experimental results show that the TTM-DF model performs much better than the TTM model. It can improve the coherence of the underlying topics and the tag-topic distributions, and capture semantic knowledge effectively. Fang Li 0003, Tingting He 0003, Xinhui Tu, Xiaohua Hu 0001 |
CIKM | 4 |
| 2012 | Author-conference topic-connection model for academic network searchabstractThis paper proposes a novel topic model, Author-Conference Topic-Connection (ACTC) Model for academic network search. The ACTC Model extends the author-conference-topic (ACT) model by adding subject of the conference and the latent mapping information between subjects and topics. It simultaneously models topical aspects of papers, authors and conferences with two latent topic layers: a subject layer corresponding to conference topic, and a topic layer corresponding to the word topic. Each author would be associated with a multinomial distribution over subjects of conference (eg., KM, DB, IR for CIKM 2012), the conference(CIKM 2012), and the topics are respectively generated from a sampled subject. Then the words are generated from the sampled topics. We conduct experiments on a data set with 8,523 authors, 22,487 papers and 1,243 conferences from the well-known Arnetminer website, and train the model with different number of subjects and topics. For a qualitative evaluation, we compare ACTC with three others models LDA, Author-Topic (AT) and ACT in academic search services. Experiments show that ACTC can effectively capture the semantic connection between different types of information in academic network and perform well in expert searching and conference searching. Xiaohua Hu 0001, Xinhui Tu, Tingting He 0003 |
CIKM | 2 |
| 2011 | Perspective hierarchical dirichlet process for user-tagged image modelingabstractIn this paper, we proposed a perspective Hierarchical Dirichlet Process (pHDP) model to deal with user-tagged image modeling. The contribution is two-fold. Firstly, we associate image features with image tags. Secondly, we incorporate the user's perspectives into the image tag generation process and introduce new latent variables to determine if an image tag is generated from user's perspectives or from the image content. Therefore, the model is able to extract both embedded semantic components and user's perspectives from user-tagged images. Based on the proposed pHDP model, we achieve automatic image tagging with users' perspective. Experimental results show that the pHDP model achieves better image tagging performance compared to state-of-the-art topic models. Xin Chen 0041, Xiaohua Hu 0001, Zunyan Xiong, Tingting He 0003, E. K. Park |
CIKM | 2 |
| 2011 | Automatically Mapping and Integrating Multiple Data Entry Forms into a Database
Ritu Khare, Il-Yeol Song, Xiaohua Hu 0001 |
ER | 4 |
| 2011 | Active Learning of Model Parameters for Influence Maximization
Tianyu Cao 0001, Xindong Wu 0001, Xiaohua Hu 0001 |
ECML/PKDD (1) | 3 |
| 2010 | Hierarchical Classification with Dynamic-Threshold SVM Ensemble for Gene Function Prediction
Yiming Chen 0002, Zhoujun Li 0001, Xiaohua Hu 0001, Junwan Liu |
ADMA (2) | 3 |
| 2010 | A probabilistic topic-connection model for automatic image annotationabstractThe explosive increase of image data on Internet has made it an important, yet very challenging task to index and automatically annotate image data. To achieve that end, sophisticated algorithms and models have been proposed to study the correlation between image content and corresponding text description. Despite the success of previous works, however, researchers are still facing two major difficulties that may undermine their effort of providing reliable and accurate annotations for images. The first difficulty is lacking of comprehensive benchmark image dataset with high quality text descriptions. The second difficulty is lacking of effective way to represent the image content and make it associate with the text descriptions. In our paper, we aim to deal with both problems. To deal with the first problem, we utilize Wikipedia as external knowledge source and enrich the ontology structure of ImageNet database with comprehensive and highly-reliable text descriptions from Wikipedia articles. To address the second problem, we develop a Probabilistic Topic-Connection (PTC) model to represent the connection between latent semantic topic in text description and latent patterns from image feature space. We compare the performance of our model with the currently popular Correspondence LDA (Corr-LDA) model under the same automatic image annotation scenario using cross-validation. Experimental results demonstrate that our model is able to well represent the connection between latent semantic topics and latent patterns in image feature space, thus facilitates knowledge organization and understanding of both image and text descriptions. Xin Chen 0041, Xiaohua Hu 0001, Zhongna Zhou, Caimei Lu, Gail L. Rosen, Tingting He 0003, E. K. Park |
CIKM | 2 |
| 2010 | The topic-perspective model for social tagging systemsabstractIn this paper, we propose a new probabilistic generative model, called Topic-Perspective Model, for simulating the generation process of social annotations. Different from other generative models, in our model, the tag generation process is separated from the content term generation process. While content terms are only generated from resource topics, social tags are generated by resource topics and user perspectives together. The proposed probabilistic model can produce more useful information than any other models proposed before. The parameters learned from this model include: (1) the topical distribution of each document, (2) the perspective distribution of each user, (3) the word distribution of each topic, (4) the tag distribution of each topic, (5) the tag distribution of each user perspective, (6) and the probabilistic of each tag being generated from resource topics or user perspectives. Experimental results show that the proposed model has better generalization performance or tag prediction ability than other two models proposed in previous research. Caimei Lu, Xiaohua Hu 0001, Xin Chen 0041, Jung-ran Park, Tingting He 0003, Zhoujun Li 0001 |
KDD | 2 |
| 2010 | Answer Diversification for Complex Question Answering on the Web
Palakorn Achananuparp, Xiaohua Hu 0001, Tingting He 0003, Christopher C. Yang, Lifan Guo |
PAKDD (1) | 2 |
| 2010 | Improving Diversity of Focused Summaries through the Negative Endorsements of Redundant FactsabstractWe present NegativeRank, a novel graph-based sentence ranking model to improve the diversity of focused summary by performing random walks over sentence graph with negative edge weights. Unlike the typical eigenvector centrality ranking, our method models the redundancy among sentence nodes as the negative edges. The negative edges can be thought of as the propagation of disapproval votes which can be used to penalize redundant sentences. As the iterative process continues, the initial ranking score of a given node will be adjusted according to a long-term negative endorsement from other sentence nodes. The evaluation results confirm that our proposed method is very effective in improving the diversity of the focused summary, compared to several well-known text summarization methods. Palakorn Achananuparp, Xiaohua Hu 0001, Lifan Guo, Tingting He 0003, Zhoujun Li 0001 |
Web Intelligence | 2 |
| 2010 | Mining hidden connections among biomedical concepts from disjoint biomedical literature sets through semantic-based association ruleabstractThe novel connection between Raynaud disease and fish oils was uncovered from two disjointed biomedical literature sets by Swanson in 1986. Since then, there have been many approaches to uncover novel connections by mining the biomedical literature. One of the popular approaches is to adapt the association rule (AR) method to automatically identify implicit novel connections between concept A and concept C from two disjointed sets of documents through intermediate B concept. Since A and C concepts do not occur together in the same data set, the mining goal is to find novel connection among A and C concepts in the disjoint data sets. It first applies association rule to the two disjointed biomedical literature sets separately to generate two rule sets (A→B, B→C), and then applies transitive law to get the novel connections A→C. However, this approach generates a huge number of possible connections among the millions of biomedical concepts and a lot of these hypothetical connections are spurious, useless, and/or biologically meaningless. Thus it is essential to develop new approach to generate highly likely novel and biologically relevant connections among the biomedical concepts. This paper presents a biomedical semantic-based association rule system (Bio-SARS) that significantly reduce spurious/useless/biologically irrelevant connections through semantic filtering. Compared to other approaches such as latent semantic indexing and traditional association rule-based approach, our approach generates much fewer rules and a lot of these rules represent relevant connections among biological concepts. © 2009 Wiley Periodicals, Inc. Xiaohua Hu 0001, Xiaodan Zhang 0001, Illhoi Yoo, Jiali Feng |
Int. J. Intell. Syst. | 1 |
| 2010 | Maintaining Mappings between Conceptual Models and Relational SchemasabstractThis paper describes a round-trip engineering approach for incrementally maintaining mappings between conceptual models and relational schemas. When either schema or conceptual model evolves to accommodate new information needs, the existing mapping must be maintained accordingly to continuously provide valid services. In this paper, the authors examine the mappings specifying “consistent” relationships between models. First, they define the consistency of a conceptual-relational mapping through “semantically compatible” instances. Next, the authors analyze the knowledge encoded in the standard database design process and develop round-trip algorithms for incrementally maintaining the consistency of conceptual-relational mappings under evolution. Finally, they conduct a set of comprehensive experiments. The results show that the proposed solution is efficient and provides significant benefits in comparison to the mapping reconstructing approach. Xiaohua Hu 0001, Il-Yeol Song |
J. Database Manag. | 2 |
| 2009 | Exploiting Wikipedia as external knowledge for document clusteringabstractIn traditional text clustering methods, documents are represented as "bags of words" without considering the semantic information of each document. For instance, if two documents use different collections of core words to represent the same topic, they may be falsely assigned to different clusters due to the lack of shared core words, although the core words they use are probably synonyms or semantically associated in other forms. The most common way to solve this problem is to enrich document representation with the background knowledge in an ontology. There are two major issues for this approach: (1) the coverage of the ontology is limited, even for WordNet or Mesh, (2) using ontology terms as replacement or additional features may cause information loss, or introduce noise. In this paper, we present a novel text clustering method to address these two issues by enriching document representation with Wikipedia concept and category information. We develop two approaches, exact match and relatedness-match, to map text documents to Wikipedia concepts, and further to Wikipedia categories. Then the text documents are clustered based on a similarity metric which combines document content information, concept information as well as category information. The experimental results using the proposed clustering framework on three datasets (20-newsgroup, TDT2, and LA Times) show that clustering performance improves significantly by enriching document representation with Wikipedia concepts and categories. Xiaohua Hu 0001, Xiaodan Zhang 0001, Caimei Lu, E. K. Park, Xiaohua Zhou |
KDD | 1 |
| 2009 | Addressing the Variability of Natural Language Expression in Sentence Similarity with Semantic Structure of the Sentences
Palakorn Achananuparp, Xiaohua Hu 0001, Christopher C. Yang |
PAKDD | 2 |
| 2009 | Spatial Weighting for Bag-of-Visual-Words and Its Application in Content-Based Image Retrieval
Xin Chen 0041, Xiaohua Hu 0001, Xiajiong Shen |
PAKDD | 2 |
| 2008 | Round-Trip Engineering for Maintaining Conceptual-Relational Mappings
Xiaohua Hu 0001, Il-Yeol Song |
CAiSE | 2 |
| 2008 | The Evaluation of Sentence Similarity Measures
Palakorn Achananuparp, Xiaohua Hu 0001, Xiajiong Shen |
DaWaK | 2 |
| 2008 | Document Clustering by Semantic Smoothing and Dynamic Growing Cell Structure (DynGCS) for Biomedical Literature
Min Song 0001, Xiaohua Hu 0001, Illhoi Yoo, Eric Koppel |
DaWaK | 2 |
| 2008 | Semantic Smoothing for Bayesian Text Classification with Small Training DataabstractBayesian text classifiers face a common issue which is referred to as data sparsity problem, especially when the size of training data is very small. The frequently used Laplacian smoothing and corpus-based background smoothing are not effective in handling it. Instead, we propose a novel semantic smoothing method to address the sparse problem. Our method extracts explicit topic signatures (e.g. words, multiword phrases, and ontology-based concepts) from a document and then statistically maps them into single-word features. We conduct comprehensive experiments on three testing collections (OHSUMED, LATimes, and 20NG) to compare semantic smoothing with other approaches. When the size of training documents is small, the bayesian classifier with semantic smoothing not only outperforms the classifiers with background smoothing and Laplacian smoothing, but also beats the state-of-the-art active learning classifiers and SVM classifiers. In this paper, we also compare three types of topic signatures with respect to their effectiveness and efficiency for semantic smoothing. Xiaohua Zhou, Xiaodan Zhang 0001, Xiaohua Hu 0001 |
SDM | 3 |
| 2008 | A comparative evaluation of different link types on enhancing document clusteringabstractWith a growing number of works utilizing link information in enhancing document clustering, it becomes necessary to make a comparative evaluation of the impacts of different link types on document clustering. Various types of links between text documents, including explicit links such as citation links and hyperlinks, implicit links such as co-authorship links, and pseudo links such as content similarity links, convey topic similarity or topic transferring patterns, which is very useful for document clustering. In this study, we adopt a Relaxation Labeling (RL)-based clustering algorithm, which employs both content and linkage information, to evaluate the effectiveness of the aforementioned types of links for document clustering on eight datasets. The experimental results show that linkage is quite effective in improving content-based document clustering. Furthermore, a series of interesting findings regarding the impacts of different link types on document clustering are discovered through our experiments. Xiaodan Zhang 0001, Xiaohua Hu 0001, Xiaohua Zhou |
SIGIR | 2 |
| 2008 | A Comparison and Scenario Analysis of Leading Data Mining SoftwareabstractFinding the right software is often hindered by different criteria as well as by technology changes. We performed an analytic hierarchy process (AHP) analysis using Expert Choice to determine which data mining package was best suitable for us. Deliberating a dozen alternatives and objectives led us to a series of pair-wise comparisons. When further synthesizing the results, Expert Choice helped us provide a clear rationale for the decision. The issue is that data mining technology is changing very rapidly. Our article focused only on the major suppliers typically available in the market place. The method and the process that we have used can be easily applied to analyze and compare other data mining software or knowledge management initiatives. John Wang 0001, Xiaohua Hu 0001, Kimberly Hollister |
Int. J. Knowl. Manag. | 2 |
| 2008 | Towards effective document clustering: A constrained K-means based approach
Guobiao Hu, Shuigeng Zhou, Jihong Guan, Xiaohua Hu 0001 |
Inf. Process. Manag. | 4 |
| 2007 | A segment-based hidden markov model for real-setting pinyin-to-chinese conversionabstractHidden markov model (HMM) is frequently used for Pinyin-to-Chinese conversion. But it only captures the dependency with the preceding character. Higher order markov models can bring higher accuracy, but are computationally unaffordable to average PC settings. We propose a segment-based hidden markov model (SHMM), which has the same magnitude of complexity as first-order HMM, but generates higher decoding accuracy. SHMM tells a word from a bigram connecting two words, and assigns a reasonable probability to words as a whole. It is more powerful than HMM to decode words containing over two characters. We conduct a comprehensive Pinyin-to-Chinese conversion evaluation on Lancaster corpus. The experiment shows the perfect sentence accuracy is improved from 34.7% (HMM) to 43.3% (SHMM). The one-error sentence accuracy is increased from 72.7% to 78.3%. Furthermore, SHMM can seamlessly integrate with pinyin typing correction, acronym pinyin input, user-defined words, and self-adaptive learning all of which are a must for a commercial Pinyin-to-Chinese conversion product in order to improve the efficiency of pinyin input. Xiaohua Zhou, Xiaohua Hu 0001, Xiaodan Zhang 0001, Xiajiong Shen |
CIKM | 2 |
| 2007 | A Comparative Study of Ontology Based Term Similarity Measures on PubMed Document Clustering
Xiaodan Zhang 0001, Liping Jing, Xiaohua Hu 0001, Michael Kwok-Po Ng, Xiaohua Zhou |
DASFAA | 3 |
| 2007 | Utilization of Global Ranking Information in Graph- Based Biomedical Literature Clustering
Xiaodan Zhang 0001, Xiaohua Hu 0001, Jiali Xia, Xiaohua Zhou, Palakorn Achananuparp |
DaWaK | 2 |
| 2007 | Integration of association rules and ontologies for semantic query expansion
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen |
Data Knowl. Eng. | 3 |
| 2007 | Topic Signature Language Models for Ad hoc RetrievalabstractSemantic smoothing, which incorporates synonym and sense information into the language models, is effective and potentially significant to improve retrieval performance. Previously implemented semantic smoothing models such as the translation model have shown good experimental results. However, these models are unable to incorporate contextual information. To overcome this limitation, we propose a novel context-sensitive semantic smoothing method that decomposes a document into a set of weighted context-sensitive topic signatures and then maps those topic signatures into query terms. The language model with such a context- sensitive semantic smoothing is referred to as the topic signature language model. In detail, we implement two types of topic signatures, depending on whether ontology exists in the application domain. One is the ontology-based concept and the other is the multiword phrase. The mapping probabilities from each topic signature to individual terms are estimated through the EM algorithm. Document models based on topic signature mapping are then derived. The new smoothing method is evaluated on the TREC 2004/ 2005 Genomics Track with ontology-based concepts, as well as the TREC Ad Hoc Track (Disks 1, 2, and 3) with multiword phrases. Both experiments show significant improvements over the two-stage language model, as well as the language model with context- insensitive semantic smoothing. Xiaohua Zhou, Xiaohua Hu 0001, Xiaodan Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2006 | Integration of cluster ensemble and EM based text mining for microarray gene cluster identification and annotationabstractIn this paper, we design and develop a unified system GE-Miner (Gene Expression Miner) to integrate cluster ensemble, text clustering and multi document summarization and provide an environment for comprehensive gene expression data analysis. We present a novel cluster ensemble approach to generate high quality gene cluster. In our text summarization module, given a gene cluster, our Expectation Maximization (EM) based algorithm can automatically identify subtopics and extract most probable terms for each topic. Then, the extracted top k topical terms from each subtopic are combined to form the biological explanation of each gene cluster. Experimental results demonstrate that our system can obtain high quality clusters and provide informative key terms for the gene clusters. Xiaohua Hu 0001, Xiaodan Zhang 0001, Xiaohua Zhou |
CIKM | 1 |
| 2006 | Relation-Based Document Retrieval for Biomedical Literature Databases
Xiaohua Zhou, Xiaohua Hu 0001, Xia Lin, Hyoil Han, Xiaodan Zhang 0001 |
DASFAA | 2 |
| 2006 | A Coherent Biomedical Literature Clustering and Summarization Approach Through Ontology-Enriched Graphical Representations
Illhoi Yoo, Xiaohua Hu 0001, Il-Yeol Song |
DaWaK | 2 |
| 2006 | Using Concept-Based Indexing to Improve Language Modeling Approach to Genomic IR
Xiaohua Zhou, Xiaodan Zhang 0001, Xiaohua Hu 0001 |
ECIR | 3 |
| 2006 | Semantic Smoothing for Model-based Document ClusteringabstractA document is often full of class-independent "general" words and short of class-specific "core " words, which leads to the difficulty of document clustering. We argue that both problems will be relieved after suitable smoothing of document models in agglomerative approaches and of cluster models in partitional approaches, and hence improve clustering quality. To the best of our knowledge, most model-based clustering approaches use Laplacian smoothing to prevent zero probability while most similarity-based approaches employ the heuristic TF*IDF scheme to discount the effect of "general" words. Inspired by a series of statistical translation language model for text retrieval, we propose in this paper a novel smoothing method referred to as context-sensitive semantic smoothing for document clustering purpose. The comparative experiment on three datasets shows that model-based clustering approaches with semantic smoothing is effective in improving cluster quality. Xiaodan Zhang 0001, Xiaohua Zhou, Xiaohua Hu 0001 |
ICDM | 3 |
| 2006 | Integration of semantic-based bipartite graph representation and mutual refinement strategy for biomedical literature clusteringabstractWe introduce a novel document clustering approach that overcomes those problems by combining a semantic-based bipartite graph representation and a mutual refinement strategy. The primary contributions of this paper are the following. First, we introduce a new representation of documents using a bipartite graph between documents and co-occurrence concepts in the documents. Second, we show how to enhance clustering quality by applying the mutual refinement strategy to the initial clustering results. Third, through the experiments on MEDLINE documents, we show that our integrated method significantly enhances cluster quality and clustering reliability compared to existing clustering methods. Our approach improves on the average 29.5 cluster quality and 26.3 clustering reliability, in terms of misclassification index, over Bisecting K-means with the best parameters. Illhoi Yoo, Xiaohua Hu 0001, Il-Yeol Song |
KDD | 2 |
| 2006 | Clustering Large Collection of Biomedical Literature Based on Ontology-Enriched Bipartite Graph Representation and Mutual Refinement Strategy
Illhoi Yoo, Xiaohua Hu 0001 |
PAKDD | 2 |
| 2006 | A Semantic Approach for Mining Hidden Links from Complementary and Non-interactive Biomedical LiteratureabstractTwo complementary and non-interactive literature sets of articles, when they are considered together, can reveal useful information of scientific interest not apparent in either of the two sets alone. Swanson called the existence of such hidden links as undiscovered public knowledge (UPK). The novel connection between Raynaud disease and fish oils was uncovered from complementary and non-interactive biomedical literature by Swanson in 1986. Since then, there have been many approaches to uncover UPK by mining the biomedical literature. These earlier works, however, required substantial manual intervention to reduce the number of possible connections. This paper proposes a semantic-based mining model for undiscovered public knowledge using the biomedical literature. Our method replaces manual ad-hoc pruning by using semantic knowledge from the biomedical ontologies. Using the semantic types and semantic relationships of the biomedical concepts, our prototype system can identify the relevant concepts collected from Medline and generate the novel hypothesis between these concepts. The system successfully replicates Swanson's two famous discoveries: Raynaud disease/fish oils and migraine/magnesium. Compared with previous approaches such as LSI-based and traditional association rule-based methods, our method generates much fewer but more relevant novel hypotheses, and requires much less human intervention in the discovery procedure. Xiaohua Hu 0001, Xiaodan Zhang 0001, Illhoi Yoo, Yan-Qing Zhang 0001 |
SDM | 1 |
| 2006 | Context-sensitive semantic smoothing for the language modeling approach to genomic IRabstractSemantic smoothing, which incorporates synonym and sense information into the language models, is effective and potentially significant to improve retrieval performance. The implemented semantic smoothing models, such as the translation model which statistically maps document terms to query terms, and a number of works that have followed have shown good experimental results. However, these models are unable to incorporate contextual information. Thus, the resulting translation might be mixed and fairly general. To overcome this limitation, we propose a novel context-sensitive semantic smoothing method that decomposes a document or a query into a set of weighted context-sensitive topic signatures and then translate those topic signatures into query terms. In detail, we solve this problem through (1) choosing concept pairs as topic signatures and adopting an ontology-based approach to extract concept pairs; (2) estimating the translation model for each topic signature using the EM algorithm; and (3) expanding document and query models based on topic signature translations. The new smoothing method is evaluated on TREC 2004/05 Genomics Track collections and significant improvements are obtained. The MAP (mean average precision) achieves a 33.6 % maximal gain over the simple language model, as well as a 7.8 % gain over the language model with context-insensitive semantic smoothing. Xiaohua Zhou, Xiaohua Hu 0001, Xiaodan Zhang 0001, Xia Lin, Il-Yeol Song |
SIGIR | 2 |
| 2005 | Mining undiscovered public knowledge from complementary and non-interactive biomedical literature through semantic pruningabstractTwo complementary and non-interactive literature sets of articles, when they are considered together, can reveal useful information of scientific interest not apparent in either of the two document sets. Swanson called the existence of such knowledge, undiscovered public knowledge (UDPK). This paper proposes a semantic-based mining model for UDPK. Our method replaces manual ad-hoc pruning with using semantic knowledge from the biomedical ontologies. Using the semantic types and semantic relationships of the biomedical concepts, our prototype system can identify the relevant concepts collected from Medline and generate the novel hypothesis between these concepts. The system successfully replicates Swanson's two famous discoveries: Raynaud disease/fish oils and migraine/magnesium. Compared with previous approaches, our methods generate much fewer but more relevant novel hypotheses, and require much less human intervention in the discovery procedure. Xiaohua Hu 0001, Illhoi Yoo, Min Song 0001, Yan-Qing Zhang 0001, Il-Yeol Song |
CIKM | 1 |
| 2005 | Semantic Query Expansion Combining Association Rules with Ontologies and Information Retrieval Techniques
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen |
DaWaK | 3 |
| 2005 | An Automatic Unsupervised Querying Algorithm for Efficient Information Extraction in Biomedical Domain
Min Song 0001, Il-Yeol Song, Xiaohua Hu 0001, Robert B. Allen |
PAKDD | 3 |
| 2004 | Managing and Mining Clinical Outcomes
Hyoil Han, Il-Yeol Song, Xiaohua Hu 0001, Ann Prestrud, Murray F. Brennan, Ari D. Brooks |
DASFAA | 3 |
| 2004 | Intelligent Query Answering Based on Neighborhood Systems and Data Mining Techniques
Tsau Young Lin, Nick Cercone, Xiaohua Hu 0001, Jianchao Han |
IDEAS | 3 |
| 2004 | A Weighted Freshness Metric for Maintaining Search Engine Local RepositoryabstractCurrent search engines maintain a local repository to improve the search efficiency. A crawler is used to periodically poll the remote web pages to update the contents of the local repository. Due to the resource limitations, some local pages may be stale. To maintain the high freshness of the repository, the crawler is expected to revisit remote web pages in optimized order and frequency. The intuitive metric of freshness of the local repository is defined as the fraction of up-to-date web pages in the repository, which is merely based on the repository content, and does not, unfortunately, reflect the perspective of the search engine users, e.g., how often is a web page queried? We propose a novel weighted metric of the repository freshness with the importance of web pages being the weights. This metric not only takes into account the local web pages themselves but also the perspectives of the search engine users. We study the repository synchronization policy under this new metric, compare this metric with others, analyze its features, and discuss how the web page importance is determined. Jianchao Han, Nick Cercone, Xiaohua Hu 0001 |
Web Intelligence | 3 |
| 2004 | Ontology-Based Scalable and Portable Information Extraction System to Extract Biological Knowledge from Huge Collection of Biomedical Web DocumentsabstractAutomated discovery and extraction of biological knowledge from biomedical web documents has become essential because of the enormous amount of biomedical literature published each year. In this paper we present an ontology-based scalable and portable information extraction system to automatically extract biological knowledge from huge collection of online biomedical web documents. Our method integrates ontology-based semantic tagging, information extraction and data mining together, automatically learns the patterns based on a few user seed tuples, and then extract new tuples from the biomedical web documents based on the discovered patterns. A novel system SPIE (Scalable and Portable Information Extraction) is implemented and tested on the PuBMed to find the chromatin protein-protein interaction and the experimental results indicate our approach is very effective in extracting biological knowledge from huge collection of biomedical web documents. Xiaohua Hu 0001, Tsau Young Lin, Il-Yeol Song, Xia Lin, Illhoi Yoo, Mark Lechner, Min Song 0001 |
Web Intelligence | 1 |
| 2004 | A data warehouse/online analytic processing framework for web usage mining and business intelligence reportingabstractWeb usage mining is the application of data mining techniques to discover usage patterns and behaviors from web data (clickstream, purchase information, customer information, etc.) in order to understand and serve e-commerce customers better and improve the online business. In this article, we present a general data warehouse/online analytic processing (OLAP) framework for web usage mining and business intelligence reporting. When we integrate the web data warehouse construction, data mining, and OLAP into the e-commerce system, this tight integration dramatically reduces the time and effort for web usage mining, business intelligence reporting, and mining deployment. Our data warehouse/OLAP framework consists of four phases: data capture, webhouse construction (clickstream marts), pattern discovery and cube construction, and pattern evaluation and deployment. We discuss data transformation operations for web usage mining and business reporting in clickstream, session, and customer levels; describe the problems and challenging issues in each phase in detail; provide plausible solutions to the issues; and demonstrate the framework with some examples from some real web sites. Our data warehouse/OLAP framework has been integrated into some commercial e-commerce systems. We believe this data warehouse/OLAP framework would be very useful for developing any real-world web usage mining and business intelligence reporting systems. © 2004 Wiley Periodicals, Inc. Xiaohua Hu 0001, Nick Cercone |
Int. J. Intell. Syst. | 1 |
| 2003 | A New Computation Model for Rough Set Theory Based on Database Systems
Jianchao Han, Xiaohua Hu 0001, Tsau Young Lin |
DaWaK | 2 |
| 2003 | Bitmap Techniques for Optimizing Decision Support Queries and Association Rule AlgorithmsabstractIn this paper, we discuss some new bitmap techniques for optimizing decision support queries and association rule algorithm. We first show how to use a new type of predefined bitmap join index (prejoin/spl I.bar/bitmap/spl I.bar/index) to efficiently execute complex decision support queries with multiple outer join operations involved and push the outer join operations from the data flow level to the bitmap level and achieve significant performance gain. Then we discuss a bitmap based association rule algorithm. Our bitmap based association rule algorithm Bit-AssocRule doesn't follow the generation-and-test strategy of a priori algorithm and adopts the divide-and-conquer strategy, thus avoids the time-consuming table scan to find and prune the item sets, all the operations of finding large item sets from the datasets are the fast bit operations. The experimental results show Bit-AssocRule is 2 to 3 orders of magnitude faster than a priori and a priori hybrid algorithms. Our results indicate that bitmap techniques can greatly improve the performance of decision support queries and association rule algorithm, and bitmap techniques are very promising for the decision support query optimization and data mining applications. Xiaohua Hu 0001, Tsau Young Lin, Eric Louie |
IDEAS | 1 |
| 2001 | Using Rough Sets Theory and Database Operations to Construct a Good Ensemble of Classifiers for Data Mining ApplicationsabstractThe article presents a novel approach to constructing a good ensemble of classifiers using rough set theory and database operations. Ensembles of classifiers are formulated precisely within the framework of rough set theory and constructed very efficiently by using set-oriented database operations. Our method first computes a set of reducts which include all the indispensable attributes required for the decision categories. For each reduct, a reduct table is generated by removing those attributes which are not in the reduct. Next, a novel rule induction algorithm is used to compute the maximal generalized rules for each reduct table and a set of reduct classifiers is formed based on the corresponding reducts. The distinctive features of our method as compared to other methods of constructing ensembles of classifiers are: (1) presents a theoretical model to explain the mechanism of constructing ensemble of classifiers; (2) each reduct is a minimum subset of attributes and has the same classification ability as the entire attributes; (3) each reduct classifier constructed from the corresponding reduct has a minimal set of classification rules, and is as accurate and complete as possible and at the same time as diverse as possible from the other classifiers; (4) the test indicates that the number of classifiers used to improve the accuracy is much less than other methods. Xiaohua Hu 0001 |
ICDM | 1 |
| 1999 | Data Mining via Discretization, Generalization and Rough Set Feature Selection
Xiaohua Hu 0001, Nick Cercone |
Knowl. Inf. Syst. | 1 |
| 1996 | Mining Knowledge Rules from Databases: A Rough Set ApproachabstractThe principle and experimental results of an attribute oriented rough set approach for knowledge discovery in databases are described. Our method integrates the database operation, rough set theory and machine learning techniques. In this method the learning procedure consists of two phases: data generalization and data reduction. In the data generalization phase, attribute oriented induction is performed attribute by attribute using attribute removal and concept ascension, some undesirable attributes to the discovery task are removed and the primitive data is generalized to the desirable level; thus a set of tuples may be generalized to the same generalized tuple. This procedure substantially reduces the computational complexity of the database learning process. Subsequently, in data reduction phase, the rough set method is applied to the generalized relation to find a minimal attribute set relevant to the learning task. The generalized relation is reduced further by removing those attributes which are irrelevant and/or unimportant to the learning task. Finally the tuples in the reduced relation are transformed into different knowledge rules based on different knowledge discovery algorithms. Based upon these principles, a prototype knowledge discovery system, DBROUGH has been constructed. In DBROUGH, a variety of knowledge discovery algorithms are incorporated and different kinds of knowledge rules, such as characteristic rules, classification rules, decision rules, maximal generalized rules can be discovered efficiently and effectively from large databases. Xiaohua Hu 0001, Nick Cercone |
ICDE | 1 |
| 1995 | Rough Sets Similarity-Based Learning from Databases
Xiaohua Hu 0001, Nick Cercone |
KDD | 1 |
| 1994 | Discovery of Decision Rules in Relational Databases: A Rough Set ApproachabstractWe develop an attribute-oriented rough set approach for the discovery of decision rules in relational databases. Our approach combines machine learning techniques and rough set theory. We consider a learning procedure to consist of the two phases data generalization and data reduction. In the data generalization phase, utilizing knowledge about concept hierarchies and relevance of the data, an attribute-oriented induction is performed attribute by attribute. Some undesirable attributes of the discovery task are removed and the primitive data in the databases are generalized to the desirable level; this process greatly decreases the number of tuples which must be examined for the discovery task and substantially reduces the computational complexity of the database learning processes. Subsequently, in data reduction phase, rough set theory is applied to the generalized relation; the cause-effect relationships among the condition and decision attributes in the databases are analyzed and the non-essential or irrelevant attributes to the discovery task are eliminated without losing information of the original database system. This process further reduces the generalized relation. Thus very concise and more accurate decision rules for each class in the decision attribute with little or no redundancy information, can be extracted automatically from the reduced relation during the learning process. Our study shows that attribute-oriented induction combined with rough set theory provide an efficient and effective mechanism for discovering decision rules in database systems. Xiaohua Hu 0001, Nick Cercone |
CIKM | 1 |