Eui-Hong Han

dblp:h/EuiHongHan · also Eui-Hong Sam Han · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
1since 2021 · last 2026
0009-0000-2988-1284ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 7 first-authorArtificial intelligence and machine learning · 10 · 4 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Web and social media mining · 71% Data mining · 29%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 80% High-performance computing · 20%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Web and social media mining › social media analysis
demographic inference
0.112010
Content-Based Methods for Predicting Web-Site Demographic Attributes · ICDM 2010
Web and social media mining › web mining
web content analysis
0.112010
Content-Based Methods for Predicting Web-Site Demographic Attributes · ICDM 2010
Data mining › pattern mining
association rule mining
0.022000
Scalable Parallel Data Mining for Association Rules · IEEE Trans. Knowl. Data Eng. 2000
Scalable Parallel Data Mining for Association Rules · SIGMOD Conference 1997
Data mining › pattern mining › association rule mining
parallel association rule mining
0.022000
Scalable Parallel Data Mining for Association Rules · IEEE Trans. Knowl. Data Eng. 2000
Scalable Parallel Data Mining for Association Rules · SIGMOD Conference 1997
Parallel and multicore computing
parallel data mining
0.012000
Scalable Parallel Data Mining for Association Rules · IEEE Trans. Knowl. Data Eng. 2000
High-performance computing › data-intensive computing
scalable data mining
0.012000
Scalable Parallel Data Mining for Association Rules · IEEE Trans. Knowl. Data Eng. 2000
Parallel and multicore computing
parallel algorithms
0.011997
Scalable Parallel Data Mining for Association Rules · SIGMOD Conference 1997

Methods — techniques the papers use, named apart from their topics

regression · 0.1matrix approximation · 0.1hash tree construction · 0.1candidate set partitioning · 0.1apriori · 0.1load balancing · 0.0candidate partitioning · 0.0
YearPublicationVenuePosition
2026 LLM-Based Content Tagging at The Washington Post
abstract
We present a production LLM-based taxonomy classification system deployed at The Washington Post that tags news content across five schemas (Subject, Person, Company, Organization, Geography) using a proprietary taxonomy of ∼ 20,400 entries across seven hierarchical levels. For the Subject schema, we employ embedding-based candidate filtering followed by LLM selection. For other schemas, we combine LLM-based named entity extraction with fuzzy n-gram matching, followed by LLM selection. Comparison of post-production F1 scores against commercial vendor baselines demonstrates significant improvements across all five schemas, with the most substantial gain in Subject schema (+29.3%, p < 0.001). The system processes hundreds to thousands of articles and news items daily with a mean latency of 3–4 seconds per request and supports zero-downtime taxonomy updates.
Meng Ling, Himanshu Jahagirdar, Janith Weerasinghe, Han Jun Yoon, Suja Thomas, Anuradha Uduwage, Eui-Hong Han
UMAP7
2019 Characterization and Early Detection of Evergreen News Articles
Yiming Liao, Shuguang Wang, Eui-Hong Han, Jongwuk Lee, Dongwon Lee 0001
ECML/PKDD (3)3
2016 Predicting the shape and peak time of news article views
abstract
Predicting the popularity of news articles - whether measured via retweets, clicks, or views - is an important problem for editors, journalists, and readers alike. In this paper, we introduce a new model to predict the shape of news article views, and use this model to determine when an article will likely reach its maximum number of views. Although volume prediction for news articles has been extensively studied predicting when a burst of views will happen, in what shape, and by how much, remains an open problem. We engineer several classes of features (metadata, contextual or content-based, temporal, and social), develop models to classify shape of views, with particular attention paid to performing online, time-updated, prediction, i.e., using data before and during the early stages of article prediction to predict its eventual peak views and update earlier predictions. The system presented here is an emerging application being developed at The Washington Post and can be used to support article placement, updating, and promotion strategies.
Yaser Keneshloo, Shuguang Wang, Eui-Hong Han, Naren Ramakrishnan
IEEE BigData3
2016 Predicting the Popularity of News Articles
abstract
Consuming news articles is an integral part of our daily lives and news agencies such as The Washington Post (WP) expend tremendous effort in providing high quality reading experiences for their readers. Journalists and editors are faced with the task of determining which articles will become popular so that they can efficiently allocate resources to support a better reading experience. The reasons behind the popularity of news articles are typically varied, and might involve contemporariness, writing quality, and other latent factors. In this paper, we cast the problem of popularity prediction problem as regression, engineer several classes of features (metadata, contextual or content-based, temporal, and social), and build models for forecasting popularity. The system presented here is deployed in a real setting at The Washington Post; we demonstrate that it is able to accurately predict article popularity with an R2 ≈ 0.8 using features harvested within 30 minutes of publication time.
Yaser Keneshloo, Shuguang Wang, Eui-Hong Han, Naren Ramakrishnan
SDM3
2015 BreakFast: Analyzing Celerity of News
abstract
In the hypercompetitive news market, news outlets race to break news first. In order to provide better breaking news service and improve the reader experience, news agencies need to understand how to identify bottlenecks and streamline their reporting and delivery processes. With that in mind, we built a system, BreakFast, to measure and compare the speed of delivery of breaking news from various news sources to readers. One of the primary challenges of this comparison is how to identify which breaking news items are about the same emerging event but reported by different news agencies with different headlines and content. To tackle this problem, we extracted keywords automatically from the content, identified important topics, and then developed a classification model. The model identifies the same breaking stories from multiple news sources with an accuracy of approximately 90%. We also proposed new metrics to evaluate the speed of breaking news services and built real-time dashboards to monitor performance over time. We deployed BreakFast into the breaking news service at The Washington Post. This integrated system narrowed in on bottlenecks in its breaking news generation and delivery process, and improved its breaking news service in terms of time by more than 50%.
Shuguang Wang, Eui-Hong Han
ICMLA2
2015 Recommending Temporally Relevant News Content from Implicit Feedback Data
abstract
News has, in this day and age, transformed primarily into a digital format with leading newspapers and news agencies having a significant online presence. The speed at which news reaches the reader notwithstanding, the proliferation of blogs and microblogs to deliver specialized content has become the order of the day. Even highly engaged users tend to disengage with a website when the content they are served is unappealing to them. While recommendation systems have been used to ensure delivery of content to the user in tune with their tastes, these systems face an unprecedented challenge - the transient nature of 'popular' news and users' changing interests. Moreover, the challenge is compounded by the absence of explicit feedback. Most recommendation systems for recommending digital news content rely on inferring user engagement through 'clicks', which is not necessarily an accurate measure as it gives us no explicit information about the degree to which a user is interested in a news article. In this paper, we introduce and study the behavior of temporal and tag-based models for news article recommendation. Our experiments indicate that incorporating temporal and taginformation improves recommendation quality and increases user engagement. We argue through experimental evaluation that the improved performance is due to recommendation of more personalized news content by the tag-based recommendation algorithms as compared to other models that do not explicitly incorporate user-tag information.
Nikhil Muralidhar, Huzefa Rangwala, Eui-Hong Han
ICTAI3
2010 Content-Based Methods for Predicting Web-Site Demographic Attributes
abstract
Demographic information plays an important role in gaining valuable insights about a web-site's user-base and is used extensively to target online advertisements and promotions. This paper investigates machine-learning approaches for predicting the demographic attributes of web-sites using information derived from their content and their hyper linked structure and not relying on any information directly or indirectly obtained from the web-site's users. Such methods are important because users are becoming increasingly more concerned about sharing their personal and behavioral information on the Internet. Regression-based approaches are developed and studied for predicting demographic attributes that utilize different content-derived features, different ways of building the prediction models, and different ways of aggregating web-page level predictions that take into account the web's hyper linked structure. In addition, a matrix-approximation based approach is developed for coupling the predictions of individual regression models into a model designed to predict the probability mass function of the attribute. Extensive experiments show that these methods are able to achieve an RMSE of 8-10% and provide insights on how to best train and apply such models.
Santosh Kabbur, Eui-Hong Han, George Karypis
ICDM2
2005 Feature-based recommendation system
abstract
The explosive growth of the world-wide-web and the emergence of e-commerce has led to the development of recommender systems--a personalized information filtering technology used to identify a set of N items that will be of interest to a certain user. User-based and model-based collaborative filtering are the most successful technology for building recommender systems to date and is extensively used in many commercial recommender systems. The basic assumption in these algorithms is that there are sufficient historical data for measuring similarity between products or users. However, this assumption does not hold in various application domains such as electronics retail, home shopping network, on-line retail where new products are introduced and existing products disappear from the catalog. Another such application domains is home improvement retail industry where a lot of products (such as window treatments, bathroom, kitchen or deck) are custom made. Each product is unique and there are very little duplicate products. In this domain, the probability of the same exact two products bought together is close to zero. In this paper, we discuss the challenges of providing recommendation in the domains where no sufficient historical data exist for measuring similarity between products or users. We present feature-based recommendation algorithms that overcome the limitations of the existing top-n recommendation algorithms. The experimental evaluation of the proposed algorithms in the real life data sets shows a great promise. The pilot project deploying the proposed feature-based recommendation algorithms in the on-line retail web site shows 75% increase in the recommendation revenue for the first 2 month period.
Eui-Hong Han, George Karypis
CIKM1
2003 Intelligent metasearch engine for knowledge management
abstract
The explosive growth of available information sources and the resulting information overload pose several problems for users in many business organizations and educational institutions. First, searching through several information sources, one at a time, is a source of enormous frustration for users. Second, top-ranked documents in search results are frequently irrelevant to what users are interested in. To address these problems, we have developed ixmeta™, a powerful metasearch engine that gathers, evaluates, ranks, and reports the most relevant results from multiple information sources, including library catalogs, proprietary databases, intranets, and Web search engines. In addition to basic metasearch capabilities, ixmetafind uses personalization and clustering techniques to find the most relevant results for users. In this paper, we briefly describe technologies used in ixmetafind and present pinpoint™ from Sagebrush Corporation, the smart research tool™ in the kindergarten through twelfth grade (K-12) school environment. Pinpoint showcases ixmetafind in the knowledge management domain of the K-12 school environment.
Eui-Hong Han, George Karypis, Doug Mewhort, Keith Hatchard
CIKM1
2001 Text Categorization Using Weight Adjusted k-Nearest Neighbor Classification
Eui-Hong Han, George Karypis, Vipin Kumar 0001
PAKDD1
2000 Fast Supervised Dimensionality Reduction Algorithm with Applications to Document Categorization & Retrieval
abstract
Retriev al techniques based on dimensionalit y reduction, such as Latent S e m a n tic Indexing (LSI), have been shown to improve the quality of the information being retrieved by c a pturing the latent meaning of the words present in the documents.Unfortunately, the high computational and memory requirements of LSI and its inabilit yto compute an eective dimensionality reduction in a supervised setting limits its applicability.In this paper we p r e s e n t a fast supervised dimensionality reduction algorithm that is derived from the recen tly dev eloped cluster-based unsupervised dimensionality reduction algorithms.We experimentally evaluate the quality of the low er dimensional spaces both in the context of document categorization and improvements in retrieval performance on a variety of dierent document collections.Our experiments sho w that the lower dimensional spaces computed by our algorithm consistently improve the performance of traditional algorithms such as C4.5, k-nearestneigh bor, and Support V ector Machines (SVM), by a n a verage of 2% to 7%.F urthermore, the supervised lower dimensional space greatly improves the retriev al performance when compared to LSI.This work w as supported
Eui-Hong Han, George Karypis
CIKM1
2000 Centroid-Based Document Classification: Analysis and Experimental Results
Eui-Hong Han, George Karypis
PKDD1
2000 Scalable Parallel Data Mining for Association Rules
abstract
The authors propose two new parallel formulations of the Apriori algorithm (R. Agrawal and R. Srikant, 1994) that is used for computing association rules. These new formulations, IDD and HD, address the shortcomings of two previously proposed parallel formulations CD and DD. Unlike the CD algorithm, the IDD algorithm partitions the candidate set intelligently among processors to efficiently parallelize the step of building the hash tree. The IDD algorithm also eliminates the redundant work inherent in DD, and requires substantially smaller communication overhead than DD. But IDD suffers from the added cost due to communication of transactions among processors. HD is a hybrid algorithm that combines the advantages of CD and DD. Experimental results on a 128-processor Cray T3E show that HD scales just as well as the CD algorithm with respect to the number of transactions, and scales as well as IDD with respect to increasing candidate set size.
Eui-Hong Han, George Karypis, Vipin Kumar 0001
IEEE Trans. Knowl. Data Eng.1
1999 Parallel Formulations of Decision-Tree Classification Algorithms
Eui-Hong Han, Vipin Kumar 0001
Data Min. Knowl. Discov.2
1999 Partitioning-based clustering for Web document categorization
Daniel Boley, Maria L. Gini, Robert Gross, Eui-Hong Han, Kyle Hastings, George Karypis, Vipin Kumar 0001, Bamshad Mobasher, Jerome Moore
Decis. Support Syst.4
1998 Parallel Formulations of Decision-Tree Classification Algorithms
abstract
Classification decision tree algorithms are used extensively for data mining in many domains such as retail target marketing, fraud detection, etc. Highly parallel algorithms for constructing classification decision trees are desirable for dealing with large data sets in reasonable amount of time. Algorithms for building classification decision trees have a natural concurrency, but are difficult to parallelize due to the inherent dynamic nature of the computation. We present parallel formulations of classification decision tree learning algorithm based on induction. We describe two basic parallel formulations. One is based on Synchronous Tree Construction Approach and the other is based on Partitioned Tree Construction Approach. We discuss the advantages and disadvantages of using these methods and propose a hybrid method that employs the good features of these methods. Experimental results on an IBM SP-2 demonstrate excellent speedups and scalability.
Eui-Hong Han, Vipin Kumar 0001
ICPP2
1997 Scalable Parallel Data Mining for Association Rules
abstract
One of the important problems in data mining is discovering association rules from databases of transactions where each transaction consists of a set of items. The most time consuming operation in this discovery process is the computation of the frequency of the occurrences of interesting subset of items (called candidates) in the database of transactions. To prune the exponentially large space of candidates, most existing algorithms, consider only those candidates that have a user defined minimum support. Even with the pruning, the task of finding all association rules requires a lot of computation power and time. Parallel computers offer a potential solution to the computation requirement of this task, provided efficient and scalable parallel algorithms can be designed. In this paper, we present two new parallel algorithms for mining association rules. The Intelligent Data Distribution algorithm efficiently uses aggregate memory of the parallel computer by employing intelligent candidate partitioning scheme and uses efficient communication mechanism to move data among the processors. The Hybrid Distribution algorithm further improves upon the Intelligent Data Distribution algorithm by dynamically partitioning the candidate set to maintain good load balance. The experimental results on a Cray T3D parallel computer show that the Hybrid Distribution algorithm scales linearly and exploits the aggregate memory better and can generate more association rules with a single scan of database per pass.
Eui-Hong Han, George Karypis, Vipin Kumar 0001
SIGMOD Conference1