Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dae Hoon Park

dblp:30/10689 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
2since 2021 · last 2026
0009-0001-9247-1736ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 6 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Information retrieval · 79% Recommender systems · 17% Data mining · 4%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
text summarization
0.522016
Retrieving Non-Redundant Questions to Summarize a Product Review · SIGIR 2016
Retrieval of Relevant Opinion Sentences for New Products · SIGIR 2015
Information retrieval › retrieval models
ad-hoc retrieval
0.412019
Adversarial Sampling and Training for Semi-Supervised Information Retrieval · WWW 2019
Recommender systems
adversarial training
0.412019
Adversarial Sampling and Training for Semi-Supervised Information Retrieval · WWW 2019
Recommender systems › collaborative filtering
implicit feedback
0.412019
Adversarial Sampling and Training for Semi-Supervised Information Retrieval · WWW 2019
Natural language and speech › Language models and text generation
neural language model
0.312017
A Neural Language Model for Query Auto-Completion · SIGIR 2017
Information retrieval › query suggestion
query auto-completion
0.312017
A Neural Language Model for Query Auto-Completion · SIGIR 2017
Information retrieval
query processing
0.312017
A Neural Language Model for Query Auto-Completion · SIGIR 2017
Information retrieval › question answering
question retrieval
0.212016
Retrieving Non-Redundant Questions to Summarize a Product Review · SIGIR 2016
Information retrieval › text summarization › user-generated content summarization
review summarization
0.212016
Retrieving Non-Redundant Questions to Summarize a Product Review · SIGIR 2016
Information retrieval
search result diversification
0.212016
Retrieving Non-Redundant Questions to Summarize a Product Review · SIGIR 2016
Information retrieval › document retrieval
opinion retrieval
0.212015
Retrieval of Relevant Opinion Sentences for New Products · SIGIR 2015
Information retrieval › text summarization
opinion summarization
0.212015
Retrieval of Relevant Opinion Sentences for New Products · SIGIR 2015
Information retrieval
retrieval models
0.212015
Leveraging User Reviews to Improve Accuracy for Mobile App Retrieval · SIGIR 2015
Data mining › text mining
topic model
0.212015
Leveraging User Reviews to Improve Accuracy for Mobile App Retrieval · SIGIR 2015
Recommender systems
item recommendation
0.112019
Adversarial Sampling and Training for Semi-Supervised Information Retrieval · WWW 2019

Methods — techniques the papers use, named apart from their topics

large language model · 1.0evolutionary search · 1.0recurrent neural network · 0.6neural network · 0.4adversarial training · 0.4adversarial sampling · 0.4submodular optimization · 0.2probabilistic retrieval · 0.2greedy algorithm · 0.2probabilistic method · 0.2
YearPublicationVenuePosition
2026 RankEvolve: Automating the Discovery of Retrieval Algorithms via LLM-Driven Evolution
Jinming Nian, Fangchen Li, Dae Hoon Park, Yi Fang 0008
SIGIR3
2023 Augmentation Robust Self-Supervised Learning for Human Activity Recognition
abstract
Human Activity Recognition (HAR) is widely applied on wearable devices in our daily lives. However, acquiring high-quality wearable sensor data set with ground-truths is challenging due to the high cost in collecting data and necessity of domain experts. In order to achieve generalization from limited data, we study augmentation-based Self-Supervised Learning (SSL) for data from wearable devices. However, there is an issue in one of the most popular SSL approaches, contrastive learning: it is sensitive to the choice of data augmentations. To resolve this, we first propose to combine contrastive learning with generative learning, which is robust to augmentations. Second, we propose an automatic augmentation policy search method to discover the most promising augmentation policy. We empirically verify our approaches on three public HAR datasets. Experimental results show that our proposed SSL approach is robust to augmentations, and delivers higher accuracy than contrastive learning. Additionally, with the searched augmentation policy we are able to further improve the accuracy of HAR task.
Yuhang Li 0001, Dae Lee, Dae Hoon Park, Hongda Mao, Huyen Do, Jonathan Chung 0001, Dinesh Nair
ICASSP4
2020 node2hash: Graph aware deep semantic text hashing
Suthee Chaidaroon, Dae Hoon Park, Yi Chang 0001, Yi Fang 0008
Inf. Process. Manag.2
2019 Adversarial Sampling and Training for Semi-Supervised Information Retrieval
abstract
Ad-hoc retrieval models with implicit feedback often have problems, e.g., the imbalanced classes in the data set. Too few clicked documents may hurt generalization ability of the models, whereas too many non-clicked documents may harm effectiveness of the models and efficiency of training. In addition, recent neural network-based models are vulnerable to adversarial examples due to the linear nature in them. To solve the problems at the same time, we propose an adversarial sampling and training framework to learn ad-hoc retrieval models with implicit feedback. Our key idea is (i) to augment clicked examples by adversarial training for better generalization and (ii) to obtain very informational non-clicked examples by adversarial sampling and training. Experiments are performed on benchmark data sets for common ad-hoc retrieval tasks such as Web search, item recommendation, and question answering. Experimental results indicate that the proposed approaches significantly outperform strong baselines especially for high-ranked documents, and they outperform IRGAN in [email protected] using only 5% of labeled data for the Web search task.
Dae Hoon Park, Yi Chang 0001
WWW1
2017 A Neural Language Model for Query Auto-Completion
abstract
Query auto-completion (QAC) systems suggest queries that complete a user's text as the user types each character. Such queries are typically selected among previously stored queries, based on specific attributes such as popularity. However, queries cannot be suggested if a user's text does not match any queries in the storage. In order to suggest queries for previously unseen text, we propose a neural language model that learns how to generate a query from a starting text, a prefix. Specifically, we employ a recurrent neural network to handle prefixes in variable length. We perform the first neural language model experiments for QAC, and we evaluate the proposed methods with a public data set. Empirical results show that the proposed methods are as effective as traditional methods for previously seen queries and are superior to the state-of-the-art QAC method for previously unseen queries.
Dae Hoon Park, Rikio Chiba
SIGIR1
2017 Which used product is more sellable? A time-aware approach
Mengwen Liu, Wanying Ding, Dae Hoon Park, Yi Fang 0008, Rui Yan 0001, Xiaohua Hu 0001
Inf. Retr. J.3
2017 Product review summarization through question retrieval and diversification
Mengwen Liu, Yi Fang 0008, Alexander G. Choulos, Dae Hoon Park, Xiaohua Hu 0001
Inf. Retr. J.4
2016 Mobile App Retrieval for Social Media Users via Inference of Implicit Intent in Social Media Text
abstract
People often implicitly or explicitly express their needs in social media in the form of "user status text". Such text can be very useful for service providers and product manufacturers to proactively provide relevant services or products that satisfy people's immediate needs. In this paper, we study how to infer a user's intent based on the user's "status text" and retrieve relevant mobile apps that may satisfy the user's needs. We address this problem by framing it as a new entity retrieval task where the query is a user's status text and the entities to be retrieved are mobile apps. We first propose a novel approach that generates a new representation for each query. Our key idea is to leverage social media to build parallel corpora that contain implicit intention text and the corresponding explicit intention text. Specifically, we model various user intentions in social media text using topic models, and we predict user intention in a query that contains implicit intention. Then, we retrieve relevant mobile apps with the predicted user intention. We evaluate the mobile app retrieval task using a new data set we create. Experiment results indicate that the proposed model is effective and outperforms the state-of-the-art retrieval models.
Dae Hoon Park, Yi Fang 0008, Mengwen Liu, ChengXiang Zhai
CIKM1
2016 Retrieving Non-Redundant Questions to Summarize a Product Review
abstract
Product reviews have become an important resource for customers before they make purchase decisions. However, the abundance of reviews makes it difficult for customers to digest them and make informed choices. In our study, we aim to help customers who want to quickly capture the main idea of a lengthy product review before they read the details. In contrast with existing work on review analysis and document summarization, we aim to retrieve a set of real-world user questions to summarize a review. In this way, users would know what questions a given review can address and they may further read the review only if they have similar questions about the product. Specifically, we design a two-stage approach which consists of question retrieval and question diversification. We first propose probabilistic retrieval models to locate candidate questions that are relevant to a review. We then design a set function to re-rank the questions with the goal of rewarding diversity in the final question set. The set function satisfies submodularity and monotonicity, which results in an efficient greedy algorithm of submodular optimization. Evaluation on product reviews from two categories shows that the proposed approach is effective for discovering meaningful questions that are representative for individual reviews.
Mengwen Liu, Yi Fang 0008, Dae Hoon Park, Xiaohua Hu 0001, Zhengtao Yu 0001
SIGIR3
2015 SpecLDA: Modeling Product Reviews and Specifications to Generate Augmented Specifications
abstract
Product specifications are often available for a product on E-commerce websites. However, novice customers often do not have enough knowledge to understand all features of a product, especially advanced features. In order to provide useful knowledge to the customers, we propose to automatically generate augmented product specifications, which contains relevant opinions for product feature values, feature importance, and product-specific words. Specifically, we propose a novel Specification Latent Dirichlet Allocation (SpecLDA) that can enable us to effectively model product reviews and specifications at the same time. It mines review texts relevant to a feature value in order to inform customers what other customers have said about the feature value in reviews of the same product and also different products. SpecLDA can also infer importance of each feature and infer which words are special for each product so that customers quickly understand products. Experiment results show that SpecLDA can effectively model product reviews with specifications. The model can be used for any text collections with specification (key-value) type prior knowledge.
Dae Hoon Park, ChengXiang Zhai, Lifan Guo
SDM1
2015 Retrieval of Relevant Opinion Sentences for New Products
abstract
With the rapid development of Internet and E-commerce, abundant product reviews have been written by consumers who bought the products. These reviews are very useful for consumers to optimize their purchasing decisions. However, since the reviews are all written by consumers who have bought and used a product, there are generally very few or even no reviews available for a new product or an unpopular product. We study the novel problem of retrieving relevant opinion sentences from the reviews of other products using specifications of a new or unpopular product as query. Our key idea is to leverage product specifications to assess product similarity between the query product and other products and extract relevant opinion sentences from the similar products where a consumer may find useful discussions. Then, we provide ranked opinion sentences for the query product that has no user-generated reviews. We first propose a popular summarization method and its modified version to solve the problem. Then, we propose our novel probabilistic methods. Experiment results show that the proposed methods can effectively retrieve useful opinion sentences for products that have no reviews.
Dae Hoon Park, Hyun Duk Kim, ChengXiang Zhai, Lifan Guo
SIGIR1
2015 Leveraging User Reviews to Improve Accuracy for Mobile App Retrieval
abstract
Smartphones and tablets with their apps pervaded our everyday life, leading to a new demand for search tools to help users find the right apps to satisfy their immediate needs. While there are a few commercial mobile app search engines available, the new task of mobile app retrieval has not yet been rigorously studied. Indeed, there does not yet exist a test collection for quantitatively evaluating this new retrieval task. In this paper, we first study the effectiveness of the state-of-the-art retrieval models for the app retrieval task using a new app retrieval test data we created. We then propose and study a novel approach that generates a new representation for each app. Our key idea is to leverage user reviews to find out important features of apps and bridge vocabulary gap between app developers and users. Specifically, we jointly model app descriptions and user reviews using topic model in order to generate app representations while excluding noise in reviews. Experiment results indicate that the proposed approach is effective and outperforms the state-of-the-art retrieval models for app retrieval.
Dae Hoon Park, Mengwen Liu, ChengXiang Zhai, Haohong Wang
SIGIR1
2011 Generate Adjective Sentiment Dictionary for Social Media Sentiment Analysis Using Constrained Nonnegative Matrix Factorization
Dae Hoon Park
ICWSM2