Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dingyi Han

dblp:56/2645 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 1 first-authorArtificial intelligence and machine learning · 9 · 1 first-authorComputer networks · 5 · 1 first-authorSystems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Data mining · 56% Recommender systems · 16% Web and social media mining · 12%
Artificial intelligence
3 papers
Information extraction and text analysis · 78% Question answering and dialogue systems · 22%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › text mining
topic modeling
0.222011
Mining topics on participations for community discovery · SIGIR 2011
Joint Emotion-Topic Modeling for Social Affective Text Mining · ICDM 2009
Recommender systems
click-through rate prediction
0.212014
Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014
Information retrieval › online advertising
display advertising
0.212014
Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014
Data mining › dimensionality reduction
feature selection
0.212014
Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014
Data mining › dimensionality reduction › feature selection
group lasso
0.212014
Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.112012
Mining Social Emotions from Affective Text · IEEE Trans. Knowl. Data Eng. 2012
Natural language and speech › Information extraction and text analysis
topic model
0.112012
Mining Social Emotions from Affective Text · IEEE Trans. Knowl. Data Eng. 2012
Natural language and speech › Question answering and dialogue systems
community question answering
0.112011
Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services · AAAI 2011
Natural language and speech › Information extraction and text analysis
text classification
0.112011
Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services · AAAI 2011
Web and social media mining
co-authorship networks
0.112011
Mining topics on participations for community discovery · SIGIR 2011
Data mining › structured data mining › graph mining
community detection
0.112011
Mining topics on participations for community discovery · SIGIR 2011
Data mining › structured data mining
graph mining
0.112011
Mining topics on participations for community discovery · SIGIR 2011
Data integration and cleaning › data extraction
web data extraction
0.112007
Homepage live: automatic block tracing for web personalization · WWW 2007
Recommender systems
web personalization
0.112007
Homepage live: automatic block tracing for web personalization · WWW 2007
Data mining › text mining
text classification
0.012012
Mining Social Emotions from Affective Text · IEEE Trans. Knowl. Data Eng. 2012
Web and social media mining › user-generated content
user-generated content analysis
0.012011
Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services · AAAI 2011
Natural language and speech › Information extraction and text analysis
emotion recognition
0.012009
Joint Emotion-Topic Modeling for Social Affective Text Mining · ICDM 2009
Software maintenance and evolution › code change analysis
change detection
0.012007
Homepage live: automatic block tracing for web personalization · WWW 2007

Methods — techniques the papers use, named apart from their topics

latent dirichlet allocation · 0.5supervised learning · 0.2feature engineering · 0.2logistic regression · 0.2feature hashing · 0.2distributed implementation · 0.2tree edit distance · 0.1topical variable modeling · 0.1probabilistic graphical model · 0.1
YearPublicationVenuePosition
2015 Topic Modeling in Semantic Space with Keywords
abstract
A common and convenient approach for user to describe his information need is to provide a set of keywords. Therefore, the technique to understand the need becomes crucial. In this paper, for the information need about a topic or category, we propose a novel method called TDCS(Topic Distilling with Compressive Sensing) for explicit and accurate modeling the topic implied by several keywords. The task is transformed as a topic reconstruction problem in the semantic space with a reasonable intuition that the topic is sparse in the semantic space. The latent semantic space could be mined from documents via unsupervised methods, e.g. LSI. Compressive sensing is leveraged to obtain a sparse representation from only a few keywords. In order to make the distilled topic more robust, an iterative learning approach is adopted. The experiment results show the effectiveness of our method. Moreover, with only a few semantic concepts remained for the topic, our method is efficient for subsequent text mining tasks.
Xiaojia Pu, Rong Jin 0001, Gangshan Wu, Dingyi Han, Gui-Rong Xue
CIKM4
2014 Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising
abstract
In display advertising, click through rate(CTR) prediction is the problem of estimating the probability that an advertisement (ad) is clicked when displayed to a user in a specific context. Due to its easy implementation and promising performance, logistic regression(LR) model has been widely used for CTR prediction, especially in industrial systems. However, it is not easy for LR to capture the nonlinear information, such as the conjunction information, from user features and ad features. In this paper, we propose a novel model, called coupled group lasso(CGL), for CTR prediction in display advertising. CGL can seamlessly integrate the conjunction information from user features and ad features for modeling. Furthermore, CGL can automatically eliminate useless features for both users and ads, which may facilitate fast online prediction. Scalability of CGL is ensured through feature hashing and distributed implementation. Experimental results on real-world data sets show that our CGL model can achieve state-of-the-art performance on web-scale CTR prediction tasks.
Wu-Jun Li, Gui-Rong Xue, Dingyi Han
ICML4
2012 Mining Social Emotions from Affective Text
abstract
This paper is concerned with the problem of mining social emotions from text. Recently, with the fast development of web 2.0, more and more documents are assigned by social users with emotion labels such as happiness, sadness, and surprise. Such emotions can provide a new aspect for document categorization, and therefore help online users to select related documents based on their emotional preferences. Useful as it is, the ratio with manual emotion labels is still very tiny comparing to the huge amount of web/enterprise documents. In this paper, we aim to discover the connections between social emotions and affective terms and based on which predict the social emotion from text content automatically. More specifically, we propose a joint emotion-topic model by augmenting Latent Dirichlet Allocation with an additional layer for emotion modeling. It first generates a set of latent topics from emotions, followed by generating affective terms from each topic. Experimental results on an online news collection show that the proposed model can effectively identify meaningful latent topics for each emotion. Evaluation on emotion prediction further verifies the effectiveness of the proposed model.
Shenghua Bao, Shengliang Xu, Li Zhang 0007, Zhong Su, Dingyi Han, Yong Yu 0001
IEEE Trans. Knowl. Data Eng.6
2011 Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services
abstract
This paper focuses on analyzing and predicting not-answered questions in Community based Question Answering (CQA) services, such as Yahoo! Answers. In CQA services, users express their information needs by submitting natural language questions and await answers from other human users. Comparing to receiving results from web search engines using keyword queries, CQA users are likely to get more specific answers, because human answerers may catch the main point of the question. However, one of the key problems of this pattern is that sometimes no one helps to give answers, while web search engines hardly fail to response. In this paper, we analyze the not-answered questions and give a first try of predicting whether questions will receive answers. More specifically, we first analyze the questions of Yahoo Answers based on the features selected from different perspectives. Then, we formalize the prediction problem as supervised learning – binary classification problem and leverage the proposed features to make predictions. Extensive experiments are made on 76,251 questions collected from Yahoo! Answers. We analyze the specific characteristics of not-answered questions and try to suggest possible reasons for why a question is not likely to be answered. As for prediction, the experimental results show that classification based on the proposed features outperforms the simple word-based approach significantly.
Lichun Yang, Shenghua Bao, Qingliang Lin, Xian Wu 0001, Dingyi Han, Zhong Su, Yong Yu 0001
AAAI5
2011 Exemplar-based Robust Coherent Biclustering
abstract
The biclustering, co-clustering, or subspace clustering problem involves simultaneously grouping the rows and columns of a data matrix to uncover biclusters or sub-matrices of the data matrix that optimize a desired objective function. In coherent biclustering, the objective function contains a coherence measure of the biclusters. We introduce a novel formulation of the coherent biclustering problem and use it to derive two algorithms. The first algorithm is based on loopy message passing; and the second relies on a greedy strategy yielding an algorithm that is significantly faster than the first. A distinguishing feature of these algorithms is that they identify an exemplar or a prototypical member of each bicluster. We note the interference from background elements in biclustering, and offer a means to circumvent such interference using additional regularization. Our experiments with synthetic as well as real-world datasets show that our algorithms are competitive with the current state-of-the-art algorithms for finding coherent biclusters.
Kewei Tu, Xixiu Ouyang, Dingyi Han, Vasant G. Honavar
SDM3
2011 Mining topics on participations for community discovery
abstract
Community discovery on large-scale linked document corpora has been a hot research topic for decades. There are two types of links. The first one, which we call d2d-link, indicates connectiveness among different documents, such as blog references and research paper citations. The other one, which we call u2u-link, represents co-occurrences or simultaneous participations of different users in one document and typically each document from u2u-link corpus has more than one user/author. Examples of u2u-link data covers email archives and research paper co-authorship networks. Community discovery in d2d-link data has achieved much success, while methods for that in u2u-link data either make no use of the textual content of the documents or make oversimplified assumptions about the users and the textual content. In this paper we propose a general approach of community discovery for u2u-link data, i.e., multiple user data, by placing topical variables on multiple authors' participations in documents. Experiments on a research proceeding co-authorship corpus and a New York Times news corpus show the effectiveness of our model.
Guoqing Zheng, Jinwen Guo, Lichun Yang, Shengliang Xu, Shenghua Bao, Zhong Su, Dingyi Han, Yong Yu 0001
SIGIR7
2011 Finding Appropriate Experts for Collaboration
Zhenjiang Zhan, Lichun Yang, Shenghua Bao, Dingyi Han, Zhong Su, Yong Yu 0001
WAIM4
2010 A topical link model for community discovery in textual interaction graph
abstract
This paper is concerned with community discovery in textual interaction graph, where the links between entities are indicated by textual documents. Specifically, we propose a Topical Link Model(TLM), which leverages Hierarchical Dirichlet Process(HDP) to introduce hidden topical variable of the links. Other than the use of links, TLM can look into the documents on the links in detail to recover sound communities. Moreover, TLM is a nonparametric model, which is able to learn the number of communities from the data. Extensive experiments on two real world corpora show TLM outperforms two state-of-the-art baseline models, which verify the effectiveness of TLM in determining the proper number of communities and generating sound communities.
Guoqing Zheng, Jinwen Guo, Lichun Yang, Shengliang Xu, Shenghua Bao, Zhong Su, Dingyi Han, Yong Yu 0001
CIKM7
2009 A study of information retrieval on accumulative social descriptions using the generation features
abstract
This paper is concerned with the study of information retrieval (IR) on Accumulative Social Descriptions (ASDs). ASDs refer to Web texts that accumulated by many Web users describing certain Web resources, such as anchor texts, search logs and social annotations. There have been some studies working on leveraging ASDs for improving search performance in both internet and intranet. However, to the best of our knowledge, no prior study has concerned the specific generation features of ASDs, which are the focus point of this paper. Specifically, we consider the generation features from two perspectives, the generation processes and the generated distributions. Further, three probabilistic IR models are derived based on them. The three models are first demonstrated with one toy dataset and then empirically evaluated with two real datasets: an internet dataset consisting of 90,295 Web pages, with 25,845,818 social annotations crawled from Del.icio.us and 31,320,005 pieces of anchor texts crawled through Yahoo! API, and an intranet dataset consisting of 179,835 Web pages with 1,245,522 annotations dumped from the intranet tagging system in IBM, named as Dogear. Extensive experimental results show that the proposed methods, which fully leverage the generation features of ASDs, improve the performance of both internet and intranet search significantly.
Lichun Yang, Shengliang Xu, Shenghua Bao, Dingyi Han, Zhong Su, Yong Yu 0001
CIKM4
2009 Joint Emotion-Topic Modeling for Social Affective Text Mining
abstract
This paper is concerned with the problem of social affective text mining, which aims to discover the connections between social emotions and affective terms based on user-generated emotion labels. We propose a joint emotion-topic model by augmenting latent Dirichlet allocation with an additional layer for emotion modeling. It first generates a set of latent topics from emotions, followed by generating affective terms from each topic. Experimental results on an online news collection show that the proposed model can effectively identify meaningful latent topics for each emotion. Evaluation on emotion prediction further verifies the effectiveness of the proposed model.
Shenghua Bao, Shengliang Xu, Li Zhang 0007, Zhong Su, Dingyi Han, Yong Yu 0001
ICDM6
2008 Understanding and Summarizing Answers in Community-Based Question Answering Services
Yuanjie Liu, Yunbo Cao, Chin-Yew Lin, Dingyi Han, Yong Yu 0001
COLING5
2008 A unified model of pollution in P2P networks
abstract
Nowadays many popular Peer-to-Peer (P2P) systems suffer from the simultaneous attacks of various pollution, including file-targeted attack and index-targeted attack. However, to our knowledge, there is no model that takes both of them into consideration. In fact, the two attacks impact the effect of each other. It makes the models considering either kind of pollution only fail to accurately illustrate the actual pollution. In this paper, we develop a unified model to remedy the defect. Through the analysis from the perspective of user behavior, the two attacks are integrated into the unified model as two factors impacting users’ choice of the files to download. The modeled file proliferation processes are consistent to those measured in real P2P systems. Moreover, the co-effect of the two attacks is also analyzed. The extremum point of co-effect is found, which corresponds to the most efficient attack of pollution. Further analysis of the model’s accuracy requires the quantitative comparison between the modeled effects of pollution and the measured ones. Nonetheless, no such metric has ever been proposed, which also causes a lot of problems in evaluating the effect of pollution and anti-pollution techniques. To fix the deficiency, we propose several metrics to assess the effect of pollution, including abort ratio, average download time of unpolluted files, etc. These metrics estimate the effect of pollution from different aspects. They are applied to the analysis of pollution emulated by our unified model. The co-effect of pollution is captured by these metrics. Furthermore, the difference between our model and the previously developed ones is also reflected by them.
Dingyi Han, Xinyao Hu, Yong Yu 0001
IPDPS2
2008 SEM: Mining Spatial Events from the Web
Kaifeng Xu, Rui Li 0049, Shenghua Bao, Dingyi Han, Yong Yu 0001
PAKDD4
2008 A dynamic routing protocol for keyword search in unstructured peer-to-peer networks
Dingyi Han, Yuanjie Liu, Shicong Meng, Yong Yu 0001
Comput. Commun.2
2007 Predicting Query Duplication with Box-Jenkins Models and Its Applications
abstract
Self-propagating worms have been terrorizing the Internet for several years and they are becoming imminent threats to large-scale Peer-to-Peer (P2P) systems featuring rich host connectivity and popular data services. In this paper, we consider topological worms, which exploit P2P host vulnerabilities and topology information to spread in an ultra-fast way. We study the feasibility of leveraging the existing P2P overlay structure for distributing automated security patches to vulnerable machines. Two approaches are examined: a partition-based approach, which utilizes immunized hosts to proactively stop worm spread in the overlay graph, and a Connected Dominating Set(CDS)-based approach, which utilizes a group of dominating nodes in the overlay to achieve fast patch dissemination in a race with the worm. We demonstrate through analysis and simulations that both methods can result in effective worm containment.
Xinyao Hu, Shicong Meng, Dingyi Han
Peer-to-Peer Computing4
2007 Homepage live: automatic block tracing for web personalization
abstract
The emergence of personalized homepage services, e.g. personalized Google Homepage and Microsoft Windows Live, has enabled Web users to select Web contents of interest and to aggregate them in a single Web page. The web contents are often predefined content blocks provided by the service providers. However, it involves intensive manual efforts to define the content blocks and maintain the information in it. In this paper, we propose a novel personalized homepage system, called .Homepage Live., to allow end users to use drag-and-drop actions to collect their favorite Web content blocks from existing Web pages and organize them in a single page. Moreover, Homepage Live automatically traces the changes of blocks with the evolvement of the container pages by measuring the tree edit distance of the selected blocks. By exploiting the immutable elements of Web pages, the tracing algorithm performance is significantly improved. The experimental results demonstrate the effectiveness and efficiency of our algorithm.
Dingyi Han, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001
WWW2
2006 A Hierarchical Model of Web Graph
Yong Yu 0001, Dingyi Han, Gui-Rong Xue
ADMA4
2006 A Statistical Study of Today's Gnutella
Shicong Meng, Dingyi Han, Yong Yu 0001
APWeb3
2006 Cuckoo Ring: BalancingWorkload for Locality Sensitive Hash
abstract
Locality sensitive hash (LSH) is widely used in peer-to-peer (P2P) systems. Although it can support range or similarity queries, it breaks the load balance mechanism of traditional distributed hash table (DHT) based system by replacing consistent hash with LSH. To solve the imbalance problem, current systems either weaken the locality preserve ability from similarity preserved to order preserved or adopt load aware peer join mechanism. The first method does not support similarity query as it loses the similarity information and the second method is greatly affected by the dynamic nature of P2P networks. In this paper, we propose a novel system, cuckoo ring, which can preserve similarity information while load balanced. It does not guide the newly joining peer to the hot areas but move the items in the hot areas to cold areas so that the short life time peers are distributed uniformly across the network instead of being guided to the hot areas. Compared to traditional DHT systems, cuckoo ring only maintains a little more information about the global light load peers and the moved indexed items
Dingyi Han, Ting Shen, Shicong Meng, Yong Yu 0001
Peer-to-Peer Computing1
2006 Reinforcement Learning for Query-Oriented Routing Indices in Unstructured Peer-to-Peer Networks
abstract
The idea of building query-oriented routing indices has changed the way of improving routing efficiency from the basis as it can learn the content distribution during the query routing process. It gradually improves routing efficiency with no excessive network overhead of the routing index construction and maintenance. However, the previously proposed mechanism is not practically effective due to the slow improvement of routing efficiency. In this paper, we propose a novel mechanism for query-oriented routing indices which quickly achieves high routing efficiency at low cost. The maintenance method employs reinforcement learning to utilize mass peer behaviors to construct and maintain routing indices. It explicitly uses the expected value of returned content number to depict the content distribution, which helps quickly approximate the real distribution. Meanwhile, the routing method is to retrieve as many contents as possible. It also helps speed up the learning process further. The experimental evaluation shows that the mechanism has high routing efficiency, quick learning ability and satisfactory performance under churn
Shicong Meng, Yuanjie Liu, Dingyi Han, Yong Yu 0001
Peer-to-Peer Computing4
2006 An Empirical Study of Data Smoothing Methods for Memory-Based and Hybrid Collaborative Filtering
Dingyi Han, Gui-Rong Xue, Yong Yu 0001
PRICAI1
2006 Exploiting Rating Behaviors for Effective Collaborative Filtering
Dingyi Han, Yong Yu 0001, Gui-Rong Xue
WISE1
2005 An Effective Resource Description Based Approach to Find Similar Peers
abstract
In text document sharing peer-to-peer (P2P) applications, similar peers are peers which share documents about the same topic. Finding similar peers can benefit document retrieval in P2P systems. In this paper, the authors suggested an effective resource description based approach to find similar peers. By combining the topic model (an extension of language model) technique and fuzzy set theory, the key problems about resource description generation and peer similarity calculation were solved. Experiments performed on the standard data sets prove that this approach is effective. Especially, comparing with the traditional approach, experimental results in the simulated P2P environment manifest that this approach works much better when only local information is available.
Dingyi Han, Weibin Zhu, Yong Yu 0001
Peer-to-Peer Computing2
2005 A Novel Resource Description Based Approach for Clustering Peers
Dingyi Han, Yong Yu 0001, Weibin Zhu
WAIM2