Yinghui Yang 0001

dblp:70/7884 · also Yinghui (Catherine) Yang · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
0since 2021 · last 2015
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-authorArtificial intelligence and machine learning · 4 · 3 first-authorTheory of computation · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 57% Data stream processing · 19% Web and social media mining · 17%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing › frequency estimation
heavy hitter detection
0.112010
Comments on "an integrated efficient solution for computing frequent and top-k elements in data streams" · ACM Trans. Database Syst. 2010
Data mining
clustering
0.122005
GHIC: A Hierarchical Pattern-Based Clustering Algorithm for Grouping Web Transactions · IEEE Trans. Knowl. Data Eng. 2005
Segmenting Customer Transactions Using a Pattern-Based Clustering Approach · ICDM 2003
Data mining › clustering
pattern-based clustering
0.122005
GHIC: A Hierarchical Pattern-Based Clustering Algorithm for Grouping Web Transactions · IEEE Trans. Knowl. Data Eng. 2005
Segmenting Customer Transactions Using a Pattern-Based Clustering Approach · ICDM 2003
Data mining › clustering › categorical data clustering
transaction clustering
0.122005
GHIC: A Hierarchical Pattern-Based Clustering Algorithm for Grouping Web Transactions · IEEE Trans. Knowl. Data Eng. 2005
Segmenting Customer Transactions Using a Pattern-Based Clustering Approach · ICDM 2003
Web and social media mining
web usage mining
0.122005
GHIC: A Hierarchical Pattern-Based Clustering Algorithm for Grouping Web Transactions · IEEE Trans. Knowl. Data Eng. 2005
Mining web logs to improve website organization · WWW 2001
Algorithms and data structures › data streams
streaming algorithms
0.012010
Comments on "an integrated efficient solution for computing frequent and top-k elements in data streams" · ACM Trans. Database Syst. 2010
Information retrieval
web navigation
0.012001
Mining web logs to improve website organization · WWW 2001
Data mining
pattern mining
0.022005
GHIC: A Hierarchical Pattern-Based Clustering Algorithm for Grouping Web Transactions · IEEE Trans. Knowl. Data Eng. 2005
Segmenting Customer Transactions Using a Pattern-Based Clustering Approach · ICDM 2003
Data mining › pattern mining
itemset mining
0.012003
Segmenting Customer Transactions Using a Pattern-Based Clustering Approach · ICDM 2003
Graph data management
backtracking search
0.012001
Mining web logs to improve website organization · WWW 2001
Web and social media mining
user behavior analysis
0.012001
Mining web logs to improve website organization · WWW 2001

Methods — techniques the papers use, named apart from their topics

empirical evaluation · 0.2counterexample construction · 0.2objective function optimization · 0.1hierarchical clustering · 0.1pattern-based clustering · 0.0mixture model · 0.0web log mining · 0.0navigation link optimization · 0.0
YearPublicationVenuePosition
2015 Selective Domain Information Acquisition to Improve Segmentation Quality
abstract
It has been well established that adding domain information about whether certain data objects (for example customers in customer segmentation application) should belong to the same segment can improve the quality of the segments. However, it can be expensive to acquire such domain knowledge. Consequently, we need to limit the number of constrains we acquire and more importantly maximize the effectiveness of these limited number of constraints we can acquire from the experts. Many of the constrained clustering methods randomly select constraints, which have been shown in our experiments to be ineffective. In this paper, we define the problem of identifying the most informative constraints. We propose an algorithm which generates the most informative constraints by maximizing the information gain from the constraints. We conducted a set of experiments on various data sets to compare our method with two other methods according to two measurements: the accuracy rate and the Vector Quantization Error (VQE), which is the objective function k-means clustering method minimizes. We illustrated that our approach not only achieves better accuracy rates, but also maintains low VQE values. Our results suggest that businesses can enhance their segmentation quality greatly by actively acquire the right type of domain information.
Yinghui Yang 0001, Zijie Qi, Hongyan Liu 0002
ICEC1
2015 Optimal Budget Allocation Across Search Advertising Markets
abstract
One critical operational decision facing online advertisers when they engage in sponsored search advertising is concerned with the allocation of a limited advertising budget. In particular, dealing with multi-keyword search markets over multiple decision periods poses significant decision-making challenges. In this paper, we develop a novel budget allocation optimization model with multiple search advertising markets and a finite time horizon. One key element of our modeling work is developing a customized advertising response function when considering distinctive features of sponsored search, including the quality score and the dynamic advertising effort. We derive a feasible solution to our budget model and study its properties. Computational experiments are conducted on real-world data to evaluate our budget model and perform parameter sensitivity analysis. Experimental results indicate that our budget allocation strategy significantly outperforms several baseline strategies. In addition, the identified properties derived from the solution process illuminate critical managerial insights for advertisers in sponsored search.
Daniel Dajun Zeng, Yinghui Yang 0001, Jie Zhang 0116
INFORMS J. Comput.3
2014 A Tree-Based Contrast Set-Mining Approach to Detecting Group Differences
abstract
Understanding differences between groups in a data set is one of the fundamental tasks in data analysis. As relevant applications accumulate, data-mining methods have been developed to specifically address the problem of group difference detection. Contrast set mining discovers group differences in the form of conjunction of feature-value pairs or items. In this paper, we incorporate absolute difference, relative difference, and statistical significance in our definition of a group difference, and develop a novel method named DIFF that uses the prefix-tree structure to compress the search space, follows a tree traversal procedure to discover the complete set of significant group differences, and employs efficient pruning strategies to expedite the search process. We conducted comprehensive experiments to compare our method with existing methods on completeness of results, pruning efficiency, and computational efficiency. The experiments demonstrate that our method guarantees completeness of results and achieves higher pruning efficiency and computational efficiency compared to STUCCO. In addition, our definition of group difference is more general than STUCCO. Our method is more effective than traditional approaches, such as classification trees, in discovering the complete set of significant group differences.
Hongyan Liu 0002, Yinghui Yang 0001, Zhuohua Chen
INFORMS J. Comput.2
2013 Discovery of Online Shopping Patterns Across Websites
abstract
In the online world, customers can easily navigate to different online stores to make purchases. The products purchased on one site are often associated with product purchases on other sites (e.g., a hotel reservation on one site and a car rental on another site). Whereas market basket analysis is often used to discover associations among products for brick-and-mortar stores, it is rarely applied in the online setting where consumers navigate among different online stores to buy products. We define online shopping patterns and develop two novel methods to perform market basket analysis across websites. While this research is motivated by online shopping applications, our contribution is mainly methodological. The two methods we develop in this paper can not only be used to identify various online shopping patterns across sites and products but can also be applied to settings where patterns exist across different dimensions. Experiments on both synthetic data and real online shopping data demonstrate the effectiveness of our methods.
Yinghui Yang 0001, Hongyan Liu 0002, Yunjue Cai
INFORMS J. Comput.1
2012 Discovery of Periodic Patterns in Sequence Data: A Variance-Based Approach
abstract
We address the discovery of periodic patterns in sequence data. Building on prior work in this area, we present definitions and new methods for characterizing and identifying four types of periodic patterns. A unifying concept across the different types of periodic patterns we consider is the use of statistical variance to define periodicity. This lends itself to efficient variance-reduction algorithms for identifying periodic patterns. We motivate and test our approach using both extensive simulated sequences and real sequence data from online clickstream data.
Yinghui Yang 0001, Balaji Padmanabhan, Hongyan Liu 0002
INFORMS J. Comput.1
2011 Product selection for promotion planning
abstract
This paper addresses a very important question—how to select the right products to promote in order to maximize promotional benefit. We set up a framework to incorporate promotion decisions into the data-mining process, formulate the profit maximization problem as an optimization problem, and propose a heuristic search solution to discover the right products to promote. Moreover, we are able to get access to real supermarket data and apply our solution to help achieve higher profits. Our experimental results on both synthetic data and real supermarket data demonstrate that our framework and method are highly effective and can potentially bring huge profit gains to a marketing campaign.
Yinghui Yang 0001, Chunhui Hao
Knowl. Inf. Syst.1
2010 Cross-Selling Optimization for Customized Promotion
abstract
The profit of a retail product not only comes from its own sales, but also comes from its influence on the sales of other products. How to promote the right products to the right customers becomes one of the key issues in marketing. In this paper, we propose a new formulation of promotion value by considering cross-selling effects within selected products and customers, which were largely ignored by existing work. We investigate the problem of customized promotion, which identifies promotional products and customers so that the promotion effect can be maximized. This problem can be decomposed into two subproblems: product selection and customer selection. The baseline methods entail an exhaustive traversal of all possible product and customer combinations, which is computationally intractable. As an alternative, we propose greedy and randomized algorithms to produce approximation solutions in an efficient manner. Experiments on both synthetic and real-world supermarket transaction data demonstrate the effectiveness and efficiency of the proposed algorithms.
Yinghui Yang 0001, Xifeng Yan
SDM2
2010 Web user behavioral profiling for user identification
Yinghui Yang 0001
Decis. Support Syst.1
2010 Toward user patterns for online security: Observation time and online user identification
Yinghui Yang 0001, Balaji Padmanabhan
Decis. Support Syst.1
2010 Comments on "an integrated efficient solution for computing frequent and top-k elements in data streams"
abstract
We investigate a well-known algorithm, Space-Saving [Metwally et al. 2006], which has been proven efficient and effective at mining frequent elements in data streams. We discovered an error in one of the theorems in Metwally et al. [2006]. Experiments are conducted to illustrate the error.
Hongyan Liu 0002, Yinghui Yang 0001
ACM Trans. Database Syst.3
2005 GHIC: A Hierarchical Pattern-Based Clustering Algorithm for Grouping Web Transactions
abstract
Grouping customer transactions into segments may help understand customers better. The marketing literature has concentrated on identifying important segmentation variables (e.g., customer loyalty) and on using cluster analysis and mixture models for segmentation. The data mining literature has provided various clustering algorithms for segmentation without focusing specifically on clustering customer transactions. Building on the notion that observable customer transactions are generated by latent behavioral traits, in this paper, we investigate using a pattern-based clustering approach to grouping customer transactions. We define an objective function that we maximize in order to achieve a good clustering of customer transactions and present an algorithm, GHIC, that groups customer transactions such that itemsets generated from each cluster, while similar to each other, are different from ones generated from others. We present experimental results from user-centric Web usage data that demonstrates that GHIC generates a highly effective clustering of transactions.
Yinghui Yang 0001, Balaji Padmanabhan
IEEE Trans. Knowl. Data Eng.1
2003 On Original Generation of Structure in Legal Documents
abstract
This position paper advocates a vision in the development of automated legal reasoning and presents evidence supporting the plausibility of that vision. The paper observes that original creation of documents of legal import in either fully formal or semistructured form offers the prospect of greatly reducing the cost and expanding the scope of knowledge engineering for legal reasoning. This, it is claimed, is most likely to be achieved via formalization of various sublanguages of legal discourse. SeaSpeak is an example of such a sublanguage and it appears to be amenable to full formalization. Short of that, much can be done with partial formalization and semistructured documents. The paper presents a tabular format for message expression, motivated by a formal agent communication language.
Steven Orla Kimbrough, Thomas Y. Lee, Balaji Padmanabhan, Yinghui Yang 0001
ICAIL4
2003 Segmenting Customer Transactions Using a Pattern-Based Clustering Approach
abstract
Grouping customer transactions into categories helps understand customers better. The marketing literature has concentrated on identifying important segmentation variables (e.g. customer loyalty) and on using clustering and mixture models for segmentation. The data mining literature has provided various clustering algorithms for segmentation. We investigate using "pattern-based" clustering approaches to grouping customer transactions. We argue that there are clusters in transaction data based on natural behavioral patterns, and present a new technique, YACA, that groups transactions such that itemsets generated from each cluster, while similar to each other, are different from ones generated from others. We present experimental results from user-centric Web usage data that demonstrates that YACA generates a highly effective clustering of transactions.
Yinghui Yang 0001, Balaji Padmanabhan
ICDM1
2001 Mining web logs to improve website organization
abstract
Many websites have a hierarchical organization of content. This organization may be quite different from the organization expected by visitors to the website. In particular, it is often unclear where a specific document is located. In this paper, we propose an algorithm to automatically find pages in a website whose location is different from where visitors expect to find them. The key insight is that visitors will backtrack if they do not find the information where they expect it: the point from where they backtrack is the expected location for the page. We present an algorithm for discovering such expected locations that can handle page caching by the browser. Expected locations with a significant number of hits are then presented to the website administrator. We also present algorithms for selecting expected locations (for adding navigation links) to optimize the benefit to the website or the visitor. We ran our algorithm on the Wharton business school website and found that even on this small website, there were many pages with expected locations different from their actual location. 1.
Ramakrishnan Srikant, Yinghui Yang 0001
WWW2