Ying Li 0040

dblp:22/1805-40 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
1since 2021 · last 2022
0000-0003-3576-5698ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 1 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 Inventory Purchase Recommendation for Merchants in Traditional FMCG Retail Business
abstract
Small and micro merchants in the traditional FastMoving Consumer Goods (FMCG) retail business play a vital role in developing economies. In Indonesia, 82% of the population gets their day-to-day groceries from traditional trade merchants. Our work aims to help these grocery merchants prosper by increasing their sales and improving their inventory turnover cycle. To this end, we conduct an in-depth analysis of merchants’ inventory-buying behaviors from a transaction record dataset and propose several item recommendation methods for merchants’ inventory based on insights gathered from the data. For analysis, we investigate the stock-keeping unit (SKU) coverage, the SKU selling speed, merchant repeat purchase patterns, the growth rate of monthly spending, as well as their correlations with total monetary spending by merchants. Our recommendation methods consider recency and utility of SKU selling speed, and are implemented using popular item-based recommendation and content-based filtering algorithms. The inventory recommendation system was deployed and tested in production. Initial results show a lift in click-through rate (CTR) and also show that merchants are accepting the recommendations of SKU items for their inventory that they had never bought before.
Ying Li 0040, Muhammad Daffa Robani, Victor Suciu 0002, Joy He-Yueya
DSAA1
2020 Acoustic Measures for Real-Time Voice Coaching
abstract
Our voices can convey many different types of thoughts and intent; how our voices carry them is often not consciously controlled and as a consequence, unintended effects may arise that negatively impact our relationships. How we say things is as important as what we say. This paper presents methodologies for computing a set of physical properties from sound waves of a speaker's voice directly, referred to as acoustic measures. Experiments are designed and conducted to establish the correlations between physical properties and auditory measures for human perception of sound waves. Based on these correlations, a voice coaching app can guide users, in real-time or deferred retrospective, to modify their speech's auditory measures, such as rate of speech, energy level, and intonation, to achieve their intended communication goals.
Ying Li 0040, Abraham Miller, Arthur Liu, Kyle Coburn, Luis J. Salazar
KDD1
2020 Domain Specific Knowledge Graphs as a Service to the Public: Powering Social-Impact Funding in the US
abstract
Web and mobile technologies enable ubiquitous access to information. Yet, it is getting harder, even for subject matter experts, to quickly identify quality, trustworthy, and reliable content available online through search engines powered by advanced knowledge graphs. This paper explores the practical applications of Domain Specific Knowledge Graphs that allow for the extraction of information from trusted published and unpublished sources, to map the extracted information to an ontology defined in collaboration with sector experts, and to enable the public to go from single queries into ongoing conversations meeting their knowledge needs reliably. We focused on Social-Impact Funding, an area of need for over one million nonprofit organizations, foundations, government entities, social entrepreneurs, impact investors, and academic institutions in the US.
Ying Li 0040, Vitalii Zakhozhyi, Daniel Zhu, Luis J. Salazar
KDD1
2018 Improved Localisation Using Spatio-Temporal Data from Cellular Network
abstract
Localisation of mobile devices has been a topic of academic research and industry practice for solving various application problems, examples can be footfall counting and profiling used for location based digital or physical advertising, crowd monitoring for public security, emergency handling, transport measurement and management, etc. Often the solutions for localisation require large scale networks hardware and/or software upgrades, which can be very costly. However, we note the fact that many commercial use cases actually do not require very high resolution of localisation and satisfactory level of accuracy may be sufficient for attaining business decision quality. A reasonable trade-off between achieving business value and minimising additional costs on network equipment purchase and maintenance is to build solutions that rely only on telco-network data and utilize data mining methods to improve the localisation accuracy of mobile devices that carried by subscribers. In this work, we aim to achieve acceptable accuracy for localisation at the resolution of region of interest (ROI), the exact shape of which is defined according to business requirements. One example is the geographical division of planning sub-zone in Singapore. We make use of the Global Positioning System (GPS) locations extracted from mobile broadband log that contain the longitude and latitude of the subscriber to annotate the telco-network data. We experimented with three learning models: maximum likelihood estimation, dominant serving ROI, and random forest, along with the baseline of localisation based on cellular tower locations. The experiment results demonstrate the effectiveness of the proposed models and demonstrate accuracy improvement from baseline of 37.8% (naive cellular tower localisation) to 78.4% (random forest classification).
Shixin Luo, Yibin Ng, Terence Zheng Wei Lim, Cliff Choon Hua Tan, Nannan He, Giuseppe Manai, Ying Li 0040
MDM7
2017 Mobility Genome™- A Framework for Mobility Intelligence from Large-Scale Spatio-Temporal Data
abstract
Massive amount of spatio-temporal data is generated as a result of user movement and data access activities, from both smart mobile devices and network infrastructure. We aim to be able to derive mobility intelligence from spatiotemporal data at scale with efficiency and yet meet the needs of diverse applications. In this paper, we propose Mobility GenomeTM, a computational framework that enables efficient and extensible discovery of mobility intelligence from large-scale spatio-temporal data. The framework is organised as multiple layers composed of fundamental and extensible computational units. We describe several algorithms and models, such as Human Daily Activity to demonstrate the derivation of mobility intelligence from these units, together with validation results. The framework has been integrated into our Mobility Intelligence platform, processing hundreds of millions of records per day in our production environment. We also show the possibility of building more advanced applications on top of the framework such as detection of home and work location, origin-destination trips and footfall analysis. With its deployment, the framework helps us avoid duplicated efforts on repeating common data processing and algorithmic tasks while serving a diverse set of mobility intelligence applications.
The Anh Dang, Jayakumaran Deepak, Shixin Luo, Yunye Jin, Yibin Ng, Aloysius Lim, Ying Li 0040
DSAA8
2016 Classification of voices that elicit soothing effect by applying a voiced vs. unvoiced feature engineering strategy
abstract
This paper introduces a novel approach of classifying voices that elicit a soothing effect on listeners from a domain knowledge inspired application of feature engineering. In particular, we utilize the characteristics of voiced vs, unvoiced speech in order to build a more accurate feature set. Large sets of training data are prepared and disciplined feature selections are conducted. Our final classifier achieved 86.84% classification accuracy of cross validation and evaluations by unknown listener population via crowdsourcing have rates of agreement with the classification model range from 80% to 90%. The technologies are deployed into Jobaline products to help service companies identify hourly-job workers whose voice can elicit soothing effect on customers.
Ying Li 0040, Kevin Mueller, Jose D. Contreras, Luis J. Salazar
ICASSP1
2015 Predicting Voice Elicited Emotions
abstract
We present the research, and product development and deployment, of Voice Analyzer' by Jobaline Inc. This is a patent pending technology that analyzes voice data and predicts human emotions elicited by the paralinguistic elements of a voice. Human voice characteristics, such as tone, complement the verbal communication. In several contexts of communication, "how" things are said is just as important as "what" is being said. This paper provides an overview of our deployed system, the raw data, the data processing steps, and the prediction algorithms we experimented with. A case study is included where, given a voice clip, our model predicts the degree in which a listener will find the voice "engaging". Our prediction results were verified through independent market research with 75% in agreement on how an average listener would feel. One application of Jobaline Voice Analyzer technology is for assisting companies to hire workers in the service industry where customers' emotional response to workers' voice may affect the service outcome. Jobaline Voice Analyzer is deployed in production as a product offer to our clients to help them identify workers who will better engage with their customers. We will also share some discoveries and lessons learned.
Ying Li 0040, Jose D. Contreras, Luis J. Salazar
KDD1
2010 Learning to rank audience for behavioral targeting
abstract
Behavioral Targeting (BT) is a recent trend of online advertising market. However, some classical BT solutions, which predefine the user segments for BT ads delivery, are sometimes too large to numerous long-tail advertisers, who cannot afford to buy any large user segments due to budget consideration. In this extend abstract, we propose to rank users according to their probability of interest in an advertisement in a learning to rank framework. We propose to extract three types of features between user behaviors such as search queries, ad click history etc and the ad content provided by advertisers. Through this way, a long-tail advertiser can select a certain number of top ranked users as needed from the user segments for ads delivery. In the experiments, we use a 30-days' ad click-through log from a commercial search engine. The results show that using our proposed features under a learning to rank framework, we can well rank users who potentially interest in an advertisement.
Ning Liu 0001, Jun Yan 0001, Dou Shen, Depin Chen, Zheng Chen 0001, Ying Li 0040
SIGIR6
2009 Product query classification
abstract
Web query classification is an effective way to understand Web user intents, which can further improve Web search and online advertising relevance. However, Web queries are usually very short which cannot fully reflect their meanings. What is more, it is quite hard to obtain enough training data for training accurate classifiers. Therefore, previous work on query classification has focused on two issues. One is how to represent Web queries through query expansion. The other is how to increase the amount of training data. In this paper, we took product query classification as an example, which is to classify Web queries into a predefined product taxonomy, and systematically studied the impact of query expansion and the size of training data. We proposed two methods of enriching Web queries and three approaches of collecting training data. Thereafter, we conducted a series of experiments to compare the classification performance of using different combinations of training data and query representations over a real data set. The data set consists of hundreds of thousands queries collected from a popular commercial search engine. From the experiments, we found some interesting observations, which were not discussed before. Finally, we proposed an effective and efficient product query classification method based on our observations.
Dou Shen, Ying Li 0040, Xiao Li 0006, Dengyong Zhou
CIKM2
2009 Exploiting term relationship to boost text classification
abstract
Document classification provides an effective way to handle the explosive online textual data. However, in practical classification settings, we face the so-called feature sparsity problem caused by a lack of training documents or the shortness of text to be classified. In this paper, we solve the sparsity problem by exploiting term relationships along with Naive Bayes classifiers. The first method is to estimate term relationships based on the co-occurrence information of two terms in a certain context. The second method estimates the term relationships based on the distribution of terms over different hierarchical categories in a publicly available document taxonomy. Thereafter, term relationship is used to augment Naive Bayes classifiers. We test our methods on two open-domain data sets to demonstrate its advantages. The experimental results show that our method can significantly improve the classification performance, especially when we do not have enough training data or the texts are Web search queries.
Dou Shen, Jianmin Wu, Bin Cao 0001, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001, Ying Li 0040
CIKM7
2008 Personal name classification in web queries
abstract
Personal names are an important kind of Web queries in Web search, and yet they are special in many ways. Strategies for retrieving information on personal names should therefore be different from the strategies for other types of queries. To improve the search quality for personal names, a first step is to detect whether a query is a personal name. Despite the importance of this problem, relatively little previous research has been done on this topic. Since Web queries are usually short, conventional supervised machine-learning algorithms cannot be applied directly. An alternative is to apply some heuristic rules coupled with name-term dictionaries. However, when the dictionaries are small, this method tends to make false negatives; when the dictionaries are large, it tends to generate false positives. A more serious problem is that this method cannot provide a good trade-off between precision and recall. To solve these problems, we propose an approach based on the construction of probabilistic name-term dictionaries and personal name grammars, and use this algorithm to predict the probability of a query to be a personal name. In this paper, we develop four different methods for building probabilistic name-term dictionaries in which a term is assigned with a probability value of the term being a name term. We compared our approach with baseline algorithms such as dictionary-based look-up methods and supervised classification algorithms including logistic regression and SVM on some manually labeled test sets. The results validate the effectiveness of our approach, whose F1 value is more than 79.8%, which outperforms the best baseline by more than 11.3%
Dou Shen, Toby Walker, Zijian Zheng 0002, Qiang Yang 0001, Ying Li 0040
WSDM5
2006 Similarity of Temporal Query Logs Based on ARIMA Model
abstract
A challenging issue faced by modern information retrieval is that of determining and satisfying users' requirements relying only on very short text queries. In this paper, we propose an algorithm to find out related queries based on Auto-Regressive Integrated Moving Average (ARIMA) Model. First, we select and estimate ARIMA model of the temporal query logs. And then each query is denoted by a sequence of coefficients. We use the correlation of ARIMA coefficients as the similarity measurement. We call it as the ARIMA Temporal Similarity (ARIMA TS). This similarity describes how strongly two time series are linearly related. On the other hand, the ARIMA model could also be treated as a dimensionality reduction procedure. It can save storage space for a large database of the query logs. In addition, ARIMA model could be used as a tool to predict the trend of a query. The experimental results on two query logs of MSN search engine 1 demonstrate that the proposed approach can achieve better similarity measurement efficiently.
Ning Liu 0001, Shuzhen Nong, Jun Yan 0001, Benyu Zhang, Zheng Chen 0001, Ying Li 0040
ICDM6
2006 Detecting online commercial intention (OCI)
abstract
Understanding goals and preferences behind a user's online activities can greatly help information providers, such as search engine and E-Commerce web sites, to personalize contents and thus improve user satisfaction. Understanding a user's intention could also provide other business advantages to information providers. For example, information providers can decide whether to display commercial content based on user's intent to purchase. Previous work on Web search defines three major types of user search goals for search queries: navigational, informational and transactional or resource [1][7]. In this paper, we focus our attention on capturing commercial intention from search queries and Web pages, i.e., when a user submits the query or browse a Web page, whether he/she is about to commit or in the middle of a commercial activity, such as purchase, auction, selling, paid service, etc. We call the commercial intentions behind a user's online activities as OCI (Online Commercial Intention). We also propose the notion of "Commercial Activity Phase" (CAP), which identifies in which phase a user is in his/her commercial activities: Research or Commit. We present the framework of building machine learning models to learn OCI based on any Web page content. Based on that framework, we build models to detect OCI from search queries and Web pages. We train machine learning models from two types of data sources for a given search query: content of algorithmic search result page(s) and contents of top sites returned by a search engine. Our experiments show that the model based on the first data source achieved better performance. We also discover that frequent queries are more likely to have commercial intention. Finally we propose our future work in learning richer commercial intention behind users' online activities.
Honghua (Kathy) Dai, Zaiqing Nie, Ji-Rong Wen, Lee Wang, Ying Li 0040
WWW6
2002 Visually Mining Web User Clickpaths
abstract
As powerful as clickpath mining methods can be, they often lead to huge incomprehensible and non-interesting result sets. Our clickpath mining practice at MSN was faced with challenges of keeping analysts closer to the data exploration process, revealing powerful insight from clickpath mining that business owners can directly act upon. These challenges stressed the importance of an interactive and visual representation of clickpath mining results. Most products today that can perform clickpath visualization do so by presenting massive cross-weaving web graphs. We present a new type of clickpath visualization which focuses only on clickpaths of interest, simplifying the visualization space while still retaining the same degree of mineable knowledge in the data. We also describe visualization techniques we have used to enhance the detection of interesting clickpath patterns from data, and provide a real-life case study that has benefited from the use of our implemented clickpath visualizer PAVE.
Teresa Mah, Ying Li 0040
ICDM2
2001 Funnel report mining for the MSN network
abstract
Data mining research has long concentrated on the five main areas: clustering, association discovery, classification, forecasting and sequential patterns. Web data mining projects are concerned mainly with text mining, user segmentation, forecasting web usage and analyzing users' clickstream patterns. We present a new type of web usage mining called funnel analysis or funnel report mining. A funnel report is a study of the retention behavior among a series of pages or sites. For example, of all hits on the home page of www.msn.com, what percentages of those are followed by hits to moneycentral.msn.com? What percentage of www.msn.com hits are followed by moneycentral.msn.com, and then www.msnbc.com? What are the most interesting funnels starting with www.msn.com? Where does the greatest drop off rate occur after a user has hit MSNBC? Funnel reports are extremely useful in e-business because they give product planners an idea of how usable and well-structured their site is. From our experience performing web usage mining for the MSN network of sites, funnel reports are requested even more than user segmentation analyses, site affiliation studies and classification exercises. In this paper, we define a framework for funnel analysis and provide a tree-based solution we have been using successfully to extract all relevant funnels using only one scan of the data file.
Teresa Mah, Hank Hoek, Ying Li 0040
KDD3