Yifan Hu 0001

dblp:92/498-1 · DBLP profile ↗
← Back
11ranked-venue papers in the field
2as first author
5since 2021 · last 2023
0000-0003-2017-924XORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 4Big Data, Cloud & Distributed Data Systems · 4Data Mining & Knowledge Discovery · 2 (2 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2023 CrypText: Database and Interactive Toolkit of Human-Written Text Perturbations in the Wild
abstract
User-generated textual contents on the Internet are often noisy, erroneous, and not in correct grammar. In fact, some online users choose to express their opinions online through carefully perturbed texts, especially in controversial topics (e.g., politics, vaccine mandate) or abusive contexts (e.g., cyberbullying, hate-speech). However, to the best of our knowledge, there is no framework that explores these online "human-written" perturbations (as opposed to algorithm-generated perturbations). Therefore, we introduce an interactive system called CrypText. CrypText is a data-intensive application that provides the users with a database and several tools to extract and interact with human-written perturbations. Specifically, CrypText helps look up, perturb, and normalize (i.e., de-perturb) texts. CrypText also provides an interactive interface to monitor and analyze text perturbations online. The demo is available at: https://lethaiq.github.io/anthro.
Thai Le, Yiran Ye, Yifan Hu 0001, Dongwon Lee 0001
ICDE3
2021 Hadoop-MTA: a system for Multi Data-center Trillion Concepts Auto-ML atop Hadoop
abstract
The ever-growing computation capability distributed infrastructure brings tremendous opportunities for mining and analysis of data that was impossible otherwise. Meanwhile, the inherent computation model of distributed system also brings unique and non-trivial challenges for traditional Auto-ML, including the explosion of data dimensions, the expected absence of features, and the heterogeneity of information. This is especially the case in modern Internet enterprises, where data in the scale of trillions are stored in multiple data centers, and the discovery of subtle signals could incur significant impact in revenue and welfare. How can we best harness the large scale distributed machine learning, but without keeping engineers constantly in the loop? In this work, we present Hadoop-MTA, a system for Multi Data-center, Trillion Concepts, Auto-ML on top of the Hadoop distributed computation environment that leverages sparsity aware heterogeneous knowledge graph representation and dimensionality agnostic parallel learning. Through multiple large scale experiments, we find that Hadoop-MTA significantly output-performs competitive state of the art distributed learning algorithms and scales well to trillion scale data-sets. Our model is rolled out to Hadoop serving infrastructure in Yahoo covering billions of unique identities and shows improvements 129.5% accuracy and 106.5 % weighted F1-score (more than 2x) on key targeting use cases.
Keqian Li, Yifan Hu 0001, Manisha Verma, Fei Tan 0002, Changwei Hu, Tejaswi Kasturi, Kevin Yen
IEEE BigData2
2021 BAN: Large Scale Brand ANonymization for Creative Recommendation via Label Light Adaptation
abstract
One of the primary component in ads creative recommendation system is the brand anonymization that removes brand-specific information from ad text for legal compliance and providing ready to use template for the advertisers to customize and consume. In our previous work [1] on ads creative recommendation system, the anonymization is done via a block list created solely based on manual reviewing, which is expensive and limits in the scale of the deployment of the ads recommendation. In this work we investigate a large scale, automated approach for brand anonymization. Such a problem presents many unique and non-trivial challenges, including the domain specificity of the brand entities, the fine-granularity requirements of structured output, the tight constraint of the limited contexts, the high level of grammatical noise in the advertisement data, and the heterogeneity of information required to perform anonymization. We propose a transformer model that leverage implicit knowledge together with a label-light adaptation procedure for this task. Our model is rolled out to ads systems in Yahoo that cover billions of impression traffic per month and improved previous production system by 68.3% F1-score on token level prediction and 61.6% on ad level prediction.
Keqian Li, Kevin Yen, Shaunak Mishra, Yifan Hu 0001, Changwei Hu, Manisha Verma
IEEE BigData4
2021 TSI: An Ad Text Strength Indicator using Text-to-CTR and Semantic-Ad-Similarity
abstract
Coming up with effective ad text is a time consuming process, and particularly challenging for small businesses with limited advertising experience. When an inexperienced advertiser onboards with a poorly written ad text, the ad platform has the opportunity to detect low performing ad text, and provide improvement suggestions. To realize this opportunity, we propose an ad text strength indicator (TSI) which: (i) predicts the click-through-rate (CTR) for an input ad text, (ii) fetches similar existing ads to create a neighborhood around the input ad, (iii) and compares the predicted CTRs in the neighborhood to declare whether the input ad is strong or weak. In addition, as suggestions for ad text improvement, TSI shows anonymized versions of superior ads (higher predicted CTR) in the neighborhood. For (i), we propose a BERT based text-to-CTR model trained on impressions and clicks associated with an ad text. For (ii), we propose a sentence-BERT based semantic-ad-similarity model trained using weak labels from ad campaign setup data. Offline experiments demonstrate that our BERT based text-to-CTR model achieves a significant lift in CTR prediction AUC for cold start (new) advertisers compared to bag-of-words based baselines. In addition, our semantic-textual-similarity model for similar ads retrieval achieves a [email protected] of 0.93 (for retrieving ads from the same product category); this is significantly higher compared to unsupervised TF-IDF, word2vec, and sentence-BERT baselines. Finally, we share promising online results from advertisers in the Yahoo (Verizon Media) ad platform where a variant of TSI was implemented with sub-second end-to-end latency.
Shaunak Mishra, Changwei Hu, Manisha Verma, Kevin Yen, Yifan Hu 0001, Maxim Sviridenko
CIKM5
2021 What's in a name? - gender classification of names with character based machine learning models
Yifan Hu 0001, Changwei Hu, Thanh Tran 0005, Tejaswi Kasturi, Elizabeth Joseph, Matt Gillingham
Data Min. Knowl. Discov.1
2019 ORIGIN: Non-Rigid Network Alignment
abstract
Network alignment is a fundamental task in many high-impact applications. Most of the existing approaches either explicitly or implicitly consider the alignment matrix as a linear transformation to map one network to another, and might overlook the complicated alignment relationship across networks. On the other hand, node representation learning based alignment methods are hampered by the incomparability among the node representations of different networks. In this paper, we propose a unified semi-supervised deep model (ORIGIN) that simultaneously finds the non-rigid network alignment and learns node representations in multiple networks in a mutually beneficial way. The key idea is to learn node representations by the effective graph convolutional networks, which subsequently enable us to formulate network alignment as a point set alignment problem. The proposed method offers two distinctive advantages. First (node representations), unlike the existing graph convolutional networks that aggregate the node information within a single network, we can effectively aggregate the auxiliary information from multiple sources, achieving far-reaching node representations. Second (network alignment), guided by the high-quality node representations, our proposed non-rigid point set alignment approach overcomes the bottleneck of the linear transformation assumption. We conduct extensive experiments that demonstrate the proposed non-rigid alignment method is (1) effective, outperforming both the state-of-the-art linear transformation-based methods and node representation based methods, and (2) efficient, with a comparable computational time between the proposed multi-network representation learning component and its single-network counterpart.
Hanghang Tong, Jiejun Xu, Yifan Hu 0001, Ross Maciejewski
IEEE BigData4
2019 Adaptive Feature Redundancy Minimization
abstract
Most existing feature selection methods select the top-ranked features according to certain criterion. However, without considering the redundancy among the features, the selected ones are frequently highly correlated with each other, which is detrimental to the performance. To tackle this problem, we propose a framework regarding adaptive redundancy minimization (ARM) for the feature selection. Unlike other feature selection methods, the proposed model has the following merits: (1) The redundancy matrix is adaptively constructed instead of presetting it as the priori information. (2) The proposed model could pick out the discriminative and non-redundant features via minimizing the global redundancy of the features. (3) ARM can reduce the redundancy of the features from both supervised and unsupervised perspectives.
Rui Zhang 0017, Hanghang Tong, Yifan Hu 0001
CIKM3
2017 Nationality Classification Using Name Embeddings
abstract
Nationality identification unlocks important demographic information, with many applications in biomedical and sociological research. Existing name-based nationality classifiers use name substrings as features and are trained on small, unrepresentative sets of labeled names, typically extracted from Wikipedia. As a result, these methods achieve limited performance and cannot support fine-grained classification.
Junting Ye, Shuchu Han, Yifan Hu 0001, Baris Coskun, Meizhu Liu, Hong Qin 0001, Steven Skiena
CIKM3
2013 CompactMap: A mental map preserving visual interface for streaming text data
abstract
As text streams become increasingly available from social media such as Facebook and Twitter, visual analysis of streaming text data is playing an important role in most business sectors. A fundamental challenge in visualizing a large amount of streaming text data is to preserve the user's mental map to enable tracking dynamic changes in topics, while simultaneously utilizing the display space efficiently. In this paper, we present CompactMap, an online visual interface that packs text clusters efficiently, with stable updates to maintain the user's mental map. It achieves spatiotemporally coherent layouts by dynamically matching clusters across time, and removing cluster overlaps according to spatial proximity and constraints. We developed a visual search engine based on CompactMaps for exploring a large amount of text streams in details on demand. We demonstrate the effectiveness of our approach in a controlled user study compared with a competing method.
Yifan Hu 0001, Stephen C. North, Han-Wei Shen
IEEE BigData2
2009 Putting recommendations on the map: visualizing clusters and relations
abstract
For users, recommendations can sometimes seem odd or counterintuitive. Visualizing recommendations can remove some of this mystery, showing how a recommendation is grouped with other choices. A drawing can also lead a user's eye to other options. Traditional 2D-embeddings of points can be used to create a basic layout, but these methods, by themselves, do not illustrate clusters and neighborhoods very well. In this paper, we propose the use of geographic maps to enhance the definition of clusters and neighborhoods, and consider the effectiveness of this approach in visualizing similarities and recommendations arising from TV shows.
Emden R. Gansner, Yifan Hu 0001, Stephen G. Kobourov, Chris Volinsky
RecSys2
2008 Collaborative Filtering for Implicit Feedback Datasets
abstract
A common task of recommender systems is to improve customer experience through personalized recommendations based on prior implicit feedback. These systems passively track different sorts of user behavior, such as purchase history, watching habits and browsing activity, in order to model user preferences. Unlike the much more extensively researched explicit feedback, we do not have any direct input from the users regarding their preferences. In particular, we lack substantial evidence on which products consumer dislike. In this work we identify unique properties of implicit feedback datasets. We propose treating the data as indication of positive and negative preference associated with vastly varying confidence levels. This leads to a factor model which is especially tailored for implicit feedback recommenders. We also suggest a scalable optimization procedure, which scales linearly with the data size. The algorithm is used successfully within a recommender system for television shows. It compares favorably with well tuned implementations of other known methods. In addition, we offer a novel way to give explanations to recommendations given by this factor model.
Yifan Hu 0001, Yehuda Koren, Chris Volinsky
ICDM1