VLDB 2026 Research / reviewers in the wild / expert
Kevin Yen
dblp:276/6575
· DBLP profile ↗
9ranked-venue papers
0as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SmartGD: A GAN-Based Graph Drawing Framework for Diverse Aesthetic GoalsabstractWhile a multitude of studies have been conducted on graph drawing, many existing methods only focus on optimizing a single aesthetic aspect of graph layouts, which can lead to sub-optimal results. There are a few existing methods that have attempted to develop a flexible solution for optimizing different aesthetic aspects measured by different aesthetic criteria. Furthermore, thanks to the significant advance in deep learning techniques, several deep learning-based layout methods were proposed recently. These methods have demonstrated the advantages of deep learning approaches for graph drawing. However, none of these existing methods can be directly applied to optimizing non-differentiable criteria without special accommodation. In this work, we propose a novel Generative Adversarial Network (GAN) based deep learning framework for graph drawing, called SmartGD, which can optimize different quantitative aesthetic goals, regardless of their differentiability. To demonstrate the effectiveness and efficiency of SmartGD, we conducted experiments on minimizing stress, minimizing edge crossing, maximizing crossing angle, maximizing shape-based metrics, and a combination of multiple aesthetics. Compared with several popular graph drawing algorithms, the experimental results show that SmartGD achieves good performance both quantitatively and qualitatively. Kevin Yen, Yifan Hu 0001, Han-Wei Shen |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Automatic labeling of Parkinson's Disease gait videos with weak supervision
Mohsen Gholami, Rabab K. Ward, Ravneet Mahal, Maryam S. Mirian, Kevin Yen, Kye Won Park, Martin J. McKeown, Z. Jane Wang 0001 |
Medical Image Anal. | 5 |
| 2023 | MGEL: Multigrained Representation Analysis and Ensemble Learning for Text ModerationabstractIn this work, we describe our efforts in addressing two typical challenges involved in the popular text classification methods when they are applied to text moderation: the representation of multibyte characters and word obfuscations. Specifically, a multihot byte-level scheme is developed to significantly reduce the dimension of one-hot character-level encoding caused by the multiplicity of instance-scarce non-ASCII characters. In addition, we introduce a simple yet effective weighting approach for fusing n-gram features to empower the classical logistic regression. Surprisingly, it outperforms well-tuned representative neural networks greatly. As a continual effort toward text moderation, we endeavor to analyze the current state-of-the-art (SOTA) algorithm bidirectional encoder representations from transformers (BERT), which works well in context understanding but performs poorly on intentional word obfuscations. To resolve this crux, we then develop an enhanced variant and remedy this drawback by integrating byte and character decomposition. It advances the SOTA performance on the largest abusive language datasets as demonstrated by our comprehensive experiments. Our work offers a feasible and effective framework to tackle word obfuscations. Fei Tan 0002, Changwei Hu, Yifan Hu 0001, Kevin Yen, Zhi Wei 0001, Aasish Pappu, Se Rim Park, Keqian Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Hadoop-MTA: a system for Multi Data-center Trillion Concepts Auto-ML atop HadoopabstractThe ever-growing computation capability distributed infrastructure brings tremendous opportunities for mining and analysis of data that was impossible otherwise. Meanwhile, the inherent computation model of distributed system also brings unique and non-trivial challenges for traditional Auto-ML, including the explosion of data dimensions, the expected absence of features, and the heterogeneity of information. This is especially the case in modern Internet enterprises, where data in the scale of trillions are stored in multiple data centers, and the discovery of subtle signals could incur significant impact in revenue and welfare. How can we best harness the large scale distributed machine learning, but without keeping engineers constantly in the loop? In this work, we present Hadoop-MTA, a system for Multi Data-center, Trillion Concepts, Auto-ML on top of the Hadoop distributed computation environment that leverages sparsity aware heterogeneous knowledge graph representation and dimensionality agnostic parallel learning. Through multiple large scale experiments, we find that Hadoop-MTA significantly output-performs competitive state of the art distributed learning algorithms and scales well to trillion scale data-sets. Our model is rolled out to Hadoop serving infrastructure in Yahoo covering billions of unique identities and shows improvements 129.5% accuracy and 106.5 % weighted F1-score (more than 2x) on key targeting use cases. Keqian Li, Yifan Hu 0001, Manisha Verma, Fei Tan 0002, Changwei Hu, Tejaswi Kasturi, Kevin Yen |
IEEE BigData | 7 |
| 2021 | BAN: Large Scale Brand ANonymization for Creative Recommendation via Label Light AdaptationabstractOne of the primary component in ads creative recommendation system is the brand anonymization that removes brand-specific information from ad text for legal compliance and providing ready to use template for the advertisers to customize and consume. In our previous work [1] on ads creative recommendation system, the anonymization is done via a block list created solely based on manual reviewing, which is expensive and limits in the scale of the deployment of the ads recommendation. In this work we investigate a large scale, automated approach for brand anonymization. Such a problem presents many unique and non-trivial challenges, including the domain specificity of the brand entities, the fine-granularity requirements of structured output, the tight constraint of the limited contexts, the high level of grammatical noise in the advertisement data, and the heterogeneity of information required to perform anonymization. We propose a transformer model that leverage implicit knowledge together with a label-light adaptation procedure for this task. Our model is rolled out to ads systems in Yahoo that cover billions of impression traffic per month and improved previous production system by 68.3% F1-score on token level prediction and 61.6% on ad level prediction. Keqian Li, Kevin Yen, Shaunak Mishra, Yifan Hu 0001, Changwei Hu, Manisha Verma |
IEEE BigData | 2 |
| 2021 | TSI: An Ad Text Strength Indicator using Text-to-CTR and Semantic-Ad-SimilarityabstractComing up with effective ad text is a time consuming process, and particularly challenging for small businesses with limited advertising experience. When an inexperienced advertiser onboards with a poorly written ad text, the ad platform has the opportunity to detect low performing ad text, and provide improvement suggestions. To realize this opportunity, we propose an ad text strength indicator (TSI) which: (i) predicts the click-through-rate (CTR) for an input ad text, (ii) fetches similar existing ads to create a neighborhood around the input ad, (iii) and compares the predicted CTRs in the neighborhood to declare whether the input ad is strong or weak. In addition, as suggestions for ad text improvement, TSI shows anonymized versions of superior ads (higher predicted CTR) in the neighborhood. For (i), we propose a BERT based text-to-CTR model trained on impressions and clicks associated with an ad text. For (ii), we propose a sentence-BERT based semantic-ad-similarity model trained using weak labels from ad campaign setup data. Offline experiments demonstrate that our BERT based text-to-CTR model achieves a significant lift in CTR prediction AUC for cold start (new) advertisers compared to bag-of-words based baselines. In addition, our semantic-textual-similarity model for similar ads retrieval achieves a [email protected] of 0.93 (for retrieving ads from the same product category); this is significantly higher compared to unsupervised TF-IDF, word2vec, and sentence-BERT baselines. Finally, we share promising online results from advertisers in the Yahoo (Verizon Media) ad platform where a variant of TSI was implemented with sub-second end-to-end latency. Shaunak Mishra, Changwei Hu, Manisha Verma, Kevin Yen, Yifan Hu 0001, Maxim Sviridenko |
CIKM | 4 |
| 2021 | BERT-Beta: A Proactive Probabilistic Approach to Text ModerationabstractText moderation for user generated content, which helps to promote healthy interaction among users, has been widely studied and many machine learning models have been proposed.In this work, we explore an alternative perspective by augmenting reactive reviews with proactive forecasting.Specifically, we propose a new concept text toxicity propensity to characterize the extent to which a text tends to attract toxic comments.Beta regression is then introduced to do the probabilistic modeling, which is demonstrated to function well in comprehensive experiments.We also propose an explanation method to communicate the model decision clearly.Both propensity scoring and interpretation benefit text moderation in a novel manner.Finally, the proposed scaling mechanism for the linear model offers useful insights beyond this work. Fei Tan 0002, Yifan Hu 0001, Kevin Yen, Changwei Hu |
EMNLP (1) | 3 |
| 2020 | TNT: Text Normalization based Pre-training of Transformers for Content ModerationabstractIn this work, we present a new language pre-training model TNT (Text Normalization based pre-training of Transformers) for content moderation.Inspired by the masking strategy and text normalization, TNT is developed to learn language representation by training transformers to reconstruct text from four operation types typically seen in text manipulation: substitution, transposition, deletion, and insertion.Furthermore, the normalization involves the prediction of both operation types and token labels, enabling TNT to learn from more challenging tasks than the standard task of masked word recovery.As a result, the experiments demonstrate that TNT outperforms strong baselines on the hate speech classification task.Additional text normalization experiments and case studies show that TNT is a new potential approach to misspelling correction. Fei Tan 0002, Yifan Hu 0001, Changwei Hu, Keqian Li, Kevin Yen |
EMNLP (1) | 5 |
| 2020 | HABERTOR: An Efficient and Effective Deep Hatespeech DetectorabstractWe present our HABERTOR model for detecting hatespeech in large scale user-generated content.Inspired by the recent success of the BERT model, we propose several modifications to BERT to enhance the performance on the downstream hatespeech classification task.HABERTOR inherits BERT's architecture, but is different in four aspects: (i) it generates its own vocabularies and is pre-trained from the scratch using the largest scale hatespeech dataset; (ii) it consists of Quaternionbased factorized components, resulting in a much smaller number of parameters, faster training and inferencing, as well as less memory usage; (iii) it uses our proposed multisource ensemble heads with a pooling layer for separate input sources, to further enhance its effectiveness; and (iv) it uses a regularized adversarial training with our proposed finegrained and adaptive noise magnitude to enhance its robustness.Through experiments on the large-scale real-world hatespeech dataset with 1.4M annotated comments, we show that HABERTOR works better than 15 state-ofthe-art hatespeech detection methods, including fine-tuning Language Models.In particular, comparing with BERT, our HABERTOR is 4∼5 times faster in the training/inferencing phase, uses less than 1/3 of the memory, and has better performance, even though we pretrain it by using less than 1% of the number of words.Our generalizability analysis shows that HABERTOR transfers well to other unseen hatespeech datasets and is a more efficient and effective alternative to BERT for the hatespeech classification. Thanh Tran 0005, Yifan Hu 0001, Changwei Hu, Kevin Yen, Fei Tan 0002, Kyumin Lee, Se Rim Park |
EMNLP (1) | 4 |