Changwei Hu

dblp:44/8760 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
6since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2023 MGEL: Multigrained Representation Analysis and Ensemble Learning for Text Moderation
abstract
In this work, we describe our efforts in addressing two typical challenges involved in the popular text classification methods when they are applied to text moderation: the representation of multibyte characters and word obfuscations. Specifically, a multihot byte-level scheme is developed to significantly reduce the dimension of one-hot character-level encoding caused by the multiplicity of instance-scarce non-ASCII characters. In addition, we introduce a simple yet effective weighting approach for fusing n-gram features to empower the classical logistic regression. Surprisingly, it outperforms well-tuned representative neural networks greatly. As a continual effort toward text moderation, we endeavor to analyze the current state-of-the-art (SOTA) algorithm bidirectional encoder representations from transformers (BERT), which works well in context understanding but performs poorly on intentional word obfuscations. To resolve this crux, we then develop an enhanced variant and remedy this drawback by integrating byte and character decomposition. It advances the SOTA performance on the largest abusive language datasets as demonstrated by our comprehensive experiments. Our work offers a feasible and effective framework to tackle word obfuscations.
Fei Tan 0002, Changwei Hu, Yifan Hu 0001, Kevin Yen, Zhi Wei 0001, Aasish Pappu, Se Rim Park, Keqian Li
IEEE Trans. Neural Networks Learn. Syst.2
2021 Hadoop-MTA: a system for Multi Data-center Trillion Concepts Auto-ML atop Hadoop
abstract
The ever-growing computation capability distributed infrastructure brings tremendous opportunities for mining and analysis of data that was impossible otherwise. Meanwhile, the inherent computation model of distributed system also brings unique and non-trivial challenges for traditional Auto-ML, including the explosion of data dimensions, the expected absence of features, and the heterogeneity of information. This is especially the case in modern Internet enterprises, where data in the scale of trillions are stored in multiple data centers, and the discovery of subtle signals could incur significant impact in revenue and welfare. How can we best harness the large scale distributed machine learning, but without keeping engineers constantly in the loop? In this work, we present Hadoop-MTA, a system for Multi Data-center, Trillion Concepts, Auto-ML on top of the Hadoop distributed computation environment that leverages sparsity aware heterogeneous knowledge graph representation and dimensionality agnostic parallel learning. Through multiple large scale experiments, we find that Hadoop-MTA significantly output-performs competitive state of the art distributed learning algorithms and scales well to trillion scale data-sets. Our model is rolled out to Hadoop serving infrastructure in Yahoo covering billions of unique identities and shows improvements 129.5% accuracy and 106.5 % weighted F1-score (more than 2x) on key targeting use cases.
Keqian Li, Yifan Hu 0001, Manisha Verma, Fei Tan 0002, Changwei Hu, Tejaswi Kasturi, Kevin Yen
IEEE BigData5
2021 BAN: Large Scale Brand ANonymization for Creative Recommendation via Label Light Adaptation
abstract
One of the primary component in ads creative recommendation system is the brand anonymization that removes brand-specific information from ad text for legal compliance and providing ready to use template for the advertisers to customize and consume. In our previous work [1] on ads creative recommendation system, the anonymization is done via a block list created solely based on manual reviewing, which is expensive and limits in the scale of the deployment of the ads recommendation. In this work we investigate a large scale, automated approach for brand anonymization. Such a problem presents many unique and non-trivial challenges, including the domain specificity of the brand entities, the fine-granularity requirements of structured output, the tight constraint of the limited contexts, the high level of grammatical noise in the advertisement data, and the heterogeneity of information required to perform anonymization. We propose a transformer model that leverage implicit knowledge together with a label-light adaptation procedure for this task. Our model is rolled out to ads systems in Yahoo that cover billions of impression traffic per month and improved previous production system by 68.3% F1-score on token level prediction and 61.6% on ad level prediction.
Keqian Li, Kevin Yen, Shaunak Mishra, Yifan Hu 0001, Changwei Hu, Manisha Verma
IEEE BigData5
2021 TSI: An Ad Text Strength Indicator using Text-to-CTR and Semantic-Ad-Similarity
abstract
Coming up with effective ad text is a time consuming process, and particularly challenging for small businesses with limited advertising experience. When an inexperienced advertiser onboards with a poorly written ad text, the ad platform has the opportunity to detect low performing ad text, and provide improvement suggestions. To realize this opportunity, we propose an ad text strength indicator (TSI) which: (i) predicts the click-through-rate (CTR) for an input ad text, (ii) fetches similar existing ads to create a neighborhood around the input ad, (iii) and compares the predicted CTRs in the neighborhood to declare whether the input ad is strong or weak. In addition, as suggestions for ad text improvement, TSI shows anonymized versions of superior ads (higher predicted CTR) in the neighborhood. For (i), we propose a BERT based text-to-CTR model trained on impressions and clicks associated with an ad text. For (ii), we propose a sentence-BERT based semantic-ad-similarity model trained using weak labels from ad campaign setup data. Offline experiments demonstrate that our BERT based text-to-CTR model achieves a significant lift in CTR prediction AUC for cold start (new) advertisers compared to bag-of-words based baselines. In addition, our semantic-textual-similarity model for similar ads retrieval achieves a [email protected] of 0.93 (for retrieving ads from the same product category); this is significantly higher compared to unsupervised TF-IDF, word2vec, and sentence-BERT baselines. Finally, we share promising online results from advertisers in the Yahoo (Verizon Media) ad platform where a variant of TSI was implemented with sub-second end-to-end latency.
Shaunak Mishra, Changwei Hu, Manisha Verma, Kevin Yen, Yifan Hu 0001, Maxim Sviridenko
CIKM2
2021 BERT-Beta: A Proactive Probabilistic Approach to Text Moderation
abstract
Text moderation for user generated content, which helps to promote healthy interaction among users, has been widely studied and many machine learning models have been proposed.In this work, we explore an alternative perspective by augmenting reactive reviews with proactive forecasting.Specifically, we propose a new concept text toxicity propensity to characterize the extent to which a text tends to attract toxic comments.Beta regression is then introduced to do the probabilistic modeling, which is demonstrated to function well in comprehensive experiments.We also propose an explanation method to communicate the model decision clearly.Both propensity scoring and interpretation benefit text moderation in a novel manner.Finally, the proposed scaling mechanism for the linear model offers useful insights beyond this work.
Fei Tan 0002, Yifan Hu 0001, Kevin Yen, Changwei Hu
EMNLP (1)4
2021 What's in a name? - gender classification of names with character based machine learning models
Yifan Hu 0001, Changwei Hu, Thanh Tran 0005, Tejaswi Kasturi, Elizabeth Joseph, Matt Gillingham
Data Min. Knowl. Discov.2
2020 Repulsive Attention: Rethinking Multi-head Attention as Bayesian Inference
abstract
Bang An, Jie Lyu, Zhenyi Wang, Chunyuan Li, Changwei Hu, Fei Tan, Ruiyi Zhang, Yifan Hu, Changyou Chen. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Bang An 0001, Jie Lyu 0004, Zhenyi Wang 0001, Chunyuan Li, Changwei Hu, Fei Tan 0002, Ruiyi Zhang 0002, Yifan Hu 0001, Changyou Chen
EMNLP (1)5
2020 TNT: Text Normalization based Pre-training of Transformers for Content Moderation
abstract
In this work, we present a new language pre-training model TNT (Text Normalization based pre-training of Transformers) for content moderation.Inspired by the masking strategy and text normalization, TNT is developed to learn language representation by training transformers to reconstruct text from four operation types typically seen in text manipulation: substitution, transposition, deletion, and insertion.Furthermore, the normalization involves the prediction of both operation types and token labels, enabling TNT to learn from more challenging tasks than the standard task of masked word recovery.As a result, the experiments demonstrate that TNT outperforms strong baselines on the hate speech classification task.Additional text normalization experiments and case studies show that TNT is a new potential approach to misspelling correction.
Fei Tan 0002, Yifan Hu 0001, Changwei Hu, Keqian Li, Kevin Yen
EMNLP (1)3
2020 HABERTOR: An Efficient and Effective Deep Hatespeech Detector
abstract
We present our HABERTOR model for detecting hatespeech in large scale user-generated content.Inspired by the recent success of the BERT model, we propose several modifications to BERT to enhance the performance on the downstream hatespeech classification task.HABERTOR inherits BERT's architecture, but is different in four aspects: (i) it generates its own vocabularies and is pre-trained from the scratch using the largest scale hatespeech dataset; (ii) it consists of Quaternionbased factorized components, resulting in a much smaller number of parameters, faster training and inferencing, as well as less memory usage; (iii) it uses our proposed multisource ensemble heads with a pooling layer for separate input sources, to further enhance its effectiveness; and (iv) it uses a regularized adversarial training with our proposed finegrained and adaptive noise magnitude to enhance its robustness.Through experiments on the large-scale real-world hatespeech dataset with 1.4M annotated comments, we show that HABERTOR works better than 15 state-ofthe-art hatespeech detection methods, including fine-tuning Language Models.In particular, comparing with BERT, our HABERTOR is 4∼5 times faster in the training/inferencing phase, uses less than 1/3 of the memory, and has better performance, even though we pretrain it by using less than 1% of the number of words.Our generalizability analysis shows that HABERTOR transfers well to other unseen hatespeech datasets and is a more efficient and effective alternative to BERT for the hatespeech classification.
Thanh Tran 0005, Yifan Hu 0001, Changwei Hu, Kevin Yen, Fei Tan 0002, Kyumin Lee, Se Rim Park
EMNLP (1)3
2019 A Deep Structural Model for Analyzing Correlated Multivariate Time Series
abstract
Multivariate time series are routinely encountered in real-world applications, and in many cases, these time series are strongly correlated. In this paper, we present a deep learning structural time series model which can (i) handle correlated multivariate time series input, and (ii) forecast the targeted temporal sequence by explicitly learning/extracting the trend, seasonality, and event components. The trend is learned via a 1D and 2D temporal CNN and LSTM hierarchical neural net. The CNN-LSTM architecture can (i) seamlessly leverage the dependency among multiple correlated time series in a natural way, (ii) extract the weighted differencing feature for better trend learning, and (iii) memorize the long-term sequential pattern. The seasonality component is approximated via a non-liner function of a set of Fourier terms, and the event components are learned by a simple linear function of regressor encoding the event dates. We compare our model with several state-of-the-art methods through a comprehensive set of experiments on a variety of time series data sets, such as forecasts of Amazon AWS Simple Storage Service (S3) and Elastic Compute Cloud (EC2) billings, and the closing prices for corporate stocks in the same category.
Changwei Hu, Yifan Hu 0001, Sungyong Seo
ICMLA1
2019 Large-Scale Gender/Age Prediction of Tumblr Users
abstract
Tumblr, as a leading content provider and social media, attracts 371 million monthly visits, 280 million blogs and 53.3 million daily posts. The popularity of Tumblr provides great opportunities for advertisers to promote their products through sponsored posts. However, it is a challenging task to target specific demographic groups for ads, since Tumblr does not require user information like gender and ages during their registration. Hence, to promote ad targeting, it is essential to predict user's demography using rich content such as posts, images and social connections. In this paper, we propose graph based and deep learning models for age and gender predictions, which take into account user activities and content features. For graph based models, we come up with two approaches, network embedding and label propagation, to generate connection features as well as directly infer user's demography. For deep learning models, we leverage convolutional neural network (CNN) and multilayer perceptron (MLP) to prediction users' age and gender. Experimental results on real Tumblr daily dataset, with hundreds of millions of active users and billions of following relations, demonstrate that our approaches significantly outperform the baseline model, by improving the accuracy relatively by 81% for age, and the AUC and accuracy by 5% for gender.
Changwei Hu, Yifan Hu 0001, Tejaswi Kasturi, Shanmugam Ramasamy, Matt Gillingham, Keith Yamamoto
ICMLA2
2017 Deep Generative Models for Relational Data with Side Information
abstract
We present a probabilistic framework for overlapping community discovery and link prediction for relational data, given as a graph. The proposed framework has: (1) a deep architecture which enables us to infer multiple layers of latent features/communities for each node, providing superior link prediction performance on more complex networks and better interpretability of the latent features; and (2) a regression model which allows directly conditioning the node latent features on the side information available in form of node attributes. Our framework handles both (1) and (2) via a clean, unified model, which enjoys full local conjugacy via data augmentation, and facilitates efficient inference via closed form Gibbs sampling. Moreover, inference cost scales in the number of edges which is attractive for massive but sparse networks. Our framework is also easily extendable to model weighted networks with count-valued edges. We compare with various state-of-the-art methods and report results, both quantitative and qualitative, on several benchmark data sets.
Changwei Hu, Piyush Rai, Lawrence Carin
ICML1
2016 Non-negative Matrix Factorization for Discrete Data with Hierarchical Side-Information
abstract
We present a probabilistic framework for efficient non-negative matrix factorization of discrete (count/binary) data with side-information. The side-information is given as a multi-level structure, taxonomy, or ontology, with nodes at each level being categorical-valued observations. For example, when modeling documents with a two-level side-information (documents being at level-zero), level-one may represent (one or more) authors associated with each document and level-two may represent affiliations of each author. The model easily generalizes to more than two levels (or taxonomy/ontology of arbitrary depth). Our model can learn embeddings of entities present at each level in the data/side-information hierarchy (e.g., documents, authors, affiliations, in the previous example), with appropriate sharing of information across levels. The model also enjoys full local conjugacy, facilitating efficient Gibbs sampling for model inference. Inference cost scales in the number of non-zero entries in the data matrix, which is especially appealing for real-world massive but sparse matrices. We demonstrate the effectiveness of the model on several real-world data sets.
Changwei Hu, Piyush Rai, Lawrence Carin
AISTATS1
2016 Topic-Based Embeddings for Learning from Large Knowledge Graphs
abstract
We present a scalable probabilistic framework for learning from multi-relational data given in form of entity-relation-entity triplets, with a potentially massive number of entities and relations (e.g., in multi-relational networks, knowledge bases, etc.). We define each triplet via a relation-specific bilinear function of the embeddings of entities associated with it (these embeddings correspond to “topics”). To handle massive number of relations and the data sparsity problem (very few observations per relation), we also extend this model to allow sharing of parameters across relations, which leads to a substantial reduction in the number of parameters to be learned. In addition to yielding excellent predictive performance (e.g., for knowledge base completion tasks), the interpretability of our topic-based embedding framework enables easy qualitative analyses. Computational cost of our models scales in the number of positive triplets, which makes it easy to scale to massive real-world multi-relational data sets, which are usually extremely sparse. We develop simple-to-implement batch as well as online Gibbs sampling algorithms and demonstrate the effectiveness of our models on tasks such as multi-relational link-prediction, and learning from large knowledge bases.
Changwei Hu, Piyush Rai, Lawrence Carin
AISTATS1
2015 Scalable Probabilistic Tensor Factorization for Binary and Count Data
Piyush Rai, Changwei Hu, Matthew Harding, Lawrence Carin
IJCAI2
2015 Large-Scale Bayesian Multi-Label Learning via Topic-Based Label Embeddings
abstract
We present a scalable Bayesian multi-label learning model based on learning low-dimensional label embeddings. Our model assumes that each label vector is generated as a weighted combination of a set of topics (each topic being a distribution over labels), where the combination weights (i.e., the embeddings) for each label vector are conditioned on the observed feature vector. This construction, coupled with a Bernoulli-Poisson link function for each label of the binary label vector, leads to a model with a computational cost that scales in the number of positive labels in the label matrix. This makes the model particularly appealing for real-world multi-label learning problems where the label matrix is usually very massive but highly sparse. Using a data-augmentation strategy leads to full local conjugacy in our model, facilitating simple and very efficient Gibbs sampling, as well as an Expectation Maximization algorithm for inference. Also, predicting the label vector at test time does not require doing an inference for the label embeddings and can be done in closed form. We report results on several benchmark data sets, comparing our model with various state-of-the art methods.
Piyush Rai, Changwei Hu, Ricardo Henao, Lawrence Carin
NIPS2
2015 Scalable Bayesian Non-negative Tensor Factorization for Massive Count Data
Changwei Hu, Piyush Rai, Changyou Chen, Matthew Harding, Lawrence Carin
ECML/PKDD (2)1
2015 Zero-Truncated Poisson Tensor Factorization for Massive Binary Tensors
Changwei Hu, Piyush Rai, Lawrence Carin
UAI1
2015 Unlocking Smart Phone through Handwaving Biometrics
abstract
Screen locking/unlocking is important for modern smart phones to avoid the unintentional operations and secure the personal stuff. Once the phone is locked, the user should take a specific action or provide some secret information to unlock the phone. The existing unlocking approaches can be categorized into four groups: motion, password, pattern, and fingerprint. Existing approaches do not support smart phones well due to the deficiency of security, high cost, and poor usability. We collect 200 users' handwaving actions with their smart phones and discover an appealing observation: the waving pattern of a person is kind of unique, stable and distinguishable. In this paper, we propose OpenSesame, which employs the users' waving patterns for locking/unlocking. The key feature of our system lies in using four fine-grained and statistic features of handwaving to verify users. Moreover, we utilize support vector machine (SVM) for accurate and fast classification. Our technique is robust compatible across different brands of smart phones, without the need of any specialized hardware. Results from comprehensive experiments show that the mean false positive rate of OpenSesame is around 15 percent, while the false negative rate is lower than 8 percent.
Lei Yang 0025, Yi Guo 0008, Jinsong Han, Yunhao Liu 0001, Cheng Wang 0001, Changwei Hu
IEEE Trans. Mob. Comput.7
2014 Latent Gaussian Models for Topic Modeling
abstract
A new approach is proposed for topic modeling, in which the latent matrix factorization employs Gaussian priors, rather than the Dirichlet-class priors widely used in such models. The use of a latent-Gaussian model permits simple and efficient approximate Bayesian posterior inference, via the Laplace approximation. On multiple datasets, the proposed approach is demonstrated to yield results as accurate as state-of-the-art approaches based on Dirichlet constructions, at a small fraction of the computation. The framework is general enough to jointly model text and binary data, here demonstrated to produce accurate and fast results for joint analysis of voting rolls and the associated legislative text. Further, it is demonstrated how the technique may be scaled up to massive data, with encouraging performance relative to alternative methods.
Changwei Hu, Eunsu Ryu, David E. Carlson, Yingjian Wang 0004, Lawrence Carin
AISTATS1
2012 A GIS-supported impact assessment of the hierarchical flood-defense systems on the plain areas of the Taihu Basin, China
abstract
The Taihu Basin is located in the east coast of China, with a total area of 36,895 km2. Low-lying floodplain areas occupy about 83% of the basin. The threat of frequent floods to this economically important area has stimulated construction of enormous flood-defense projects along the complex system of rivers and lakes. Digital modeling of flooding processes and quantitative assessment of flood damages in this basin remain challenging due to the complexity. This article reports on an approach to simulate the flooding processes, which integrates hydrological and hydraulic modeling with dike-reliability analysis and socioeconomic information within a GIS platform. A new algorithm is introduced to calculate the influence of the flood-defense systems on spatial distributions of floodwater and consequential damages. Scenario analysis indicates that the modeling is particularly sensitive to the assumed rainfall, dike reliability, and the pump capacities within local polders. The model is validated by comparison with observations from historical flood records. The analysis reveals that the defense systems have significantly reduced the basin-wide flood risk and changed the spatial distributions of floodwater. Such a GIS-based approach can be potentially used to assess the benefit from construction of flood defenses and to avoid unintended spatial redistribution of flooding.
Chaoqing Yu, Xiaotao Cheng, Jim W. Hall, Edward P. Evans, Changwei Hu, Haoyun Wu, Jon Wicks, Mathew Scott, Minglei Ren, Zongxue Xu
Int. J. Geogr. Inf. Sci.6
2010 Compressed sensing MRI with combined sparsifying transforms and smoothed l0 norm minimization
abstract
Undersampling the k-space is an efficient way to speed up the magnetic resonance imaging (MRI). Recently emerged compressed sensing MRI shows promising results. However, most of them only enforce the sparsity of images in single transform, e.g. total variation, wavelet, etc. In this paper, based on the principle of basis pursuit, we propose a new framework to combine sparsifying transforms in compressed sensing MRI. Each transform can efficiently represent specific feature that the other can not. This framework is implemented via the state-of-art smoothed l0norm in overcomplete sparse decomposition. Simulation results demonstrate that the proposed method can improve image quality when comparing to single sparsifying transform.
Xiaobo Qu 0001, Xue Cao, Di Guo 0003, Changwei Hu, Zhong Chen 0005
ICASSP4