VLDB 2026 Research / reviewers in the wild / expert
Fei Tan 0002
dblp:80/8352-2
· DBLP profile ↗
21ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-3232-1912ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced PolicyabstractLarge language models have achieved substantial progress in mathematical reasoning, yet their advancement is limited by the scarcity of high-quality, high-difficulty training data. Existing synthesis methods largely rely on transforming human-written templates, limiting both diversity and scalability. We propose MathSmith, a novel framework for synthesizing challenging mathematical problems to enhance LLM reasoning. Rather than modifying existing problems, MathSmith constructs new ones from scratch by randomly sampling concept–explanation pairs from PlanetMath, ensuring data independence and avoiding contamination. To increase difficulty, we design nine predefined strategies as soft constraints during rationales. We further adopts reinforcement learning to jointly optimize structural validity, reasoning complexity, and answer consistency. The length of the reasoning trace generated under autoregressive prompting is used to reflect cognitive complexity, encouraging the creation of more demanding problems aligned with long-chain-of-thought reasoning. Experiments across five benchmarks, categorized as easy & medium (GSM8K, MATH-500) and hard (AIME2024, AIME2025, OlympiadBench), show that MathSmith consistently outperforms existing baselines under both short and long CoT settings. Additionally, a weakness-focused variant generation module enables targeted improvement on specific concepts. Overall, MathSmith exhibits strong scalability, generalization, and transferability, highlighting the promise of high-difficulty synthetic data in advancing LLM reasoning capabilities. Shaoxiong Zhan, Yanlin Lai, Dahua Lin, Ziqing Yang 0006, Fei Tan 0002 |
AAAI | 6 |
| 2025 | ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuningabstractFine-tuning multi-turn dialogue systems requires high-quality supervision but often suffers from degraded performance when exposed to low-quality data.Supervision errors in early turns can propagate across subsequent turns, undermining coherence and response quality.Existing methods typically address data quality via static prefiltering, which decouples quality control from training and fails to mitigate turn-level error propagation.In this context, we propose ReSURE (Regularizing Supervision UnREliability), an adaptive learning method that dynamically downweights unreliable supervision without explicit filtering.ReSURE estimates per-turn loss distributions using Welford's online statistics and reweights sample losses on the fly accordingly.Experiments on both singlesource and mixed-quality datasets show improved stability and response quality.Notably, ReSURE enjoys positive Spearman correlations (0.21 ∼ 1.0 across multiple benchmarks) between response scores and number of samples regardless of data quality, which potentially paves the way for utilizing large-scale data effectively. Yiming Du, Yifan Xiang, Bin Liang 0004, Dahua Lin, Kam-Fai Wong, Fei Tan 0002 |
EMNLP | 6 |
| 2024 | SDA: Simple Discrete Augmentation for Contrastive Sentence Representation LearningabstractContrastive learning has recently achieved compelling performance in unsupervised sentence representation. As an essential element, data augmentation protocols, however, have not been well explored. The pioneering work SimCSE resorting to a simple dropout mechanism (viewed as continuous augmentation) surprisingly dominates discrete augmentations such as cropping, word deletion, and synonym replacement as reported. To understand the underlying rationales, we revisit existing approaches and attempt to hypothesize the desiderata of reasonable data augmentation methods: balance of semantic consistency and expression diversity. We then develop three simple yet effective discrete sentence augmentation schemes: punctuation insertion, modal verbs, and double negation. They act as minimal noises at lexical level to produce diverse forms of sentences. Furthermore, standard negation is capitalized on to generate negative samples for alleviating feature suppression involved in contrastive learning. We experimented extensively with semantic textual similarity on diverse datasets. The results support the superiority of the proposed methods consistently. Our key code is available at https://github.com/Zhudongsheng75/SDA Dongsheng Zhu, Zhenyu Mao, Jinghui Lu, Fei Tan 0002 |
LREC/COLING | 5 |
| 2024 | CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsabstractLarge Language Models (LLMs) excel in diverse tasks but often underperform in specialized fields due to limited domain-specific or proprietary corpus.Continual pre-training (CPT) enhances LLM capabilities by imbuing new domain-specific or proprietary knowledge while replaying general corpus to prevent catastrophic forgetting.The data mixture ratio of general corpus and domain-specific corpus, however, has been chosen heuristically, leading to sub-optimal training efficiency in practice.In this context, we attempt to re-visit the scaling behavior of LLMs under the hood of CPT, and discover a power-law relationship between loss, mixture ratio, and training tokens scale.We formalize the trade-off between general and domain-specific capabilities, leading to a well-defined Critical Mixture Ratio (CMR) of general and domain data.By striking the balance, CMR maintains the model's general ability and achieves the desired domain transfer, ensuring the highest utilization of available resources.Considering the balance between efficiency and effectiveness, CMR can be regarded as the optimal mixture ratio.Through extensive experiments, we ascertain the predictability of CMR, propose CMR scaling law and have substantiated its generalization.These findings offer practical guidelines for optimizing LLM training in specialized domains, ensuring both general and domain-specific performance while efficiently managing training resources. Jiawei Gu, Zacc Yang, Chuanghao Ding, Fei Tan 0002 |
EMNLP | 5 |
| 2023 | PUnifiedNER: A Prompting-Based Unified NER System for Diverse DatasetsabstractMuch of named entity recognition (NER) research focuses on developing dataset-specific models based on data from the domain of interest, and a limited set of related entity types. This is frustrating as each new dataset requires a new model to be trained and stored. In this work, we present a ``versatile'' model---the Prompting-based Unified NER system (PUnifiedNER)---that works with data from different domains and can recognise up to 37 entity types simultaneously, and theoretically it could be as many as possible. By using prompt learning, PUnifiedNER is a novel approach that is able to jointly train across multiple corpora, implementing intelligent on-demand entity recognition. Experimental results show that PUnifiedNER leads to significant prediction benefits compared to dataset-specific models with impressively reduced model deployment costs. Furthermore, the performance of PUnifiedNER can achieve competitive or even better performance than state-of-the-art domain-specific methods for some datasets. We also perform comprehensive pilot and ablation studies to support in-depth analysis of each component in PUnifiedNER. Jinghui Lu, Brian Mac Namee, Fei Tan 0002 |
AAAI | 4 |
| 2023 | What Makes Pre-trained Language Models Better Zero-shot Learners?abstractCurrent methods for prompt learning in zeroshot scenarios widely rely on a development set with sufficient human-annotated data to select the best-performing prompt template a posteriori.This is not ideal because in a real-world zero-shot scenario of practical relevance, no labelled data is available.Thus, we propose a simple yet effective method for screening reasonable prompt templates in zero-shot text classification: Perplexity Selection (Perplection).We hypothesize that language discrepancy can be used to measure the efficacy of prompt templates, and thereby develop a substantiated perplexity-based scheme allowing for forecasting the performance of prompt templates in advance.Experiments show that our method leads to improved prediction performance in a realistic zero-shot setting, eliminating the need for any labelled examples.PPL Acc.(%) PPL Acc.(%) PPL Acc.(%) PPL Acc.(%) DOUBAN 24.61 57.12 40.93 50.98 28.80 56.68 71.01 51.31 WEIBO 19.78 61.79 30.37 51.16 22.34 58.35 44.45 50.92WAIMAI 16.44 67.80 23.34 53.15 19.68 69.72 36.07 48.49ECOMMERCE 14.07 73.12 18.45 55.68 16.88 67. Jinghui Lu, Dongsheng Zhu, Weidong Han 0002, Brian Mac Namee, Fei Tan 0002 |
ACL (1) | 6 |
| 2023 | MGEL: Multigrained Representation Analysis and Ensemble Learning for Text ModerationabstractIn this work, we describe our efforts in addressing two typical challenges involved in the popular text classification methods when they are applied to text moderation: the representation of multibyte characters and word obfuscations. Specifically, a multihot byte-level scheme is developed to significantly reduce the dimension of one-hot character-level encoding caused by the multiplicity of instance-scarce non-ASCII characters. In addition, we introduce a simple yet effective weighting approach for fusing n-gram features to empower the classical logistic regression. Surprisingly, it outperforms well-tuned representative neural networks greatly. As a continual effort toward text moderation, we endeavor to analyze the current state-of-the-art (SOTA) algorithm bidirectional encoder representations from transformers (BERT), which works well in context understanding but performs poorly on intentional word obfuscations. To resolve this crux, we then develop an enhanced variant and remedy this drawback by integrating byte and character decomposition. It advances the SOTA performance on the largest abusive language datasets as demonstrated by our comprehensive experiments. Our work offers a feasible and effective framework to tackle word obfuscations. Fei Tan 0002, Changwei Hu, Yifan Hu 0001, Kevin Yen, Zhi Wei 0001, Aasish Pappu, Se Rim Park, Keqian Li |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Hadoop-MTA: a system for Multi Data-center Trillion Concepts Auto-ML atop HadoopabstractThe ever-growing computation capability distributed infrastructure brings tremendous opportunities for mining and analysis of data that was impossible otherwise. Meanwhile, the inherent computation model of distributed system also brings unique and non-trivial challenges for traditional Auto-ML, including the explosion of data dimensions, the expected absence of features, and the heterogeneity of information. This is especially the case in modern Internet enterprises, where data in the scale of trillions are stored in multiple data centers, and the discovery of subtle signals could incur significant impact in revenue and welfare. How can we best harness the large scale distributed machine learning, but without keeping engineers constantly in the loop? In this work, we present Hadoop-MTA, a system for Multi Data-center, Trillion Concepts, Auto-ML on top of the Hadoop distributed computation environment that leverages sparsity aware heterogeneous knowledge graph representation and dimensionality agnostic parallel learning. Through multiple large scale experiments, we find that Hadoop-MTA significantly output-performs competitive state of the art distributed learning algorithms and scales well to trillion scale data-sets. Our model is rolled out to Hadoop serving infrastructure in Yahoo covering billions of unique identities and shows improvements 129.5% accuracy and 106.5 % weighted F1-score (more than 2x) on key targeting use cases. Keqian Li, Yifan Hu 0001, Manisha Verma, Fei Tan 0002, Changwei Hu, Tejaswi Kasturi, Kevin Yen |
IEEE BigData | 4 |
| 2021 | BERT-Beta: A Proactive Probabilistic Approach to Text ModerationabstractText moderation for user generated content, which helps to promote healthy interaction among users, has been widely studied and many machine learning models have been proposed.In this work, we explore an alternative perspective by augmenting reactive reviews with proactive forecasting.Specifically, we propose a new concept text toxicity propensity to characterize the extent to which a text tends to attract toxic comments.Beta regression is then introduced to do the probabilistic modeling, which is demonstrated to function well in comprehensive experiments.We also propose an explanation method to communicate the model decision clearly.Both propensity scoring and interpretation benefit text moderation in a novel manner.Finally, the proposed scaling mechanism for the linear model offers useful insights beyond this work. Fei Tan 0002, Yifan Hu 0001, Kevin Yen, Changwei Hu |
EMNLP (1) | 1 |
| 2020 | DeepVar: An End-to-End Deep Learning Approach for Genomic Variant Recognition in Biomedical LiteratureabstractWe consider the problem of Named Entity Recognition (NER) on biomedical scientific literature, and more specifically the genomic variants recognition in this work. Significant success has been achieved for NER on canonical tasks in recent years where large data sets are generally available. However, it remains a challenging problem on many domain-specific areas, especially the domains where only small gold annotations can be obtained. In addition, genomic variant entities exhibit diverse linguistic heterogeneity, differing much from those that have been characterized in existing canonical NER tasks. The state-of-the-art machine learning approaches heavily rely on arduous feature engineering to characterize those unique patterns. In this work, we present the first successful end-to-end deep learning approach to bridge the gap between generic NER algorithms and low-resource applications through genomic variants recognition. Our proposed model can result in promising performance without any hand-crafted features or post-processing rules. Our extensive experiments and results may shed light on other similar low-resource NER applications. Chaoran Cheng, Fei Tan 0002, Zhi Wei 0001 |
AAAI | 2 |
| 2020 | Repulsive Attention: Rethinking Multi-head Attention as Bayesian InferenceabstractBang An, Jie Lyu, Zhenyi Wang, Chunyuan Li, Changwei Hu, Fei Tan, Ruiyi Zhang, Yifan Hu, Changyou Chen. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Bang An 0001, Jie Lyu 0004, Zhenyi Wang 0001, Chunyuan Li, Changwei Hu, Fei Tan 0002, Ruiyi Zhang 0002, Yifan Hu 0001, Changyou Chen |
EMNLP (1) | 6 |
| 2020 | TNT: Text Normalization based Pre-training of Transformers for Content ModerationabstractIn this work, we present a new language pre-training model TNT (Text Normalization based pre-training of Transformers) for content moderation.Inspired by the masking strategy and text normalization, TNT is developed to learn language representation by training transformers to reconstruct text from four operation types typically seen in text manipulation: substitution, transposition, deletion, and insertion.Furthermore, the normalization involves the prediction of both operation types and token labels, enabling TNT to learn from more challenging tasks than the standard task of masked word recovery.As a result, the experiments demonstrate that TNT outperforms strong baselines on the hate speech classification task.Additional text normalization experiments and case studies show that TNT is a new potential approach to misspelling correction. Fei Tan 0002, Yifan Hu 0001, Changwei Hu, Keqian Li, Kevin Yen |
EMNLP (1) | 1 |
| 2020 | HABERTOR: An Efficient and Effective Deep Hatespeech DetectorabstractWe present our HABERTOR model for detecting hatespeech in large scale user-generated content.Inspired by the recent success of the BERT model, we propose several modifications to BERT to enhance the performance on the downstream hatespeech classification task.HABERTOR inherits BERT's architecture, but is different in four aspects: (i) it generates its own vocabularies and is pre-trained from the scratch using the largest scale hatespeech dataset; (ii) it consists of Quaternionbased factorized components, resulting in a much smaller number of parameters, faster training and inferencing, as well as less memory usage; (iii) it uses our proposed multisource ensemble heads with a pooling layer for separate input sources, to further enhance its effectiveness; and (iv) it uses a regularized adversarial training with our proposed finegrained and adaptive noise magnitude to enhance its robustness.Through experiments on the large-scale real-world hatespeech dataset with 1.4M annotated comments, we show that HABERTOR works better than 15 state-ofthe-art hatespeech detection methods, including fine-tuning Language Models.In particular, comparing with BERT, our HABERTOR is 4∼5 times faster in the training/inferencing phase, uses less than 1/3 of the memory, and has better performance, even though we pretrain it by using less than 1% of the number of words.Our generalizability analysis shows that HABERTOR transfers well to other unseen hatespeech datasets and is a more efficient and effective alternative to BERT for the hatespeech classification. Thanh Tran 0005, Yifan Hu 0001, Changwei Hu, Kevin Yen, Fei Tan 0002, Kyumin Lee, Se Rim Park |
EMNLP (1) | 5 |
| 2019 | User Response Driven Content Understanding with Causal InferenceabstractContent understanding with many potential industrial applications, is spurring interest by researchers in many areas in artificial intelligence. We propose to revisit the content understanding problem in digital marketing from three novel perspectives. First, our problem is to explore the way how user experience is delivered with divergent key multimedia elements. Second, we treat understanding as to elucidate their causal implications in driving user responses. Third, we propose to understand content based on observational audience visit logs. To approach this problem, we measure and generate heterogeneous content features and model them as binary, multivalued or continuous genres. Multiple key performance indicators (KPIs) are introduced to quantify user responses. We then develop a flexible and adaptive doubly robust estimator to identify the causality between these features and user responses from observational data. The comprehensive experiments are performed on real-world data sets. We show that the further analysis of the experimental results can shed actionable insights on how to improve KPIs. Our work will benefit content distribution and optimization in digital marketing. Fei Tan 0002, Zhi Wei 0001, Abhishek Pani, Zhenyu Yan 0001 |
ICDM | 1 |
| 2019 | Success Prediction on Crowdfunding with Multimodal Deep LearningabstractWe consider the problem of project success prediction on crowdfunding platforms. Despite the information in a project profile can be of different modalities such as text, images, and metadata, most existing prediction approaches leverage only the text dominated modality. Nowadays rich visual images have been utilized in more and more project profiles for attracting backers, little work has been conducted to evaluate their effects towards success prediction. Moreover, meta information has been exploited in many existing approaches for improving prediction accuracy. However, such meta information is usually limited to the dynamics after projects are posted, e.g., funding dynamics such as comments and updates. Such a requirement of using after-posting information makes both project creators and platforms not able to predict the outcome in a timely manner. In this work, we designed and evaluated advanced neural network schemes that combine information from different modalities to study the influence of sophisticated interactions among textual, visual, and metadata on project success prediction. To make pre-posting prediction possible, our approach requires only information collected from the pre-posting profile. Our extensive experimental results show that the image features could improve success prediction performance significantly, particularly for project profiles with little text information. Furthermore, we identified contributing elements. Chaoran Cheng, Fei Tan 0002, Xiurui Hou, Zhi Wei 0001 |
IJCAI | 2 |
| 2019 | Modeling and elucidation of housing price
Fei Tan 0002, Chaoran Cheng, Zhi Wei 0001 |
Data Min. Knowl. Discov. | 1 |
| 2019 | A Deep Learning Approach to Competing Risks Representation in Peer-to-Peer LendingabstractOnline peer-to-peer (P2P) lending is expected to benefit both investors and borrowers due to their low transaction cost and the elimination of expensive intermediaries. From the lenders' perspective, maximizing their return on investment is an ultimate goal during their decision-making procedure. In this paper, we explore and address a fundamental problem underlying such a goal: how to represent the two competing risks, charge-off and prepayment, in funded loans. We propose to model both potential risks simultaneously, which remains largely unexplored until now. We first develop a hierarchical grading framework to integrate two risks of loans both qualitatively and quantitatively. Afterward, we introduce an end-to-end deep learning approach to solve this problem by breaking it down into multiple binary classification subproblems that are amenable to both feature representation and risks learning. Particularly, we leverage deep neural networks to jointly solve these subtasks, which leads to the in-depth exploration of the interaction involved in these tasks. To the best of our knowledge, this is the first attempt to characterize competing risks for loans in P2P lending via deep neural networks. The comprehensive experiments on real-world loan data show that our methodology is able to achieve an appealing investment performance by modeling the competition within and between risks explicitly and properly. The feature analysis based on saliency maps provides useful insights into payment dynamics of loans for potential investors intuitively. Fei Tan 0002, Xiurui Hou, Jie Zhang 0049, Zhi Wei 0001, Zhenyu Yan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | A Blended Deep Learning Approach for Predicting User Intended ActionsabstractUser intended actions are widely seen in many areas. Forecasting these actions and taking proactive measures to optimize business outcome is a crucial step towards sustaining the steady business growth. In this work, we focus on predicting attrition, which is one of typical user intended actions. Conventional attrition predictive modeling strategies suffer a few inherent drawbacks. To overcome these limitations, we propose a novel end-to-end learning scheme to keep track of the evolution of attrition patterns for the predictive modeling. It integrates user activity logs, dynamic and static user profiles based on multi-path learning. It exploits historical user records by establishing a decaying multi-snapshot technique. And finally it employs the precedent user intentions via guiding them to the subsequent learning procedure. As a result, it addresses all disadvantages of conventional methods. We evaluate our methodology on two public data repositories and one private user usage dataset provided by Adobe Creative Cloud. The extensive experiments demonstrate that it can offer the appealing performance in comparison with several existing approaches as rated by different popular metrics. Furthermore, we introduce an advanced interpretation and visualization strategy to effectively characterize the periodicity of user activity logs. It can help to pinpoint important factors that are critical to user attrition and retention and thus suggests actionable improvement targets for business practice. Our work will provide useful insights into the prediction and elucidation of other user intended actions as well. Fei Tan 0002, Zhi Wei 0001, Zhenyu Yan 0001 |
ICDM | 1 |
| 2018 | Modeling Item-specific Effects for Video ClickabstractPrediction is widely employed to improve the number of video clicks and views, which are the key important indicators (KPIs) due to their contribution to revenue. The available predictive features, however, are generally limited as compared to the expected prediction capability from the algorithm side. Inspired by the intrinsic dependence among multiple clicks for the same video, we hypothesize that there exist some consistent effects involved in grouped click records. We then propose to recover such effects from the associated hidden features, which are likely to alleviate the insufficiency of features. The simulation studies are performed to elucidate how the derived grouped effects empower a model with additional discriminating capacity compared with the original one. The proposed methodology is further examined on the repository of PPTV (a leading video service provider in China) click records comprehensively. The results confirm the existence of the hypothesized effects and demonstrate their critical role in the performance improvement of video click prediction. Fei Tan 0002, Kuang Du, Zhi Wei 0001, Chenguang Qin |
SDM | 1 |
| 2017 | Time-Aware Latent Hierarchical Model for Predicting House PricesabstractIt is widely acknowledged that the value of a house is the mixture of a large number of characteristics. House price prediction thus presents a unique set of challenges in practice. While a large body of works are dedicated to this task, their performance and applications have been limited by the shortage of long time span of transaction data, the absence of real-world settings and the insufficiency of housing features. To this end, a time-aware latent hierarchical model is introduced to capture underlying spatiotemporal interactions behind the evolution of house prices. The hierarchical perspective obviates the need for historical transaction data of exactly same houses when temporal effects are considered. The proposed framework is examined on a large-scale dataset of the property transaction in Beijing. The whole experimental procedure strictly complies with the real-world scenario. The empirical evaluation results demonstrate the outperformance of our approach over alternative competitive methods. Fei Tan 0002, Chaoran Cheng, Zhi Wei 0001 |
ICDM | 1 |
| 2016 | Modeling Real Estate for School District IdentificationabstractThe affiliated school district of a real estate property is often a crucial concern. How to automate the identification of residential homes located in a favorable educational environment, however, is largely unexplored until now. The availability of heterogeneous estate-related data offers a great opportunity for this task. Nevertheless, it is such heterogeneity that poses significant challenges to their amalgamation in a unified fashion. To this end, we develop G-LRMM model to integrate digital price, textual comments, and geographical location information together. The proposed approach is able to capture the in-depth interaction among multi-type data greatly. The evaluation on the dataset of Beijing property market justifies the benefits of our approach over baselines. The further comparison among different components is also conducted and demonstrates their important roles. Moreover, the proposed model can offer useful insights into modeling heterogeneous data sources. Fei Tan 0002, Chaoran Cheng, Zhi Wei 0001 |
ICDM | 1 |