VLDB 2026 Research / reviewers in the wild / expert
Haiqin Yang
dblp:63/3939
· DBLP profile ↗
65ranked-venue papers
18as first author
19since 2021 · last 2026
0000-0001-5453-476XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 15 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 12 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQLabstractYaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang, Ziyi Liu, Haiqin Yang, Yiying Yang, Yuhang Zhao, Pingfu Chao, Wenhao Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang 0003, Ziyi Liu 0005, Haiqin Yang, Pingfu Chao |
ACL (1) | 6 |
| 2025 | Intrinsic Test of Unlearning Using Parametric Knowledge TracesabstractThe task of "unlearning" certain concepts in large language models (LLMs) has gained attention for its role in mitigating harmful, private, or incorrect outputs.Current evaluations mostly rely on behavioral tests, without monitoring residual knowledge in model parameters, which can be adversarially exploited to recover erased information.We argue that unlearning should also be assessed internally by tracking changes in the parametric traces of unlearned concepts.To this end, we propose a general evaluation methodology that uses vocabulary projections to inspect concepts encoded in model parameters.We apply this approach to localize "concept vectors" -parameter vectors encoding concrete conceptsand construct CONCEPTVECTORS, a benchmark of hundreds of such concepts and their parametric traces in two open-source LLMs.Evaluation on CONCEPTVECTORS shows that existing methods minimally alter concept vectors, mostly suppressing them at inference time, while direct ablation of these vectors removes the associated knowledge and reduces adversarial susceptibility.Our findings reveal limitations of behavior-only evaluations and advocate for parameter-based assessments.We release our code and benchmark at https://github. com/yihuaihong/ConceptVectors.* Bias term is omitted for brevity. Yihuai Hong, Haiqin Yang, Shauli Ravfogel, Mor Geva |
EMNLP | 3 |
| 2025 | Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of SportsabstractPunit Kumar Singh, Nishant Kumar, Akash Ghosh, Kunal Pasad, Khushi Soni, Manisha Jaishwal, Sriparna Saha, Syukron Abu Ishaq Alfarozi, Asres Temam Abagissa, Kitsuchart Pasupa, Haiqin Yang, Jose G Moreno. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Punit Kumar Singh, Akash Ghosh, Kunal Pasad, Khushi Soni, Manisha Jaishwal, Sriparna Saha 0001, Syukron Abu Ishaq Alfarozi, Asres Temam Abagissa, Kitsuchart Pasupa, Haiqin Yang, José G. Moreno 0001 |
EMNLP | 11 |
| 2025 | Natural Language Interfaces for Tabular Data Querying and Visualization: A Survey (Extended Abstract)abstractNatural Language Interfaces (NLIs) have transformed data interaction by enabling natural language querying and visualization of tabular data. Despite the growing importance of NLIs, prior research has examined querying and visualization tasks separately, lacking a unified perspective, especially in the era of Large Language Models (LLMs). To fill this gap, this survey provides a comprehensive analysis of NLIs for tabular data, examining their evolution and fundamental components: datasets, evaluation metrics, and architectural designs. By analyzing over 60 approaches and 38 datasets, we explore recent advancements in Text-to-SQL and Text-to-Vis tasks, focusing on semantic parsing techniques for natural language translation to SQL queries and visualization specifications. We evaluate the impact of LLMs on these systems, discussing their capabilities and limitations. Our systematic review serves as a roadmap for developing NLIs in the foundation model era. Weixu Zhang, Yuanfeng Song, Victor Junqiu Wei, Yuxing Tian, Yiyan Qi, Jonathan H. Chan, Raymond Chi-Wing Wong, Haiqin Yang |
ICDE | 9 |
| 2025 | Boosting Long-Tailed Recognition With Label Descriptor and BeyondabstractLong-Tailed Recognition (LTR) poses significant challenges due to the heavily imbalanced nature of real-world data, which severely skews data-driven deep neural networks. Despite the rapid progress of Vision-Language Models (VLMs), they still face challenges in effectively learning from long-tailed visual data. In this paper, we present a comprehensive analysis of the reasons behind the underperformance of VLMs and propose a hierarchical inference framework to address this issue. Specifically, we prompt the large language models to generatesentence-leveldescriptors for class labels and conduct the open vocabulary classification by computing the average similarity between the image and each descriptor. Areweightingmechanism is further proposed to filter out uninformative descriptors. To mitigate model bias incurred by the long-tail distribution, we propose a feature adapter with the logit adjustment technique and fine-tune the CLIP model via visual prompt tokens. We introduce the Shared Feature space Mixup (SFM) to enhance the interaction between modalities to address tail visual feature insufficiency. Finally, we propose a hierarchical inference manner to combine the aforementioned proposals. Extensive evaluations demonstrate that our approach achieves state-of-the-art performance by fine-tuning only a few parameters on the Places-LT, ImageNet-LT, and iNaturalist 2018 benchmarks. Zhengzhuo Xu, Ruikang Liu, Zenghao Chai, Yiyan Qi, Lei Li 0051, Haiqin Yang, Chun Yuan 0003 |
IEEE Trans. Multim. | 6 |
| 2024 | Dissecting Fine-Tuning Unlearning in Large Language ModelsabstractFine-tuning-based unlearning methods prevail for preventing targeted harmful, sensitive, or copyrighted information within large language models while preserving overall capabilities.However, the true effectiveness of these methods is unclear.In this work, we delve into the limitations of fine-tuning-based unlearning through activation patching and parameter restoration experiments.Our findings reveal that these methods alter the model's knowledge retrieval process, providing further evidence that they do not genuinely erase the problematic knowledge embedded in the model parameters.Instead, the coefficients generated by the MLP components in the model's final layer are the primary contributors to these seemingly positive unlearning effects, playing a crucial role in controlling the model's behaviors.Furthermore, behavioral tests demonstrate that this unlearning mechanism inevitably impacts the global behavior of the models, affecting unrelated knowledge or capabilities.The code is released at https://github.com/yihuaihong/ Dissecting-FT-Unlearning. Yihuai Hong, Yuelin Zou, Lijie Hu, Ziqian Zeng, Di Wang 0015, Haiqin Yang |
EMNLP | 6 |
| 2024 | Unleashing Trigger-Free Event Detection: Revealing Event Correlations Via a Contrastive Derangement FrameworkabstractEvent detection (ED), detecting events with specified types observed in given texts, is critical to many downstream applications. Existing ED methods generally require high-quality triggers annotated by human experts, which is labor-intensive, especially for those nontrivial texts about breaking events. In this paper, we propose a novel trigger-free ED framework that detects multiple events from a given text without pre-defined triggers. Specifically, we first shed light on the event correlations with input texts using a joint embedding paradigm. Next, we devise derangement-based contrastive learning to model fine-grained correlations between multi-event instances. Since events in training benchmarks are usually imbalanced, we further design a simple yet effective event derangement module for balanced training. Experimental results on two benchmarks show that our trigger-free method is remarkably competitive to state-of-the-art trigger-based baselines. Hongzhan Lin 0001, Haiqin Yang, Jing Ma 0004 |
ICASSP | 2 |
| 2024 | Dirichlet Continual Learning: Tackling Catastrophic Forgetting in NLPabstractCatastrophic forgetting poses a significant challenge in continual learning (CL). In the context of Natural Language Processing, generative-based rehearsal CL methods have made progress in avoiding expensive retraining. However, generating pseudo samples that accurately capture the task-specific distribution remains a daunting task. In this paper, we propose Dirichlet Continual Learning (DCL), a novel generative-based rehearsal strategy designed specifically for CL. Different from the conventional use of Gaussian latent variable in Conditional Variational Autoencoder, DCL employs the flexibility of the Dirichlet distribution to model the latent variable. This allows DCL to effectively capture sentence-level features from previous tasks and guide the generation of pseudo samples. Additionally, we introduce Jensen-Shannon Knowledge Distillation, a robust logit-based knowledge distillation method that enhances knowledge transfer during pseudo-sample generation. Our extensive experiments show that DCL outperforms state-of-the-art methods in two typical tasks of task-oriented dialogue systems, demonstrating its efficacy. Haiqin Yang, Wei Xue 0002, Yike Guo |
UAI | 2 |
| 2024 | Natural Language Interfaces for Tabular Data Querying and Visualization: A SurveyabstractThe emergence of natural language processing has revolutionized the way users interact with tabular data, enabling a shift from traditional query languages and manual plotting to more intuitive, language-based interfaces. The rise of large language models (LLMs) such as ChatGPT and its successors has further advanced this field, opening new avenues for natural language processing techniques. This survey presents a comprehensive overview of natural language interfaces for tabular data querying and visualization, which allow users to interact with data using natural language queries. We introduce the fundamental concepts and techniques underlying these interfaces with a particular emphasis on semantic parsing, the key technology facilitating the translation from natural language to SQL queries or data visualization commands. We then delve into the recent advancements in Text-to-SQL and Text-to-Vis problems from the perspectives of datasets, methodologies, metrics, and system designs. This includes a deep dive into the influence of LLMs, highlighting their strengths, limitations, and potential for future improvements. Through this survey, we aim to provide a roadmap for researchers and practitioners interested in developing and applying natural language interfaces for data interaction in the era of large language models. Weixu Zhang, Yuanfeng Song, Victor Junqiu Wei, Yuxing Tian, Yiyan Qi, Jonathan H. Chan, Raymond Chi-Wing Wong, Haiqin Yang |
IEEE Trans. Knowl. Data Eng. | 9 |
| 2024 | Towards Effective Collaborative Learning in Long-Tailed RecognitionabstractReal-world data usually suffers from severe class imbalance and long-tailed distributions, where minority classes are significantly underrepresented compared to the majority ones. Recent research prefers to utilize multi-expert architectures to mitigate the model uncertainty on the minority, where collaborative learning is employed to aggregate the knowledge of experts, i.e., online distillation. In this article, we observe that the knowledge transfer between experts is imbalanced in terms of class distribution, which results in limited performance improvement of the minority classes. To address it, we propose a re-weighted distillation loss by comparing two classifiers' predictions, which are supervised by online distillation and label annotations, respectively. We also emphasize that feature-level distillation will significantly improve model performance and increase feature robustness. Finally, we propose an Effective Collaborative Learning (ECL) framework that integrates a contrastive proxy task branch to further improve feature quality. Quantitative and qualitative experiments on four standard datasets demonstrate that ECL achieves state-of-the-art performance and the detailed ablation studies manifest the effectiveness of each component in ECL. Zhengzhuo Xu, Zenghao Chai, Chengyin Xu, Chun Yuan 0003, Haiqin Yang |
IEEE Trans. Multim. | 5 |
| 2023 | D2Match: Leveraging Deep Learning and Degeneracy for Subgraph MatchingabstractSubgraph matching is a fundamental building block for graph-based applications and is challenging due to its high-order combinatorial nature. Existing studies usually tackle it by combinatorial optimization or learning-based methods. However, they suffer from exponential computational costs or searching the matching without theoretical guarantees. In this paper, we develop $D^2$Match by leveraging the efficiency of Deep learning and Degeneracy for subgraph matching. More specifically, we first prove that subgraph matching can degenerate to subtree matching, and subsequently is equivalent to finding a perfect matching on a bipartite graph. We can then yield an implementation of linear time complexity by the built-in tree-structured aggregation mechanism on graph neural networks. Moreover, circle structures and node attributes can be easily incorporated in $D^2$Match to boost the matching performance. Finally, we conduct extensive experiments to show the superior performance of our $D^2$Match and confirm that our $D^2$Match indeed exploits the subtrees and differs from existing GNNs-based subgraph matching methods that depend on memorizing the data distribution divergence. Xuanzhou Liu, Yujiu Yang 0001, Haiqin Yang |
ICML | 5 |
| 2023 | Do Not Train It: A Linear Neural Architecture Search of Graph Neural NetworksabstractNeural architecture search (NAS) for Graph neural networks (GNNs), called NAS-GNNs, has achieved significant performance over manually designed GNN architectures. However, these methods inherit issues from the conventional NAS methods, such as high computational cost and optimization difficulty. More importantly, previous NAS methods have ignored the uniqueness of GNNs, where GNNs possess expressive power without training. With the randomly-initialized weights, we can then seek the optimal architecture parameters via the sparse coding objective and derive a novel NAS-GNNs method, namely neural architecture coding (NAC). Consequently, our NAC holds a no-update scheme on GNNs and can efficiently compute in linear time. Empirical evaluations on multiple GNN benchmark datasets demonstrate that our approach leads to state-of-the-art performance, which is up to $200\times$ faster and $18.8%$ more accurate than the strong baselines. Peng Xu 0052, Xuanzhou Liu, Yue Zhao 0016, Haiqin Yang, Bei Yu 0001 |
ICML | 6 |
| 2022 | Vision-and-Language Pretrained Models: A SurveyabstractPretrained models have produced great success in both Computer Vision (CV) and Natural Language Processing (NLP). This progress leads to learning joint representations of vision and language pretraining by feeding visual and linguistic contents into a multi-layer transformer, Visual-Language Pretrained Models (VLPMs). In this paper, we present an overview of the major advances achieved in VLPMs for producing joint representations of vision and language. As the preliminaries, we briefly describe the general task definition and genetic architecture of VLPMs. We first discuss the language and vision data encoding methods and then present the mainstream VLPM structure as the core content. We further summarise several essential pretraining and fine-tuning strategies. Finally, we highlight three future directions for both CV and NLP researchers to provide insightful guidance. Siqu Long, Feiqi Cao, Soyeon Caren Han, Haiqin Yang |
IJCAI | 4 |
| 2021 | KGSynNet: A Novel Entity Synonyms Discovery Framework with Knowledge Graph
Xi Yin 0007, Haiqin Yang, Xingjian Fei, Hao Peng 0001, Kaijie Zhou, Kunfeng Lai, Jianping Shen |
DASFAA (1) | 3 |
| 2021 | Progressive Open-Domain Response Generation with Multiple Controllable AttributesabstractIt is desirable to include more controllable attributes to enhance the diversity of generated responses in open-domain dialogue systems. However, existing methods can generate responses with only one controllable attribute or lack a flexible way to generate them with multiple controllable attributes. In this paper, we propose a Progressively trained Hierarchical Encoder-Decoder (PHED) to tackle this task. More specifically, PHED deploys Conditional Variational AutoEncoder (CVAE) on Transformer to include one aspect of attributes at one stage. A vital characteristic of the CVAE is to separate the latent variables at each stage into two types: a global variable capturing the common semantic features and a specific variable absorbing the attribute information at that stage. PHED then couples the CVAE latent variables with the Transformer encoder and is trained by minimizing a newly derived ELBO and controlled losses to produce the next stage's input and produce responses as required. Finally, we conduct extensive evaluations to show that PHED significantly outperforms the state-of-the-art neural generation models and produces more diverse responses as expected. Haiqin Yang, Xiaoyuan Yao, Yiqun Duan, Jianping Shen, Kun Zhang 0001 |
IJCAI | 1 |
| 2021 | RefBERT: Compressing BERT by Referencing to Pre-computed RepresentationsabstractRecently developed large pre-trained language models, e.g., BERT, have achieved remarkable performance in many downstream natural language processing applications. These pre-trained language models often contain hundreds of millions of parameters and suffer from high computation and latency in real-world applications. It is desirable to reduce the computation overhead of the models for fast training and inference while keeping the model performance in downstream applications. Several lines of work utilize knowledge distillation to compress the teacher model to a smaller student model. However, they usually discard the teacher's knowledge when in inference. Differently, in this paper, we propose RemERT to leverage the knowledge learned from the teacher, i.e., facilitating the pre-computed BERT representation on the reference sample and compressing BERT into a smaller student model. To guarantee our proposal, we provide theoretical justification on the loss function and the usage of reference samples. Significantly, the theoretical result shows that including the pre-computed teacher's representations on the reference samples indeed increases the mutual information in learning the student model. Finally, we conduct the empirical evaluation and show that our RemERT can beat the vanilla TinyBERT over 8.1 % and achieves more than 94% of the performance of$\mathbf{BERT}_{\mathbf{BASE}}$on the GLUE benchmark. Meanwhile, RemERT is$\mathbf{7.4x}$smaller and$\mathbf{9.5x}$faster on inference than$\mathbf{BERT}_{\mathbf{BASE}}$. Xinyi Wang 0003, Haiqin Yang, Yang Mo, Jianping Shen |
IJCNN | 2 |
| 2021 | Emotion Dynamics Modeling via BERTabstractEmotion dynamics modeling is a significant task in emotion recognition in conversation. It aims to predict conversational emotions when building empathetic dialogue systems. Existing studies mainly develop models based on Recurrent Neural Networks (RNNs). They cannot benefit from the power of the recently-developed pre-training strategies for better token representation learning in conversations. More seriously, it is hard to distinguish the dependency of interlocutors and the emotional influence among interlocutors by simply assembling the features on top of RNNs. In this paper, we develop a series of BERT-based models to specifically capture the inter-interlocutor and intra-interlocutor dependencies of the conversational emotion dynamics. Concretely, we first substitute BERT for RNNs to enrich the token representations. Then, a Flat-structured BERT (F-BERT) is applied to link up utterances in a conversation directly, and a Hierarchically-structured BERT (H-BERT) is employed to distinguish the interlocutors when linking up utterances. More importantly, a Spatial-Temporal-structured BERT, namely ST-BERT, is proposed to further determine the emotional influence among interlocutors. Finally, we conduct extensive experiments on two popular emotion recognition in conversation benchmark datasets and demonstrate that our proposed models can attain around 5% and 10% improvement over the state-of-the-art baselines, respectively. Haiqin Yang, Jianping Shen |
IJCNN | 1 |
| 2021 | Automatic Intent-Slot Induction for Dialogue SystemsabstractAutomatically and accurately identifying user intents and filling the associated slots from their spoken language are critical to the success of dialogue systems. Traditional methods require manually defining the DOMAIN-INTENT-SLOT schema and asking many domain experts to annotate the corresponding utterances, upon which neural models are trained. This procedure brings the challenges of information sharing hindering, out-of-schema, or data sparsity in open domain dialogue systems. To tackle these challenges, we explore a new task of automatic intent-slot induction and propose a novel domain-independent tool. That is, we design a coarse-to-fine three-step procedure including Role-labeling, Concept-mining, And Pattern-mining (RCAP): (1) role-labeling: extracting key phrases from users’ utterances and classifying them into a quadruple of coarsely-defined intent-roles via sequence labeling; (2) concept-mining: clustering the extracted intent-role mentions and naming them into abstract fine-grained concepts; (3) pattern-mining: applying the Apriori algorithm to mine intent-role patterns and automatically inferring the intent-slot using these coarse-grained intent-role labels and fine-grained concepts. Empirical evaluations on both real-world in-domain and out-of-domain datasets show that: (1) our RCAP can generate satisfactory SLU schema and outperforms the state-of-the-art supervised learning method; (2) our RCAP can be directly applied to out-of-domain datasets and gain at least 76% improvement of F1-score on intent detection and 41% improvement of F1-score on slot filling; (3) our RCAP exhibits its power in generic intent-slot extractions with less manual effort, which opens pathways for schema induction on new domains and unseen intent-slot discovery for generalizable dialogue systems. Zengfeng Zeng, Haiqin Yang, Zhen Gou, Jianping Shen |
WWW | 3 |
| 2021 | Making Online Sketching Hashing Even FasterabstractData-dependent hashing methods have demonstrated good performance in various machine learning applications to learn a low-dimensional representation from the original data. However, they still suffer from several obstacles: First, most of existing hashing methods are trained in a batch mode, yielding inefficiency for training streaming data. Second, the computational cost and the memory consumption increase extraordinarily in the big data setting, which perplexes the training procedure. Third, the lack of labeled data hinders the improvement of the model performance. To address these difficulties, we utilize online sketching hashing (OSH) and present a FasteR Online Sketching Hashing (FROSH) algorithm to sketch the data in a more compact form via an independent transformation. We provide theoretical justification to guarantee that our proposed FROSH consumes less time and achieves a comparable sketching precision under the same memory cost of OSH. We also extend FROSH to its distributed implementation, namely DFROSH, to further reduce the training time cost of FROSH while deriving the theoretical bound of the sketching precision. Finally, we conduct extensive experiments on both synthetic and real datasets to demonstrate the attractive merits of FROSH and DFROSH. Haiqin Yang, Shenglin Zhao, Irwin King, Michael R. Lyu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Block-term tensor neural networks
Jinmian Ye, Guangxi Li, Haiqin Yang, Shandian Zhe, Zenglin Xu |
Neural Networks | 4 |
| 2020 | Effective Data-Aware Covariance Estimator From Compressed DataabstractEstimating covariance matrix from massive high-dimensional and distributed data is significant for various real-world applications. In this paper, we propose a data-aware weighted sampling-based covariance matrix estimator, namely DACE, which can provide an unbiased covariance matrix estimation and attain more accurate estimation under the same compression ratio. Moreover, we extend our proposed DACE to tackle multiclass classification problems with theoretical justification and conduct extensive experiments on both synthetic and real-world data sets to demonstrate the superior performance of our DACE. Haiqin Yang, Shenglin Zhao, Michael R. Lyu, Irwin King |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Neural Machine Translation for Financial Listing Documents
Linkai Luo, Haiqin Yang, Sai Cheong Siu, Francis Y. L. Chin |
ICONIP (5) | 2 |
| 2018 | TreeNet: Learning Sentence Representations with Unconstrained Tree StructureabstractRecursive neural network (RvNN) has been proved to be an effective and promising tool to learn sentence representations by explicitly exploiting the sentence structure. However, most existing work can only exploit simple tree structure, e.g., binary trees, or ignore the order of nodes, which yields suboptimal performance. In this paper, we proposed a novel neural network, namely TreeNet, to capture sentences structurally over the raw unconstrained constituency trees, where the number of child nodes can be arbitrary. In TreeNet, each node is learning from its left sibling and right child in a bottom-up left-to-right order, thus enabling the net to learn over any tree. Furthermore, multiple soft gates and a memory cell are employed in implementing the TreeNet to determine to what extent it should learn, remember and output, which proves to be a simple and efficient mechanism for semantic synthesis. Moreover, TreeNet significantly suppresses convolutional neural networks (CNN) and Long Short-Term Memory (LSTM) with fewer parameters. It improves the classification accuracy by 2%-5% with 42% of the best CNN’s parameters or 94% of standard LSTM’s. Extensive experiments demonstrate TreeNet achieves the state-of-the-art performance on all four typical text classification tasks. Zhou Cheng, Chun Yuan 0003, Jiancheng Li, Haiqin Yang |
IJCAI | 4 |
| 2018 | A Broad Neural Network Structure for Class Incremental Learning
Wenzhang Liu, Haiqin Yang, Yuewen Sun, Changyin Sun 0001 |
ISNN | 2 |
| 2018 | Predicting the quality of online health expert question-answering services with temporal features in a deep learning framework
Ze Hu, Zhan Zhang 0002, Haiqin Yang, De-Cheng Zuo |
Neurocomputing | 3 |
| 2018 | Factorization machines and deep views-based co-training for improving answer quality prediction in online health expert question-answering services
Zhan Zhang 0002, Ze Hu, Haiqin Yang, De-Cheng Zuo |
J. Biomed. Informatics | 3 |
| 2018 | Online Nonlinear AUC Maximization for Imbalanced Data SetsabstractClassifying binary imbalanced streaming data is a significant task in both machine learning and data mining. Previously, online area under the receiver operating characteristic (ROC) curve (AUC) maximization has been proposed to seek a linear classifier. However, it is not well suited for handling nonlinearity and heterogeneity of the data. In this paper, we propose the kernelized online imbalanced learning (KOIL) algorithm, which produces a nonlinear classifier for the data by maximizing the AUC score while minimizing a functional regularizer. We address four major challenges that arise from our approach. First, to control the number of support vectors without sacrificing the model performance, we introduce two buffers with fixed budgets to capture the global information on the decision boundary by storing the corresponding learned support vectors. Second, to restrict the fluctuation of the learned decision function and achieve smooth updating, we confine the influence on a new support vector to its -nearest opposite support vectors. Third, to avoid information loss, we propose an effective compensation scheme after the replacement is conducted when either buffer is full. With such a compensation scheme, the performance of the learned model is comparable to the one learned with infinite budgets. Fourth, to determine good kernels for data similarity representation, we exploit the multiple kernel learning framework to automatically learn a set of kernels. Extensive experiments on both synthetic and real-world benchmark data sets demonstrate the efficacy of our proposed approach. Junjie Hu 0001, Haiqin Yang, Michael R. Lyu, Irwin King, Anthony Man-Cho So |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Efficient Non-Oblivious Randomized Reduction for Risk Minimization with Improved Excess Risk GuaranteeabstractIn this paper, we address learning problems for high dimensional data. Previously, oblivious random projection based approaches that project high dimensional features onto a random subspace have been used in practice for tackling high-dimensionality challenge in machine learning. Recently, various non-oblivious randomized reduction methods have been developed and deployed for solving many numerical problems such as matrix product approximation, low-rank matrix approximation, etc. However, they are less explored for the machine learning tasks, e.g., classification. More seriously, the theoretical analysis of excess risk bounds for risk minimization, an important measure of generalization performance, has not been established for non-oblivious randomized reduction methods. It therefore remains an open problem what is the benefit of using them over previous oblivious random projection based approaches. To tackle these challenges, we propose an algorithmic framework for employing non-oblivious randomized reduction method for general empirical risk minimizing in machine learning tasks, where the original high-dimensional features are projected onto a random subspace that is derived from the data with a small matrix approximation error. We then derive the first excess risk bound for the proposed non-oblivious randomized reduction approach without requiring strong assumptions on the training data. The established excess risk bound exhibits that the proposed approach provides much better generalization performance and it also sheds more insights about different randomized reduction approaches. Finally, we conduct extensive experiments on both synthetic and real-world benchmark datasets, whose dimension scales to O(10^7), to demonstrate the efficacy of our proposed approach. Yi Xu 0008, Haiqin Yang, Lijun Zhang 0005, Tianbao Yang |
AAAI | 2 |
| 2017 | Heterogeneous Features Integration in Deep Knowledge Tracing
Lap Pong Cheung, Haiqin Yang |
ICONIP (2) | 2 |
| 2017 | A deep learning approach for predicting the quality of online health expert question-answering services
Ze Hu, Zhan Zhang 0002, Haiqin Yang, De-Cheng Zuo |
J. Biomed. Informatics | 3 |
| 2016 | STELLAR: Spatial-Temporal Latent Ranking for Successive Point-of-Interest RecommendationabstractSuccessive point-of-interest (POI) recommendation in location-based social networks (LBSNs) becomes a significant task since it helps users to navigate a number of candidate POIs and provides the best POI recommendations based on users’ most recent check-in knowledge. However, all existing methods for successive POI recommendation only focus on modeling the correlation between POIs based on users’ check-in sequences, but ignore an important fact that successive POI recommendation is a time-subtle recommendation task. In fact, even with the same previous check-in information, users would prefer different successive POIs at different time. To capture the impact of time on successive POI recommendation, in this paper, we propose a spatial-temporal latent ranking (STELLAR) method to explicitly model the interactions among user, POI, and time. In particular, the proposed STELLAR model is built upon a ranking-based pairwise tensor factorization framework with a fine-grained modeling of user-POI, POI-time, and POI-POI interactions for successive POI recommendation. Moreover, we propose a new interval-aware weight utility function to differentiate successive check-ins’ correlations, which breaks the time interval constraint in prior work. Evaluations on two real-world datasets demonstrate that the STELLAR model outperforms state-of-the-art successive POI recommendation model about 20% in Precision@5 and Recall@5. Shenglin Zhao, Tong Zhao 0002, Haiqin Yang, Michael R. Lyu, Irwin King |
AAAI | 3 |
| 2016 | Distributed Information-Theoretic Metric Learning in Apache SparkabstractDistance metric learning (DML) is an effective similarity learning tool to learn a distance function from examples to enhance the model performance in applications of classification, regression, and ranking, etc. Most DML algorithms need to learn a Mahalanobis matrix, a positive semidefinite matrix that scales quadratically with the number of dimensions of input data. This brings huge computational cost in the learning procedure, and makes all proposed algorithms infeasible for extremely high-dimensional data even with the low-rank approximation. Differently, in this paper, we take advantage of the power of parallel computation and propose a novel distributed distance metric learning algorithm based on a state-of-the-art DML algorithm, Information-Theoretic Metric Learning (ITML).More specifically, we utilize the property that each positive semidefinite matrix can be decomposed into a combination of rank-one and trace-one matrices and convert the original sequential training procedure into a parallel one. In most cases, the communication demands of the proposed method are also reduced from O(d2) to O(cd), where d is the number of dimensions of the data and c is the number of constraints in DML and can be smaller than d by appropriate selection. Moreover importantly, we present a rigorous theoretical analysis to upper bound the Bregman divergence between the sequential algorithm and the parallel algorithm, which guarantees the correctness and performance of the proposed algorithm. Our experiments on datasets with O(105) features demonstrate the competitive scalability and the performance compared with the original ITML algorithm. Yuxin Su 0001, Haiqin Yang, Irwin King, Michael R. Lyu |
IJCNN | 2 |
| 2016 | Online non-negative dictionary learning via moment information for sparse Poisson codingabstractOnline dictionary learning for sparse coding is an effective tool for data analysis. It incrementally learns a set of basis vectors with sparse linear combinations of these vectors when new samples appear. Previous work assumes that the samples embed Gaussian noises, which weaken the power of these methods in handling real applications with non-negative data (e.g., frequency data in word counts). Differently, in this paper, we concentrate on online learning for non-negative dictionary by using moment information for sparse Poisson coding. We exploit the non-negativity of Poisson models to learn a set of non-negative basis vectors and a non-negative sparse linear combination for the moment information of samples. Specifically, we first formulate the online learning problem via the maximum-a-posteriori (MAP) framework. We then propose a novel online algorithm which alternatively updates the sparse-coefficient vector and the basis vectors with non-negativity constraints when a new sample arrives. More importantly, we present sufficient convergence analyses to guarantee the performance of the proposed algorithm, which leads to convergence of a stable dictionary for characterizing the moment information of samples. We finally conduct a series of experiments on word-counts data and image data to show merits of the proposed online algorithm. Xiaotian Yu, Haiqin Yang, Irwin King, Michael R. Lyu |
IJCNN | 2 |
| 2016 | A Unified Point-of-Interest Recommendation Framework in Location-Based Social NetworksabstractLocation-based social networks (LBSNs), such as Gowalla, Facebook, Foursquare, Brightkite, and so on, have attracted millions of users to share their social friendship and their locations via check-ins in the past few years. Plenty of valuable information is accumulated based on the check-in behaviors, which makes it possible to learn users’ moving patterns as well as their preferences. In LBSNs, point-of-interest (POI) recommendation is one of the most significant tasks because it can help targeted users explore their surroundings as well as help third-party developers provide personalized services. Matrix factorization is a promising method for this task because it can capture users’ preferences to locations and is widely adopted in traditional recommender systems such as movie recommendation. However, the sparsity of the check-in data makes it difficult to capture users’ preferences accurately. Geographical influence can help alleviate this problem and have a large impact on the final recommendation result. By studying users’ moving patterns, we find that users tend to check in around several centers and different users have different numbers of centers. Based on this, we propose a Multi-center Gaussian Model (MGM) to capture this pattern via modeling the probability of a user’s check-in on a location. Moreover, users are usually more interested in the top 20 or even top 10 recommended POIs, which makes personalized ranking important in this task. From previous work, directly optimizing for pairwise ranking like Bayesian Personalized Ranking (BPR) achieves better performance in the top- k recommendation than directly using matrix matrix factorization that aims to minimize the point-wise rating error. To consider users’ preferences, geographical influence and personalized ranking, we propose a unified POI recommendation framework, which unifies all of them together. Specifically, we first fuse MGM with matrix factorization methods and further with BPR using two different approaches. We conduct experiments on Gowalla and Foursquare datasets, which are two large-scale real-world LBSN datasets publicly available online. The results on both datasets show that our unified POI recommendation framework can produce better performance. Haiqin Yang, Irwin King, Michael R. Lyu |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2015 | Kernelized Online Imbalanced Learning with Fixed BudgetsabstractOnline learning from imbalanced streaming data to capture the nonlinearity and heterogeneity of the data is significant in machine learning and data mining. To tackle this problem, we propose a kernelized online imbalanced learning (KOIL) algorithm to directly maximize the area under the ROC curve (AUC). We address two more challenges: 1) How to control the number of support vectors without sacrificing model performance; and 2) how to restrict the fluctuation of the learned decision function to attain smooth updating. To this end, we introduce two buffers with fixed budgets (buffer sizes) for positive class and negative class, respectively, to store the learned support vectors, which can allow us to capture the global information of the decision boundary. When determining the weight of a new support vector, we confine its influence only to its $k$-nearest opposite support vectors. This can restrict the effect of new instances and prevent the harm of outliers. More importantly, we design a sophisticated scheme to compensate the model after replacement is conducted when either buffer is full. With this compensation, the learned model approaches the one learned with infinite budgets. We present both theoretical analysis and extensive experimental comparison to demonstrate the effectiveness of our proposed KOIL. Junjie Hu 0001, Haiqin Yang, Irwin King, Michael R. Lyu, Anthony Man-Cho So |
AAAI | 2 |
| 2015 | Training-Efficient Feature Map for Shift-Invariant Kernels
Haiqin Yang, Irwin King, Michael R. Lyu |
IJCAI | 2 |
| 2015 | WSDM'15 Workshop Summary / Scalable Data Analytics: Theory and ApplicationsabstractThe SDA workshop at WSDM 2015 is the fifth International Workshop on Scalable Data Analytics, following the previous four workshops of SDA respectively held at IEEE Big Data 2013, PAKDD 2014, IEEE Big Data 2014, and IEEE ICDM 2014. This series of workshops aims to provide professionals, researchers, and technologists with a single forum where they can discuss and share the state-of-the-art theories and applications of scalable data analytics technologies. In particular, in the era of information explosion, the scientific, biomedical, and engineering research communities are undergoing a profound transformation where discoveries and innovations increasingly rely on massive amounts of data. The characteristics of volume, velocity, variety and veracity originated in the massive big data then bring challenges to current data analytics techniques. The focus of the fifth SDA is to discuss how we can scale up data analytics techniques for modeling and analyzing big data from various domains. Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu |
WSDM | 2 |
| 2015 | Maximum margin semi-supervised learning with irrelevant data
Haiqin Yang, Kaizhu Huang, Irwin King, Michael R. Lyu |
Neural Networks | 1 |
| 2015 | Budget constrained non-monotonic feature selection
Haiqin Yang, Zenglin Xu, Michael R. Lyu, Irwin King |
Neural Networks | 1 |
| 2015 | Boosting Response Aware Model-Based Collaborative FilteringabstractRecommender systems are promising for providing personalized favorite services. Collaborative filtering (CF) technologies, making prediction of users' preference based on users' previous behaviors, have become one of the most successful techniques to build modern recommender systems. Several challenging issues occur in previously proposed CF methods: (1) most CF methods ignore users' response patterns and may yield biased parameter estimation and suboptimal performance; (2) some CF methods adopt heuristic weight settings, which lacks a systematical implementation; and (3) the multinomial mixture models may weaken the computational ability of matrix factorization for generating the data matrix, thus increasing the computational cost of training. To resolve these issues, we incorporate users' response models into the probabilistic matrix factorization (PMF), a popular matrix factorization CF model, to establish the response aware probabilistic matrix factorization (RAPMF) framework. More specifically, we make the assumption on the user response as a Bernoulli distribution which is parameterized by the rating scores for the observed ratings while as a step function for the unobserved ratings. Moreover, we speed up the algorithm by a mini-batch implementation and a crafting scheduling policy. Finally, we design different experimental protocols and conduct systematical empirical evaluation on both synthetic and real-world datasets to demonstrate the merits of the proposed RAPMF and its mini-batch implementation. Haiqin Yang, Guang Ling, Yuxin Su 0001, Michael R. Lyu, Irwin King |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Non-monotonic Feature Selection for Regression
Haiqin Yang, Zenglin Xu, Irwin King, Michael R. Lyu |
ICONIP (2) | 1 |
| 2013 | Sparse Poisson coding for high dimensional document clusteringabstractDocument clustering plays an important role in large scale textual data analysis, which generally faces with great challenge of the high dimensional textual data. One remedy is to learn the high-level sparse representation by the sparse coding techniques. In contrast to traditional Gaussian noise-based sparse coding methods, in this paper, we employ a Poisson distribution model to represent the word-count frequency feature of a text for sparse coding. Moreover, a novel sparse-constrained Poisson regression algorithm is proposed to solve the induced optimization problem. Different from previous Poisson regression with the family of ℓ1-regularization to enhance the sparse solution, we introduce a sparsity ratio measure which make use of both ℓ1-norm and ℓ2-norm on the learned weight. An important advantage of the sparsity ratio is that it bounded in the range of 0 and 1. This makes it easy to set for practical applications. To further make the algorithm trackable for the high dimensional textual data, a projected gradient descent algorithm is proposed to solve the regression problem. Extensive experiments have been conducted to show that our proposed approach can achieve effective representation for document clustering compared with state-of-the-art regression methods. Chenxia Wu, Haiqin Yang, Jianke Zhu, Jiemi Zhang, Irwin King, Michael R. Lyu |
IEEE BigData | 2 |
| 2013 | Where You Like to Go Next: Successive Point-of-Interest Recommendation
Haiqin Yang, Michael R. Lyu, Irwin King |
IJCAI | 2 |
| 2013 | Efficient online learning for multitask feature selectionabstractLearning explanatory features across multiple related tasks, or MultiTask Feature Selection (MTFS), is an important problem in the applications of data mining, machine learning, and bioinformatics. Previous MTFS methods fulfill this task by batch-mode training. This makes them inefficient when data come sequentially or when the number of training data is so large that they cannot be loaded into the memory simultaneously. In order to tackle these problems, we propose a novel online learning framework to solve the MTFS problem. A main advantage of the online algorithm is its efficiency in both time complexity and memory cost. The weights of the MTFS models at each iteration can be updated by closed-form solutions based on the average of previous subgradients. This yields the worst-case bounds of the time complexity and memory cost at each iteration, both in the order ofO(d×Q), wheredis the number of feature dimensions andQis the number of tasks. Moreover, we provide theoretical analysis for the average regret of the online learning algorithms, which also guarantees the convergence rate of the algorithms. Finally, we conduct detailed experiments to show the characteristics and merits of the online learning algorithms in solving several MTFS problems. Haiqin Yang, Michael R. Lyu, Irwin King |
ACM Trans. Knowl. Discov. Data | 1 |
| 2012 | Fused Matrix Factorization with Geographical and Social Influence in Location-Based Social NetworksabstractRecently, location-based social networks (LBSNs), such as Gowalla, Foursquare, Facebook, and Brightkite, etc., have attracted millions of users to share their social friendship and their locations via check-ins. The available check-in information makes it possible to mine users’ preference on locations and to provide favorite recommendations. Personalized Point-of-interest (POI) recommendation is a significant task in LBSNs since it can help targeted users explore their surroundings as well as help third-party developers to provide personalized services. To solve this task, matrix factorization is a promising tool due to its success in recommender systems. However, previously proposed matrix factorization (MF) methods do not explore geographical influence, e.g., multi-center check-in property, which yields suboptimal solutions for the recommendation. In this paper, to the best of our knowledge, we are the first to fuse MF with geographical and social influence for POI recommendation in LBSNs. We first capture the geographical influence via modeling the probability of a user’s check-in on a location as a Multi-center Gaussian Model (MGM). Next, we include social information and fuse the geographical influence into a generalized matrix factorization framework. Our solution to POI recommendation is efficient and scales linearly with the number of observations. Finally, we conduct thorough experiments on a large-scale real-world LBSNs dataset and demonstrate that the fused matrix factorization framework with MGM utilizes the distance information sufficiently and outperforms other state-of-the-art methods significantly. Haiqin Yang, Irwin King, Michael R. Lyu |
AAAI | 2 |
| 2012 | Online learning for collaborative filteringabstractCollaborative filtering (CF), aiming at predicting users' unknown preferences based on observational preferences from some users, has become one of the most successful methods to building recommender systems. Various approaches to CF have been proposed in this area, but seldom do they consider the dynamic scenarios: 1) new items arriving in the system, 2) new users joining the system; or 3) new rating updating the system are all dynamically obtained with respect to time. To capture these changes, in this paper, we develop an online learning framework for collaborative filtering. Specifically, we construct this framework consisting of two state-of-the-art matrix factorization based CF methods: the probabilistic matrix factorization and the top-one probability based ranking matrix factorization. Moreover, we demonstrate that the proposed online algorithms bring several attractive advantages: 1) they scale linearly with the number of observed ratings and the size of latent features; 2) they obviate the need to load all ratings in memory; 3) they can adapt to new ratings easily. Finally, we conduct a series of detailed experiments on real-world datasets to demonstrate the merits of the proposed online learning algorithms under various settings. Guang Ling, Haiqin Yang, Irwin King, Michael R. Lyu |
IJCNN | 2 |
| 2012 | Response Aware Model-Based Collaborative Filtering
Guang Ling, Haiqin Yang, Michael R. Lyu, Irwin King |
UAI | 2 |
| 2011 | Can irrelevant data help semi-supervised learning, why and how?abstractPrevious semi-supervised learning (SSL) techniques usually assume unlabeled data are relevant to the target task. That is, they follow the same distribution as the targeted labeled data. In this paper, we address a different and very difficult scenario in SSL, where the unlabeled data may be a mixture of data relevant or irrelevant to the target binary classification task. In our framework, we do not require explicitly prior knowledge on the relatedness of the unlabeled data to the target data. In order to alleviate the effect of the irrelevant unlabeled data and utilize the implicit knowledge among all available data, we develop a novel maximum margin classifier, named the tri-class support vector machine (3C-SVM), to seek an inductive rule to separate the target binary classification task well while finding out the irrelevant data by-product. To attain this goal, we introduce a new min loss function, which can relieve the impact of the irrelevant data while relying more on the labeled data and the relevant unlabeled data. This loss function can therefore achieve the maximum entropy principle. The 3C-SVM can then generalize standard SVMs, Semi-supervised SVMs, and SVMs learned from the universum as its special cases. We further analyze the property of 3C-SVM on why the irrelevant data can help to improve the model performance. For implementation, we make relaxation and approximate the objective by the convex-concave procedure, which turns the original optimization from integral programming problem to a problem by just solving a finite number of quadratic programming problems. Empirical results are reported to demonstrate the advantages of our 3C-SVM model. Haiqin Yang, Shenghuo Zhu, Irwin King, Michael R. Lyu |
CIKM | 1 |
| 2011 | Efficient Sparse Generalized Multiple Kernel LearningabstractKernel methods have been successfully applied in various applications. To succeed in these applications, it is crucial to learn a good kernel representation, whose objective is to reveal the data similarity precisely. In this paper, we address the problem of multiple kernel learning (MKL), searching for the optimal kernel combination weights through maximizing a generalized performance measure. Most MKL methods employ the L(1)-norm simplex constraints on the kernel combination weights, which therefore involve a sparse but non-smooth solution for the kernel weights. Despite the success of their efficiency, they tend to discard informative complementary or orthogonal base kernels and yield degenerated generalization performance. Alternatively, imposing the L(p)-norm (p > 1) constraint on the kernel weights will keep all the information in the base kernels. This leads to non-sparse solutions and brings the risk of being sensitive to noise and incorporating redundant information. To tackle these problems, we propose a generalized MKL (GMKL) model by introducing an elastic-net-type constraint on the kernel weights. More specifically, it is an MKL model with a constraint on a linear combination of the L(1)-norm and the squared L(2)-norm on the kernel weights to seek the optimal kernel combination weights. Therefore, previous MKL problems based on the L(1)-norm or the L(2)-norm constraints can be regarded as special cases. Furthermore, our GMKL enjoys the favorable sparsity property on the solution and also facilitates the grouping effect. Moreover, the optimization of our GMKL is a convex optimization problem, where a local solution is the global optimal solution. We further derive a level method to efficiently solve the optimization problem. A series of experiments on both synthetic and real-world datasets have been conducted to show the effectiveness and efficiency of our GMKL. Haiqin Yang, Zenglin Xu, Jieping Ye, Irwin King, Michael R. Lyu |
IEEE Trans. Neural Networks | 1 |
| 2010 | Online learning for multi-task feature selectionabstractMulti-task feature selection (MTFS) is an important tool to learn the explanatory features across multiple related tasks. Previous MTFS methods fulfill this task in batch-mode training. This makes them inefficient when data come in sequence or when the number of training data is so large that they cannot be loaded into the memory simultaneously. To tackle these problems, we propose the first online learning framework for MTFS. A main advantage of the online algorithms is the efficiency in both time complexity and memory cost due to the closed-form solutions in updating the model weights at each iteration. Experimental results on a real-world dataset attest to the merits of the proposed algorithms. Haiqin Yang, Irwin King, Michael R. Lyu |
CIKM | 1 |
| 2010 | Simple and Efficient Multiple Kernel Learning by Group Lasso
Zenglin Xu, Rong Jin 0001, Haiqin Yang, Irwin King, Michael R. Lyu |
ICML | 3 |
| 2010 | Online Learning for Group Lasso
Haiqin Yang, Zenglin Xu, Irwin King, Michael R. Lyu |
ICML | 1 |
| 2010 | Multi-task Learning for one-class classificationabstractIn this paper, we address the problem of one-class classification. Taking into account the fact that in some applications, the given training samples are rather limited, we attempt to utilize the advantages of Multi-task Learning (MTL), where the data of related tasks may share similar structure and helpful information. We then propose an MTL framework for one-class classification. The framework derives from the one-class v-SVM and makes use of related tasks by constraining them to have similar solutions. This formulation can be cast into a second-order cone program, which achieves a global solution and is solved efficiently. Further, the framework also maintains the favorable property of the v parameter in the v-SVM, which can control the fraction of outliers and support vectors, in one-class classification. This framework also connects with several existing models. Experimental results on both synthetic and real-world datasets demonstrate the properties and advantages of our proposed model. Haiqin Yang, Irwin King, Michael R. Lyu |
IJCNN | 1 |
| 2009 | Ensemble Learning for Imbalanced E-commerce Transaction Anomaly Classification
Haiqin Yang, Irwin King |
ICONIP (1) | 1 |
| 2009 | Localized support vector regression for time series prediction
Haiqin Yang, Kaizhu Huang, Irwin King, Michael R. Lyu |
Neurocomputing | 1 |
| 2008 | Sprinkled Latent Semantic Indexing for Text Classification with Background Knowledge
Haiqin Yang, Irwin King |
ICONIP (2) | 1 |
| 2008 | Efficient Minimax Clustering Probability Machine by Generalized Probability Product KernelabstractMinimax Probability Machine (MPM), learning a decision function by minimizing the maximum probability of misclassification, has demonstrated very promising performance in classification and regression. However, MPM is often challenged for its slow training and test procedures. Aiming to solve this problem, we propose an efficient model named Minimax Clustering Probability Machine (MCPM). Following many traditional methods, we represent training data points by several clusters. Different from these methods, a Generalized Probability Product Kernel is appropriately defined to grasp the inner distributional information over the clusters. Incorporating clustering information via a non-linear kernel, MCPM can fast train and test in classification problem with promising performance. Another appealing property of the proposed approach is that MCPM can still derive an explicit worst-case accuracy bound for the decision boundary. Experimental results on synthetic and real data validate the effectiveness of MCPM for classification while attaining high accuracy. Haiqin Yang, Kaizhu Huang, Irwin King, Michael R. Lyu |
IJCNN | 1 |
| 2008 | Maxi-Min Margin Machine: Learning Large Margin Classifiers Locally and GloballyabstractIn this paper, we propose a novel large margin classifier, called the maxi-min margin machine M(4). This model learns the decision boundary both locally and globally. In comparison, other large margin classifiers construct separating hyperplanes only either locally or globally. For example, a state-of-the-art large margin classifier, the support vector machine (SVM), considers data only locally, while another significant model, the minimax probability machine (MPM), focuses on building the decision hyperplane exclusively based on the global information. As a major contribution, we show that SVM yields the same solution as M(4) when data satisfy certain conditions, and MPM can be regarded as a relaxation model of M(4). Moreover, based on our proposed local and global view of data, another popular model, the linear discriminant analysis, can easily be interpreted and extended as well. We describe the M(4) model definition, provide a geometrical interpretation, present theoretical justifications, and propose a practical sequential conic programming method to solve the optimization problem. We also show how to exploit Mercer kernels to extend M(4) for nonlinear classifications. Furthermore, we perform a series of evaluations on both synthetic data sets and real-world benchmark data sets. Comparison with SVM and MPM demonstrates the advantages of our new model. Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu |
IEEE Trans. Neural Networks | 2 |
| 2006 | Local Support Vector Regression for Financial Time Series PredictionabstractWe consider the regression problem for financial time series. Typically, financial time series are non-stationary and volatile in nature. Because of its good generalization power and the tractability of the problem, the Support Vector Regression (SVR) has been extensively applied in financial time series prediction. The standard SVR adopts the lp-norm (p = 1 or 2) to model the functional complexity of the whole data set and employs a fixed ε-tube to tolerate noise. Although this approach has proved successful both theoretically and empirically, it considers data in a global fashion only. Therefore it may lack the flexibility to capture the local trend of data; this is a critical aspect of volatile data, especially financial time series data. Aiming to address this issue, we propose the Local Support Vector Regression (LSVR) model. This novel model is demonstrated to provide a systematic and automatic scheme to adapt the margin locally and flexibly; the margin is fixed globally in the standard SVR. Therefore, the LSVR can tolerate noise adaptively. We provide both theoretical justifications and empirical evaluations for this novel model. The experimental results on synthetic data and real financial data demonstrate its advantages over the standard SVR. Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu |
IJCNN | 2 |
| 2006 | Imbalanced learning with a biased minimax probability machineabstractImbalanced learning is a challenged task in machine learning. In this context, the data associated with one class are far fewer than those associated with the other class. Traditional machine learning methods seeking classification accuracy over a full range of instances are not suitable to deal with this problem, since they tend to classify all the data into a majority class, usually the less important class. In this correspondence, the authors describe a new approach named the biased minimax probability machine (BMPM) to deal with the problem of imbalanced learning. This BMPM model is demonstrated to provide an elegant and systematic way for imbalanced learning. More specifically, by controlling the accuracy of the majority class under all possible choices of class-conditional densities with a given mean and covariance matrix, this model can quantitatively and systematically incorporate a bias for the minority class. By establishing an explicit connection between the classification accuracy and the bias, this approach distinguishes itself from the many current imbalanced-learning methods; these methods often impose a certain bias on the minority data by adapting intermediate factors via the trial-and-error procedure. The authors detail the theoretical foundation, prove its solvability, propose an efficient optimization algorithm, and perform a series of experiments to evaluate the novel model. The comparison with other competitive methods demonstrates the effectiveness of this new model. Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2004 | Learning Classifiers from Imbalanced Data Based on Biased Minimax Probability Machine
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu |
CVPR (2) | 2 |
| 2004 | Learning large margin classifiers locally and globallyabstractA new large margin classifier, named Maxi-Min Margin Machine (M4) is proposed in this paper. This new classifier is constructed based on both a "local: and a "global" view of data, while the most popular large margin classifier, Support Vector Machine (SVM) and the recently-proposed important model, Minimax Probability Machine (MPM) consider data only either locally or globally. This new model is theoretically important in the sense that SVM and MPM can both be considered as its special case. Furthermore, the optimization of M4 can be cast as a sequential conic programming problem, which can be solved efficiently. We describe the M4 model definition, provide a clear geometrical interpretation, present theoretical justifications, propose efficient solving methods, and perform a series of evaluations on both synthetic data sets and real world benchmark data sets. Its comparison with SVM and MPM also demonstrates the advantages of our new model. Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu |
ICML | 2 |
| 2004 | Outliers Treatment in Support Vector Regression for Financial Time Series Prediction
Haiqin Yang, Kaizhu Huang, Lai-Wan Chan, Irwin King, Michael R. Lyu |
ICONIP | 1 |
| 2004 | The Minimum Error Minimax Probability Machine
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu, Lai-Wan Chan |
J. Mach. Learn. Res. | 2 |
| 2002 | Support Vector Machine Regression for Volatile Stock Market Prediction
Haiqin Yang, Lai-Wan Chan, Irwin King |
IDEAL | 1 |