VLDB 2026 Research / reviewers in the wild / expert
Yanjun Qi
dblp:81/3848
· DBLP profile ↗
56ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-5796-7453ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 18 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 12 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Security and privacy · 3Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Preference Optimization via Contrastive Divergence: Your Policy Is Secretly an NLL EstimatorabstractExisting studies on preference optimization (PO) have been focused on constructing pairwise preference data following simple heuristics, such as maximizing the margin between chosen and rejected responses based on human (or AI) ratings. In this work, we develop a novel PO framework that provides theoretical guidance to effectively sample rejected responses. To achieve this, we formulate PO as minimizing the negative log-likelihood (NLL) of a probability model and propose a sampling-based solution to estimate its normalization constant via contrastive divergence. We show that these estimative samples can act as rejected responses in PO. Leveraging the connection established between PO and NLL estimation, we propose a novel PO algorithm, called Monte-Carlo-based PO (MC-PO), that applies a MC kernel to sample *hard negatives* w.r.t.~the log-likelihood of the target policy. Intuitively, these hard negatives represent the rejected samples that are most difficult for the current policy to differentiate. We show that MC-PO outperforms existing SOTA baselines on popular alignment benchmarks. Zhoutong Chen, Xuan Zhu 0002, Haozhu Wang, Yanjun Qi, Mohammad Ghavamzadeh |
AAAI | 6 |
| 2024 | Launchpad: Learning to Schedule Using Offline and Online RL Methods
Vanamala Venkataswamy, Jake Grigsby, Andrew Grimshaw, Yanjun Qi |
JSSPP | 4 |
| 2023 | Improving Interpretability via Explicit Word Interaction Graph LayerabstractRecent NLP literature has seen growing interest in improving model interpretability. Along this direction, we propose a trainable neural network layer that learns a global interaction graph between words and then selects more informative words using the learned word interactions. Our layer, we call WIGRAPH, can plug into any neural network-based NLP text classifiers right after its word embedding layer. Across multiple SOTA NLP models and various NLP datasets, we demonstrate that adding the WIGRAPH layer substantially improves NLP models' interpretability and enhances models' prediction performance at the same time. Arshdeep Sekhon, Aman Shrivastava, Zhe Wang 0025, Yangfeng Ji, Yanjun Qi |
AAAI | 6 |
| 2023 | PGrad: Learning Principal Gradients For Domain Generalization
Zhe Wang 0025, Jake Grigsby, Yanjun Qi |
ICLR | 3 |
| 2022 | Beyond Data Samples: Aligning Differential Networks Estimation with Scientific KnowledgeabstractLearning the differential statistical dependency network between two contexts is essential for many real-life applications, mostly in the high dimensional low sample regime. In this paper, we propose a novel differential network estimator that allows integrating various sources of knowledge beyond data samples. The proposed estimator is scalable to a large number of variables and achieves a sharp asymptotic convergence rate. Empirical experiments on extensive simulated data and four real-world applications (one on neuroimaging and three from functional genomics) show that our approach achieves improved differential network estimation and provides better supports to downstream tasks like classification. Our results highlight significant benefits of integrating group, spatial and anatomic knowledge during differential genetic network identification and brain connectome change discovery. Arshdeep Sekhon, Zhe Wang 0025, Yanjun Qi |
AISTATS | 3 |
| 2022 | RARE: Renewable Energy Aware Resource Management in Datacenters
Vanamala Venkataswamy, Jake Grigsby, Andrew Grimshaw, Yanjun Qi |
JSSPP | 4 |
| 2022 | ST-MAML : A stochastic-task based method for task-heterogeneous meta-learningabstractOptimization-based meta-learning typically assumes tasks are sampled from a single distribution - an assumption that oversimplifies and limits the diversity of tasks that meta-learning can model. Handling tasks from multiple distributions is challenging for meta-learning because it adds ambiguity to task identities. This paper proposes a novel method, ST-MAML, that empowers model-agnostic meta-learning (MAML) to learn from multiple task distributions. ST-MAML encodes tasks using a stochastic neural network module, that summarizes every task with a stochastic representation. The proposed Stochastic Task (ST) strategy learns a distribution of solutions for an ambiguous task and allows a meta-model to self-adapt to the current task. ST-MAML also propagates the task representation to enhance input variable encodings. Empirically, we demonstrate that ST-MAML outperforms the state-of-the-art on two few-shot image classification tasks, one curve regression benchmark, one image completion problem, and a real-world temperature prediction application. Zhe Wang 0025, Jake Grigsby, Arshdeep Sekhon, Yanjun Qi |
UAI | 4 |
| 2022 | Forecasting Cloud Application Workloads With CloudInsight for Predictive Resource ManagementabstractPredictive cloud resource management has been widely adopted to overcome the limitations of reactive cloud autoscaling. The predictive resource management is highly relying on workload predictors, which estimate short-/long-term fluctuations of cloud application workloads. These predictors tend to be pre-optimized for specific workload patterns. However, such predictors are still insufficient to handle real-world cloud workloads whose patterns may be unknown a priori, may dynamically change over time and may be irregular. As a result, these predictors often cause over-/under-provisioning of cloud resources. To address this problem, we have created CloudInsight, a novel cloud workload prediction framework, leveraging the combined power of multiple workload predictors. CloudInsight creates an ensemble model using multiple predictors to make accurate predictions for real workloads. The weights of the predictors in CloudInsight are determined at runtime with their accuracy for the current workload using multi-class regression. The ensemble model is periodically optimized to handle sudden changes in the workload. We evaluated CloudInsight with various real workload traces. The results show that CloudInsight has 13–27 percent higher accuracy than state-of-the-art predictors. Moreover, the results from trace-based simulations with a cloud resource manager show that CloudInsight has 15–20 percent less under-/over-provisioning periods, resulting in high cost-efficiency and low SLA violations. In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
IEEE Trans. Cloud Comput. | 3 |
| 2021 | Curriculum Labeling: Revisiting Pseudo-Labeling for Semi-Supervised LearningabstractIn this paper we revisit the idea of pseudo-labeling in the context of semi-supervised learning where a learning algorithm has access to a small set of labeled samples and a large set of unlabeled samples. Pseudo-labeling works by applying pseudo-labels to samples in the unlabeled set by using a model trained on combination of the labeled samples and any previously pseudo-labeled samples, and iteratively repeating this process in a self-training cycle. Current methods seem to have abandoned this approach in favor of consistency regularization methods that train models under a combination of different styles of self-supervised losses on the unlabeled samples and standard supervised losses on the labeled samples. We empirically demonstrate that pseudo-labeling can in fact be competitive with the state-of-the-art, while being more resilient to out-of-distribution samples in the unlabeled set. We identify two key factors that allow pseudo-labeling to achieve such remarkable results (1) applying curriculum learning principles and (2) avoiding concept drift by restarting model parameters before each self-training cycle. We obtain 94.91% accuracy on CIFAR-10 using only 4,000 labeled samples, and 68.87% top-1 accuracy on Imagenet-ILSVRC using only 10% of the labeled samples. Paola Cascante-Bonilla, Fuwen Tan, Yanjun Qi, Vicente Ordonez |
AAAI | 3 |
| 2021 | Evolving Image Compositions for Feature Representation Learning
Paola Cascante-Bonilla, Arshdeep Sekhon, Yanjun Qi, Vicente Ordonez |
BMVC | 3 |
| 2021 | General Multi-Label Image Classification With TransformersabstractMulti-label image classification is the task of predicting a set of labels corresponding to objects, attributes or other entities present in an image. In this work we propose the Classification Transformer (C-Tran), a general framework for multi-label image classification that leverages Transformers to exploit the complex dependencies among visual features and labels. Our approach consists of a Transformer encoder trained to predict a set of target labels given an input set of masked labels, and visual features from a convolutional neural network. A key ingredient of our method is a label mask training objective that uses a ternary encoding scheme to represent the state of the labels as positive, negative, or unknown during training. Our model shows state-of-the-art performance on challenging datasets such as COCO and Visual Genome. Moreover, because our model explicitly represents the label state during training, it is more general by allowing us to produce improved results for images with partial or extra label annotations during inference. We demonstrate this additional capability in the COCO, Visual Genome, News-500, and CUB image datasets. Jack Lanchantin, Vicente Ordonez, Yanjun Qi |
CVPR | 4 |
| 2020 | FastSK: fast sequence analysis with gapped string kernelsabstractMOTIVATION: Gapped k-mer kernels with support vector machines (gkm-SVMs) have achieved strong predictive performance on regulatory DNA sequences on modestly sized training sets. However, existing gkm-SVM algorithms suffer from slow kernel computation time, as they depend exponentially on the sub-sequence feature length, number of mismatch positions, and the task's alphabet size. RESULTS: In this work, we introduce a fast and scalable algorithm for calculating gapped k-mer string kernels. Our method, named FastSK, uses a simplified kernel formulation that decomposes the kernel calculation into a set of independent counting operations over the possible mismatch positions. This simplified decomposition allows us to devise a fast Monte Carlo approximation that rapidly converges. FastSK can scale to much greater feature lengths, allows us to consider more mismatches, and is performant on a variety of sequence analysis tasks. On multiple DNA transcription factor binding site prediction datasets, FastSK consistently matches or outperforms the state-of-the-art gkmSVM-2.0 algorithms in area under the ROC curve, while achieving average speedups in kernel computation of ∼100× and speedups of ∼800× for large feature lengths. We further show that FastSK outperforms character-level recurrent and convolutional neural networks while achieving low variance. We then extend FastSK to 7 English-language medical named entity recognition datasets and 10 protein remote homology detection datasets. FastSK consistently matches or outperforms these baselines. AVAILABILITY AND IMPLEMENTATION: Our algorithm is available as a Python package and as C++ source code at https://github.com/QData/FastSK. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Derrick Blakely, Eamon Collins, Ritambhara Singh, Andrew P. Norton, Jack Lanchantin, Yanjun Qi |
Bioinform. | 6 |
| 2020 | Graph convolutional networks for epigenetic state prediction using both sequence and 3D genome dataabstractMOTIVATION: Predictive models of DNA chromatin profile (i.e. epigenetic state), such as transcription factor binding, are essential for understanding regulatory processes and developing gene therapies. It is known that the 3D genome, or spatial structure of DNA, is highly influential in the chromatin profile. Deep neural networks have achieved state of the art performance on chromatin profile prediction by using short windows of DNA sequences independently. These methods, however, ignore the long-range dependencies when predicting the chromatin profiles because modeling the 3D genome is challenging. RESULTS: In this work, we introduce ChromeGCN, a graph convolutional network for chromatin profile prediction by fusing both local sequence and long-range 3D genome information. By incorporating the 3D genome, we relax the independent and identically distributed assumption of local windows for a better representation of DNA. ChromeGCN explicitly incorporates known long-range interactions into the modeling, allowing us to identify and interpret those important long-range dependencies in influencing chromatin profiles. We show experimentally that by fusing sequential and 3D genome data using ChromeGCN, we get a significant improvement over the state-of-the-art deep learning methods as indicated by three metrics. Importantly, we show that ChromeGCN is particularly useful for identifying epigenetic effects in those DNA windows that have a high degree of interactions with other DNA windows. AVAILABILITY AND IMPLEMENTATION: https://github.com/QData/ChromeGCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jack Lanchantin, Yanjun Qi |
Bioinform. | 2 |
| 2019 | Neural Message Passing for Multi-label Classification
Jack Lanchantin, Arshdeep Sekhon, Yanjun Qi |
ECML/PKDD (2) | 3 |
| 2019 | Transfer String Kernel for Cross-Context DNA-Protein Binding PredictionabstractThrough sequence-based classification, this paper tries to accurately predict the DNA binding sites of transcription factors (TFs) in an unannotated cellular context. Related methods in the literature fail to perform such predictions accurately, since they do not consider sample distribution shift of sequence segments from an annotated (source) context to an unannotated (target) context. We, therefore, propose a method called "Transfer String Kernel" (TSK) that achieves improved prediction of transcription factor binding site (TFBS) using knowledge transfer via cross-context sample adaptation. TSK maps sequence segments to a high-dimensional feature space using a discriminative mismatch string kernel framework. In this high-dimensional space, labeled examples of the source context are re-weighted so that the revised sample distribution matches the target context more closely. We have experimentally verified TSK for TFBS identifications on 14 different TFs under a cross-organism setting. We find that TSK consistently outperforms the state-of-the-art TFBS tools, especially when working with TFs whose binding sequences are not conserved across contexts. We also demonstrate the generalizability of TSK by showing its cutting-edge performance on a different set of cross-context tasks for the MHC peptide binding predictions. Ritambhara Singh, Jack Lanchantin, Gabriel Robins, Yanjun Qi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | CloudInsight: Utilizing a Council of Experts to Predict Future Cloud Application WorkloadsabstractMany predictive approaches have been proposed to overcome the limitations of reactive autoscaling on clouds. These approaches leverage workload predictors that are usually targeted for a particular workload pattern and can fail to handle real-world cloud workloads whose patterns may be unknown a priori, may dynamically change over time, or may be irregular. The result is that resources are frequently under-and overprovisioned. To address this problem, we create a novel cloud workload prediction framework called CloudInsight, leveraging the combined power of multiple workload predictors that collectively provide a "council of experts". The weights of the predictors in this ensemble model are determined in real-time based on their accuracy for current workload using multi-class regression. Under real workload traces, CloudInsight has 13% - 27% better accuracy than state-of-the-art predictors. It also has low overhead for predicting future workload changes (<; 100 ms) and creating a new ensemble workload predictor (<; 1.1 sec.). In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
IEEE CLOUD | 3 |
| 2018 | Fast and Scalable Learning of Sparse Changes in High-Dimensional Gaussian Graphical Model StructureabstractWe focus on the problem of estimating the change in the dependency structures of two $p$-dimensional Gaussian Graphical models (GGMs). Previous studies for sparse change estimation in GGMs involve expensive and difficult non-smooth optimization. We propose a novel method, DIFFEE for estimating DIFFerential networks via an Elementary Estimator under a high-dimensional situation. DIFFEE is solved through a faster and closed form solution that enables it to work in large-scale settings. We conduct a rigorous statistical analysis showing that surprisingly DIFFEE achieves the same asymptotic convergence rates as the state-of-the-art estimators that are much more difficult to compute. Our experimental results on multiple synthetic datasets and one real-world data about brain connectivity show strong performance improvements over baselines, as well as significant computational benefits. Beilun Wang, Arshdeep Sekhon, Yanjun Qi |
AISTATS | 3 |
| 2018 | A Fast and Scalable Joint Estimator for Integrating Additional Knowledge in Learning Multiple Related Sparse Gaussian Graphical ModelsabstractWe consider the problem of including additional knowledge in estimating sparse Gaussian graphical models (sGGMs) from aggregated samples, arising often in bioinformatics and neuroimaging applications. Previous joint sGGM estimators either fail to use existing knowledge or cannot scale-up to many tasks (large $K$) under a high-dimensional (large $p$) situation. In this paper, we propose a novel \underline{J}oint \underline{E}lementary \underline{E}stimator incorporating additional \underline{K}nowledge (JEEK) to infer multiple related sparse Gaussian Graphical models from large-scale heterogeneous data. Using domain knowledge as weights, we design a novel hybrid norm as the minimization objective to enforce the superposition of two weighted sparsity constraints, one on the shared interactions and the other on the task-specific structural patterns. This enables JEEK to elegantly consider various forms of existing knowledge based on the domain at hand and avoid the need to design knowledge-specific optimization. JEEK is solved through a fast and entry-wise parallelizable solution that largely improves the computational efficiency of the state-of-the-art $O(p^5K^4)$ to $O(p^2K^4)$. We conduct a rigorous statistical analysis showing that JEEK achieves the same convergence rate $O(\log(Kp)/n_{tot})$ as the state-of-the-art estimators that are much harder to compute. Empirically, on multiple synthetic datasets and one real-world data from neuroscience, JEEP outperforms the speed of the state-of-arts significantly while achieving the same level of prediction accuracy. Beilun Wang, Arshdeep Sekhon, Yanjun Qi |
ICML | 3 |
| 2018 | Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks
Weilin Xu, David Evans 0001, Yanjun Qi |
NDSS | 3 |
| 2018 | DeepDiff: DEEP-learning for predicting DIFFerential gene expression from histone modificationsabstractMotivation: Computational methods that predict differential gene expression from histone modification signals are highly desirable for understanding how histone modifications control the functional heterogeneity of cells through influencing differential gene regulation. Recent studies either failed to capture combinatorial effects on differential prediction or primarily only focused on cell type-specific analysis. In this paper we develop a novel attention-based deep learning architecture, DeepDiff, that provides a unified and end-to-end solution to model and to interpret how dependencies among histone modifications control the differential patterns of gene regulation. DeepDiff uses a hierarchy of multiple Long Short-Term Memory (LSTM) modules to encode the spatial structure of input signals and to model how various histone modifications cooperate automatically. We introduce and train two levels of attention jointly with the target prediction, enabling DeepDiff to attend differentially to relevant modifications and to locate important genome positions for each modification. Additionally, DeepDiff introduces a novel deep-learning based multi-task formulation to use the cell-type-specific gene expression predictions as auxiliary tasks, encouraging richer feature embeddings in our primary task of differential expression prediction. Results: Using data from Roadmap Epigenomics Project (REMC) for ten different pairs of cell types, we show that DeepDiff significantly outperforms the state-of-the-art baselines for differential gene expression prediction. The learned attention weights are validated by observations from previous studies about how epigenetic mechanisms connect to differential gene expression. Availability and implementation: Codes and results are available at deepchrome.org. Supplementary information: Supplementary data are available at Bioinformatics online. Arshdeep Sekhon, Ritambhara Singh, Yanjun Qi |
Bioinform. | 3 |
| 2017 | A Fast and Scalable Joint Estimator for Learning Multiple Related Sparse Gaussian Graphical ModelsabstractEstimating multiple sparse Gaussian Graphical Models (sGGMs) jointly for many related tasks (large $K$) under a high-dimensional (large $p$) situation is an important task. Most previous studies for the joint estimation of multiple sGGMs rely on penalized log-likelihood estimators that involve expensive and difficult non-smooth optimizations. We propose a novel approach, FASJEM for \underlinefast and \underlinescalable \underlinejoint structure-\underlineestimation of \underlinemultiple sGGMs at a large scale. As the first study of joint sGGM using the M-estimator framework, our work has three major contributions: (1) We solve FASJEM through an entry-wise manner which is parallelizable. (2) We choose a proximal algorithm to optimize FASJEM. This improves the computational efficiency from $O(Kp^3)$ to $O(Kp^2)$ and reduces the memory requirement from $O(Kp^2)$ to $O(K)$. (3) We theoretically prove that FASJEM achieves a consistent estimation with a convergence rate of $O(\log(Kp)/n_tot)$. On several synthetic and four real-world datasets, FASJEM shows significant improvements over baselines on accuracy, computational complexity and memory costs. Beilun Wang, Ji Gao, Yanjun Qi |
AISTATS | 3 |
| 2017 | Attend and Predict: Understanding Gene Regulation by Selective Attention on ChromatinabstractThe past decade has seen a revolution in genomic technologies that enabled a flood of genome-wide profiling of chromatin marks. Recent literature tried to understand gene regulation by predicting gene expression from large-scale chromatin measurements. Two fundamental challenges exist for such learning tasks: (1) genome-wide chromatin signals are spatially structured, high-dimensional and highly modular; and (2) the core aim is to understand what are the relevant factors and how they work together. Previous studies either failed to model complex dependencies among input signals or relied on separate feature analysis to explain the decisions. This paper presents an attention-based deep learning approach; AttentiveChrome, that uses a unified architecture to model and to interpret dependencies among chromatin factors for controlling gene regulation. AttentiveChrome uses a hierarchy of multiple Long Short-Term Memory (LSTM) modules to encode the input signals and to model how various chromatin marks cooperate automatically. AttentiveChrome trains two levels of attention jointly with the target prediction, enabling it to attend differentially to relevant marks and to locate important positions per mark. We evaluate the model across 56 different cell types (tasks) in human. Not only is the proposed architecture more accurate, but its attention scores also provide a better interpretation than state-of-the-art feature visualization methods such as saliency map. Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, Yanjun Qi |
NIPS | 4 |
| 2017 | GaKCo: A Fast Gapped k-mer String Kernel Using Counting
Ritambhara Singh, Arshdeep Sekhon, Kamran Kowsari, Jack Lanchantin, Beilun Wang, Yanjun Qi |
ECML/PKDD (1) | 6 |
| 2017 | Adversarial-Playground: A visualization suite showing how adversarial examples fool deep learningabstractRecent studies have shown that attackers can force deep learning models to misclassify so-called “adversarial examples:” maliciously generated images formed by making imperceptible modifications to pixel values. With growing interest in deep learning for security applications, it is important for security experts and users of machine learning to recognize how learning systems may be attacked. Due to the complex nature of deep learning, it is challenging to understand how deep models can be fooled by adversarial examples. Thus, we present a web-based visualization tool, Adversarial-Playground, to demonstrate the efficacy of common adversarial methods against a convolutional neural network (CNN) system. Adversarial-Playground is educational, modular and interactive. (1) It enables non-experts to compare examples visually and to understand why an adversarial example can fool a CNN-based image classifier. (2) It can help security experts explore more vulnerability of deep learning as a software module. (3) Building an interactive visualization is challenging in this domain due to the large feature space of image classification (generating adversarial examples is slow in general and visualizing images are costly). Through multiple novel design choices, our tool can provide fast and accurate responses to user requests. Empirically, we find that our client-server division strategy reduced the response time by an average of 1.5 seconds per sample. Our other innovation, a faster variant of JSMA evasion algorithm, empirically performed twice as fast as JSMA and yet maintains a comparable evasion rate1. Andrew P. Norton, Yanjun Qi |
VizSEC | 2 |
| 2017 | A constrained $$\ell $$ ℓ 1 minimization approach for estimating multiple sparse Gaussian or nonparanormal graphical modelsabstractIdentifying context-specific entity networks from aggregated data is an important task, arising often in bioinformatics and neuroimaging applications. Computationally, this task can be formulated as jointly estimating multiple different, but related, sparse undirected graphical models (UGM) from aggregated samples across several contexts. Previous joint-UGM studies have mostly focused on sparse Gaussian graphical models (sGGMs) and can’t identify context-specific edge patterns directly. We, therefore, propose a novel approach, SIMULE (detecting Shared and Individual parts of MULtiple graphs Explicitly) to learn multi-UGM via a constrained $$\ell $$ 1 minimization. SIMULE automatically infers both specific edge patterns that are unique to each context and shared interactions preserved among all the contexts. Through the $$\ell $$ 1 constrained formulation, this problem is cast as multiple independent subtasks of linear programming that can be solved efficiently in parallel. In addition to Gaussian data, SIMULE can also handle multivariate Nonparanormal data that greatly relaxes the normality assumption that many real-world applications do not follow. We provide a novel theoretical proof showing that SIMULE achieves a consistent result at the rate $$O(\log (Kp)/n_{tot})$$ . On multiple synthetic datasets and two biomedical datasets, SIMULE shows significant improvement over state-of-the-art multi-sGGM and single-UGM baselines (SIMULE implementation and the used datasets @ https://github.com/QData/SIMULE ). Beilun Wang, Ritambhara Singh, Yanjun Qi |
Mach. Learn. | 3 |
| 2016 | Empirical Evaluation of Workload Forecasting Techniques for Predictive Cloud Resource ScalingabstractMany predictive resource scaling approaches have been proposed to overcome the limitations of the conventional reactive approaches most often used in clouds today. In general, due to the complexity of clouds, these reactive approaches were often forced to make significant limiting assumptions in either the operating conditions/requirements or expected workload patterns. As such, it is extremely difficult for cloud users to know which - if any - existing workload predictor will work best for their particular cloud activity, especially when considering highly-variable workload patterns, non-trivial billing models, variety of resources to add/subtract, etc. To solve this problem, we conduct comprehensive evaluations for a variety of workload predictors under real-world cloud configurations. The workload predictors cover four classes of 21 predictors: naive, regression, temporal, and non-temporal methods. We simulate a cloud application under four realistic workload patterns, two different cloud billing models, and three different styles of predictive scaling. Our evaluation confirms that no workload predictor is universally best for all workload patterns, and shows that Predictive Scaling-out + Predictive Scaling-in has the best cost efficiency and the lowest job deadline miss rate in cloud resource management, on average providing 30% better cost efficiency and 80% less job deadline miss rate compared to other styles of predictive scaling. In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
CLOUD | 3 |
| 2016 | MUST-CNN: A Multilayer Shift-and-Stitch Deep Convolutional Architecture for Sequence-Based Protein Structure PredictionabstractPredicting protein properties such as solvent accessibility and secondary structure from its primary amino acid sequence is an important task in bioinformatics. Recently, a few deep learning models have surpassed the traditional window based multilayer perceptron. Taking inspiration from the image classification domain we propose a deep convolutional neural network architecture, MUST-CNN, to predict protein properties. This architecture uses a novel multilayer shift-and-stitch (MUST) technique to generate fully dense per-position predictions on protein sequences. Our model is significantly simpler than the state-of-the-art, yet achieves better results. By combining MUST and the efficient convolution operation, we can consider far more parameters while retaining very fast prediction speeds. We beat the state-of-the-art performance on two large protein property prediction datasets. Zeming Lin, Jack Lanchantin, Yanjun Qi |
AAAI | 3 |
| 2016 | Automatically Evading Classifiers: A Case Study on PDF Malware Classifiers
Weilin Xu, Yanjun Qi, David Evans 0001 |
NDSS | 2 |
| 2016 | DeepChrome: deep-learning for predicting gene expression from histone modificationsabstractMOTIVATION: Histone modifications are among the most important factors that control gene regulation. Computational methods that predict gene expression from histone modification signals are highly desirable for understanding their combinatorial effects in gene regulation. This knowledge can help in developing 'epigenetic drugs' for diseases like cancer. Previous studies for quantifying the relationship between histone modifications and gene expression levels either failed to capture combinatorial effects or relied on multiple methods that separate predictions and combinatorial analysis. This paper develops a unified discriminative framework using a deep convolutional neural network to classify gene expression using histone modification data as input. Our system, called DeepChrome, allows automatic extraction of complex interactions among important features. To simultaneously visualize the combinatorial interactions among histone modifications, we propose a novel optimization-based technique that generates feature pattern maps from the learnt deep model. This provides an intuitive description of underlying epigenetic mechanisms that regulate genes. RESULTS: We show that DeepChrome outperforms state-of-the-art models like Support Vector Machines and Random Forests for gene expression classification task on 56 different cell-types from REMC database. The output of our visualization technique not only validates the previous observations but also allows novel insights about combinatorial interactions among histone modification marks, some of which have recently been observed by experimental studies. AVAILABILITY AND IMPLEMENTATION: Codes and results are available at www.deepchrome.org CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ritambhara Singh, Jack Lanchantin, Gabriel Robins, Yanjun Qi |
Bioinform. | 4 |
| 2016 | Piecewise Linear Dynamical Model for Action Clustering from Real-World Deployments of Inertial Body SensorsabstractHuman motion has been reported as having great relevance to various disease, disorder, injuries and emotional state. Therefore, motion assessment using inertial body sensor networks (BSNs) is gaining popularity as an outcome measure in clinical study and neuroscience research. The efficacy of motion assessment heavily relies on the accurate temporal clustering of human motion into actions on various time scales. However, two human factors in real-world deployments of inertial BSNs make such motion assessment challenging: mounting errors (where sensor displacement and orientation do not match what is assumed by processing algorithms) and insecure mounting (where sensors are loosely worn causing them to shake during operations). In order to enhance the robustness of human actions clustering from real-world BSN data, this work leverages dynamical systems modeling with the considerations of human factors. By proposing a computational body-model framework called the piecewise linear dynamical model (PLDM), we derive a robust method to segment time series data of inertial BSNs in real-world deployment with human factors into motion primitives and actions. We test the proposed method on three different inertial BSN datasets, extract actions on different temporal scales and recognize the actions into clusters. The experimental results demonstrate the effectiveness of our approach. Jiaqi Gong, Philip Asare, Yanjun Qi, John C. Lach |
IEEE Trans. Affect. Comput. | 3 |
| 2016 | Causality Analysis of Inertial Body Sensors for Multiple Sclerosis Diagnostic EnhancementabstractInertial body sensors have emerged in recent years as an effective tool for evaluating mobility impairment resulting from various diseases, disorders, and injuries. For example, body sensors have been used in 6-min walk (6 MW) tests for multiple sclerosis (MS) patients to identify gait features useful in the study, diagnosis, and tracking of the disease. However, most studies to date have focused on features localized to the lower or upper extremities and do not provide a holistic assessment of mobility. This paper presents a causality analysis method focused on the coordination between extremities to identify subtle whole-body mobility impairment that may aid disease diagnosis. This method was developed for and utilized in an MS pilot study with 41 subjects (28 persons with MS (PwMS) and 13 healthy controls) performing 6 MW tests. Compared with existing methods, the causality analysis provided better discrimination between healthy controls and PwMS and a deeper understanding of MS disease impact on mobility. Jiaqi Gong, Yanjun Qi, Myla D. Goldman, John C. Lach |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | Kernelized Information-Theoretic Metric Learning for Cancer Diagnosis Using High-Dimensional Molecular Profiling DataabstractWith the advancement of genome-wide monitoring technologies, molecular expression data have become widely used for diagnosing cancer through tumor or blood samples. When mining molecular signature data, the process of comparing samples through an adaptive distance function is fundamental but difficult, as such datasets are normally heterogeneous and high dimensional. In this article, we present kernelized information-theoretic metric learning (KITML) algorithms that optimize a distance function to tackle the cancer diagnosis problem and scale to high dimensionality. By learning a nonlinear transformation in the input space implicitly through kernelization, KITML permits efficient optimization, low storage, and improved learning of distance metric. We propose two novel applications of KITML for diagnosing cancer using high-dimensional molecular profiling data: (1) for sample-level cancer diagnosis, the learned metric is used to improve the performance of k -nearest neighbor classification; and (2) for estimating the severity level or stage of a group of samples, we propose a novel set-based ranking approach to extend KITML. For the sample-level cancer classification task, we have evaluated on 14 cancer gene microarray datasets and compared with eight other state-of-the-art approaches. The results show that our approach achieves the best overall performance for the task of molecular-expression-driven cancer sample diagnosis. For the group-level cancer stage estimation, we test the proposed set-KITML approach using three multi-stage cancer microarray datasets, and correctly estimated the stages of sample groups for all three studies. Feiyu Xiong, Moshe Kam, Leonid Hrebien, Beilun Wang, Yanjun Qi |
ACM Trans. Knowl. Discov. Data | 5 |
| 2015 | Causal analysis of inertial body sensors for enhancing gait assessment separability towards multiple sclerosis diagnosisabstractGait assessment is a common method for diagnosing various diseases, disorders, and injuries, studying their impact on mobility, and evaluating the efficacy of various therapeutic interventions. The recent emergence of inertial body sensors for gait assessment addresses the limitations of visual observation and subjective clinical evaluation by providing more precise and objective measures. Inertial sensors have been included in an ongoing study at the University of Virginia Medical Center on Multiple Sclerosis (MS), a chronic autoimmune disorder of the central nervous system (CNS) that produces neurologic impairment and functional disability over time, with the goal of improving the ability to assess MS-affected gait and to distinguish between subjects with MS and those without MS. This work presents a gait assessment technique based on causal modeling to distinguish MS-affected gait and healthy gait. The approach in this work is based on the hypothesis that the strength of interaction between body parts during walking is greater in healthy controls that in MS subjects. The strength of interaction was quantified using a causality index based on the pairwise causal relationships between body parts as characterized by the Phase Slope Index (PSI) of inertial signals from pairs of body parts. In a pilot study with 41 subjects (28 MS subjects and 13 healthy controls), the approach developed in this paper provided better separability (p <; 0.0001) compared with existing methods. Jiaqi Gong, John C. Lach, Yanjun Qi, Myla D. Goldman |
BSN | 3 |
| 2015 | MAPer: A Multi-scale Adaptive Personalized Model for Temporal Human Behavior PredictionabstractThe primary objective of this research is to develop a simple and interpretable predictive framework to perform temporal modeling of individual user's behavior traits based on each person's past observed traits/behavior. Individual-level human behavior patterns are possibly influenced by various temporal features (e.g., lag, cycle) and vary across temporal scales (e.g., hour of the day, day of the week). Most of the existing forecasting models do not capture such multi-scale adaptive regularity of human behavior or lack interpretability due to relying on hidden variables. Hence, we build a multi-scale adaptive personalized (MAPer) model that quantifies the effect of both lag and behavior cycle for predicting future behavior. MAper includes a novel basis vector to adaptively learn behavior patterns and capture the variation of lag and cycle across multi-scale temporal contexts. We also extend MAPer to capture the interaction among multiple behaviors to improve the prediction performance. Sarah Masud Preum, John A. Stankovic, Yanjun Qi |
CIKM | 3 |
| 2015 | Association Rule Mining with the Micron Automata ProcessorabstractAssociation rule mining (ARM) is a widely used data mining technique for discovering sets of frequently associated items in large databases. As datasets grow in size and real-time analysis becomes important, the performance of ARM implementation can impede its applicability. We accelerate ARM by using Micron's Automata Processor (AP), a hardware implementation of non-deterministic finite automata (NFAs), with additional features that significantly expand the APs capabilities beyond those of traditional NFAs. The Apriori algorithm that ARM uses for discovering item sets maps naturally to the massive parallelism of the AP. We implement the multipass pruning strategy used in the Apriori ARM through the APs symbol replacement capability, a form of lightweight reconfigurability. Up to 129X and 49X speedups are achieved by the AP-accelerated Apriori on seven synthetic and real-world datasets, when compared with the Apriori single-core CPU implementation and Eclat, a more efficient ARM algorithm, 6-core multicourse CPU implementation, respectively. The AP-accelerated Apriori solution also outperforms GPU implementations of Eclat especially for large datasets. Technology scaling projections suggest even better speedups from future generations of AP. Ke Wang 0011, Yanjun Qi, Jeffrey J. Fox, Mircea R. Stan, Kevin Skadron |
IPDPS | 2 |
| 2014 | Deep Learning for Character-Based Information Extraction
Yanjun Qi, Sujatha G. Das, Ronan Collobert, Jason Weston |
ECIR | 1 |
| 2014 | Extracting Researcher Metadata with Labeled FeaturesabstractProfessional homepages of researchers contain metadata that provides crucial evidence in several digital library tasks such as academic network extraction, record linkage and expertise search. Due to inherent diversity in values for certain metadata fields (e.g., affiliation) supervised algorithms require a large number of labeled examples for accurately identifying values for these fields. We address this issue with feature labeling, a recent semi-supervised machine learning technique. We apply feature labeling to researcher metadata extraction from homepages by combining a small set of expert-provided feature distributions with few fully-labeled examples. We study two types of labeled features: (1) Dictionary features provide unigram hints related to specific metadata fields, whereas, (2) Proximity features capture the layout information between metadata fields on a homepage in a second stage. We experimentally show that this two-stage approach along with labeled features provides significant improvements in the tagging performance. In one experiment with only ten labeled homepages and 22 expert-specified labeled features, we obtained a 45% relative increase in the F1 value for the affiliation field, while the overall F1 improves by 9%. Sujatha Das Gollapalli, Yanjun Qi, Prasenjit Mitra 0001, C. Lee Giles |
SDM | 2 |
| 2014 | Unsupervised Feature Learning by Deep Sparse CodingabstractIn this paper, we propose a new unsupervised feature learning framework, namely Deep Sparse Coding (DeepSC), that extends sparse coding to a multi-layer architecture for visual object recognition tasks. The main innovation of the framework is that it connects the sparse-encoders from different layers by a sparse-to-dense module. The sparse-to-dense module is a composition of a local spatial pooling step and a low-dimensional embedding process, which takes advantage of the spatial smoothness information in the image. As a result, the new method is able to learn multiple layers of sparse representations of the image which capture features at a variety of abstraction levels and simultaneously preserve the spatial smoothness between the neighboring image patches. Combining the feature representations from multiple layers, DeepSC achieves the state-of-the-art performance on multiple object recognition tasks. Koray Kavukcuoglu, Arthur Szlam, Yanjun Qi |
SDM | 5 |
| 2012 | Large-scale image classification using supervised spatial encoder
Dmitriy Bespalov, Yanjun Qi, Ali Shokoufandeh |
ICPR | 2 |
| 2012 | Learning the Dependency Structure of Latent FactorsabstractIn this paper, we study latent factor models with the dependency structure in the latent space. We propose a general learning framework which induces sparsity on the undirected graphical model imposed on the vector of latent factors. A novel latent factor model SLFA is then proposed as a matrix factorization problem with a special regularization term that encourages collaborative reconstruction. The main benefit (novelty) of the model is that we can simultaneously learn the lower-dimensional representation for data and model the pairwise relationships between latent factors explicitly. An on-line learning algorithm is devised to make the model feasible for large-scale learning problems. Experimental results on two synthetic data and two real-world data sets demonstrate that pairwise relationships and latent factors learned by our model provide a more structured way of exploring high-dimensional data, and the learned representations achieve the state-of-the-art classification performance. Yanjun Qi, Koray Kavukcuoglu, Haesun Park |
NIPS | 2 |
| 2012 | Sentiment Classification with Supervised Sequence Embedding
Dmitriy Bespalov, Yanjun Qi, Ali Shokoufandeh |
ECML/PKDD (1) | 2 |
| 2011 | Sentiment classification based on supervised latent n-gram analysisabstractIn this paper, we propose an efficient embedding for modeling higher-order (n-gram) phrases that projects the n-grams to low-dimensional latent semantic space, where a classification function can be defined. We utilize a deep neural network to build a unified discriminative framework that allows for estimating the parameters of the latent space as well as the classification function with a bias for the target classification task at hand. We apply the framework to large-scale sentimental classification task. We present comparative evaluation of the proposed method on two (large) benchmark data sets for online product reviews. The proposed method achieves superior performance in comparison to the state of the art. Dmitriy Bespalov, Yanjun Qi, Ali Shokoufandeh |
CIKM | 3 |
| 2011 | Sparse Latent Semantic AnalysisabstractLatent semantic analysis (LSA), as one of the most popular unsupervised dimension reduction tools, has a wide range of applications in text mining and information retrieval. The key idea of LSA is to learn a projection matrix that maps the high dimensional vector space representations of documents to a lower dimensional latent space, i.e. so called latent topic space. In this paper, we propose a new model called Sparse LSA, which produces a sparse projection matrix via the ℓ1 regularization. Compared to the traditional LSA, Sparse LSA selects only a small number of relevant words for each topic and hence provides a compact representation of topic-word relationships. Moreover, Sparse LSA is computationally very efficient with much less memory usage for storing the projection matrix. Furthermore, we propose two important extensions of Sparse LSA: group structured Sparse LSA and non-negative Sparse LSA. We conduct experiments on several benchmark datasets and compare Sparse LSA and its extensions with several widely used methods, e.g. LSA, Sparse Coding and LDA. Empirical results suggest that Sparse LSA achieves similar performance gains to LSA, but is more efficient in projection computation, storage, and also well explain the topic-word relationships. Xi Chen 0010, Yanjun Qi, Qihang Lin, Jaime G. Carbonell |
SDM | 2 |
| 2011 | Semi-Supervised Convolution Graph Kernels for Relation ExtractionabstractExtracting semantic relations between entities is an important step towards automatic text understanding. In this paper, we propose a novel Semi-supervised Convolution Graph Kernel (SCGK) method for semantic Relation Extraction (RE) from natural language. By encoding English sentences as dependence graphs among words, SCGK computes kernels (similarities) between sentences using a convolution strategy, i.e., calculating similarities over all possible short single paths from two dependence graphs. Furthermore, SCGK adds three semi-supervised strategies in the kernel calculation to incorporate soft-matches between (1) words, (2) grammatical dependencies, and (3) entire sentences, respectively. From a large unannotated corpus, these semi-supervision steps learn to capture contextual semantic patterns of elements in natural sentences, which therefore alleviate the lack of annotated examples in most RE corpora. Through convolutions and multi-level semi-supervisions, SCGK provides a powerful model to encode both syntactic and semantic evidence existing in natural English sentences, which effectively recovers the target relational patterns of interest. We perform extensive experiments on five RE benchmark datasets which aim to identify interaction relations from biomedical literature. Our results demonstrate that SCGK achieves the state-of-the-art performance on the task of semantic relation extraction. Xia Ning, Yanjun Qi |
SDM | 2 |
| 2010 | Learning Preferences with Millions of Parameters by Enforcing SparsityabstractWe study the retrieval task that ranks a set of objects for a given query in the pair wise preference learning framework. Recently researchers found out that raw features (e.g. words for text retrieval) and their pair wise features which describe relationships between two raw features (e.g. word synonymy or polysemy) could greatly improve the retrieval precision. However, most existing methods can not scale up to problems with many raw features (e.g. English vocabulary), due to the prohibitive computational cost on learning and the memory requirement to store a quadratic number of parameters. In this paper, we propose to learn a sparse representation of the pair wise features under the preference learning framework using the L1 regularization. Based on stochastic gradient descent, an online algorithm is devised to enforce the sparsity using a mini-batch shrinkage strategy. On multiple benchmark datasets, we show that our method achieves better performance with fast convergence, and takes much less memory on models with millions of parameters. Xi Chen 0010, Yanjun Qi, Qihang Lin, Jaime G. Carbonell |
ICDM | 3 |
| 2010 | Semi-supervised Abstraction-Augmented String Kernel for Multi-level Bio-Relation Extraction
Pavel P. Kuksa, Yanjun Qi, Ronan Collobert, Jason Weston, Vladimir Pavlovic 0001, Xia Ning |
ECML/PKDD (2) | 2 |
| 2010 | Semi-supervised Bio-named Entity Recognition with Word-Codebook LearningabstractWe describe a novel semi-supervised method called Word-Codebook Learning (WCL), and apply it to the task of bio-named entity recognition (bioNER). Typical bioNER systems can be seen as tasks of assigning labels to words in bio-literature text. To improve supervised tagging, WCL learns a class of word-level feature embeddings to capture word semantic meanings or word label patterns from a large unlabeled corpus. Words are then clustered according to their embedding vectors through a vector quantization step, where each word is assigned into one of the codewords in a codebook. Finally codewords are treated as new word attributes and are added for entity labeling. Two types of word-codebook learning are proposed: (1) General WCL, where an unsupervised method uses contextual semantic similarity of words to learn accurate word representations; (2) Task-oriented WCL, where for every word a semi-supervised method learns target-class label patterns from unlabeled data using supervised signals from trained bioNER model. Without the need for complex linguistic features, we demonstrate utility of WCL on the BioCreativeII gene name recognition competition data, where WCL yields state-of-the-art performance and shows great improvements over supervised baselines and semi-supervised counter peers. Pavel P. Kuksa, Yanjun Qi |
SDM | 2 |
| 2010 | Semi-supervised multi-task learning for predicting interactions between HIV-1 and human proteinsabstractMOTIVATION: Protein-protein interactions (PPIs) are critical for virtually every biological function. Recently, researchers suggested to use supervised learning for the task of classifying pairs of proteins as interacting or not. However, its performance is largely restricted by the availability of truly interacting proteins (labeled). Meanwhile, there exists a considerable amount of protein pairs where an association appears between two partners, but not enough experimental evidence to support it as a direct interaction (partially labeled). RESULTS: We propose a semi-supervised multi-task framework for predicting PPIs from not only labeled, but also partially labeled reference sets. The basic idea is to perform multi-task learning on a supervised classification task and a semi-supervised auxiliary task. The supervised classifier trains a multi-layer perceptron network for PPI predictions from labeled examples. The semi-supervised auxiliary task shares network layers of the supervised classifier and trains with partially labeled examples. Semi-supervision could be utilized in multiple ways. We tried three approaches in this article, (i) classification (to distinguish partial positives with negatives); (ii) ranking (to rate partial positive more likely than negatives); (iii) embedding (to make data clusters get similar labels). We applied this framework to improve the identification of interacting pairs between HIV-1 and human proteins. Our method improved upon the state-of-the-art method for this task indicating the benefits of semi-supervised multi-task learning using auxiliary information. AVAILABILITY: http://www.cs.cmu.edu/~qyj/HIVsemi. Yanjun Qi, Öznur Tastan, Jaime G. Carbonell, Judith Klein-Seetharaman, Jason Weston |
Bioinform. | 1 |
| 2010 | Learning to rank with (a lot of) word features
Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Olivier Chapelle, Kilian Q. Weinberger |
Inf. Retr. | 6 |
| 2009 | Supervised semantic indexingabstractIn this article we propose Supervised Semantic Indexing (SSI), an algorithm that is trained on (query, document) pairs of text documents to predict the quality of their match. Like Latent Semantic Indexing (LSI), our models take account of correlations between words (synonymy, polysemy). However, unlike LSI our models are trained with a supervised signal directly on the ranking task of interest, which we argue is the reason for our superior results. As the query and target texts are modeled separately, our approach is easily generalized to different retrieval tasks, such as online advertising placement. Dealing with models on all pairs of words features is computationally challenging. We propose several improvements to our basic model for addressing this issue, including low rank (but diagonal preserving) representations, and correlated feature hashing (CFH). We provide an empirical study of all these methods on retrieval tasks based on Wikipedia documents as well as an Internet advertisement task. We obtain state-of-the-art performance while providing realistically scalable methods. Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Olivier Chapelle, Kilian Q. Weinberger |
CIKM | 6 |
| 2009 | Combining labeled and unlabeled data with word-class distribution learningabstractWe describe a novel simple and highly scalable semi-supervised method called Word-Class Distribution Learning (WCDL), and apply it task of information extraction (IE) by utilizing unlabeled sentences to improve supervised classification methods. WCDL iteratively builds class label distributions for each word in the dictionary by averaging predicted labels over all cases in the unlabeled corpus, and re-training a base classifier adding these distributions as word features. In contrast, traditional self-training or co-training methods self-labeled examples (rather than features) which can degrade performance due to incestuous learning bias. WCDL exhibits robust behavior, and has no difficult parameters to tune. We applied our method on German and English name entity recognition (NER) tasks. WCDL shows improvements over self-training, multi-task semi-supervision or supervision alone, in particular yielding a state-of-the art 75.72 F1 score on the German NER task. Yanjun Qi, Ronan Collobert, Pavel P. Kuksa, Koray Kavukcuoglu, Jason Weston |
CIKM | 1 |
| 2009 | Semi-Supervised Sequence Labeling with Self-Learned FeaturesabstractTypical information extraction (IE) systems can be seen as tasks assigning labels to words in a natural language sequence. The performance is restricted by the availability of labeled words. To tackle this issue, we propose a semi-supervised approach to improve the sequence labeling procedure in IE through a class of algorithms with self-learned features (SLF). A supervised classifier can be trained with annotated text sequences and used to classify each word in a large set of unannotated sentences. By averaging predicted labels over all cases in the unlabeled corpus, SLF training builds class label distribution patterns for each word (or word attribute) in the dictionary and re-trains the current model iteratively adding these distributions as extra word features. Basic SLF models how likely a word could be assigned to target class types. Several extensions are proposed, such as learning words' class boundary distributions. SLF exhibits robust and scalable behaviour and is easy to tune. We applied this approach on four classical IE tasks: named entity recognition (German and English), part-of-speech tagging (English) and one gene name recognition corpus. Experimental results show effective improvements over the supervised baselines on all tasks. In addition, when compared with the closely related self-training idea, this approach shows favorable advantages. Yanjun Qi, Pavel P. Kuksa, Ronan Collobert, Kunihiko Sadamasa, Koray Kavukcuoglu, Jason Weston |
ICDM | 1 |
| 2009 | Polynomial Semantic IndexingabstractWe present a class of nonlinear (polynomial) models that are discriminatively trained to directly map from the word content in a query-document or document-document pair to a ranking score. Dealing with polynomial models on word features is computationally challenging. We propose a low rank (but diagonal preserving) representation of our polynomial models to induce feasible memory and computation requirements. We provide an empirical study on retrieval tasks based on Wikipedia documents, where we obtain state-of-the-art performance while providing realistically scalable methods. Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Corinna Cortes, Mehryar Mohri |
NIPS | 6 |
| 2008 | Protein complex identification by supervised graph local clusteringabstractMOTIVATION: Protein complexes integrate multiple gene products to coordinate many biological functions. Given a graph representing pairwise protein interaction data one can search for subgraphs representing protein complexes. Previous methods for performing such search relied on the assumption that complexes form a clique in that graph. While this assumption is true for some complexes, it does not hold for many others. New algorithms are required in order to recover complexes with other types of topological structure. RESULTS: We present an algorithm for inferring protein complexes from weighted interaction graphs. By using graph topological patterns and biological properties as features, we model each complex subgraph by a probabilistic Bayesian network (BN). We use a training set of known complexes to learn the parameters of this BN model. The log-likelihood ratio derived from the BN is then used to score subgraphs in the protein interaction graph and identify new complexes. We applied our method to protein interaction data in yeast. As we show our algorithm achieved a considerable improvement over clique based algorithms in terms of its ability to recover known complexes. We discuss some of the new complexes predicted by our algorithm and determine that they likely represent true complexes. AVAILABILITY: Matlab implementation is available on the supporting website: www.cs.cmu.edu/~qyj/SuperComplex. Yanjun Qi, Fernanda Balem, Christos Faloutsos, Judith Klein-Seetharaman, Ziv Bar-Joseph |
ISMB | 1 |
| 2007 | A mixture of feature experts approach for protein-protein interaction predictionabstractBACKGROUND: High-throughput methods can directly detect the set of interacting proteins in model species but the results are often incomplete and exhibit high false positive and false negative rates. A number of researchers have recently presented methods for integrating direct and indirect data for predicting interactions. These methods utilize a common classifier for all pairs. However, due to missing data and high redundancy among the features used, different protein pairs may benefit from different features based on the set of attributes available. In addition, in many cases it is hard to directly determine which of the data sources contributed to a prediction. This information is important for biologists using these predications in the design of new experiments. RESULTS: To address these challenges we propose a Mixture-of-Feature-Experts method for protein-protein interaction prediction. We split the features into roughly homogeneous sets of feature experts. The individual experts use logistic regression and their scores are combined using another logistic regression. When combining the scores the weighting of each expert depends on the set of input attributes available for that pair. Thus, different experts will have different influence on the prediction depending on the available features. CONCLUSION: We applied our method to predict the set of interacting proteins in yeast and human cells. Our method improved upon the best previous methods for this task. In addition, the weighting of the experts provides means to evaluate the prediction based on the high scoring features. Yanjun Qi, Judith Klein-Seetharaman, Ziv Bar-Joseph |
BMC Bioinform. | 1 |
| 2003 | Supervised classification for video shot segmentationabstractIn this paper, we explore supervised classification methods for video shot segmentation. We transform the temporal segmentation problem into a multi-class categorization issue. This approach provides a uniform framework for using different kinds of features extracted from the video and for detecting various types of shot boundaries. The approach utilizes manual labeled training data and a simple classification structure, which eliminates arbitrary thresholds and achieves more reliable estimation than previous threshold-based methods. Contrastive experiments on 13 videos (/spl sim/4 hours) show excellent performance on the 2001 TREC video track shot classification task in terms of precision and recall. Yanjun Qi, Alex Hauptmann 0001, Ting Liu 0005 |
ICME | 1 |