VLDB 2026 Research / reviewers in the wild / expert
Ning Chen 0002
dblp:56/1670-2
· DBLP profile ↗
42ranked-venue papers
9as first author
5since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-authorDatabases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Towards Effective Adversarial Textured 3D Meshes on Physical Face RecognitionabstractFace recognition is a prevailing authentication solution in numerous biometric applications. Physical adversarial attacks, as an important surrogate, can identify the weak-nesses of face recognition systems and evaluate their ro-bustness before deployed. However, most existing physical attacks are either detectable readily or ineffective against commercial recognition systems. The goal of this work is to develop a more reliable technique that can carry out an end-to-end evaluation of adversarial robustness for commercial systems. It requires that this technique can simultaneously deceive black-box recognition models and evade defensive mechanisms. To fulfill this, we design adversarial textured 3D meshes (AT3D) with an elaborate topology on a human face, which can be 3D-printed and pasted on the attacker's face to evade the defenses. However, the mesh-based op-timization regime calculates gradients in high-dimensional mesh space, and can be trapped into local optima with un-satisfactory transferability. To deviate from the mesh-based space, we propose to perturb the low-dimensional coefficient space based on 3D Morphable Model, which signifi-cantly improves black-box transferability meanwhile enjoying faster search efficiency and better visual quality. Exten-sive experiments in digital and physical scenarios show that our method effectively explores the security vulnerabilities of multiple popular commercial services, including three recognition A PIs, four anti-spoofing A PIs, two prevailing mobile phones and two automated access control systems. Xiao Yang 0028, Chang Liu 0077, Longlong Xu, Yikai Wang 0001, Yinpeng Dong, Ning Chen 0002, Hang Su 0006, Jun Zhu 0001 |
CVPR | 6 |
| 2023 | Towards Viewpoint-Invariant Visual Recognition via Adversarial TrainingabstractVisual recognition models are not invariant to viewpoint changes in the 3D world, as different viewing directions can dramatically affect the predictions given the same object. Compared to 2D transformations, the exploration of 3D viewpoint invariance deserves more attention for its greater practical significance. Motivated by the success of adversarial training in promoting model robustness, we propose Viewpoint-Invariant Adversarial Training (VIAT) to improve viewpoint robustness of common image classifiers. By regarding viewpoint transformation as an attack, VIAT is formulated as a minimax optimization problem, where the inner maximization characterizes diverse adversarial viewpoints by learning a Gaussian mixture distribution based on a new attack GMVFool, while the outer minimization trains a viewpoint-invariant classifier by minimizing the expected loss over the worst-case adversarial viewpoint distributions. To further improve the generalization performance, a distribution sharing strategy is introduced leveraging the transferability of adversarial viewpoints across objects. Experiments validate the effectiveness of VIAT in improving the viewpoint robustness of various image classifiers based on the diversity of adversarial viewpoints generated by GMVFool. Shouwei Ruan, Yinpeng Dong, Hang Su 0006, Jianteng Peng, Ning Chen 0002, Xingxing Wei 0001 |
ICCV | 5 |
| 2022 | Towards Safe Reinforcement Learning via Constraining Conditional Value-at-RiskabstractThough deep reinforcement learning (DRL) has obtained substantial success, it may encounter catastrophic failures due to the intrinsic uncertainty of both transition and observation. Most of the existing methods for safe reinforcement learning can only handle transition disturbance or observation disturbance since these two kinds of disturbance affect different parts of the agent; besides, the popular worst-case return may lead to overly pessimistic policies. To address these issues, we first theoretically prove that the performance degradation under transition disturbance and observation disturbance depends on a novel metric of Value Function Range (VFR), which corresponds to the gap in the value function between the best state and the worst state. Based on the analysis, we adopt conditional value-at-risk (CVaR) as an assessment of risk and propose a novel reinforcement learning algorithm of CVaR-Proximal-Policy-Optimization (CPPO) which formalizes the risk-sensitive constrained optimization problem by keeping its CVaR under a given threshold. Experimental results show that CPPO achieves a higher cumulative reward and is more robust against both observation and transition disturbances on a series of continuous control tasks in MuJoCo. Chengyang Ying, Xinning Zhou, Hang Su 0006, Ning Chen 0002, Jun Zhu 0001 |
IJCAI | 5 |
| 2021 | Two Birds with One Stone: Series Saliency for Accurate and Interpretable Multivariate Time Series ForecastingabstractIt is important yet challenging to perform accurate and interpretable time series forecasting. Though deep learning methods can boost forecasting accuracy, they often sacrifice interpretability. In this paper, we present a new scheme of series saliency to boost both accuracy and interpretability. By extracting series images from sliding windows of the time series, we design series saliency as a mixup strategy with a learnable mask between the series images and their perturbed versions. Series saliency is model agnostic and performs as an adaptive data augmentation method for training deep models. Moreover, by slightly changing the objective, we optimize series saliency to find a mask for interpretable forecasting in both feature and time dimensions. Experimental results on several real datasets demonstrate that series saliency is effective to produce accurate time-series forecasting results as well as generate temporal interpretations. Qingyi Pan, Wenbo Hu 0001, Ning Chen 0002 |
IJCAI | 3 |
| 2021 | Combining Tree Search and Action Prediction for State-of-the-Art Performance in DouDiZhuabstractAlphaZero has achieved superhuman performance on various perfect-information games, such as chess, shogi and Go. However, directly applying AlphaZero to imperfect-information games (IIG) is infeasible, due to the fact that traditional MCTS methods cannot handle missing information of other players. Meanwhile, there have been several extensions of MCTS for IIGs, by implicitly or explicitly sampling a state of other players. But, due to the inability to handle private and public information well, the performance of these methods is not satisfactory. In this paper, we extend AlphaZero to multiplayer IIGs by developing a new MCTS method, Action-Prediction MCTS (AP-MCTS). In contrast to traditional MCTS extensions for IIGs, AP-MCTS first builds the search tree based on public information, adopts the policy-value network to generalize between hidden states, and finally predicts other players' actions directly. This design bypasses the inefficiency of sampling and the difficulty of predicting the state of other players. We conduct extensive experiments on the popular 3-player poker game DouDiZhu to evaluate the performance of AP-MCTS combined with the framework AlphaZero. When playing against experienced human players, AP-MCTS achieved a 65.65\% winning rate, which is almost twice the human's winning rate. When comparing with state-of-the-art DouDiZhu AIs, the Elo rating of AP-MCTS is 50 to 200 higher than them. The ablation study shows that accurate action prediction is the key to AP-MCTS winning. Bei Shi, Haobo Fu, Qiang Fu 0016, Hang Su 0006, Jun Zhu 0001, Ning Chen 0002 |
IJCAI | 8 |
| 2020 | Rethinking Softmax Cross-Entropy Loss for Adversarial Robustness
Tianyu Pang, Kun Xu 0004, Yinpeng Dong, Ning Chen 0002, Jun Zhu 0001 |
ICLR | 5 |
| 2020 | Automatic Emergency Diagnosis with Knowledge-Based Tree DecodingabstractAutomatic diagnosis based on clinical notes is critical especially in the emergency department, where a fast and professional result is vital in assuring proper and timely treatment. Previous works formalize this task as plain text classification and fail to utilize the medically significant tree structure of International Classification of Diseases (ICD) coding system. Besides, external medical knowledge is rarely used before, and we explore it by extracting relevant materials from Wikipedia or Baidupedia. In this paper, we propose a knowledge-based tree decoding model (K-BTD), and the inference procedure is a top-down decoding process from the root node to leaf nodes. The stepwise inference procedure enables the model to give support for decision at each step, which visualizes the diagnosis procedure and adds to the interpretability of final predictions. Experiments on real-world data from the emergency department of a large-scale hospital indicate that the proposed model outperforms all baselines in both micro-F1 and macro-F1, and reduce the semantic distance dramatically. Xuyan Chen, Ning Chen 0002, Ting Chen 0006 |
IJCAI | 3 |
| 2020 | Joint Medical Ontology Representation Learning for Healthcare PredictionsabstractHealthcare predictions aim at predicting diseases of the next visit to hospital with historical Electronic Health Records (EHR), which is a key research field in personalized healthcare. Previous research has demonstrated that learning meaningful medical ontology representations within the healthcare prediction model can alleviate the data insufficiency problem and thus is beneficial to this task. There are two main pathways of learning medical ontology representations. The first is through pre-defined knowledge graph such as the ICD tree, and the second is through the co-occurrence of diseases within each visit. Majority of existing works formalize their model under only one pathway, and fail to utilize the mutual benefits between them. To exploit these benefits, we propose JMRL, an end-to-end and accurate model for healthcare predictions with Joint Medical ontology Representation Learning. JMRL not only utilizes the joint information from both knowledge graph and co-occurrence statistics, but also make use of the mutual benefits between them in an advanced way with two explicit feedback strategies. Experimental results on the MIMIC-III dataset demonstrate the superiority of our model over all existing state-of-the-art approaches. Ning Chen 0002, Ting Chen 0006 |
IJCNN | 2 |
| 2019 | Bone Age Assessment by Deep Convolutional Neural Networks Combined with Clinical TW3-RUSabstractBone age assessment is critical to diagnosis of various growth disorders in children, such as endocrine, nutritional disorders and dysplasia. X-rays of hand and wrist are the most common modality used to calculate bone age. In this paper, we propose a novel approach called DeepTW3 for automatic bone age assessment from X-ray images. DeepTW3 integrates Convolutional Neural Networks (CNNs) with expertise knowledge of TW3(Tanner-Whitehouse 3nd edition)-RUS(radius, ulna and short bones) bone age assessment system. The proposed method is tested on a dataset containing 1,100 hand bone X-ray images, all of which were manually annotated with selected region of interests(ROIs). Our method achieved mean absolute errors (MAE) of 0.2685, outperforming all state-of-the-art methods. For the task of grading skeletal maturity, our method using continuous stage distribution is complementary to using the clinical TW3-RUS categorical stages when interpreting critical cases of intermediate bone stage. Xiaohong Liu 0007, Yiming Xu 0010, Ning Chen 0002, Ting Chen 0006 |
BIBM | 4 |
| 2019 | DeepTriager: A Neural Attention Model for Emergency Triage with Electronic Health RecordsabstractAs the first pass for emergency patients, triage is the most important factor affecting emergency department (ED) overcrowding. So it is crucial to develop a data-driven and evidence-based triage method to quickly identify acute and severe patients, and prevent the limited emergency resources from over-diagnosis. To address these challenges, we propose an attention based deep learning framework, named DeepTriager. Trained and tested on 70,918 clinical records, DeepTriager achieved highly accurate performance on assessment of acuity level I (endangered patients), with AUC of 0.98, which was 0.16 higher than the clinical scale method MEWS and NEWS, and 0.04 higher than traditional machine learning methods. In summary, we presented a new approach for clinical evidence based discovery using a cohort of Electronic Health Records (EHRs). This approach not only outperforms the traditional word segmentation methods but also provides evidence for interpreting the results. Xiaohong Liu 0007, Ken Xie, Ning Chen 0002, Ting Chen 0006 |
BIBM | 4 |
| 2019 | How Robust is Your Automatic Diagnosis Model?abstractAutomatic diagnosis based on clinical notes has become a popular research field recently, and many proposed deep learning models have achieved competitive performance in diseases inference. However, previous research reveals that deep learning models are susceptible to negligibly perturbed inputs named adversarial examples, which contradicts with the safety and reliability requirements of the medical domain. To analyze the vulnerability and robustness of current automatic diagnosis models, we investigate in the generation of adversarial text examples. The main challenges for generating adversarial text examples are divided into three parts. First, the word embedding space is discrete, which makes it hard to perturb as small as adversarial image examples generation. Second, previous adversarial example generation methods focus mainly on multi-class classification models, while automatic diagnosis is a multi-label classification task. Third, the semantic and medical meaning of clinical notes are vital in disease inference, and even small perturbations can change them to a large extent. In this paper, we address the three main challenges and propose Clinical-Attacker, a general framework for both white-box and black-box adversarial text examples generation against automatic diagnosis models. Experimental results on MIMIC-III dataset demonstrate that our framework can easily alter the predictions of automatic diagnosis models with the semantic and medical meaning preserved. Ning Chen 0002, Ting Chen 0006 |
BIBM | 3 |
| 2019 | DCMN: Double Core Memory Network for Patient Outcome Prediction with Multimodal DataabstractMore and more healthcare data are becoming readily available nowadays. These data can help the healthcare professionals and patient themselves to better understand the patient status and potentially lead to improved care quality. However, the analysis of these data are challenging because they are large-scale and heterogeneous, high-dimensional and sparse, temporal but irregularly sampled. In this paper, we propose a method called Double Core Memory Networks (DCMN) to integrate information from different modalities of the longitudinal patient data and learn a joint patient representation effective for downstream analytical tasks such as risk prediction. DCMN is designed not only to disentangle the temporal and non-linear intra-modal dependencies for the data within each modality but also to capture the long-term inter-modal interactions. DCMN models are the end-to-end memory networks with two external memory cores where each modality of data is compressed and stored. Each memory core has an information-flow controller named query to interact with an external memory module. In addition, we incorporate a gating mechanism into basic DCMN model to perform dynamic regulation of memory interaction. DCMN models have multiple computational layers (hops) allowing data of different modalities interacting with each other recurrently along with a mechanism of alternating access of external memory for each memory core hop-by-hop. We evaluate DCMN models on two outcome prediction tasks, including a mortality prediction on the public Medical Information Mart for Intensive Care III (MIMIC-III) database and a cost prediction on the Hospital Quality Monitoring System (HQMS) dataset. Experimental results demonstrate that our DCMN models are more competitive over the baseline methods in the multimodal prediction setting. Yujuan Feng, Ning Chen 0002, Ting Chen 0006, Fei Wang 0001 |
ICDM | 4 |
| 2019 | Improving Adversarial Robustness via Promoting Ensemble DiversityabstractThough deep neural networks have achieved significant progress on various tasks, often enhanced by model ensemble, existing high-performance models can be vulnerable to adversarial attacks. Many efforts have been devoted to enhancing the robustness of individual networks and then constructing a straightforward ensemble, e.g., by directly averaging the outputs, which ignores the interaction among networks. This paper presents a new method that explores the interaction among individual networks to improve robustness for ensemble models. Technically, we define a new notion of ensemble diversity in the adversarial setting as the diversity among non-maximal predictions of individual members, and present an adaptive diversity promoting (ADP) regularizer to encourage the diversity, which leads to globally better robustness for the ensemble by making adversarial examples difficult to transfer among individual members. Our method is computationally efficient and compatible with the defense methods acting on individual networks. Empirical results on various datasets verify that our method can improve adversarial robustness while maintaining state-of-the-art accuracy on normal examples. Tianyu Pang, Kun Xu 0004, Ning Chen 0002, Jun Zhu 0001 |
ICML | 4 |
| 2019 | Transferable Adversarial Attacks for Image and Video Object DetectionabstractIdentifying adversarial examples is beneficial for understanding deep networks and developing robust models. However, existing attacking methods for image object detection have two limitations: weak transferability---the generated adversarial examples often have a low success rate to attack other kinds of detection methods, and high computation cost---they need much time to deal with video data, where many frames need polluting. To address these issues, we present a generative method to obtain adversarial images and videos, thereby significantly reducing the processing time. To enhance transferability, we manipulate the feature maps extracted by a feature network, which usually constitutes the basis of object detectors. Our method is based on the Generative Adversarial Network (GAN) framework, where we combine a high-level class loss and a low-level feature loss to jointly train the adversarial example generator. Experimental results on PASCAL VOC and ImageNet VID datasets show that our method efficiently generates image and video adversarial examples, and more importantly, these adversarial examples have better transferability, therefore being able to simultaneously attack two kinds of representative object detection models: proposal based models like Faster-RCNN and regression based models like SSD. Xingxing Wei 0001, Siyuan Liang 0004, Ning Chen 0002, Xiaochun Cao |
IJCAI | 3 |
| 2019 | An adaptive PNN-DS approach to classification using multi-sensor information fusion
Ning Chen 0002, Fuchun Sun 0001, Linge Ding |
Neural Comput. Appl. | 1 |
| 2018 | Ontology-based Venous Thromboembolism Risk Factors Mining and Model Developing from Medical Records
Xin Wang 0226, Ning Chen 0002, Juhong Shi, Ting Chen 0006 |
BIBM | 4 |
| 2018 | Message Passing Stein Variational Gradient DescentabstractStein variational gradient descent (SVGD) is a recently proposed particle-based Bayesian inference method, which has attracted a lot of interest due to its remarkable approximation ability and particle efficiency compared to traditional variational inference and Markov Chain Monte Carlo methods. However, we observed that particles of SVGD tend to collapse to modes of the target distribution, and this particle degeneracy phenomenon becomes more severe with higher dimensions. Our theoretical analysis finds out that there exists a negative correlation between the dimensionality and the repulsive force of SVGD which should be blamed for this phenomenon. We propose Message Passing SVGD (MP-SVGD) to solve this problem. By leveraging the conditional independence structure of probabilistic graphical models (PGMs), MP-SVGD converts the original high-dimensional global inference problem into a set of local ones over the Markov blanket with lower dimensions. Experimental results show its advantages of preventing vanishing repulsive force in high-dimensional space over SVGD, and its particle efficiency and approximation flexibility over other inference methods on graphical models. Jingwei Zhuo, Chang Liu 0030, Jiaxin Shi, Jun Zhu 0001, Ning Chen 0002, Bo Zhang 0010 |
ICML | 5 |
| 2018 | Dropout training for SVMs with data augmentation
Ning Chen 0002, Jun Zhu 0001, Jianfei Chen 0001, Ting Chen 0006 |
Frontiers Comput. Sci. | 1 |
| 2017 | Learning Attributes from the Crowdsourced Relative LabelsabstractFinding semantic attributes to describe related concepts is typically a hard problem. The commonly used attributes in most fields are designed by domain experts, which is expensive and time-consuming. In this paper we propose an efficient method to learn human comprehensible attributes with crowdsourcing. We first design an analogical interface to collect relative labels from the crowds. Then we propose a hierarchical Bayesian model, as well as an efficient initialization strategy, to aggregate labels and extract concise attributes. Our experimental results demonstrate promise on discovering diverse and convincing attributes, which significantly improve the performance of the challenging zero-shot learning tasks. Tian Tian 0001, Ning Chen 0002, Jun Zhu 0001 |
AAAI | 2 |
| 2017 | Patient outcome prediction via convolutional neural networks based on multi-granularity medical concept embeddingabstractThe large availability of biomedical data brings opportunities and challenges to health care. Representation of medical concepts has been well studied in many applications, such as medical informatics, cohort selection, risk prediction, and health care quality measurement. In this paper, we propose an efficient multichannel convolutional neural network (CNN) model based on multi-granularity embeddings of medical concepts named MG-CNN, to examine the effect of individual patient characteristics including demographic factors and medical comorbidities on total hospital costs and length of stay (LOS) by using the Hospital Quality Monitoring System (HQMS) data. The proposed embedding method leverages prior medical hierarchical ontology and improves the quality of embedding for rare medical concepts. The embedded vectors are further visualized by the t-Distributed Stochastic Neighbor Embedding (t-SNE) technique to demonstrate the effectiveness of grouping related medical concepts. Experimental results demonstrate that our MG-CNN model outperforms traditional regression methods based on the one-hot representation of medical concepts, especially in the outcome prediction tasks for patients with low-frequency medical events. In summary, MG-CNN model is capable of mining potential knowledge from the clinical data and will be broadly applicable in medical research and inform clinical decisions. Yujuan Feng, Xu Min, Ning Chen 0002, Xiaolei Xie, Ting Chen 0006 |
BIBM | 3 |
| 2017 | DACE: a scalable DP-means algorithm for clustering extremely large sequence dataabstractMotivation: Advancements in next-generation sequencing technology have produced large amounts of reads at low cost in a short time. In metagenomics, 16S and 18S rRNA gene have been widely used as marker genes to profile diversity of microorganisms in environmental samples. Through clustering of sequencing reads we can determine both number of OTUs and their relative abundance. In many applications, clustering of very large sequencing data with high efficiency and accuracy is essential for downstream analysis. Results: Here, we report a scalable D irichlet Process Means (DP-means) a lgorithm for c lustering e xtremely large sequencing data, termed . With an efficient random projection partition strategy for parallel clustering, DACE can cluster billions of sequences within a couple of hours. Experimental results show that DACE runs between 6 and 80 times faster than state-of-the-art programs, while maintaining overall better clustering accuracy. Using 80 cores, DACE clustered the Lake Taihu 16S rRNA gene sequencing data (∼316M reads, 30 GB) in 25 min, and the Ocean TARA Eukaryotic 18S rRNA gene sequencing data (∼500M reads, 88 GB) into ∼100 000 clusters within an hour. When applied to the IGC gene catalogs in human gut microbiome (∼10M genes), DACE produced 9.8M clusters with 52K redundant genes in 1.5 hours of running time. Availability and Implementation: DACE is available at https://github.com/tinglab/DACE . Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Linhao Jiang, Yichao Dong, Ning Chen 0002, Ting Chen 0006 |
Bioinform. | 3 |
| 2017 | Chromatin accessibility prediction via convolutional long short-term memory networks with k-mer embeddingabstractMOTIVATION: Experimental techniques for measuring chromatin accessibility are expensive and time consuming, appealing for the development of computational approaches to predict open chromatin regions from DNA sequences. Along this direction, existing methods fall into two classes: one based on handcrafted k -mer features and the other based on convolutional neural networks. Although both categories have shown good performance in specific applications thus far, there still lacks a comprehensive framework to integrate useful k -mer co-occurrence information with recent advances in deep learning. RESULTS: We fill this gap by addressing the problem of chromatin accessibility prediction with a convolutional Long Short-Term Memory (LSTM) network with k -mer embedding. We first split DNA sequences into k -mers and pre-train k -mer embedding vectors based on the co-occurrence matrix of k -mers by using an unsupervised representation learning approach. We then construct a supervised deep learning architecture comprised of an embedding layer, three convolutional layers and a Bidirectional LSTM (BLSTM) layer for feature learning and classification. We demonstrate that our method gains high-quality fixed-length features from variable-length sequences and consistently outperforms baseline methods. We show that k -mer embedding can effectively enhance model performance by exploring different embedding strategies. We also prove the efficacy of both the convolution and the BLSTM layers by comparing two variations of the network architecture. We confirm the robustness of our model to hyper-parameters by performing sensitivity analysis. We hope our method can eventually reinforce our understanding of employing deep learning in genomic studies and shed light on research regarding mechanisms of chromatin accessibility. AVAILABILITY AND IMPLEMENTATION: The source code can be downloaded from https://github.com/minxueric/ismb2017_lstm . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. Xu Min, Wanwen Zeng, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001 |
Bioinform. | 3 |
| 2017 | Predicting enhancers with deep convolutional neural networksabstractBACKGROUND: With the rapid development of deep sequencing techniques in the recent years, enhancers have been systematically identified in such projects as FANTOM and ENCODE, forming genome-wide landscapes in a series of human cell lines. Nevertheless, experimental approaches are still costly and time consuming for large scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. RESULTS: To facilitate the identification of enhancers, we propose a computational framework, named DeepEnhancer, to distinguish enhancers from background genomic sequences. Our method purely relies on DNA sequences to predict enhancers in an end-to-end manner by using a deep convolutional neural network (CNN). We train our deep learning model on permissive enhancers and then adopt a transfer learning strategy to fine-tune the model on enhancers specific to a cell line. Results demonstrate the effectiveness and efficiency of our method in the classification of enhancers against random sequences, exhibiting advantages of deep learning over traditional sequence-based classifiers. We then construct a variety of neural networks with different architectures and show the usefulness of such techniques as max-pooling and batch normalization in our method. To gain the interpretability of our approach, we further visualize convolutional kernels as sequence logos and successfully identify similar motifs in the JASPAR database. CONCLUSIONS: DeepEnhancer enables the identification of novel enhancers using only DNA sequences via a highly accurate deep learning model. The proposed computational framework can also be applied to similar problems, thereby prompting the use of machine learning methods in life sciences. Xu Min, Wanwen Zeng, Shengquan Chen, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001 |
BMC Bioinform. | 4 |
| 2016 | Discriminative Nonparametric Latent Feature Relational Models with Data AugmentationabstractWe present a discriminative nonparametric latent feature relational model (LFRM) for link prediction to automatically infer the dimensionality of latent features. Under the generic RegBayes (regularized Bayesian inference) framework, we handily incorporate the prediction loss with probabilistic inference of a Bayesian model; set distinct regularization parameters for different types of links to handle the imbalance issue in real networks; and unify the analysis of both the smooth logistic log-loss and the piecewise linear hinge loss. For the nonconjugate posterior inference, we present a simple Gibbs sampler via data augmentation, without making restricting assumptions as done in variational methods. We further develop an approximate sampler using stochastic gradient Langevin dynamics to handle large networks with hundreds of thousands of entities and millions of links, orders of magnitude larger than what existing LFRM models can process. Extensive studies on various real networks show promising performance. Ning Chen 0002, Jun Zhu 0001, Jiaming Song, Bo Zhang 0010 |
AAAI | 2 |
| 2016 | DeepEnhancer: Predicting enhancers by convolutional neural networksabstractEnhancers are crucial to the understanding of mechanisms underlying gene transcriptional regulation. Although having been successfully applied in such projects as ENCODE and Roadmap to generate landscape of enhancers in human cell lines, high-throughput biological experimental techniques are still costly and time consuming for even larger scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. In this paper, we propose a computational framework, named DeepEnhancer, to classify enhancers from background genomic sequences. We construct convolutional neural networks of various architectures and compare the classification performance with traditional sequence-based classifiers. We first train the deep learning model on the FANTOM5 permissive enhancer dataset, and then fine-tune the model on ENCODE cell type-specific enhancer datasets by adopting the transfer learning strategy. Experimental results demonstrate that DeepEnhancer has superior efficiency and effectiveness in classification tasks, and the use of max-pooling and batch normalization is beneficial to higher accuracy. To make our approach more understandable, we propose a strategy to visualize the convolutional kernels as sequence logos and compare them against the JASPAR database using TOMTOM. In summary, DeepEnhancer allows researchers to train highly accurate deep models and will be broadly applicable in computational biology. Xu Min, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001 |
BIBM | 2 |
| 2016 | mLDM: A New Hierarchical Bayesian Statistical Model for Sparse Microbial Association Discovery
Ning Chen 0002, Ting Chen 0006 |
RECOMB | 2 |
| 2015 | Discriminative Relational Topic ModelsabstractRelational topic models (RTMs) provide a probabilistic generative process to describe both the link structure and document contents for document networks, and they have shown promise on predicting network structures and discovering latent topic representations. However, existing RTMs have limitations in both the restricted model expressiveness and incapability of dealing with imbalanced network data. To expand the scope and improve the inference accuracy of RTMs, this paper presents three extensions: 1) unlike the common link likelihood with a diagonal weight matrix that allows the-same-topic interactions only, we generalize it to use a full weight matrix that captures all pairwise topic interactions and is applicable to asymmetric networks; 2) instead of doing standard Bayesian inference, we perform regularized Bayesian inference (RegBayes) with a regularization parameter to deal with the imbalanced link structure issue in real networks and improve the discriminative ability of learned latent representations; and 3) instead of doing variational approximation with strict mean-field assumptions, we present collapsed Gibbs sampling algorithms for the generalized relational topic models by exploring data augmentation without making restricting assumptions. Under the generic RegBayes framework, we carefully investigate two popular discriminative loss functions, namely, the logistic log-loss and the max-margin hinge loss. Experimental results on several real network datasets demonstrate the significance of these extensions on improving prediction performance. Ning Chen 0002, Jun Zhu 0001, Fei Xia 0005, Bo Zhang 0010 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Dropout Training for Support Vector MachinesabstractDropout and other feature noising schemes have shown promising results in controlling over-fitting by artificially corrupting the training data. Though extensive theoretical and empirical studies have been performed for generalized linear models, little work has been done for support vector machines (SVMs), one of the most successful approaches for supervised learning. This paper presents dropout training for linear SVMs. To deal with the intractable expectation of the non-smooth hinge loss under corrupting distributions, we develop an iteratively re-weighted least square (IRLS) algorithm by exploring data augmentation techniques. Our algorithm iteratively minimizes the expectation of a re-weighted least square problem, where the re-weights have closed-form solutions. The similar ideas are applied to develop a new IRLS algorithm for the expected logistic loss under corrupting distributions. Our algorithms offer insights on the connection and difference between the hinge loss and logistic loss in dropout training. Empirical results on several real datasets demonstrate the effectiveness of dropout training on significantly boosting the classification accuracy of linear SVMs. Ning Chen 0002, Jun Zhu 0001, Jianfei Chen 0001, Bo Zhang 0010 |
AAAI | 1 |
| 2014 | Max-margin latent feature relational models for entity-attribute networksabstractLink prediction is a fundamental task in statistical analysis of network data. Though much research has concentrated on predicting entity-entity relationships in homogeneous networks, it has attracted increasing attentions to predict relationships in heterogeneous networks, which consist of multiple types of nodes and relational links. Existing work on heterogeneous network link prediction mainly focuses on using input features that are explicitly extracted by humans. This paper presents an approach to automatically learn latent features from partially observed heterogeneous networks, with a particular focus on entity-attribute networks (EANs), and making predictions for unseen pairs. To make the latent features discriminative, we adopt the max-margin idea under the framework of maximum entropy discrimination (MED). Our maximum entropy discrimination joint relational model (MED-JRM) can jointly predict entity-entity relationships as well as the missing attributes of entities in EANs. Experimental results on several real networks demonstrate that our model has improved performance over state-of-the-art homogeneous and heterogeneous network link prediction algorithms. Fei Xia 0005, Ning Chen 0002, Jun Zhu 0001, Aonan Zhang, Xiaoming Jin |
IJCNN | 2 |
| 2014 | Gibbs max-margin topic models with data augmentation
Jun Zhu 0001, Ning Chen 0002, Hugh Perkins, Bo Zhang 0010 |
J. Mach. Learn. Res. | 2 |
| 2014 | Bayesian inference with posterior regularization and applications to infinite latent SVMs
Jun Zhu 0001, Ning Chen 0002, Eric P. Xing |
J. Mach. Learn. Res. | 2 |
| 2014 | Learning Harmonium Models With Infinite Latent FeaturesabstractUndirected latent variable models represent an important class of graphical models that have been successfully developed to deal with various tasks. One common challenge in learning such models is to determine the number of hidden units that are unknown a priori. Although Bayesian nonparametrics have provided promising results in bypassing the model selection problem in learning directed Bayesian Networks, very little effort has been made toward applying Bayesian nonparametrics to learn undirected latent variable models. In this paper, we present the infinite exponential family Harmonium (iEFH), a bipartite undirected latent variable model that automatically determines the number of latent units from an unbounded pool. We also present two important extensions of iEFH to 1) multiview iEFH for dealing with heterogeneous data, and 2) infinite maximum-margin Harmonium (iMMH) for incorporating supervising side information to learn predictive latent features. We develop variational inference algorithms to learn model parameters. Our methods are computationally competitive because of the avoidance of selecting the number of latent units. Our extensive experiments on real image datasets and text datasets appear to demonstrate the benefits of iEFH and iMMH inherited from Bayesian nonparametrics and max-margin learning. Such results were not available until now and contribute to expanding the scope of Bayesian nonparametrics to learn the structures of undirected latent variable models. Ning Chen 0002, Jun Zhu 0001, Fuchun Sun 0001, Bo Zhang 0010 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | Gibbs Max-Margin Topic Models with Fast Sampling AlgorithmsabstractExisting max-margin supervised topic models rely on an iterative procedure to solve multiple latent SVM subproblems with additional mean-field assumptions on the desired posterior distributions. This paper presents Gibbs max-margin supervised topic models by minimizing an expected margin loss, an upper bound of the existing margin loss derived from an expected prediction rule. By introducing augmented variables, we develop simple and fast Gibbs sampling algorithms with no restricting assumptions and no need to solve SVM subproblems for both classification and regression. Empirical results demonstrate significant improvements on time efficiency. The classification performance is also significantly improved over competitors. Jun Zhu 0001, Ning Chen 0002, Hugh Perkins, Bo Zhang 0010 |
ICML (1) | 2 |
| 2013 | Generalized Relational Topic Models with Data Augmentation
Ning Chen 0002, Jun Zhu 0001, Fei Xia 0005, Bo Zhang 0010 |
IJCAI | 1 |
| 2012 | Large-Margin Predictive Latent Subspace Learning for Multiview Data AnalysisabstractLearning salient representations of multiview data is an essential step in many applications such as image classification, retrieval, and annotation. Standard predictive methods, such as support vector machines, often directly use all the features available without taking into consideration the presence of distinct views and the resultant view dependencies, coherence, and complementarity that offer key insights to the semantics of the data, and are therefore offering weak performance and are incapable of supporting view-level analysis. This paper presents a statistical method to learn a predictive subspace representation underlying multiple views, leveraging both multiview dependencies and availability of supervising side-information. Our approach is based on a multiview latent subspace Markov network (MN) which fulfills a weak conditional independence assumption that multiview observations and response variables are conditionally independent given a set of latent variables. To learn the latent subspace MN, we develop a large-margin approach which jointly maximizes data likelihood and minimizes a prediction loss on training data. Learning and inference are efficiently done with a contrastive divergence method. Finally, we extensively evaluate the large-margin latent MN on real image and hotel review datasets for classification, regression, image annotation, and retrieval. Our results demonstrate that the large-margin approach can achieve significant improvements in terms of prediction performance and discovering predictive latent subspace representations. Ning Chen 0002, Jun Zhu 0001, Fuchun Sun 0001, Eric P. Xing |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Infinite SVM: a Dirichlet Process Mixture of Large-margin Kernel Machines
Jun Zhu 0001, Ning Chen 0002, Eric P. Xing |
ICML | 2 |
| 2011 | Conditional topical coding: an efficient topic model conditioned on rich featuresabstractProbabilistic topic models have shown remarkable success in many application domains. However, a probabilistic conditional topic model can be extremely inefficient when considering a rich set of features because it needs to define a normalized distribution, which usually involves a hard-to-compute partition function. This paper presents conditional topical coding (CTC), a novel formulation of conditional topic models which is non-probabilistic. CTC relaxes the normalization constraints as in probabilistic models and learns non-negative document codes and word codes. CTC does not need to define a normalized distribution and can efficiently incorporate a rich set of features for improved topic discovery and prediction tasks. Moreover, CTC can directly control the sparsity of inferred representations by using appropriate regularization. We develop an efficient and easy-to-implement coordinate descent learning algorithm, of which each coding substep has a closed-form solution. Finally, we demonstrate the advantages of CTC on online review analysis datasets. Our results show that conditional topical coding can achieve state-of-the-art prediction performance and is much more efficient in training (one order of magnitude faster) and testing (two orders of magnitude faster) than probabilistic conditional topic models. Jun Zhu 0001, Ni Lao, Ning Chen 0002, Eric P. Xing |
KDD | 3 |
| 2011 | Infinite Latent SVM for Classification and Multi-task LearningabstractUnlike existing nonparametric Bayesian models, which rely solely on specially conceived priors to incorporate domain knowledge for discovering improved latent representations, we study nonparametric Bayesian inference with regularization on the desired posterior distributions. While priors can indirectly affect posterior distributions through Bayes' theorem, imposing posterior regularization is arguably more direct and in some cases can be much easier. We particularly focus on developing infinite latent support vector machines (iLSVM) and multi-task infinite latent support vector machines (MT-iLSVM), which explore the large-margin idea in combination with a nonparametric Bayesian model for discovering predictive latent features for classification and multi-task learning, respectively. We present efficient inference methods and report empirical studies on several benchmark datasets. Our results appear to demonstrate the merits inherited from both large-margin learning and Bayesian nonparametrics. Jun Zhu 0001, Ning Chen 0002, Eric P. Xing |
NIPS | 2 |
| 2010 | Predictive Subspace Learning for Multi-view Data: a Large Margin ApproachabstractLearning from multi-view data is important in many applications, such as image classification and annotation. In this paper, we present a large-margin learning framework to discover a predictive latent subspace representation shared by multiple views. Our approach is based on an undirected latent space Markov network that fulfills a weak conditional independence assumption that multi-view observations and response variables are independent given a set of latent variables. We provide efficient inference and parameter estimation methods for the latent subspace model. Finally, we demonstrate the advantages of large-margin learning on real video and web image data for discovering predictive latent representations and improving the performance on image classification, annotation and retrieval. Ning Chen 0002, Jun Zhu 0001, Eric P. Xing |
NIPS | 1 |
| 2010 | An unbiased LSSVM model for classification and regression
Fuchun Sun 0001, Yan-Ning Cai, Linge Ding, Ning Chen 0002 |
Soft Comput. | 5 |
| 2009 | An adaptive PNN-DS approach to classification using multi-sensor information fusion
Ning Chen 0002, Fuchun Sun 0001, Linge Ding |
Neural Comput. Appl. | 1 |
| 2008 | A Sparse Sampling Method for Classification Based on Likelihood Factor
Linge Ding, Fuchun Sun 0001, Ning Chen 0002 |
ISNN (2) | 4 |