VLDB 2026 Research / reviewers in the wild / expert
Kai Zhang 0001
dblp:55/957-1
· DBLP profile ↗
71ranked-venue papers
21as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 19 first-author · 16 since 2021Databases, data management, data science and information retrieval · 23 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COAT-GNN: Cooperative Attribute Learning and Topological Optimization for Protein-Protein Interaction Sites Prediction
Rongfan Tang, Chenglin Wang 0010, Danlin Liu, Jie Zhang 0012, Honglin Li 0003, Kai Zhang 0001 |
DASFAA (3) | 9 |
| 2025 | Harnessing Multimodal Large Language Models for Multimodal Sequential RecommendationabstractRecent advances in Large Language Models (LLMs) have demonstrated significant potential in the field of Recommendation Systems (RSs). Most existing studies have focused on converting user behavior logs into textual prompts and leveraging techniques such as prompt tuning to enable LLMs for recommendation tasks. Meanwhile, research interest has recently grown in multimodal recommendation systems that integrate data from images, text, and other sources using modality fusion techniques. This introduces new challenges to the existing LLM-based recommendation paradigm which relies solely on text modality information. Moreover, although Multimodal Large Language Models (MLLMs) capable of processing multi-modal inputs have emerged, how to equip MLLMs with multi-modal recommendation capabilities remains largely unexplored. To this end, in this paper, we propose the Multimodal Large Language Model-enhanced Sequential Multimodal Recommendation (MLLM-MSR) model. To capture the dynamic user preference, we design a two-stage user preference summarization method. Specifically, we first utilize an MLLM-based item-summarizer to extract image feature given an item and convert the image into text. Then, we employ a recurrent user preference summarization generation paradigm to capture the dynamic changes in user preferences based on an LLM-based user-summarizer. Finally, to enable the MLLM for multi-modal recommendation task, we propose to fine-tune a MLLM-based recommender using Supervised Fine-Tuning (SFT) techniques. Extensive evaluations across various datasets validate the effectiveness of MLLM-MSR, showcasing its superior ability to capture and adapt to the evolving dynamics of user preferences. Yuyang Ye 0002, Zhi Zheng 0008, Yishan Shen, Tianshu Wang 0004, Hengruo Zhang, Peijun Zhu, Runlong Yu, Kai Zhang 0001, Hui Xiong 0001 |
AAAI | 8 |
| 2025 | RolePlot: A Systematic Framework for Evaluating and Enhancing the Plot-Progression Capabilities of Role-Playing AgentsabstractPinyi Zhang, Siyu An, Lingfeng Qiao, Yifei Yu, Jingyang Chen, Jie Wang, Di Yin, Xing Sun, Kai Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Pinyi Zhang, Siyu An, Lingfeng Qiao, Jie Wang 0146, Xing Sun 0001, Kai Zhang 0001 |
ACL (1) | 9 |
| 2025 | Decoupling and Reconstructing: A Multimodal Sentiment Analysis Framework Towards RobustnessabstractMultimodal sentiment analysis (MSA) has shown promising results but often poses significant challenges in real-world applications due to its dependence on the complete and aligned multimodal sequences. While existing approaches attempt to address missing modalities through feature reconstruction, they often neglect the complex interplay between homogeneous and heterogeneous relationships in multimodal features. To address this problem, we propose Decoupled-Adaptive Reconstruction (DAR), a novel framework that explicitly addresses these limitations through two key components: (1) a mutual information-based decoupling module that decomposes features into common and independent representations, and (2) a reconstruction module that independently processes these decoupled features before fusion for downstream tasks. Extensive experiments on two benchmark datasets demonstrate that DAR significantly outperforms existing methods in both modality reconstruction and sentiment analysis tasks, particularly in scenarios with missing or unaligned modalities. Our results show improvements of 2.21% in bi-classification accuracy and 3.9% in regression error compared to state-of-the-art baselines on the MOSEI dataset. Mingzheng Yang, Kai Zhang 0001, Yuyang Ye 0002, Yanghai Zhang, Runlong Yu, Min Hou 0004 |
IJCAI | 2 |
| 2025 | ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional DependenciesabstractText-driven image editing has achieved remarkable success in following single instructions. However, real-world scenarios often involve complex, multi-step instructions, particularly ''chain'' instructions where operations are interdependent. Current models struggle with these intricate directives, and existing benchmarks inadequately evaluate such capabilities. Specifically, they often overlook multi-instruction and chain-instruction complexities, and common consistency metrics are flawed. To address this, we introduce ComplexBench-Edit, a novel benchmark designed to systematically assess model performance on complex, multi-instruction, and chain-dependent image editing tasks. ComplexBench-Edit also features a new vision consistency evaluation method that accurately assesses non-modified regions by excluding edited areas. Furthermore, we propose a simple yet powerful Chain-of-Thought (CoT)-based approach that significantly enhances the ability of existing models to follow complex instructions. Our extensive experiments demonstrate ComplexBench-Edit's efficacy in differentiating model capabilities and highlight the superior performance of our CoT-based method in handling complex edits. The data and code are released at https://github.com/llllly26/ComplexBench-Edit. Chenglin Wang 0010, Yucheng Zhou 0001, Qianning Wang, Kai Zhang 0001 |
ACM Multimedia | 5 |
| 2025 | Alternate Geometric and Semantic Denoising Diffusion for Protein Inverse Folding
Chenglin Wang 0010, Yucheng Zhou 0001, Zijie Zhai, Jianbing Shen, Kai Zhang 0001 |
ECML/PKDD (3) | 6 |
| 2025 | Denoising Structure against Adversarial Attacks on Graph Representation LearningabstractDespite their excellent performance in graph representation learning, graph convolutional networks have been proved to be vulnerable to adversarial perturbations on the connectivity between nodes in an unnoticed manner. In this work, by looking into the impacts of adversarial attacks on graph data, we empirically find that the dominant edge-addition attacks generally increase the heterophily between connected nodes, which will fool the transductive inference models on node classification task. To defend against such attacks, we develop a Two-Stage Denoising (TSD) method that aims at removing possible malicious edges so as to mitigate the heterophily issue introduced by attacks. In particular, after a rough removal of the links that have quite low feature similarity, our method further spots the potentially heterophilous links by predicting node labels with a multi-view labeling consensus. This design is based on assumption that if the label predictions for the same node from two different views of a graph data are consistent, then we have a high chance to acquire the reliable labeling. The experiments demonstrate that by denoising a graph this way, the robustness of graph convolutional networks on node classification task is remarkably improved, compared to several strong competitive robust graph neural network models. Ping Li 0024, Jincheng Huang 0005, Kai Zhang 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2025 | Learning Temporal Features With Alternated Similarity and Proximity Attention for Time-Series PredictionabstractTime-series prediction is a fundamental problem in various scientific and engineering domains. Recently, attention-based models have shown great promise in long-term time-series forecasting. However, we prove that vanilla attention is equivalent to a one-step random walk on a bipartite graph between the query and the keys, in which the limited number of walks and simplified graph structure could make it less powerful in capturing complex, high-order featural and temporal dependencies. Inspired by how human brains iteratively reactivate memories through reminding, we propose "Alternated Similarity And Proximity Attention," or ASAP-attention. ASAP-attention employs a random walk on two concurrent views (graphs) that, respectively, capture the featural similarity and the temporal proximity between time points. In particular, the random walk alternately visits the two graphs, each time remembering the previous probability configuration to build a coherent chain of distributions to retrieve useful historical data. This dynamic interplay between temporal and featural clues enhances the model's ability to capture implicit and heterogeneous data dependencies without using positional encoding. When incorporating ASAP-attention with encoder-only Transformer architecture, we observed highly promising results against a wide collection of state-of-the-art methods on various benchmark datasets for long time-series forecasts (e.g., weather, electricity, illness, and exchange-rate data). Our source code is available at https://github.com/jychen01/ASAP-attention. Ping Li 0024, Jiancheng Lv 0001, Hongyuan Zha, Kai Zhang 0001, Jie Zhang 0012 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Message Passing on Semantic-Anchor-Graphs for Fine-grained Emotion Representation Learning and ClassificationabstractEmotion classification has wide applications in education, robotics, virtual reality, etc.However, identifying subtle differences between fine-grained emotion categories remains challenging.Current methods typically aggregate numerous token embeddings of a sentence into a single vector, which, while being an efficient compressor, may not fully capture their complex semantic and temporal distributions.To solve this problem, we propose SEmantic ANchor Graph Neural Networks (SEAN-GNN) for fine-grained emotion classification.It learns a group of representative, multi-faceted semantic anchors in the token embedding space: using these anchors as global reference, any sentence can be projected onto them to form a "semantic-anchor graph", with node attributes and edge weights quantifying semantic and temporal information, respectively.The graph structure is well aligned across sentences and, importantly, allows for generating comprehensive emotion representations regarding K different anchors.Message passing on the anchor graph can further integrate the semantic and temporal information and refine the learned features.Empirically, SEAN-GNN produces meaningful semantic anchors and discriminative graph patterns, with promising classification results on 6 popular benchmark datasets against state-of-the-arts. Pinyi Zhang, Junchen Shen, Zijie Zhai, Ping Li 0024, Jie Zhang 0012, Kai Zhang 0001 |
EMNLP | 7 |
| 2024 | High-Order Contrastive Learning with Fine-grained Comparative Levels for Sparse Ordinal Tensor CompletionabstractContrastive learning is a powerful paradigm for representation learning with prominent success in computer vision and NLP, but how to extend its success to high-dimensional tensors remains a challenge. This is because tensor data often exhibit high-order mode-interactions that are hard to profile and with negative samples growing combinatorially faster than second-order contrastive learning; furthermore, many real-world tensors have ordinal entries that necessitate more delicate comparative levels. To solve the challenge, we propose High-Order Contrastive Tensor Completion (HOCTC), an innovative network to extend contrastive learning to sparse ordinal tensor data. HOCTC employs a novel attention-based strategy with query-expansion to capture high-order mode interactions even in case of very limited tokens, which transcends beyond second-order learning scenarios. Besides, it extends two-level comparisons (positive-vs-negative) to fine-grained contrast-levels using ordinal tensor entries as a natural guidance. Efficient sampling scheme is proposed to enforce such delicate comparative structures, generating comprehensive self-supervised signals for high-order representation learning. Extensive experiments show that HOCTC has promising results in sparse tensor completion in traffic/recommender applications. Junchen Shen, Zijie Zhai, Danlin Liu, Yu Sun 0076, Ping Li 0024, Jie Zhang 0012, Kai Zhang 0001 |
ICML | 9 |
| 2024 | Knowledge Graph Information Bottleneck for Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) prediction is an important but challenging task in drug safety surveillance. With the accumulation of biological data, biomedical knowledge graphs (KGs) become available to model DDIs and related biological mechanisms. However, the presence of substantial noise in large-scale KGs hampers prediction performance and the identification of interpretable biological pathways. To fill the gaps, this paper proposes an information bottleneck-based (IB-based) framework that simultaneously denoises the KG and identifies key entities around drug pairs. Moreover, KG-based prediction methods rarely exploit the structural information of drug molecules. To this end, the proposed framework relates drug structures to IB objectives, together with a unique drug pair-centered readout to fuse molecular information into KG subgraph embeddings. Extensive experimental results and case studies demonstrate the effectiveness and interpretability of the framework. Gaoqi He, Kai Zhang 0001, Honglin Li 0003 |
IJCNN | 3 |
| 2024 | Prototype-based contrastive substructure identification for molecular property predictionabstractSubstructure-based representation learning has emerged as a powerful approach to featurize complex attributed graphs, with promising results in molecular property prediction (MPP). However, existing MPP methods mainly rely on manually defined rules to extract substructures. It remains an open challenge to adaptively identify meaningful substructures from numerous molecular graphs to accommodate MPP tasks. To this end, this paper proposes Prototype-based cOntrastive Substructure IdentificaTion (POSIT), a self-supervised framework to autonomously discover substructural prototypes across graphs so as to guide end-to-end molecular fragmentation. During pre-training, POSIT emphasizes two key aspects of substructure identification: firstly, it imposes a soft connectivity constraint to encourage the generation of topologically meaningful substructures; secondly, it aligns resultant substructures with derived prototypes through a prototype-substructure contrastive clustering objective, ensuring attribute-based similarity within clusters. In the fine-tuning stage, a cross-scale attention mechanism is designed to integrate substructure-level information to enhance molecular representations. The effectiveness of the POSIT framework is demonstrated by experimental results from diverse real-world datasets, covering both classification and regression tasks. Moreover, visualization analysis validates the consistency of chemical priors with identified substructures. The source code is publicly available at https://github.com/VRPharmer/POSIT. Gaoqi He, Changbo Wang, Kai Zhang 0001, Honglin Li 0003 |
Briefings Bioinform. | 5 |
| 2024 | Chain-aware graph neural networks for molecular property predictionabstractMOTIVATION: Predicting the properties of molecules is a fundamental problem in drug design and discovery, while how to learn effective feature representations lies at the core of modern deep-learning-based prediction methods. Recent progress shows expressive power of graph neural networks (GNNs) in capturing structural information for molecular graphs. However, we find that most molecular graphs exhibit low clustering along with dominating chains. Such topological characteristics can induce feature squashing during message passing and thus impair the expressivity of conventional GNNs. RESULTS: Aiming at improving node features' expressiveness, we develop a novel chain-aware graph neural network model, wherein the chain structures are captured by learning the representation of the center node along the shortest paths starting from it, and the redundancy between layers are mitigated via initial residual difference connection (IRDC). Then the molecular graph is represented by attentive pooling of all node representations. Compared to standard graph convolution, our chain-aware learning scheme offers a more straightforward feature interaction between distant nodes, thus it is able to capture the information about long-range dependency. We provide extensive empirical analysis on real-world datasets to show the outperformance of the proposed method. AVAILABILITY AND IMPLEMENTATION: The MolPath code is publicly available at https://github.com/Assassinswhh/Molpath. Honghao Wang 0001, Acong Zhang, Junlei Tang, Kai Zhang 0001, Ping Li 0024 |
Bioinform. | 5 |
| 2024 | DPGCL: Dual pass filtering based graph contrastive learning
Ping Li 0024, Kai Zhang 0001 |
Neural Networks | 3 |
| 2024 | A Transformative Topological Representation for Link Modeling, Prediction and Cross-Domain Network AnalysisabstractMany complex social, biological, or physical systems are characterized as networks, and recovering the missing links of a network could shed important lights on its structure and dynamics. A good topological representation is crucial to accurate link modeling and prediction, yet how to account for the kaleidoscopic changes in link formation patterns remains a challenge, especially for analysis in cross-domain studies. We propose a new link representation scheme by projecting the local environment of a link into a "dipole plane", where neighboring nodes of the link are positioned via their relative proximity to the two anchors of the link, like a dipole. By doing this, complex and discrete topology arising from link formation is turned to differentiable point-cloud distribution, opening up new possibilities for topological feature-engineering with desired expressiveness, interpretability and generalization. Our approach has comparable or even superior results against state-of-the-art GNNs, meanwhile with a model up to hundreds of times smaller and running much faster. Furthermore, it provides a universal platform to systematically profile, study, and compare link-patterns from miscellaneous real-world networks. This allows building a global link-pattern atlas, based on which we have uncovered interesting common patterns of link formation, i.e., the bridge-style, the radiation-style, and the community-style across a wide collection of networks with highly different nature. Kai Zhang 0001, Junchen Shen, Gaoqi He, Yu Sun 0076, Haibin Ling, Hongyuan Zha, Honglin Li 0003, Jie Zhang 0012 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Building Shortcuts between Distant Nodes with Biaffine Mapping for Graph Convolutional NetworksabstractMultiple recent studies show a paradox in graph convolutional networks (GCNs)—that is, shallow architectures limit the capability of learning information from high-order neighbors, whereas deep architectures suffer from over-smoothing or over-squashing. To enjoy the simplicity of shallow architectures and overcome their limits of neighborhood extension, in this work we introduce a biaffine technique to improve the expressiveness of GCNs with a shallow architecture. The core design of our method is to learn direct dependency on long-distance neighbors for nodes, with which only 1-hop message passing is capable of capturing rich information for node representation. Besides, we propose a multi-view contrastive learning method to exploit the representations learned from long-distance dependencies. Extensive experiments on nine graph benchmark datasets suggest that the shallow biaffine graph convolutional networks (BAGCN) significantly outperform state-of-the-art GCNs (with deep or shallow architectures) on semi-supervised node classification. We further verify the effectiveness of biaffine design in node representation learning and the performance consistency on different sizes of training data. Acong Zhang, Jincheng Huang 0005, Ping Li 0024, Kai Zhang 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | Fast Convolutional Factorization Machine With Enhanced RobustnessabstractRecently, factorization machine and its variants have shown promising results for context-aware recommender systems (CARS), especially when combined with deep neural networks. Among them, convolutional factorization machine (CFM) is a prominent example. The key to the success of CFM is its 3D convolutional architecture for capturing complex interactions on top of embedded features. However, the resultant computational cost can also be demanding. Moreover, the feature embedding scheme of CFM and other factorization models can be potentially vulnerable to noise. To tackle these issues, in this study we propose two models, namely, the fast convolutional factorization machine (FCFM) that slims down the complete pairwise feature interaction for higher computational efficiency, and adversarial fast convolutional factorization machine (AFCFM) that further enhances the robustness of the model by introducing adversarial noise to the feature interaction image generated by the model. Experimental results on four benchmark datasets prove that the proposed FCFM is nearly five times faster than CFM with competitive performance, while AFCFM improves the performance of the state-of-the-art models by about 8\% with higher efficiency than CFM. Jie Zhang 0012, Kai Zhang 0001, Ping Li 0024 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Diagnostic Sparse Connectivity Networks With Regularization TemplateabstractDynamic systems are often monitored with multivariate time series where each dimension represents a local component measured through a (virtual) sensor. Performing accurate diagnostic for dynamic systems while simultaneously taking into account their similarities/distinctions, is a non-trivial task. To this end, we develop an adaptive regularization approach to learning sparse connectivity structures in complex dynamic systems. The learned connectivity networks shed lights on the structural compositions of the system and hence can serve as highly informative inputs for various machine learning tasks such as classification. In particular, we focus on high-dimensional and semi-supervised learning scenarios and present a joint learning approach to recover system-wise connectivity patterns by adaptively constructing a shared, sparsity-inducing regularization template across all systems. The shared template can be physically interpreted and used as a modeling template for analyzing new systems. Moreover, our approach has the flexibility to incorporate supervising information such as must-links and cannot-links for constructing regularization templates. Overall, our approach, named sparse adaptive regularization (SAR), can extract structure-related connectivity features efficiently and effectively, and result in significant improvements for machine learning tasks in dynamic systems. We benchmark our approach against the state-of-the-art methods with real-world data. Our results demonstrate the superiority of our approach. Chuanren Liu, Kai Zhang 0001, Keli Xiao, Bo Jin 0001, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Structural Landmarking and Interaction Modelling: A "SLIM" Network for Graph ClassificationabstractGraph neural networks are a promising architecture for learning and inference with graph-structured data. Yet, how to generate informative, fixed dimensional features for graphs with varying size and topology can still be challenging. Typically, this is achieved through graph-pooling, which summarizes a graph by compressing all its nodes into a single vector. Is such a “collapsing-style” graph-pooling the only choice for graph classification? From complex system’s point of view, properties of a complex system arise largely from the interaction among its components. Therefore, we speculate that preserving the interacting relation between parts, instead of pooling them together, could benefit system level prediction. To verify this, we propose SLIM, a graph neural network model for Structural Landmarking and Interaction Modelling. The main idea is to compute a set of end-to-end optimizable sub-structure landmarks, so that any input graph can be projected onto these (spatially) local structural representatives for a faithful, global characterization. By doing so, explicit interaction between component parts of a graph can be leveraged directly in generating discriminative graph representation. Encouraging results are observed on benchmark datasets for graph classification, demonstrating the value of interaction modelling in the design of graph neural networks. Yaokang Zhu, Kai Zhang 0001, Jun Wang 0006, Haibin Ling, Jie Zhang 0012, Hongyuan Zha |
AAAI | 2 |
| 2022 | Semantic consistency for graph representation learningabstractIn graph learning, it is fundamental to integrate the features from graph structure and node attributes. Towards this end, graph convolution technique has been devised based on the premise that the similarity of node attributes between two nodes is semantically consistent with their topological proximity. However, many real-networks are found to exhibit the semantic inconsistency, i.e., the phenomenon that directly connected nodes are dissimilar in their attributes. This work is concerned with two related issues: how do we quantitatively measure the semantic consistency between node attributes and graph structure? can we leverage this information to facilitate graph representation? To answer those questions, we first introduce a novel metric to evaluate the semantic consistency in a graph, and then we identify a set of key designs to encode the local semantic consistency information into a type of ego's node feature. Then, we fuse this new node feature with the original node attributes by concatenating the two parts using the semantic consistency metric as weight factor. Experiments on real-world datasets show that linear classifier (e.g. multilayer perceptrons) based on our unsupervised feature learning scheme achieves strong performance across the datasets, especially on the datasets with low semantic consistency, compared to the popular supervised GCNs and other competitive unsupervised graph representation learning models. Jincheng Huang 0005, Ping Li 0024, Kai Zhang 0001 |
IJCNN | 3 |
| 2022 | APG: Adaptive Parameter Generation Network for Click-Through Rate PredictionabstractIn many web applications, deep learning-based CTR prediction models (deep CTR models for short) are widely adopted. Traditional deep CTR models learn patterns in a static manner, i.e., the network parameters are the same across all the instances. However, such a manner can hardly characterize each of the instances which may have different underlying distributions. It actually limits the representation power of deep CTR models, leading to sub-optimal results. In this paper, we propose an efficient, effective, and universal module, named as Adaptive Parameter Generation network (APG), which can dynamically generate parameters for deep CTR models on-the-fly based on different instances. Extensive experimental evaluation results show that APG can be applied to a variety of deep CTR models and significantly improve their performance. Meanwhile, APG can reduce the time cost by 38.7\% and memory usage by 96.6\% compared to a regular deep CTR model.We have deployed APG in the industrial sponsored search system and achieved 3\% CTR gain and 1\% RPM gain respectively. Bencheng Yan, Pengjie Wang 0002, Kai Zhang 0001, Feng Li 0067, Hongbo Deng, Jian Xu 0015, Bo Zheng 0007 |
NeurIPS | 3 |
| 2022 | Multi-modal chemical information reconstruction from images and texts for exploring the near-drug spaceabstractIdentification of new chemical compounds with desired structural diversity and biological properties plays an essential role in drug discovery, yet the construction of such a potential space with elements of 'near-drug' properties is still a challenging task. In this work, we proposed a multimodal chemical information reconstruction system to automatically process, extract and align heterogeneous information from the text descriptions and structural images of chemical patents. Our key innovation lies in a heterogeneous data generator that produces cross-modality training data in the form of text descriptions and Markush structure images, from which a two-branch model with image- and text-processing units can then learn to both recognize heterogeneous chemical entities and simultaneously capture their correspondence. In particular, we have collected chemical structures from ChEMBL database and chemical patents from the European Patent Office and the US Patent and Trademark Office using keywords 'A61P, compound, structure' in the years from 2010 to 2020, and generated heterogeneous chemical information datasets with 210K structural images and 7818 annotated text snippets. Based on the reconstructed results and substituent replacement rules, structural libraries of a huge number of near-drug compounds can be generated automatically. In quantitative evaluations, our model can correctly reconstruct 97% of the molecular images into structured format and achieve an F1-score around 97-98% in the recognition of chemical entities, which demonstrated the effectiveness of our model in automatic information extraction from chemical patents, and hopefully transforming them to a user-friendly, structured molecular database enriching the near-drug space to realize the intelligent retrieval technology of chemical knowledge. Jie Wang 0146, Zihao Shen, Yichen Liao, Shiliang Li, Gaoqi He, Man Lan, Xuhong Qian, Kai Zhang 0001, Honglin Li 0003 |
Briefings Bioinform. | 9 |
| 2022 | Node Embedding and Classification with Adaptive Structural Fingerprint
Yaokang Zhu, Jun Wang 0006, Jie Zhang 0012, Kai Zhang 0001 |
Neurocomputing | 4 |
| 2021 | Learning Effective and Efficient Embedding via an Adaptively-Masked Twins-based LayerabstractEmbedding learning for categorical features is crucial for the deep learning-based recommendation models (DLRMs). Each feature value is mapped to an embedding vector via an embedding learning process. Conventional methods configure a fixed and uniform embedding size to all feature values from the same feature field. However, such a configuration is not only sub-optimal for embedding learning but also memory costly. Existing methods that attempt to resolve these problems, either rule-based or neural architecture search (NAS)-based, need extensive efforts on the human design or network training. They are also not flexible in embedding size selection or in warm-start-based applications. In this paper, we propose a novel and effective embedding size selection scheme. Specifically, we design an Adaptively-Masked Twins-based Layer (AMTL) behind the standard embedding layer. AMTL generates a mask vector to mask the undesired dimensions for each embedding vector. The mask vector brings flexibility in selecting the dimensions and the proposed layer can be easily added to either untrained or trained DLRMs. Extensive experimental evaluations show that the proposed scheme outperforms competitive baselines on all the benchmark tasks, and is also memory-efficient, saving 60% memory usage without compromising any performance metrics. Bencheng Yan, Pengjie Wang 0002, Kai Zhang 0001, Wei Lin 0016, Kuang-chih Lee, Jian Xu 0015, Bo Zheng 0007 |
CIKM | 3 |
| 2021 | GMOT-40: A Benchmark for Generic Multiple Object TrackingabstractMultiple Object Tracking (MOT) has witnessed remarkable advances in recent years. However, existing studies dominantly request prior knowledge of the tracking target (eg, pedestrians), and hence may not generalize well to unseen categories. In contrast, Generic Multiple Object Tracking (GMOT), which requires little prior information about the target, is largely under-explored. In this paper, we make contributions to boost the study of GMOT in three aspects. First, we construct the first publicly available dense GMOT dataset, dubbed GMOT-40, which contains 40 carefully annotated sequences evenly distributed among 10 object categories. In addition, two tracking protocols are adopted to evaluate different characteristics of tracking algorithms. Second, by noting the lack of devoted tracking algorithms, we have designed a series of baseline GMOT algorithms. Third, we perform a thorough evaluations on GMOT-40, involving popular MOT algorithms (with necessary modifications) and the proposed baselines. The GMOT-40 benchmark is publicly available at https://github.com/Spritea/GMOT40. Hexin Bai, Wensheng Cheng, Peng Chu, Juehuan Liu, Kai Zhang 0001, Haibin Ling |
CVPR | 5 |
| 2020 | Online Bayesian Sparse Learning with Spike and Slab PriorsabstractIn many applications, a parsimonious model is often preferred for better interpretability and predictive performance. Online algorithms have been studied extensively for building such models in big data and fast evolving environments, with a prominent example, FTRL-proximal [1]. However, existing methods typically do not provide confidence levels, and with the usage of L1 regularization, the model estimation can be undermined by the uniform shrinkage on both relevant and irrelevant features. To address these issues, we developed OLSS, a Bayesian online sparse learning algorithm based on the spike-and-slab prior. OLSS achieves the same scalability as FTRL-proximal, but realizes appealing selective shrinkage and produces rich uncertainty information, such as posterior inclusion probabilities and feature weight variances. On the tasks of text classification and click-through-rate (CTR) prediction for Yahoo!'s display and search advertisement platforms, OLSS often demonstrates superior predictive performance to the state-of-the-art methods in industry, including Vowpal Wabbit [2] and FTRL-proximal. Shikai Fang, Shandian Zhe, Kuang-chih Lee, Kai Zhang 0001, Jennifer Neville |
ICDM | 4 |
| 2020 | Probabilistic Neural-Kernel Tensor DecompositionabstractTensor decomposition is a fundamental framework to model and analyze multiway data, which are ubiquitous in realworld applications. A critical challenge of tensor decomposition is to capture a variety of complex relationships/interactions while avoiding overfitting the data that are usually very sparse. Although numerous tensor decomposition methods have been proposed, they are mostly based on a multilinear form and hence are incapable of estimating more complex, nonlinear relationships. To address the challenge, we propose POND, PrObabilistic Neural-kernel tensor Decomposition that unifies the self-adaptation of Bayes nonparametric function learning and the expressive power of neural networks. POND uses Gaussian processes (GPs) to model the hidden relationships and can automatically detect their complexity in tensors, preventing both underfitting and overfitting. POND then incorporates convolutional neural networks to construct the GP kernel to greatly promote the capability of estimating highly nonlinear relationships. To scale POND to large data, we use the sparse variational GP framework and reparameterization trick to develop an efficient stochastic variational learning algorithm. On both synthetic and real-world benchmark datasets, POND often exhibits better predictive performance than the state-of-the-art nonlinear tensor decomposition methods. In addition, as a Bayesian approach, POND provides the posterior distribution of the latent factors, and hence can conveniently quantify their uncertainty and the confidence levels for predictions. Conor Tillinghast, Shikai Fang, Kai Zhang 0001, Shandian Zhe |
ICDM | 3 |
| 2020 | Adaptive Structural Fingerprints for Graph Attention Networks
Kai Zhang 0001, Yaokang Zhu, Jun Wang 0006, Jie Zhang 0012 |
ICLR | 1 |
| 2019 | Brain annotation toolbox: exploring the functional and genetic associations of neuroimaging resultsabstractMOTIVATION: Advances in neuroimaging and sequencing techniques provide an unprecedented opportunity to map the function of brain regions and identify the roots of psychiatric diseases. However, the results from most neuroimaging studies, i.e. activated clusters/regions or functional connectivities between brain regions, frequently cannot be conveniently and systematically interpreted, rendering the biological meaning unclear. RESULTS: We describe a brain annotation toolbox that generates functional and genetic annotations for neuroimaging results. The voxel-level functional description from the Neurosynth database and gene expression profile from the Allen Human Brain Atlas are used to generate functional/genetic information for region-level neuroimaging results. The validity of the approach is demonstrated by showing that the functional and genetic annotations for specific brain regions are consistent with each other; and further the region by region functional similarity network and genetic similarity network are highly correlated for major brain atlases. One application of brain annotation toolbox is to help provide functional/genetic annotations for newly discovered regions with unknown functions, e.g. the 97 new regions identified in the Human Connectome Project. Importantly, this toolbox can help understand differences between psychiatric patients and controls, and this is demonstrated using schizophrenia and autism data, for which the functional and genetic annotations for the neuroimaging changes in patients are consistent with each other and help interpret the results. AVAILABILITY AND IMPLEMENTATION: BAT is implemented as a free and open-source MATLAB toolbox and is publicly available at http://123.56.224.61:1313/post/bat. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhaowen Liu, Edmund T. Rolls, Zhi Liu 0004, Kai Zhang 0001, Jingnan Du, Weikang Gong, Wei Cheng 0011, He Wang 0016, Kâmil Ugurbil, Jie Zhang 0012, Jianfeng Feng |
Bioinform. | 4 |
| 2019 | Augmented label propagation for seed set expansion
Xinyu Peng, Ping Li 0024, Kai Zhang 0001, Yan Chen 0057 |
Knowl. Based Syst. | 4 |
| 2019 | A Feature Sampling Strategy for Analysis of High Dimensional Genomic DataabstractWith the development of high throughput technology, it has become feasible and common to profile tens of thousands of gene activities simultaneously. These genomic data typically have sample size of hundreds or fewer, which is much less than the feature size (number of genes). In addition, the genes, in particular the ones from the same pathway, are often highly correlated. These issues impose a great challenge for selecting meaningful genes from a large number of (correlated) candidates in many genomic studies. Quite a few methods have been proposed to attack this challenge. Among them, regularization-based techniques, e.g., lasso, become much more appealing, because they can do model fitting and variable selection at the same time. However, the lasso regression has its known limitations. One is that the number of genes selected by the lasso couldn't exceed the number of samples. Another limitation is that, if causal genes are highly correlated, the lasso tends to select only one or few genes from them. Biologists, however, desire to identify them all. To overcome these limitations, we present here a novel, robust, and stable variable selection method. Through simulation studies and a real application to the transcriptome data, we demonstrate the superiority of the proposed method in selecting highly correlated causal genes. We also provide some theoretical justifications for this feature sampling strategy based on the mean and variance analyses. Jie Zhang 0049, Zhigen Zhao, Kai Zhang 0001, Zhi Wei 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | Scaling Up Kernel SVM on Limited Resources: A Low-Rank Linearization ApproachabstractKernel support vector machines (SVMs) deliver state-of-the-art results in many real-world nonlinear classification problems, but the computational cost can be quite demanding in order to maintain a large number of support vectors. Linear SVM, on the other hand, is highly scalable to large data but only suited for linearly separable problems. In this paper, we propose a novel approach called low-rank linearized SVM to scale up kernel SVM on limited resources. Our approach transforms a nonlinear SVM to a linear one via an approximate empirical kernel map computed from efficient kernel low-rank decompositions. We theoretically analyze the gap between the solutions of the approximate and optimal rank- k kernel map, which in turn provides guidance on the sampling scheme of the Nyström approximation. Furthermore, we extend it to a semisupervised metric learning scenario in which partially labeled samples can be exploited to further improve the quality of the low-rank embedding. Our approach inherits rich representability of kernel SVM and high efficiency of linear SVM. Experimental results demonstrate that our approach is more robust and achieves a better tradeoff between model representability and scalability against state-of-the-art algorithms for large-scale SVMs. Liang Lan, Shandian Zhe, Wei Cheng 0002, Jun Wang 0006, Kai Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2018 | Collaborative Alert Ranking for Anomaly DetectionabstractGiven a large number of low-quality heterogeneous categorical alerts collected from an anomaly detection system, how to characterize the complex relationships between different alerts and deliver trustworthy rankings to end users? While existing techniques focus on either mining alert patterns or filtering out false positive alerts, it can be more advantageous to consider the two perspectives simultaneously in order to improve detection accuracy and better understand abnormal system behaviors. In this paper, we propose CAR, a collaborative alert ranking framework that exploits both temporal and content correlations from heterogeneous categorical alerts. CAR first builds a hierarchical Bayesian model to capture both short-term and long-term dependencies in each alert sequence. Then, an entity embedding-based model is proposed to learn the content correlations between alerts via their heterogeneous categorical attributes. Finally, by incorporating both temporal and content dependencies into a unified optimization framework, CAR ranks both alerts and their corresponding alert patterns. Our experiments-using both synthetic and real-world enterprise security alert data-show that CAR can accurately identify true positive alerts and successfully reconstruct the attack scenarios at the same time. Zhengzhang Chen, Lu-An Tang, Kai Zhang 0001, Wei Cheng 0002, Zhichun Li |
CIKM | 5 |
| 2018 | DisTenC: A Distributed Algorithm for Scalable Tensor Completion on SparkabstractHow can we efficiently recover missing values for very large-scale real-world datasets that are multi-dimensional even when the auxiliary information is regularized at certain mode? Tensor completion is a useful tool to recover a low-rank tensor that best approximates partially observed data and further predicts the unobserved data by this low-rank tensor, which has been successfully used for many applications such as location-based recommender systems, link prediction, targeted advertising, social media search, and event detection. Due to the curse of dimensionality, existing algorithms for tensor completion that integrate auxiliary information do not scale for tensors with billions of elements. In this paper, we propose DisTenC, a new distributed large-scale tensor completion algorithm that can be distributed on Spark. Our key insights are to (i) efficiently handle trace-based regularization terms; (ii) update factor matrices with caching; and (iii) optimize the update of the new tensor via residuals. In this way, we can tackle the high computational costs of traditional approaches and minimize intermediate data, leading to order-of-magnitude improvements in tensor completion. Experimental results demonstrate that DisTenC is capable of handling up to 10~1000X larger tensors than existing methods with much faster convergence rate, shows better linearity on machine scalability, and achieves up to an average improvement of 23.5% in accuracy in applications. Hancheng Ge, Kai Zhang 0001, Majid Alfifi, Xia Ben Hu, James Caverlee |
ICDE | 2 |
| 2018 | NetWalk: A Flexible Deep Embedding Approach for Anomaly Detection in Dynamic NetworksabstractMassive and dynamic networks arise in many practical applications such as social media, security and public health. Given an evolutionary network, it is crucial to detect structural anomalies, such as vertices and edges whose "behaviors'' deviate from underlying majority of the network, in a real-time fashion. Recently, network embedding has proven a powerful tool in learning the low-dimensional representations of vertices in networks that can capture and preserve the network structure. However, most existing network embedding approaches are designed for static networks, and thus may not be perfectly suited for a dynamic environment in which the network representation has to be constantly updated. In this paper, we propose a novel approach, NetWalk, for anomaly detection in dynamic networks by learning network representations which can be updated dynamically as the network evolves. We first encode the vertices of the dynamic network to vector representations by clique embedding, which jointly minimizes the pairwise distance of vertex representations of each walk derived from the dynamic networks, and the deep autoencoder reconstruction error serving as a global regularization. The vector representations can be computed with constant space requirements using reservoir sampling. On the basis of the learned low-dimensional vertex representations, a clustering-based technique is employed to incrementally and dynamically detect network anomalies. Compared with existing approaches, NetWalk has several advantages: 1) the network embedding can be updated dynamically, 2) streaming network nodes and edges can be encoded efficiently with constant memory space usage, 3). flexible to be applied on different types of networks, and 4) network anomalies can be detected in real-time. Extensive experiments on four real datasets demonstrate the effectiveness of NetWalk. Wenchao Yu, Wei Cheng 0002, Charu C. Aggarwal, Kai Zhang 0001, Wei Wang 0010 |
KDD | 4 |
| 2018 | Network Inference from Contrastive Groups Using Discriminative Structural RegularizationabstractGaussian graphical models (GGMs) are a popular tool for exploring conditional dependence among high dimensional data. We consider developing an estimator for GGMs for multiple graph analysis, wherein the graphs are assumed to come from two (or more) contrastive groups, and exhibit not only major global similarity, but also substantial between-group disparity. Under this setting, inferring each group of networks separately ignores the common structure, while simply assuming a global common network structure would mask the critical disparity. We propose a novel approach to pursue simultaneous network inference using discriminative and adaptive structural regularizations. We introduce a heterogeneity ratio parameter to balance the within group similarity and the between group disparity. This formulation for the first time, to our knowledge, generalizes the existing single-group network analysis to multiple-group network analysis. In other words, our proposed multiple-group network analysis reduces to single-group network analysis, when the heterogeneity ratio equal to 1. By iteratively updating a global regularization template with individual network structures, together with a feature screening module specifying relevant dimensions to satisfy the group-level constraints, our generalized approach can recover the underlying conditional independence with greater flexibility and improved accuracy. Theoretically, we show the asymptotic consistency for the proposed method in joint reconstruction of multiple network structures. We demonstrate its superior performance via extensive simulation studies. We also illustrate its practical usage in an application to polychromatic flow cytometry data sets for protein interactions under different conditions. Ruihua Cheng, Zhi Wei 0001, Kai Zhang 0001 |
SDM | 3 |
| 2017 | Efficient Discovery of Abnormal Event Sequences in Enterprise Security SystemsabstractIntrusion detection system (IDS) is an important part of enterprise security system architecture. In particular, anomaly-based IDS has been widely applied to detect single abnormal process events that deviate from the majority. However, intrusion activity usually consists of a series of low-level heterogeneous events. The gap between low-level process events and high-level intrusion activities makes it particularly challenging to identify process events that are truly involved in a real malicious activity, and especially considering the massive 'noisy' events filling the event sequences. Hence, the existing work that focus on detecting single events can hardly achieve high detection accuracy. In this work, we formulate a novel problem in intrusion detection - suspicious event sequence discovery, and propose GID, an efficient graph-based intrusion detection technique that can identify abnormal event sequences from massive heterogeneous process traces with high accuracy. We fully implement GID and deploy it into a real-world enterprise security system, and it greatly helps detect the advanced threats and optimize the incident response. Executing GID on both static and streaming data shows that GID is efficient (processes about 2 million records per minute) and accurate for intrusion detection. Boxiang Dong, Zhengzhang Chen, Wendy Hui Wang, Lu-An Tang, Kai Zhang 0001, Zhichun Li |
CIKM | 5 |
| 2017 | Ranking Causal Anomalies by Modeling Local Propagations on Networked SystemsabstractComplex systems are prevalent in many fields such as finance, security and industry. A fundamental problem in system management is to perform diagnosis in case of system failure such that the causal anomalies, i.e., root causes, can be identified for system debugging and repair. Recently, invariant network has proven a powerful tool in characterizing complex system behaviors. In an invariant network, a node represents a system component, and an edge indicates a stable interaction between two components. Recent approaches have shown that by modeling fault propagation in the invariant network, causal anomalies can be effectively discovered. Despite their success, the existing methods have a major limitation: they typically assume there is only a single and global fault propagation in the entire network. However, in real-world large-scale complex systems, it's more common for multiple fault propagations to grow simultaneously and locally within different node clusters and jointly define the system failure status. Inspired by this key observation, we propose a two-phase framework to identify and rank causal anomalies. In the first phase, a probabilistic clustering is performed to uncover impaired node clusters in the invariant network. Then, in the second phase, a low-rank network diffusion model is designed to backtrack causal anomalies in different impaired clusters. Extensive experimental results on real-life datasets demonstrate the effectiveness of our method. Jingchao Ni, Wei Cheng 0002, Kai Zhang 0001, Dongjin Song, Tan Yan, Xiang Zhang 0001 |
ICDM | 3 |
| 2017 | Randomization or Condensation?: Linear-Cost Matrix Sketching Via Cascaded Compression SamplingabstractMatrix sketching is aimed at finding compact representations of a matrix while simultaneously preserving most of its properties, which is a fundamental building block in modern scientific computing. Randomized algorithms represent state-of-the-art and have attracted huge interest from the fields of machine learning, data mining, and theoretic computer science. However, it still requires the use of the entire input matrix in producing desired factorizations, which can be a major computational and memory bottleneck in truly large problems. In this paper, we uncover an interesting theoretic connection between matrix low-rank decomposition and lossy signal compression, based on which a cascaded compression sampling framework is devised to approximate an m-by-n matrix in only O(m+n) time and space. Indeed, the proposed method accesses only a small number of matrix rows and columns, which significantly improves the memory footprint. Meanwhile, by sequentially teaming two rounds of approximation procedures and upgrading the sampling strategy from a uniform probability to more sophisticated, encoding-orientated sampling, significant algorithmic boosting is achieved to uncover more granular structures in the data. Empirical results on a wide spectrum of real-world, large-scale matrices show that by taking only linear time and space, the accuracy of our method rivals those state-of-the-art randomized algorithms consuming a quadratic, O(mn), amount of resources. Kai Zhang 0001, Chuanren Liu, Jie Zhang 0012, Hui Xiong 0001, Eric P. Xing, Jieping Ye |
KDD | 1 |
| 2017 | Low-rank decomposition meets kernel learning: A generalized Nyström method
Liang Lan, Kai Zhang 0001, Hancheng Ge, Wei Cheng 0002, Jun Liu 0003, Andreas Rauber, Xiaoli Li 0001, Jun Wang 0006, Hongyuan Zha |
Artif. Intell. | 2 |
| 2017 | Ranking Causal Anomalies for System Fault Diagnosis via Temporal and Dynamical Analysis on Vanishing CorrelationsabstractDetecting system anomalies is an important problem in many fields such as security, fault management, and industrial optimization. Recently, invariant network has shown to be powerful in characterizing complex system behaviours. In the invariant network, a node represents a system component and an edge indicates a stable, significant interaction between two components. Structures and evolutions of the invariance network, in particular the vanishing correlations, can shed important light on locating causal anomalies and performing diagnosis. However, existing approaches to detect causal anomalies with the invariant network often use the percentage of vanishing correlations to rank possible casual components, which have several limitations: (1) fault propagation in the network is ignored, (2) the root casual anomalies may not always be the nodes with a high percentage of vanishing correlations, (3) temporal patterns of vanishing correlations are not exploited for robust detection, and (4) prior knowledge on anomalous nodes are not exploited for (semi-)supervised detection. To address these limitations, in this article we propose a network diffusion based framework to identify significant causal anomalies and rank them. Our approach can effectively model fault propagation over the entire invariant network and can perform joint inference on both the structural and the time-evolving broken invariance patterns. As a result, it can locate high-confidence anomalies that are truly responsible for the vanishing correlations and can compensate for unstructured measurement noise in the system. Moreover, when the prior knowledge on the anomalous status of some nodes are available at certain time points, our approach is able to leverage them to further enhance the anomaly inference accuracy. When the prior knowledge is noisy, our approach also automatically learns reliable information and reduces impacts from noises. By performing extensive experiments on synthetic datasets, bank information system datasets, and coal plant cyber-physical system datasets, we demonstrate the effectiveness of our approach. Wei Cheng 0002, Jingchao Ni, Kai Zhang 0001, Guofei Jiang, Yu Shi 0002, Xiang Zhang 0001, Wei Wang 0010 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2016 | Entity Embedding-Based Anomaly Detection for Heterogeneous Categorical Events
Ting Chen 0007, Lu-An Tang, Yizhou Sun, Zhengzhang Chen, Kai Zhang 0001 |
IJCAI | 5 |
| 2016 | Ranking Causal Anomalies via Temporal and Dynamical Analysis on Vanishing CorrelationsabstractModern world has witnessed a dramatic increase in our ability to collect, transmit and distribute real-time monitoring and surveillance data from large-scale information systems and cyber-physical systems. Detecting system anomalies thus attracts significant amount of interest in many fields such as security, fault management, and industrial optimization. Recently, invariant network has shown to be a powerful way in characterizing complex system behaviours. In the invariant network, a node represents a system component and an edge indicates a stable, significant interaction between two components. Structures and evolutions of the invariance network, in particular the vanishing correlations, can shed important light on locating causal anomalies and performing diagnosis. However, existing approaches to detect causal anomalies with the invariant network often use the percentage of vanishing correlations to rank possible casual components, which have several limitations: 1) fault propagation in the network is ignored; 2) the root casual anomalies may not always be the nodes with a high-percentage of vanishing correlations; 3) temporal patterns of vanishing correlations are not exploited for robust detection. To address these limitations, in this paper we propose a network diffusion based framework to identify significant causal anomalies and rank them. Our approach can effectively model fault propagation over the entire invariant network, and can perform joint inference on both the structural, and the time-evolving broken invariance patterns. As a result, it can locate high-confidence anomalies that are truly responsible for the vanishing correlations, and can compensate for unstructured measurement noise in the system. Extensive experiments on synthetic datasets, bank information system datasets, and coal plant cyber-physical system datasets demonstrate the effectiveness of our approach. Wei Cheng 0002, Kai Zhang 0001, Guofei Jiang, Zhengzhang Chen, Wei Wang 0010 |
KDD | 2 |
| 2016 | Annealed Sparsity via Adaptive and Dynamic ShrinkingabstractSparse learning has received tremendous amount of interest in high-dimensional data analysis due to its model interpretability and the low-computational cost. Among the various techniques, adaptive l1-regularization is an effective framework to improve the convergence behaviour of the LASSO, by using varying strength of regularization across different features. In the meantime, the adaptive structure makes it very powerful in modelling grouped sparsity patterns as well, being particularly useful in high-dimensional multi-task problems. However, choosing an appropriate, global regularization weight is still an open problem. In this paper, inspired by the annealing technique in material science, we propose to achieve "annealed sparsity" by designing a dynamic shrinking scheme that simultaneously optimizes the regularization weights and model coefficients in sparse (multi-task) learning. The dynamic structures of our algorithm are twofold. Feature-wise (spatially), the regularization weights are updated interactively with model coefficients, allowing us to improve the global regularization structure. Iteration-wise (temporally), such interaction is coupled with gradually boosted l1-regularization by adjusting an equality norm-constraint, achieving an annealing effect to further improve model selection. This renders interesting shrinking behaviour in the whole solution path. Our method competes favorably with state-of-the-art methods in sparse (multi-task) learning. We also apply it in expression quantitative trait loci analysis (eQTL), which gives useful biological insights in human cancer (melanoma) study. Kai Zhang 0001, Shandian Zhe, Chaoran Cheng, Zhi Wei 0001, Zhengzhang Chen, Guofei Jiang, Yuan Qi 0001, Jieping Ye |
KDD | 1 |
| 2016 | Distributed Flexible Nonlinear Tensor FactorizationabstractTensor factorization is a powerful tool to analyse multi-way data. Recently proposed nonlinear factorization methods, although capable of capturing complex relationships, are computationally quite expensive and may suffer a severe learning bias in case of extreme data sparsity. Therefore, we propose a distributed, flexible nonlinear tensor factorization model, which avoids the expensive computations and structural restrictions of the Kronecker-product in the existing TGP formulations, allowing an arbitrary subset of tensor entries to be selected for training. Meanwhile, we derive a tractable and tight variational evidence lower bound (ELBO) that enables highly decoupled, parallel computations and high-quality inference. Based on the new bound, we develop a distributed, key-value-free inference algorithm in the MapReduce framework, which can fully exploit the memory cache mechanism in fast MapReduce systems such as Spark. Experiments demonstrate the advantages of our method over several state-of-the-art approaches, in terms of both predictive performance and computational efficiency. Shandian Zhe, Kai Zhang 0001, Pengyuan Wang 0001, Kuang-chih Lee, Zenglin Xu, Yuan Qi 0001, Zoubin Ghahramani |
NIPS | 2 |
| 2016 | Enhancing semi-supervised learning through label-aware base kernels
Qiaojun Wang, Kai Zhang 0001, Zhengzhang Chen, Dequan Wang, Guofei Jiang, Ivan Marsic |
Neurocomputing | 2 |
| 2016 | Temporal Skeletonization on Sequential Data: Patterns, Categorization, and VisualizationabstractSequential pattern analysis aims at finding statistically relevant temporal structures where the values are delivered in a sequence. With the growing complexity of real-world dynamic scenarios, more and more symbols are often needed to encode the sequential values. This is so-called “curse of cardinality”, which can impose significant challenges to the design of sequential analysis methods in terms of computational efficiency and practical use. Indeed, given the overwhelming scale and the heterogeneous nature of the sequential data, new visions and strategies are needed to face the challenges. To this end, in this paper, we propose a “temporal skeletonization” approach to proactively reduce the cardinality of the representation for sequences by uncovering significant, hidden temporal structures. The key idea is to summarize the temporal correlations in an undirected graph, and use the “skeleton” of the graph as a higher granularity on which hidden temporal patterns are more likely to be identified. As a consequence, the embedding topology of the graph allows us to translate the rich temporal content into a metric space. This opens up new possibilities to explore, quantify, and visualize sequential data. Our approach has shown to greatly alleviate the curse of cardinality in challenging tasks of sequential pattern mining and clustering. Evaluation on a business-to-business (B2B) marketing application demonstrates that our approach can effectively discover critical buying paths from noisy customer event data. Chuanren Liu, Kai Zhang 0001, Hui Xiong 0001, Guofei Jiang, Qiang Yang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | A Quality Control Engine for Complex Physical SystemsabstractThis paper proposes a novel framework to automatically pinpoint suspicious sensors that lead to the quality change in physical systems such as manufacture plants. Our framework treats sensor readings as time series, and contains three main stages: time series transformation to feature series, feature ranking, and ranking score fusion. In the first step, we transform time series into a number of different feature series to describe the underlying dynamics of each sensor data. After that, the importance scores of all feature series are computed by utilizing several feature selection and ranking techniques, each of which discovers specific aspects of feature importance and their dependencies in the feature space. Finally we combine importance scores from all the rankers and all the features to obtain the final ranking of each sensor with respect to the system quality change. Our experiments based on synthetic time series as well as sensor data from a real system demonstrate the effectiveness of proposed method. In addition, we have implemented our framework as a production engine, and successfully applied it to several real physical systems. Takehiko Mizoguchi, Kai Zhang 0001, Geoff Jiang |
DSN | 4 |
| 2015 | Efficient Long-Term Degradation Profiling in Time Series for Complex Physical SystemsabstractThe long term operation of physical systems inevitably leads to their wearing out, and may cause degradations in performance or the unexpected failure of the entire system. To reduce the possibility of such unanticipated failures, the system must be monitored for tell-tale symptoms of degradation that are suggestive of imminent failure. In this work, we introduce a novel time series analysis technique that allows the decomposition of the time series into trend and fluctuation components, providing the monitoring software with actionable information about the changes of the system's behavior over time. We analyze the underlying problem and formulate it to a Quadratic Programming (QP) problem that can be solved with existing QP-solvers. However, when the profiling resolution is high, as generally required by real-world applications, such a decomposition becomes intractable to general QP-solvers. To speed up the problem solving, we further transform the problem and present a novel QP formulation, Non-negative QP, for the problem and demonstrate a tractable solution that bypasses the use of slow general QP-solvers. We demonstrate our ideas on both synthetic and real datasets, showing that our method allows us to accurately extract the degradation phenomenon of time series. We further demonstrate the generality of our ideas by applying them beyond classic machine prognostics to problems in identifying the influence of news events on currency exchange rates and stock prices. We fully implement our profiling system and deploy it into several physical systems, such as chemical plants and nuclear power plants, and it greatly helps detect the degradation phenomenon, and diagnose the corresponding components. Liudmila Ulanova, Tan Yan, Guofei Jiang, Eamonn J. Keogh, Kai Zhang 0001 |
KDD | 6 |
| 2015 | From Categorical to Numerical: Multiple Transitive Distance Learning and EmbeddingabstractCategorical data are ubiquitous in real-world databases. However, due to the lack of an intrinsic proximity measure, many powerful algorithms for numerical data analysis may not work well on their categorical counterparts, making it a bottleneck in practical applications. In this paper, we propose a novel method to transform categorical data to numerical representations, so that abundant numerical learning methods can be exploited in categorical data mining. Our key idea is to learn a pairwise dissimilarity among categorical symbols, henceforth a continuous embedding, which can then be used for subsequent numerical treatment. There are two important criteria for learning the dissimilarities. First, it should capture the important “transitivity” which has shown to be particularly useful in measuring the proximity relation in categorical data. Second, the pairwise sample geometry arising from the learned symbol distances should be maximally consistent with prior knowledge (e.g., class labels) to obtain a good generalization performance. We achieve them through multiple transitive distance learning and embedding. Encouraging results are observed on a number of benchmark classification tasks against state-of-the-art. Kai Zhang 0001, Qiaojun Wang, Zhengzhang Chen, Ivan Marsic, Vipin Kumar 0001, Guofei Jiang, Jie Zhang 0012 |
SDM | 1 |
| 2015 | Scaling Up Graph-Based Semisupervised Learning via Prototype Vector MachinesabstractWhen the amount of labeled data are limited, semisupervised learning can improve the learner's performance by also using the often easily available unlabeled data. In particular, a popular approach requires the learned function to be smooth on the underlying data manifold. By approximating this manifold as a weighted graph, such graph-based techniques can often achieve state-of-the-art performance. However, their high time and space complexities make them less attractive on large data sets. In this paper, we propose to scale up graph-based semisupervised learning using a set of sparse prototypes derived from the data. These prototypes serve as a small set of data representatives, which can be used to approximate the graph-based regularizer and to control model complexity. Consequently, both training and testing become much more efficient. Moreover, when the Gaussian kernel is used to define the graph affinity, a simple and principled method to select the prototypes can be obtained. Experiments on a number of real-world data sets demonstrate encouraging performance and scaling properties of the proposed approach. It also compares favorably with models learned via l1 -regularization at the same level of model sparsity. These results demonstrate the efficacy of the proposed approach in producing highly parsimonious and accurate models for semisupervised learning. Kai Zhang 0001, Liang Lan, James T. Kwok, Slobodan Vucetic, Bahram Parvin |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Improving Semi-Supervised Target Alignment via Label-Aware Base Kernels
Qiaojun Wang, Kai Zhang 0001, Guofei Jiang, Ivan Marsic |
AAAI | 2 |
| 2014 | Temporal skeletonization on sequential data: patterns, categorization, and visualizationabstractSequential pattern analysis targets on finding statistically relevant temporal structures where the values are delivered in a sequence. With the growing complexity of real-world dynamic scenarios, more and more symbols are often needed to encode a meaningful sequence. This is so-called 'curse of cardinality', which can impose significant challenges to the design of sequential analysis methods in terms of computational efficiency and practical use. Indeed, given the overwhelming scale and the heterogeneous nature of the sequential data, new visions and strategies are needed to face the challenges. To this end, in this paper, we propose a 'temporal skeletonization' approach to proactively reduce the representation of sequences to uncover significant, hidden temporal structures. The key idea is to summarize the temporal correlations in an undirected graph. Then, the 'skeleton' of the graph serves as a higher granularity on which hidden temporal patterns are more likely to be identified. In the meantime, the embedding topology of the graph allows us to translate the rich temporal content into a metric space. This opens up new possibilities to explore, quantify, and visualize sequential data. Our approach has shown to greatly alleviate the curse of cardinality in challenging tasks of sequential pattern mining and clustering. Evaluation on a Business-to-Business (B2B) marketing application demonstrates that our approach can effectively discover critical buying paths from noisy customer event data. Chuanren Liu, Kai Zhang 0001, Hui Xiong 0001, Geoff Jiang, Qiang Yang 0001 |
KDD | 2 |
| 2014 | Sparse semi-supervised learning on low-rank kernel
Kai Zhang 0001, Qiaojun Wang, Liang Lan, Yu Sun 0076, Ivan Marsic |
Neurocomputing | 1 |
| 2013 | Covariate Shift in Hilbert Space: A Solution via Sorrogate KernelsabstractCovariate shift is a unconventional learning scenario in which training and testing data have different distributions. A general principle to solve the problem is to make the training data distribution similar to the test one, such that classifiers computed on the former generalizes well to the latter. Current approaches typically target on the sample distribution in the input space, however, for kernel-based learning methods, the algorithm performance depends directly on the geometry of the kernel-induced feature space. Motivated by this, we propose to match data distributions in the Hilbert space, which, given a pre-defined empirical kernel map, can be formulated as aligning kernel matrices across domains. In particular, to evaluate similarity of kernel matrices defined on arbitrarily different samples, the novel concept of surrogate kernel is introduced based on the Mercer's theorem. Our approach caters the model adaptation specifically to kernel-based learning mechanism, and demonstrates promising results on several real-world applications. Kai Zhang 0001, Vincent Wenchen Zheng, Qiaojun Wang, James T. Kwok, Qiang Yang 0001, Ivan Marsic |
ICML (3) | 1 |
| 2012 | Improved Nystrom Low-rank Decomposition with Priors
Kai Zhang 0001, Liang Lan, Jun Liu 0003, Andreas Rauber |
ICML | 1 |
| 2010 | Sparse multitask regression for identifying common mechanism of response to therapeutic targetsabstractMOTIVATION: Molecular association of phenotypic responses is an important step in hypothesis generation and for initiating design of new experiments. Current practices for associating gene expression data with multidimensional phenotypic data are typically (i) performed one-to-one, i.e. each gene is examined independently with a phenotypic index and (ii) tested with one stress condition at a time, i.e. different perturbations are analyzed separately. As a result, the complex coordination among the genes responsible for a phenotypic profile is potentially lost. More importantly, univariate analysis can potentially hide new insights into common mechanism of response. RESULTS: In this article, we propose a sparse, multitask regression model together with co-clustering analysis to explore the intrinsic grouping in associating the gene expression with phenotypic signatures. The global structure of association is captured by learning an intrinsic template that is shared among experimental conditions, with local perturbations introduced to integrate effects of therapeutic agents. We demonstrate the performance of our approach on both synthetic and experimental data. Synthetic data reveal that the multi-task regression has a superior reduction in the regression error when compared with traditional L(1)-and L(2)-regularized regression. On the other hand, experiments with cell cycle inhibitors over a panel of 14 breast cancer cell lines demonstrate the relevance of the computed molecular predictors with the cell cycle machinery, as well as the identification of hidden variables that are not captured by the baseline regression analysis. Accordingly, the system has identified CLCA2 as a hidden transcript and as a common mechanism of response for two therapeutic agents of CI-1040 and Iressa, which are currently in clinical use. Kai Zhang 0001, Joe W. Gray, Bahram Parvin |
Bioinform. | 1 |
| 2010 | Fast and accurate kernel density approximation using a divide-and-conquer approachabstractDensity-based nonparametric clustering techniques, such as the mean shift algorithm, are well known for their flexibility and effectiveness in real-world vision-based problems. The underlying kernel density estimation process can be very expensive on large datasets. In this paper, the divide-and-conquer method is proposed to reduce these computational requirements. The dataset is first partitioned into a number of small, compact clusters. Components of the kernel estimator in each local cluster are then fit to a single, representative density function. The key novelty presented here is the efficient derivation of the representative density function using concepts from function approximation, such that the expensive kernel density estimator can be easily summarized by a highly compact model with very few basis functions. The proposed method has a time complexity that is only linear in the sample size and data dimensionality. Moreover, the bandwidth of the resultant density model is adaptive to local data distribution. Experiments on color image filtering/segmentation show that, the proposed method is dramatically faster than both the standard mean shift and fast mean shift implementations based on kd-trees while producing competitive image segmentation results. Yan-xia Jin, Kai Zhang 0001, James T. Kwok, Han-chang Zhou |
J. Zhejiang Univ. Sci. C | 2 |
| 2010 | Rhythmic Dynamics and Synchronization via Dimensionality Reduction: Application to Human GaitabstractReliable characterization of locomotor dynamics of human walking is vital to understanding the neuromuscular control of human locomotion and disease diagnosis. However, the inherent oscillation and ubiquity of noise in such non-strictly periodic signals pose great challenges to current methodologies. To this end, we exploit the state-of-the-art technology in pattern recognition and, specifically, dimensionality reduction techniques, and propose to reconstruct and characterize the dynamics accurately on the cycle scale of the signal. This is achieved by deriving a low-dimensional representation of the cycles through global optimization, which effectively preserves the topology of the cycles that are embedded in a high-dimensional Euclidian space. Our approach demonstrates a clear advantage in capturing the intrinsic dynamics and probing the subtle synchronization patterns from uni/bivariate oscillatory signals over traditional methods. Application to human gait data for healthy subjects and diabetics reveals a significant difference in the dynamics of ankle movements and ankle-knee coordination, but not in knee movements. These results indicate that the impaired sensory feedback from the feet due to diabetes does not influence the knee movement in general, and that normal human walking is not critically dependent on the feedback from the peripheral nervous system. Jie Zhang 0012, Kai Zhang 0001, Jianfeng Feng, Michael Small |
PLoS Comput. Biol. | 2 |
| 2010 | Simplifying mixture models through function approximationabstractThe finite mixture model is widely used in various statistical learning problems. However, the model obtained may contain a large number of components, making it inefficient in practical applications. In this paper, we propose to simplify the mixture model by minimizing an upper bound of the approximation error between the original and the simplified model, under the use of the L (2) distance measure. This is achieved by first grouping similar components together and then performing local fitting through function approximation. The simplified model obtained can then be used as a replacement of the original model to speed up various algorithms involving mixture models during training (e.g., Bayesian filtering, belief propagation) and testing [e.g., kernel density estimation, support vector machine (SVM) testing]. Encouraging results are observed in the experiments on density estimation, clustering-based image segmentation, and simplification of SVM decision functions. Kai Zhang 0001, James T. Kwok |
IEEE Trans. Neural Networks | 1 |
| 2010 | Clustered Nyström method for large scale manifold learning and dimension reductionabstractKernel (or similarity) matrix plays a key role in many machine learning algorithms such as kernel methods, manifold learning, and dimension reduction. However, the cost of storing and manipulating the complete kernel matrix makes it infeasible for large problems. The Nyström method is a popular sampling-based low-rank approximation scheme for reducing the computational burdens in handling large kernel matrices. In this paper, we analyze how the approximating quality of the Nyström method depends on the choice of landmark points, and in particular the encoding powers of the landmark points in summarizing the data. Our (non-probabilistic) error analysis justifies a "clustered Nyström method" that uses the k-means clustering centers as landmark points. Our algorithm can be applied to scale up a wide variety of algorithms that depend on the eigenvalue decomposition of kernel matrix (or its variant), such as kernel principal component analysis, Laplacian eigenmap, spectral clustering, as well as those involving kernel matrix inverse such as least-squares support vector machine and Gaussian process regression. Extensive experiments demonstrate the competitive performance of our algorithm in both accuracy and efficiency. Kai Zhang 0001, James T. Kwok |
IEEE Trans. Neural Networks | 1 |
| 2009 | Prototype vector machine for large scale semi-supervised learningabstractPractical data mining rarely falls exactly into the supervised learning scenario. Rather, the growing amount of unlabeled data poses a big challenge to large-scale semi-supervised learning (SSL). We note that the computational intensiveness of graph-based SSL arises largely from the manifold or graph regularization, which in turn lead to large models that are difficult to handle. To alleviate this, we proposed the prototype vector machine (PVM), a highly scalable, graph-based algorithm for large-scale SSL. Our key innovation is the use of "prototypes vectors" for efficient approximation on both the graph-based regularizer and model representation. The choice of prototypes are grounded upon two important criteria: they not only perform effective low-rank approximation of the kernel matrix, but also span a model suffering the minimum information loss compared with the complete model. We demonstrate encouraging performance and appealing scaling properties of the PVM on a number of machine learning benchmark data sets. Kai Zhang 0001, James T. Kwok, Bahram Parvin |
ICML | 1 |
| 2009 | Density-Weighted Nyström Method for Computing Large Kernel EigensystemsabstractThe Nyström method is a well-known sampling-based technique for approximating the eigensystem of large kernel matrices. However, the chosen samples in the Nyström method are all assumed to be of equal importance, which deviates from the integral equation that defines the kernel eigenfunctions. Motivated by this observation, we extend the Nyström method to a more general, density-weighted version. We show that by introducing the probability density function as a natural weighting scheme, the approximation of the eigensystem can be greatly improved. An efficient algorithm is proposed to enforce such weighting in practice, which has the same complexity as the original Nyström method and hence is notably cheaper than several other alternatives. Experiments on kernel principal component analysis, spectral clustering, and image segmentation demonstrate the encouraging performance of our algorithm. Kai Zhang 0001, James T. Kwok |
Neural Comput. | 1 |
| 2009 | Maximum Margin Clustering Made PracticalabstractMotivated by the success of large margin methods in supervised learning, maximum margin clustering (MMC) is a recent approach that aims at extending large margin methods to unsupervised learning. However, its optimization problem is nonconvex and existing MMC methods all rely on reformulating and relaxing the nonconvex optimization problem as semidefinite programs (SDP). Though SDP is convex and standard solvers are available, they are computationally very expensive and only small data sets can be handled. To make MMC more practical, we avoid SDP relaxations and propose in this paper an efficient approach that performs alternating optimization directly on the original nonconvex problem. A key step to avoid premature convergence in the resultant iterative procedure is to change the loss function from the hinge loss to the Laplacian/square loss so that overconfident predictions are penalized. Experiments on a number of synthetic and real-world data sets demonstrate that the proposed approach is more accurate, much faster (hundreds to tens of thousands of times faster), and can handle data sets that are hundreds of times larger than the largest data set reported in the MMC literature. Kai Zhang 0001, Ivor W. Tsang, James T. Kwok |
IEEE Trans. Neural Networks | 1 |
| 2008 | Improved Nyström low-rank approximation and error analysisabstractLow-rank matrix approximation is an effective tool in alleviating the memory and computational burdens of kernel methods and sampling, as the mainstream of such algorithms, has drawn considerable attention in both theory and practice. This paper presents detailed studies on the Nyström sampling scheme and in particular, an error analysis that directly relates the Nyström approximation quality with the encoding powers of the landmark points in summarizing the data. The resultant error bound suggests a simple and efficient sampling scheme, the k-means clustering algorithm, for Nyström low-rank approximation. We compare it with state-of-the-art approaches that range from greedy schemes to probabilistic sampling. Our algorithm achieves significant performance gains in a number of supervised/unsupervised learning tasks including kernel PCA and least squares SVM. Kai Zhang 0001, Ivor W. Tsang, James T. Kwok |
ICML | 1 |
| 2007 | Maximum margin clustering made practicalabstractMaximum margin clustering (MMC) is a recent large margin unsupervised learning approach that has often outperformed conventional clustering methods. Computationally, it involves non-convex optimization and has to be relaxed to different semidefinite programs (SDP). However, SDP solvers are computationally very expensive and only small data sets can be handled by MMC so far. To make MMC more practical, we avoid SDP relaxations and propose in this paper an efficient approach that performs alternating optimization directly on the original non-convex problem. A key step to avoid premature convergence is on the use of SVR with the Laplacian loss, instead of SVM with the hinge loss, in the inner optimization subproblem. Experiments on a number of synthetic and real-world data sets demonstrate that the proposed approach is often more accurate, much faster and can handle much larger data sets. Kai Zhang 0001, Ivor W. Tsang, James T. Kwok |
ICML | 1 |
| 2006 | Accelerated Convergence Using Dynamic Mean Shift
Kai Zhang 0001, James T. Kwok, Ming Tang 0001 |
ECCV (2) | 1 |
| 2006 | Fast Speaker Adaption Via Maximum Penalized Likelihood Kernel RegressionabstractMaximum likelihood linear regression (MLLR) has been a popular speaker adaptation method for many years. In this paper, we investigate a generalization of MLLR using nonlinear regression. Specifically, kernel regression is applied with appropriate regularization to determine the transformation matrix in MLLR for fast speaker adaptation. The proposed method, called maximum penalized likelihood kernel regression adaptation (MPLKR), is computationally simple and the mean vectors of the speaker adapted acoustic model can be obtained analytically by simply solving a linear system. Since no nonlinear optimization is involved, the obtained solution is always guaranteed to be globally optimal. The new adaptation method was evaluated on the resource management task with 5s and 10s of adaptation speech. Results show that MPLKR outperforms the standard MLLR method Ivor W. Tsang, James T. Kwok, Brian Kan-Wing Mak, Kai Zhang 0001, Jeffrey Junfeng Pan |
ICASSP (1) | 4 |
| 2006 | Block-quantized kernel matrix for fast spectral embeddingabstractEigendecomposition of kernel matrix is an indispensable procedure in many learning and vision tasks. However, the cubic complexity O(N3) is impractical for large problem, where N is the data size. In this paper, we propose an efficient approach to solve the eigendecomposition of the kernel matrix W. The idea is to approximate W with W that is composed of m2 constant blocks. The eigenvectors of W, which can be solved in O(m3) time, is then used to recover the eigenvectors of the original kernel matrix. The complexity of our method is only O(mN + m3), which scales more favorably than state-of-the-art low rank approximation and sampling based approaches (O(m2N + m3)), and the approximation quality can be controlled conveniently. Our method demonstrates encouraging scaling behaviors in experiments of image segmentation (by spectral clustering) and kernel principal component analysis. Kai Zhang 0001, James T. Kwok |
ICML | 1 |
| 2006 | Simplifying Mixture Models through Function ApproximationabstractFinite mixture model is a powerful tool in many statistical learning problems. In this paper, we propose a general, structure-preserving approach to reduce its model complexity, which can bring significant computational benefits in many applications. The basic idea is to group the original mixture components into compact clusters, and then minimize an upper bound on the approximation error between the original and simplified models. By adopting the L2 norm as the dis- tance measure between mixture models, we can derive closed-form solutions that are more robust and reliable than using the KL-based distance measure. Moreover, the complexity of our algorithm is only linear in the sample size and dimensional- ity. Experiments on density estimation and clustering-based image segmentation demonstrate its outstanding performance in terms of both speed and accuracy. Kai Zhang 0001, James T. Kwok |
NIPS | 1 |
| 2005 | Applying Neighborhood Consistency for Fast Clustering and Kernel Density EstimationabstractNearest neighborhood consistency is an important concept in statistical pattern recognition, which underlies the well-known k-nearest neighbor method. In this paper, we combine this idea with kernel density estimation based clustering, and derive the fast mean shift algorithm (FMS). FMS greatly reduces the complexity of feature space analysis, resulting satisfactory precision of classification. More importantly, we show that with FMS algorithm, we are in fact relying on a conceptually novel approach of density estimation, the fast kernel density estimation (FKDE) for clustering. The FKDE combines smooth and non-smooth estimators and thus inherits advantages from both. Asymptotic analysis reveals the approximation of the FKDE to standard kernel density estimator. Data clustering and image segmentation experiments demonstrate the efficiency of FMS. Kai Zhang 0001, Ming Tang 0001, James T. Kwok |
CVPR (2) | 1 |