VLDB 2026 Research / reviewers in the wild / expert
Zhe Quan
dblp:37/8296
· DBLP profile ↗
38ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-2669-9190ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9Systems, architecture and hardware · 8 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural Causal Graph for Interpretable and Intervenable ClassificationabstractAdvancements in neural networks have significantly enhanced the performance of classification models, achieving remarkable accuracy across diverse datasets. However, these models often lack transparency and do not support interactive reasoning with human users, which are essential attributes for applications that require trust and user engagement. To overcome these limitations, we introduce an innovative framework, Neural Causal Graph (NCG), that integrates causal inference with neural networks to enable interpretable and intervenable reasoning. We then propose an intervention training method to model the intervention probability of the prediction, serving as a contextual prompt to facilitate the fine-grained reasoning and human-AI interaction abilities of NCG. Our experiments show that the proposed framework significantly enhances the performance of traditional classification baselines. Furthermore, NCG achieves nearly 95\% top-1 accuracy on the ImageNet dataset by employing a test-time intervention method. This framework not only supports sophisticated post-hoc interpretation but also enables dynamic human-AI interactions, significantly improving the model's transparency and applicability in real-world scenarios. Jiawei Wang 0025, Shaofei Lu, Da Cao, Yuquan Le, Zhe Quan, Tat-Seng Chua |
ICLR | 6 |
| 2025 | NtNDet: Hardware Trojan detection based on pre-trained language models
Shijie Kuang, Zhe Quan, Guoqi Xie, Xiaoqian Chen, Keqin Li 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Graph Reasoning With Supervised Contrastive Learning for Legal Judgment PredictionabstractGiven the fact descriptions of legal cases, the legal judgment prediction (LJP) problem aims to determine three judgment tasks of law articles, charges, and the term of penalty. Most existing studies have considered task dependencies while neglecting the prior dependencies of labels among different tasks. Therefore, how to make better use of the information on the relation dependencies among tasks and labels becomes a crucial issue. To this end, we transform the text classification problem into a node classification framework based on graph reasoning and supervised contrastive learning (SCL) techniques, named GraSCL. Specifically, we first design a graph reasoning network to model the potential dependency structures and facilitate relational learning under various graph topologies. Then, we introduce the SCL method for the LJP task to further leverage the label relation on the graph. To accommodate the node classification settings, we extend the traditional SCL method to novel variants for SCL at the node level, which allows the GraSCL framework to be trained efficiently even with small batches. Furthermore, to recognize the importance of hard negative samples in contrastive learning, we introduce a simple yet effective technique called online hard negative mining (OHNM) to enhance our SCL approach. This technique complements our SCL method and enables us to control the number and complexity of negative samples, leading to further improvements in the model's performance. Finally, extensive experiments are conducted on two well-known benchmarks, demonstrating the effectiveness and rationality of our proposed SCL approach as compared to the state-of-the-art competitors. Jiawei Wang 0025, Yuquan Le, Da Cao, Shaofei Lu, Zhe Quan, Meng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | $\boldsymbol{R}^{2}$: A Novel Recall & Ranking Framework for Legal Judgment PredictionabstractThe legal judgment prediction (LJP) task is to automatically decide appropriate law articles, charges, and term of penalty for giving the fact description of a law case. It considerably influences many real legal applications and has thus attracted the attention of legal practitioners and AI researchers in recent years. In real scenarios, many confusing charges are encountered, which makes LJP challenging. Intuitively, for a controversial legal case, legal practitioners usually first obtain various possible judgment results as candidates based on the fact description of the case; then these candidates generally need to be carefully considered based on the facts and the rationality of the candidates. Inspired by this observation, this paper presents a novelRecall &Ranking framework, dubbed as$\boldsymbol{R}^{2}$, which attempts to formalize LJP as a two-stage problem. The recall stage is designed to collect high-likelihood judgment results for a given case; these results are regarded as candidates for the ranking stage. The ranking stage introduces a verification technique to learn the relationships between the fact description and the candidates. It treats the partially correct candidates as semi-negative samples, and thus has a certain ability to distinguish confusing candidates. Moreover, we devise a comprehensive judgment strategy to refine the final judgment results by comprehensively considering the rationality of multiple probable candidates. We carry out numerous experiments on two widely used benchmark datasets. The experimental results demonstrate our proposed approach's effectiveness compared to the other competitive baselines. Yuquan Le, Zhe Quan, Jiawei Wang 0025, Da Cao, Kenli Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | XHYPRE: a reliable parallel numerical algorithm library for solving large-scale sparse linear equations
Chuanying Li, Stef Graillat, Zhe Quan, Tongxiang Gu, Hao Jiang 0001, Kenli Li 0001 |
CCF Trans. High Perform. Comput. | 3 |
| 2023 | Effectively Identifying Compound-Protein Interaction Using Graph Neural RepresentationabstractEffectively identifying compound-protein interactions (CPIs) is crucial for new drug design, which is an important step in silico drug discovery. Current machine learning methods for CPI prediction mainly use one-demensional (1D) compound/protein strings and/or the specific descriptors. However, they often ignore the fact that molecules are essentially modeled by the molecular graph. We observe that in real-world scenarios, the topological structure information of the molecular graph usually provides an overview of how the atoms are connected, and the local chemical context reveals the functionality of the protein sequence in CPI. These two types of information are complementary to each other and they are both significant for modeling compound-protein pairs. Motivated by this, we propose an end-to-end deep learning framework named GraphCPI, which captures the structural information of compounds and leverages the chemical context of protein sequences for solving the CPI prediction task. Our framework can integrate any popular graph neural networks for learning compounds, and it combines with a convolutional neural network for embedding sequences. To compare our method with classic and state-of-the-art deep learning methods, we conduct extensive experiments based on several widely-used CPI datasets. The experimental results show the feasibility and competitiveness of our proposed method. Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Xiangxiang Zeng, Philip S. Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Cover Trees Revisited: Exploiting Unused Distance and Direction InformationabstractThe cover tree (CT) and its improved version are hierarchical data structures that simplified navigating nets while maintaining good runtime guarantees. They can perform nearest neighbor search in logarithmic time and provide efficient computation in practice. In this article, we revisit cover trees for nearest neighbor search, and propose a more competitive method. The central idea of our method is to fully exploit the unused distance and direction information. More specially, our method introduces three novel concepts/techniques: (i) range list, (ii) quadrant information, and (iii) vectorial angle cosine. These techniques are seamlessly integrated into our suggested data structure and search algorithms. As an extra bonus, we explore approximate nearest neighbor and$k$nearest neighbor based on the proposed techniques, and present algorithms for handling updates. Extensive experimental results, based on both real and synthetic datasets, consistently demonstrate that our method is attractive and competitive, compared against existing cover tree structures for nearest neighbor search and its variants. Zhi-Jie Wang 0009, Mengdie Nie, Kaiqi Zhao 0001, Zhe Quan, Bin Yao 0002 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | LPV: A Log Parsing Framework Based on VectorizationabstractLogs are pervasive in modern computing systems, and are valuable to service and system management. Nevertheless, with the rapidly growing size and complexity of computing systems, the log volume is exploding, which makes automatic log analysis imperative. Generally, in automatic log analysis, the first and fundamental step is log parsing, to which a lot of effort has been devoted. However, in most existing log parsing methods, log messages are merely treated as plain text. In natural language processing (NLP) area, it is a common practice to represent words and sentences with vectors, then the similarity between two words or sentences can be measured by the distance between their vectors. Inspired by these, we put forward a novel log parsing framework, named LPV (LogParser based onVectorization), which performs log parsing by converting log messages and log templates into vectors, with the help of a vectorization method in NLP. LPV incorporates offline and online log parsing. In the offline log parsing, the central idea is to first represent log messages with vectors, so that the similarity between two log messages can be measured by the distance between their vectors, then we cluster log messages via clustering the vectors, and finally we extract log templates from the resultant clusters. By the end of the offline log parsing, each log template is assigned with an average vector, so that in the online log parsing, the similarity between an incoming log message and each log template can also be measured by the distance between their vectors. Extensive experiments have been conducted based on several public log datasets to evaluate LPV with three different vectorization methods. The results demonstrate that, with a proper vectorization method, LPV performs competitive with state-of-the-art log parsing methods, in both effectiveness and efficiency. Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Kaiqi Zhao 0001, Xiangke Liao, Yunfei Du 0001, Kenli Li 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | Loader: A Log Anomaly Detector Based on TransformerabstractDetecting anomalies in logs is crucial for service and system management, since logs are widely used to record the runtime status, and are often the only data available for postmortem analysis. Since anomalies are usually rare in real-world services and systems, a common and feasible practice is to mine or learn normal patterns from logs, and deem those violating the normal patterns as anomalies. As log sequences are a kind of time series data, RNN (Recurrent Neural Network) and its variants have been extensively employed to capture the normal patterns. Nevertheless, the sequential nature of RNN and its variants makes them hard to parallelize and capture long-term dependencies, which may hinder their performance. To address this issue, in this paper we propose Loader, a novel semi-supervisedloganomalydetector based on Transformer, because the Transformer architecture eschews recurrence and is able to draw global dependencies. Loader leverages the Transformer encoder to capture normal patterns from normal log sequences. When detecting, it gives a set of candidate log templates, that may appear after the input log substring under normal conditions. If the template of the actual next log message is not within the candidate set, this implies an anomaly. Previous similar methods select the most possible$k$log templates as candidates in any case, so the performance is sensitive to$k$, and it is nontrivial to pick a proper$k$. To alleviate this, we design a more flexible and robust ‘top-$p$’ algorithm, which determines the candidate set based on the cumulative probability of the most possible log templates. Extensive experiments are conducted based on three public log datasets, the experimental results validate the effectiveness and competitiveness of our approach. Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Yuquan Le, Yunfei Du 0001, Xiangke Liao, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Unified program cross-architecture migration framework modelabstractThe program cross-architecture migration technology can solve the problem that the software resources of the ARM architecture are not rich enough, so that the processors of the ARM architecture can obtain the abundant software resources of the x86 architecture. Since there are different migration schemes for different source programs, program migration technology involves many fields, and there is no unified method. This paper proposes a unified program migration framework model, and designs two technical modules of source code program migration and no-source code program migration according to whether the program has source code or not. The source code program migration module provides migration guidance for developers, reducing the difficulty of migration and improving the success rate of code migration; the no-source code migration module is based on binary translation technology, and integrates excellent open source tools QEMU and BOX64. Finally, experimental verification is carried out. The correctness and validity of the framework model. Minhao Zhou, Zhe Quan |
APSEC | 2 |
| 2022 | Legal Charge Prediction via Bilinear Attention NetworkabstractThe legal charge prediction task aims to judge appropriate charges according to the given fact description in cases. Most existing methods formulate it as a multi-class text classification problem and have achieved tremendous progress. However, the performance on low-frequency charges is still unsatisfactory. Previous studies indicate leveraging the charge label information can facilitate this task, but the approaches to utilizing the label information are not fully explored. In this paper, inspired by the vision-language information fusion techniques in the multi-modal field, we propose a novel model (denoted as LeapBank) by fusing the representations of text and labels to enhance the legal charge prediction task. Specifically, we devise a representation fusion block based on the bilinear attention network to interact the labels and text tokens seamlessly. Extensive experiments are conducted on three real-world datasets to compare our proposed method with state-of-the-art models. Experimental results show that LeapBank obtains up to 8.5% Macro-F1 improvements on the low-frequency charges, demonstrating our model's superiority and competitiveness. Yuquan Le, Meng Chen 0006, Zhe Quan, Xiaodong He 0001, Kenli Li 0001 |
CIKM | 4 |
| 2022 | RIRCNN: A Fault Diagnosis Method for Aviation Turboprop EngineabstractAero-engine is the 'heart' of the aviation aircraft.Practical failure prediction of aero-engines is difficult due to the performance degradation covered by the continuous switching between various operating conditions.In order to solve the above problem, we propose a new type of aero-engine fault diagnosis model-RIRCNN (Residual Independently Reccurent and Convolutional Neural Network).It can process long sequences, and has superior feature extraction effect.We gather flight data sets through ground bench experiment of the aviation turboprop engine, and intensively conduct comparative experiments to evaluate the effectiveness of our model.The verification results demonstrate that our model can achieve excellent performance compared with other available baseline models. Zhe Quan, Tong Xiao 0002, Xiaofei Jiang, Xinjian Hu, Peibing Du |
SEKE | 2 |
| 2020 | A Novel Model for Imbalanced Data ClassificationabstractRecently, imbalanced data classification has received much attention due to its wide applications. In the literature, existing researches have attempted to improve the classification performance by considering various factors such as the imbalanced distribution, cost-sensitive learning, data space improvement, and ensemble learning. Nevertheless, most of the existing methods focus on only part of these main aspects/factors. In this work, we propose a novel imbalanced data classification model that considers all these main aspects. To evaluate the performance of our proposed model, we have conducted experiments based on 14 public datasets. The results show that our model outperforms the state-of-the-art methods in terms of recall, G-mean, F-measure and AUC. Jian Yin 0001, Chunjing Gan, Kaiqi Zhao 0001, Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009 |
AAAI | 5 |
| 2020 | DeepGS: Deep Representation Learning of Graphs and Sequences for Drug-Target Binding Affinity PredictionabstractAccurately predicting drug-target binding affinity (DTA) in silico is a key task in drug discovery. Most of the conventional DTA prediction methods are simulation-based, which rely heavily on domain knowledge or the assumption of having the 3D structure of the targets, which are often difficult to obtain. Meanwhile, traditional machine learning-based methods apply various features and descriptors, and simply depend on the similarities between drug-target pairs. Recently, with the increasing amount of affinity data available and the success of deep representation learning models on various domains, deep learning techniques have been applied to DTA prediction. However, these methods consider either label/one-hot encodings or the topological structure of molecules, without considering the local chemical context of amino acids and SMILES sequences. Motivated by this, we propose a novel end-to-end learning framework, called DeepGS, which uses deep neural networks to extract the local chemical context from amino acids and SMILES sequences, as well as the molecular structure from the drugs. To assist the operations on the symbolic data, we propose to use advanced embedding techniques (i.e., Smi2Vec and Prot2Vec) to encode the amino acids and SMILES sequences to a distributed representation. Meanwhile, we suggest a new molecular structure modeling approach that works well under our framework. We have conducted extensive experiments to compare our proposed method with state-of-the-art models including KronRLS, SimBoost, DeepDTA and DeepCPI. Extensive experimental results demonstrate the superiorities and competitiveness of DeepGS. Xuan Lin, Kaiqi Zhao 0001, Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Philip S. Yu |
ECAI | 4 |
| 2020 | LPV: A Log Parser Based on Vectorization for Offline and Online Log ParsingabstractAs the first and foremost step of typical automatic log analysis, log parsing has attracted a lot of interest. Most of existing studies treat log messages as pure strings and rely on string matching or string distance. In NLP, word2vec has shown very efficient and effective in representing words with low dimensional vectors. Inspired by this, in this paper we propose a novel method, called LPV (Log Parser based on Vectorization), for both offline and online log parsing. The central idea of our method in offline log parsing is to first convert log messages into vectors, and measure the similarity between two log messages by the distance between two vectors, then log messages can be clustered via clustering the vectors, and log templates can be extracted from the resulting clusters. For online log parsing, we also assign log templates with some kind of average vectors, so that the similarity between an incoming log message and each log template can also be measured by the distance between two vectors. We have conducted extensive experiments based on three widely used log datasets, and the results demonstrate that our proposed method LPV can achieve a competitive performance, compared against state-of-the-art log parsing methods. Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Kaiqi Zhao 0001, Xiangke Liao |
ICDM | 2 |
| 2020 | Text to Image Synthesis With Bidirectional Generative Adversarial NetworkabstractGenerating realistic images from text descriptions is a challenging problem in computer vision. Although previous works have shown remarkable progress, guaranteeing semantic consistency between text descriptions and images remains challenging. To generate semantically consistent images, we propose two semantics-enhanced modules and a novel Textual-Visual Bidirectional Generative Adversarial Network (TVBi-GAN). Specifically, this paper proposes a semanticsenhanced attention module and a semantics-enhanced batch normalization module. These modules improve consistency of synthesized images by involving precisely semantic features. What's more, an encoder network is proposed to extract semantic features from images. During the adversarial process, the encoder could guide our generator to explore corresponding features behind descriptions. With extensive experiments on CUB and COCO datasets, we demonstrate that our TVBi-GAN outperforms state-of-the-art methods. Zhe Quan, Zhi-Jie Wang 0009, Xinjian Hu, Yangyang Chen 0006 |
ICME | 2 |
| 2020 | KGNN: Knowledge Graph Neural Network for Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) prediction is a challenging problem in pharmacology and clinical application, and effectively identifying potential DDIs during clinical trials is critical for patients and society. Most of existing computational models with AI techniques often concentrate on integrating multiple data sources and combining popular embedding methods together. Yet, researchers pay less attention to the potential correlations between drug and other entities such as targets and genes. Moreover, recent studies also adopted knowledge graph (KG) for DDI prediction. Yet, this line of methods learn node latent embedding directly, but they are limited in obtaining the rich neighborhood information of each entity in the KG. To address the above limitations, we propose an end-to-end framework, called Knowledge Graph Neural Network (KGNN), to resolve the DDI prediction. Our framework can effectively capture drug and its potential neighborhoods by mining their associated relations in KG. To extract both high-order structures and semantic relations of the KG, we learn from the neighborhoods for each entity in the KG as their local receptive, and then integrate neighborhood information with bias from representation of the current entity. This way, the receptive field can be naturally extended to multiple hops away to model high-order topological information and to obtain drugs potential long-distance correlations. We have implemented our method and conducted experiments based on several widely-used datasets. Empirical results show that KGNN outperforms the classic and state-of-the-art models. Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Tengfei Ma 0002, Xiangxiang Zeng |
IJCAI | 2 |
| 2020 | A novel molecular representation with BiGRU neural networks for learning atomabstractMolecular representations play critical roles in researching drug design and properties, and effective methods are beneficial to assisting in the calculation of molecules and solving related problem in drug discovery. In previous years, most of the traditional molecular representations are based on hand-crafted features and rely heavily on biological experimentations, which are often costly and time consuming. However, recent researches achieve promising results using machine learning on various domains. In this article, we present a novel method named Smi2Vec-BiGRU that is designed for learning atoms and solving the single- and multitask binary classification problems in the field of drug discovery, which are the basic and also key problems in this field. Specifically, our approach transforms the molecule data in the SMILES format into a set of sample vectors and then feeds them into the bidirectional gated recurrent unit neural networks for training, which learns low-dimensional vector representations for molecular drug. We conduct extensive experiments on several widely used benchmarks including Tox21, SIDER and ClinTox. The experimental results show that our approach can achieve state-of-the-art performance on these benchmarking datasets, demonstrating the feasibility and competitiveness of our proposed approach. Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Xiangxiang Zeng |
Briefings Bioinform. | 2 |
| 2020 | Task Scheduling for Energy Consumption Constrained Parallel Applications on Heterogeneous Computing SystemsabstractPower-aware task scheduling on processors has been a research hotspot in computing systems. Given an application G containing a set N of tasks {n1,...,n|N|}, and a system containing a set U of processors {u1,..., u|U|}, the power-aware task scheduling generally refers to finding the appropriate processor and frequency for each task ni, so as to make sure that all the tasks can be finished efficiently and the overall energy consumption is guaranteed. In this article, we study the problem of minimizing the schedule length for energy consumption constrained parallel applications on heterogeneous computing systems, where the schedule length refers to the time interval between starting the first task and finishing the last task. For this problem, existing work adopts a policy that preassigns the minimum energy consumption for each unassigned task. Nevertheless, our analysis reveals that, such a pre-assignment policy could be unfair for the low priority tasks, and it may not achieve an optimistic schedule length. Thereby, we propose a new task scheduling algorithm that suggests a weight-based mechanism to preassign energy consumption for unassigned tasks, and we provide the rigorous proof to show its feasibility. Further, we show that this idea can be extended to solve reliability maximization problems with energy consumption constraint or with both deadline and energy consumption constraints, where the reliability refers to the probability of executing application G without failures, and the deadline constraint refers to the “allowable” maximum schedule length. We have conducted extensive experiments based on real parallel applications. The experimental results consistently demonstrate that our proposed algorithms can achieve favourable performance, compared to state-of-the-art algorithms. Zhe Quan, Zhi-Jie Wang 0009, Ting Ye, Song Guo 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | An Improved Hierarchical Datastructure for Nearest Neighbor Search
Mengdie Nie, Zhi-Jie Wang 0009, Chunjing Gan, Zhe Quan, Bin Yao 0002, Jian Yin 0001 |
AAAI | 4 |
| 2019 | GraphCPI: Graph Neural Representation Learning for Compound-Protein InteractionabstractAccurately predicting compound-protein interactions (CPIs) is of great help to increase the efficiency and reduce costs in drug development. Most of existing machine learning models for CPI prediction often represent compounds and proteins in one-dimensional strings, or use the descriptor-based methods. These models might ignore the fact that molecules are essentially structured by the chemical bond of atoms. However, in real-world scenarios, the topological structure information usually provides an overview of how the atoms are connected, and the local chemical context reveals the functionality of the protein sequence in CPI. These two types of information are complementary to each other and they are both important for modeling compounds and proteins. Motivated by this, this paper suggests an end-to-end deep learning framework called GraphCPI, which captures the structural information of compounds and leverages the chemical context of protein sequences for solving the CPI prediction task. Our framework can integrate any popular graph nerual networks for learning compounds, and it combines with a convolutional neural network for embedding sequences. We conduct extensive experiments based on two benchmark CPI datasets. The experimental results demonstrate that our proposed framework is feasible and also competitive, comparing against classic and state-of-the-art methods. Zhe Quan, Xuan Lin, Zhi-Jie Wang 0009, Xiangxiang Zeng |
BIBM | 1 |
| 2019 | Iteratively Reweighted Penalty Alternating Minimization Methods with Continuation for Image DeblurringabstractIn this paper, we consider a class of nonconvex problems with linear constraints appearing frequently in the area of image processing. We solve this problem by the penalty method and propose the iteratively reweighted alternating minimization algorithm. To speed up the algorithm, we also apply the continuation strategy to the penalty parameter. A convergence result is proved for the algorithm. Compared with the nonconvex ADMM, the proposed algorithm enjoys both theoretical and computational advantages like weaker convergence requirements and faster speed. Numerical results demonstrate the efficiency of the proposed algorithm. Tao Sun 0005, Dongsheng Li 0001, Hao Jiang 0001, Zhe Quan |
ICASSP | 4 |
| 2019 | Heavy-ball Algorithms Always Escape Saddle PointsabstractNonconvex optimization algorithms with random initialization have attracted increasing attention recently. It has been showed that many first-order methods always avoid saddle points with random starting points. In this paper, we answer a question: can the nonconvex heavy-ball algorithms with random initialization avoid saddle points? The answer is yes! Direct using the existing proof technique for the heavy-ball algorithms is hard due to that each iteration of the heavy-ball algorithm consists of current and last points. It is impossible to formulate the algorithms as iteration like xk+1= g(xk) under some mapping g. To this end, we design a new mapping on a new space. With some transfers, the heavy-ball algorithm can be interpreted as iterations after this mapping. Theoretically, we prove that heavy-ball gradient descent enjoys larger stepsize than the gradient descent to escape saddle points to escape the saddle point. And the heavy-ball proximal point algorithm is also considered; we also proved that the algorithm can always escape the saddle point. Tao Sun 0005, Dongsheng Li 0001, Zhe Quan, Hao Jiang 0001, Shengguo Li, Yong Dou |
IJCAI | 3 |
| 2019 | Implementation and optimization of a data protecting model on the Sunway TaihuLight supercomputer with heterogeneous many-core processorsabstractSummary With the rapid development of information technology, the security of massive amounts of digital data has attracted huge attention in recent years. The Advanced Encryption Standard (AES) algorithm and the Security Hash Algorithm 3 (SHA3) are extensively used as cryptographic algorithms for protecting the security of information. The Sunway TaihuLight, with massive heterogeneous many‐core SW26010 processors, has the peak performance of over 100 PFlops. To achieve high efficiency of data encryption/decryption and guarantee the data integrity for large‐scale applications, this paper proposes a fast and secure data protecting model using the parallel AES algorithm and the SHA3 on the Sunway TaihuLight. According to the particular computing architecture and memory hierarchy of the Sunway TaihuLight, we propose a fine‐grained software design for the data protecting model to fully exploit the parallelism and properly arranges the data on the Sunway. Furthermore, we propose optimization strategies for the parallel AES algorithm of the data protecting model. It is proved that our data protecting model has high security, good scalability, and excellent encryption/decryption efficiency on the Sunway. Our data protecting model achieves a high throughput of 269.95 Gbits/s, and the optimized parallel AES algorithm achieves 511.28 Gbits/s on the Sunway. Yuedan Chen, Kenli Li 0001, Xiongwei Fei, Zhe Quan, Keqin Li 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | An efficient real-time data collection framework on petascale systems
Li-Qian Zhou, Yutong Lu, Tong Xiao 0002, Can Leng, Chuanying Li, Zhe Quan |
Neurocomputing | 7 |
| 2019 | An Efficient Framework for Sentence Similarity ModelingabstractSentence similarity modeling lies at the core of many natural language processing applications, and thus has received much attention. Owing to the success of word embeddings, recently, popular neural network methods achieved sentence embedding. Most of them focused on learning semantic information and modeling it as a continuous vector, yet the syntactic information of sentences has not been fully exploited. On the other hand, prior works have shown the benefits of structured trees that include syntactic information, while few methods in this branch utilized the advantages of word embeddings and another powerful technique-attention weight mechanism. This paper suggests to absorb their advantages by merging these techniques in a unified structure, dubbed as attention constituency vector tree (ACVT). Meanwhile, this paper develops a new tree kernel, known as ACVT kernel, which is tailored for sentence similarity measure based on the proposed structure. The experimental results, based on 19 widely used semantic textual similarity datasets, demonstrate that our model is effective and competitive, when compared against state-of-the-art models. Additionally, the experimental results validate that many attention weight mechanisms and word embedding techniques can be seamlessly integrated into our model, demonstrating the robustness and universality of our model. Zhe Quan, Zhi-Jie Wang 0009, Yuquan Le, Bin Yao 0002, Kenli Li 0001, Jian Yin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | GMOD: a dynamic GPU memory overflow detectorabstractRich thread-level parallelism in GPU has motivated co-running GPU kernels on a single GPU. However, when GPU kernels co-run, it is possible that a kernel can leverage buffer overflow to attack another kernel running on the same GPU. There is very limited work aiming to detect buffer overflow for GPU. The existing work has either large performance overhead or limited capability to detect buffer overflow. Bang Di, Jianhua Sun 0002, Dong Li 0001, Hao Chen 0002, Zhe Quan |
PACT | 5 |
| 2018 | A System for Learning Atoms Based on Long Short-Term Memory Recurrent Neural Networks
Zhe Quan, Xuan Lin, Zhi-Jie Wang 0009, Yan Liu 0032, Kenli Li 0001 |
BIBM | 1 |
| 2018 | ISAECC: An Improved Scheduling Approach for Energy Consumption Constrained Parallel Applications on Heterogeneous Distributed SystemsabstractPower-aware task scheduling on processors has been a hot topic. In this paper, we study the problem of minimizing the schedule length for energy consumption constrained parallel applications on heterogeneous distributed systems. Previous work (solving this problem) adopts a policy that preassigns the minimum energy consumption for each unassigned task. Nevertheless, our analysis reveals that such a preassignment policy could be unfair, and it may not achieve an optimistic schedule length. Motivated by this, we propose a new task scheduling algorithm that suggests a weight-based mechanism to preassign energy consumption for unassigned tasks. We theoretically prove that our preassignment mechanism can guarantee the energy consumption constraint. Also, we have conducted extensive experiments based on two real parallel applications. The results consistently demonstrate that, compared to state-of-the-art algorithms, our approach can achieve smaller schedule length while satisfying the energy consumption constraint. Ting Ye, Zhi-Jie Wang 0009, Zhe Quan, Song Guo 0001, Kenli Li 0001, Keqin Li 0001 |
ICPADS | 3 |
| 2018 | ACV-tree: A New Method for Sentence Similarity ModelingabstractSentence similarity modeling lies at the core of many natural language processing applications, and thus has received much attention. Owing to the success of word embeddings, recently, popular neural network methods have achieved sentence embedding, obtaining attractive performance. Nevertheless, most of them focused on learning semantic information and modeling it as a continuous vector, while the syntactic information of sentences has not been fully exploited. On the other hand, prior works have shown the benefits of structured trees that include syntactic information, while few methods in this branch utilized the advantages of word embeddings and another powerful technique ? attention weight mechanism. This paper makes the first attempt to absorb their advantages by merging these techniques in a unified structure, dubbed as ACV-tree. Meanwhile, this paper develops a new tree kernel, known as ACVT kernel, that is tailored for sentence similarity measure based on the proposed structure. The experimental results, based on 19 widely-used datasets, demonstrate that our model is effective and competitive, compared against state-of-the-art models. Yuquan Le, Zhi-Jie Wang 0009, Zhe Quan, Jiawei He 0003, Bin Yao 0002 |
IJCAI | 3 |
| 2018 | Implementing molecular dynamics simulation on the Sunway TaihuLight system with heterogeneous many-core processorsabstractSummary In various research of atom and molecule physical movements, molecular dynamics (MD) simulation is a common tool to simulate and investigate the real molecular motion. However, it introduces a significant penalty in performance, power, electricity, and running time. Consequently, once the simulation size scales up and computing demands keep growing, it comes at substantial costs in performance and energy usage. In this paper, an optimized MD implementation on the Sunway TaihuLight supercomputer with heterogeneous many‐core processors is developed to address the abovementioned issues. The Sunway TaihuLight is a totally independently designed and developed Chinese supercomputer with a profusely customized integration approach and a brand new many‐core processor, the SW26010. The computing power mainly supported by the homegrown many‐core SW26010 processors differs from other existing heterogeneous supercomputers. The Sunway TaihuLight is a heterogeneous supercomputer that ranks first in the world with its peak performance over 100 PFLOPS. Firstly, we propose a new algorithm based on cluster particles to fit the special architecture of SW26010. Then, three optimization methods of the MD simulation are implemented step by step: parallelized extensions to SW26010, memory‐access optimizations, and vectorization. After these optimization processes, a 14x speedup is achieved on a single computing node and yields significant performance and energy improvements. Superiorly, almost 24,000 computing nodes (6,000,000 cores) are applied in our experiments, which gain an almost linear speedup. Besides, the proposed methods also can be adapted and fit to other molecular dynamics codes, even other similar scientific applications. Wenqian Dong, Kenli Li 0001, Letian Kang, Zhe Quan, Keqin Li 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2018 | Optimizing power consumption of mobile devices for video streaming over 4G LTE networks
Zhi-Jie Wang 0009, Zhe Quan, Jian Yin 0001, Minyi Guo |
Peer-to-Peer Netw. Appl. | 3 |
| 2017 | Exploring Synchronization in Cache Coherent Manycore Systems: A Case Study with Xeon PhiabstractIntel Xeon Phi is a many-core architecture, featuring more than 50 cores and 200 hardware threads. Given this scale and its other distinctive architectural features, highly-concurrent applications on Xeon Phi may behave differently than on tradi- tional multi-core systems. Yet, concurrency issues especially for synchronization intensive applications on this platform have not been thoroughly analyzed. In this paper, we conduct an extensive analysis at multiple layers, from the underlying hardware cache- coherence protocol up to the user-level applications, aiming to present the most exhaustive study of synchronization on Xeon Phi. Through a range of benchmarks, we testify the feasibility and advantage of accelerating concurrent applications with Xeon Phi. Meanwhile, we identify severe scalability issues relevant to synchronization, and solutions to these issues are discussed. We believe this work can be used as guidelines both for designing better synchronization mechanisms and in optimizing concurrent applications in order to fully exploit the capability of Xeon Phi. Xin He 0054, Zhiwen Chen 0006, Jianhua Sun 0002, Hao Chen 0002, Dong Li 0001, Zhe Quan |
ICPADS | 6 |
| 2017 | Design and evaluation of a parallel neighbor algorithm for the disjunctively constrained knapsack problemabstractSummary We investigate the use of a parallel computing model for solving the disjunctively constrained knapsack problem. This parallel approach is based on a multi‐neighborhood search. In this approach, search threads asynchronously exchange information about the best solutions and use the information to guide the search. The performance of the proposed method was evaluated on the set of the standard benchmark instances. We show encouraging results and compare them to the state‐of‐the‐art solutions. Copyright © 2016 John Wiley & Sons, Ltd. Zhe Quan, Lei Wu 0005 |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Implementation and Optimization of AES Algorithm on the Sunway TaihuLightabstractWith the rapid development of information technology, the security of massive amounts of digital data has attracted huge attention in recent years. In this paper, we provide an efficient parallel implementation of the Advanced Encryption Standard (AES) algorithm, a widely used symmetrical block encryption algorithm, based on the Sunway TaihuLight. The Sunway TaihuLight is a China's independently developed heterogeneous supercomputer with peak performance over 100 PFlops. We also optimize the parallel implementation of the AES algorithm based on the Sunway TaihuLight to achieve more optimized performance. The optimization of the parallel AES algorithm in a single SW26010 node is provided. Specifically, we expand the scale to 1024 nodes and achieve the throughput of about 63.91 GB/s (511.28 Gbits/s). Our parallel implementation of the AES algorithm has great parallel scalability and the speedup ratio can be very high with the number of nodes increasing. Yuedan Chen, Kenli Li 0001, Xiongwei Fei, Zhe Quan, Keqin Li 0001 |
PDCAT | 4 |
| 2010 | An Efficient Branch-and-Bound Algorithm Based on MaxSAT for the Maximum Clique ProblemabstractState-of-the-art branch-and-bound algorithms for the maximum clique problem (Maxclique) frequently use an upper bound based on a partition P of a graph into independent sets for a maximum clique of the graph, which cannot be very tight for imperfect graphs. In this paper we propose a new encoding from Maxclique into MaxSAT and use MaxSAT technology to improve the upper bound based on the partition P. In this way, the strength of specific algorithms for Maxclique in partitioning a graph and the strength of MaxSAT technology in propositional reasoning are naturally combined to solve Maxclique. Experimental results show that the approach is very effective on hard random graphs and on DIMACS Maxclique benchmarks, and allows to close an open DIMACS problem. Chu Min Li 0001, Zhe Quan |
AAAI | 2 |
| 2010 | Combining Graph Structure Exploitation and Propositional Reasoning for the Maximum Clique ProblemabstractState-of-the-art branch-and-bound algorithms for the maximum clique problem (Maxclique) generally exploit the structural information of a graph G to partition G into independent sets, in order to derive an upper bound for the cardinality of a maximum clique of G, which cannot be very tight for imperfect graphs. On the other hand, while Maxclique can be easily encoded into MaxSAT to be solved using a MaxSAT solver, general-purpose MaxSAT solvers are not competitive for solving Maxclique, because they do not exploit the structural information of the graph. Recently, we have shown that propositional reasoning developed for the MaxSAT solvers can be used to improve the upper bound based on the partition. In this paper, we propose and study several improvements to this approach by better combining graph structure exploitation and propositional reasoning to solve Maxclique. Experimental results show that the improvements are very effective on hard random graphs and on DIMACS Maxclique benchmarks which are widely used to evaluate branch-and-bound algorithms for Maxclique. Chu Min Li 0001, Zhe Quan |
ICTAI (1) | 2 |
| 2010 | Exact MinSAT Solving
Chu Min Li 0001, Felip Manyà, Zhe Quan |
SAT | 3 |