EDBT 2026 Demo / reviewers in the wild / expert
Zhi-Jie Wang 0009
dblp:26/11197 · also Zhijie Wang 0009
· DBLP profile ↗
48ranked-venue papers
8as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 2 since 2021Systems, architecture and hardware · 5Computer networks · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TPTCD: A prompt tuning based two-Step framework for cross-Domain text classification
Zhi-Jie Wang 0009, Pei-Pei Li 0001 |
Expert Syst. Appl. | 1 |
| 2025 | External Knowledge-Enhanced Semi-Supervised Multi-Label Short Text Classification
Zhi-Jie Wang 0009, Yirui Li, Shuhui Cao, Pei-Pei Li 0001 |
ICIC (23) | 1 |
| 2025 | Few-Shot Fine-Grained Image Classification via Vision TransformerabstractFew-shot fine-grained image classification (FS-FGIC) is to classify images of the same class into fine-grained subclasses, where only a very limited number of labeled samples in each subclass are available (e.g., 5 or even 1 labeled sample). For most of existing methods, the feature representation capabilities are insufficient, which may harm the performance. Vision Transformers (ViTs) have shown strong feature representation capabilities in many research fields. In this paper, we attempt to solve FSFGIC problem via ViT or its variants. Generally, we use an enhanced ViT as the backbone and adopt a three-stage training strategy. More specifically, (i) we utilize an image matting module to enhance the focus of ViT on the main subjects of images; (ii) we introduce a part selection module to better focus on local image details; and (iii) we incorporate an AdaptMLP module into each Transformer Encoder to reduce the number of parameters that require fine-tuning. Extensive experiments based on three benchmark datasets show us that the proposed model is highly competitive, compared against state-of-the-art models. Tong Xiao 0002, Zeao Chen, Zhi-Jie Wang 0009 |
SMC | 5 |
| 2025 | MDEval: Evaluating and Enhancing Markdown Awareness in Large Language ModelsabstractLarge language models (LLMs) are expected to offer structured Markdown responses for the sake of readability in web chatbots (e.g., ChatGPT). Although there are a myriad of metrics to evaluate LLMs, they fail to evaluate the readability from the view of output content structure. To this end, we focus on an overlooked yet important metric --- Markdown Awareness, which directly impacts the readability and structure of the content generated by these language models. In this paper, we introduce MDEval, a comprehensive benchmark to assess Markdown Awareness for LLMs, by constructing a dataset with 20K instances covering 10 subjects in English and Chinese. Unlike traditional model-based evaluations, MDEval provides excellent interpretability by combining model-based generation tasks and statistical methods. Our results demonstrate that MDEval achieves a Spearman correlation of 0.791 and an accuracy of 84.1% with human, outperforming existing methods by a large margin. Extensive experimental results also show that through fine-tuning over our proposed dataset, less performant open-source models are able to achieve comparable performance to GPT-4o in terms of Markdown Awareness. To ensure reproducibility and transparency, MDEval is open sourced at https://github.com/SWUFE-DB-Group/MDEval-Benchmark. Zhongpu Chen, Yinfeng Liu, Long Shi 0002, Zhi-Jie Wang 0009, Xingyan Chen, Yu Zhao 0019, Fuji Ren |
WWW | 4 |
| 2023 | Image deraining based on dual-channel component decomposition
Xiao Lin 0012, Duojiu Xu, Peiwen Tan, Lizhuang Ma, Zhi-Jie Wang 0009 |
Comput. Graph. | 5 |
| 2023 | Effectively Identifying Compound-Protein Interaction Using Graph Neural RepresentationabstractEffectively identifying compound-protein interactions (CPIs) is crucial for new drug design, which is an important step in silico drug discovery. Current machine learning methods for CPI prediction mainly use one-demensional (1D) compound/protein strings and/or the specific descriptors. However, they often ignore the fact that molecules are essentially modeled by the molecular graph. We observe that in real-world scenarios, the topological structure information of the molecular graph usually provides an overview of how the atoms are connected, and the local chemical context reveals the functionality of the protein sequence in CPI. These two types of information are complementary to each other and they are both significant for modeling compound-protein pairs. Motivated by this, we propose an end-to-end deep learning framework named GraphCPI, which captures the structural information of compounds and leverages the chemical context of protein sequences for solving the CPI prediction task. Our framework can integrate any popular graph neural networks for learning compounds, and it combines with a convolutional neural network for embedding sequences. To compare our method with classic and state-of-the-art deep learning methods, we conduct extensive experiments based on several widely-used CPI datasets. The experimental results show the feasibility and competitiveness of our proposed method. Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Xiangxiang Zeng, Philip S. Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Semantic-Aware Dehazing Network With Adaptive Feature FusionabstractDespite that convolutional neural networks (CNNs) have shown high-quality reconstruction for single image dehazing, recovering natural and realistic dehazed results remains a challenging problem due to semantic confusion in the hazy scene. In this article, we show that it is possible to recover textures faithfully by incorporating semantic prior into dehazing network since objects in haze-free images tend to show certain shapes, textures, and colors. We propose a semantic-aware dehazing network (SDNet) in which the semantic prior is taken as a color constraint for dehazing, benefiting the acquisition of a reasonable scene configuration. In addition, we design a densely connected block to capture global and local information for dehazing and semantic prior estimation. To eliminate the unnatural appearance of some objects, we propose to fuse the features from shallow and deep layers adaptively. Experimental results demonstrate that our proposed model performs favorably against the state-of-the-art single image dehazing approaches. Shengdong Zhang, Wenqi Ren, Xin Tan 0002, Zhi-Jie Wang 0009, Yong Liu 0018, Jingang Zhang, Xiaoqin Zhang 0002, Xiaochun Cao |
IEEE Trans. Cybern. | 4 |
| 2023 | Cover Trees Revisited: Exploiting Unused Distance and Direction InformationabstractThe cover tree (CT) and its improved version are hierarchical data structures that simplified navigating nets while maintaining good runtime guarantees. They can perform nearest neighbor search in logarithmic time and provide efficient computation in practice. In this article, we revisit cover trees for nearest neighbor search, and propose a more competitive method. The central idea of our method is to fully exploit the unused distance and direction information. More specially, our method introduces three novel concepts/techniques: (i) range list, (ii) quadrant information, and (iii) vectorial angle cosine. These techniques are seamlessly integrated into our suggested data structure and search algorithms. As an extra bonus, we explore approximate nearest neighbor and$k$nearest neighbor based on the proposed techniques, and present algorithms for handling updates. Extensive experimental results, based on both real and synthetic datasets, consistently demonstrate that our method is attractive and competitive, compared against existing cover tree structures for nearest neighbor search and its variants. Zhi-Jie Wang 0009, Mengdie Nie, Kaiqi Zhao 0001, Zhe Quan, Bin Yao 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | LPV: A Log Parsing Framework Based on VectorizationabstractLogs are pervasive in modern computing systems, and are valuable to service and system management. Nevertheless, with the rapidly growing size and complexity of computing systems, the log volume is exploding, which makes automatic log analysis imperative. Generally, in automatic log analysis, the first and fundamental step is log parsing, to which a lot of effort has been devoted. However, in most existing log parsing methods, log messages are merely treated as plain text. In natural language processing (NLP) area, it is a common practice to represent words and sentences with vectors, then the similarity between two words or sentences can be measured by the distance between their vectors. Inspired by these, we put forward a novel log parsing framework, named LPV (LogParser based onVectorization), which performs log parsing by converting log messages and log templates into vectors, with the help of a vectorization method in NLP. LPV incorporates offline and online log parsing. In the offline log parsing, the central idea is to first represent log messages with vectors, so that the similarity between two log messages can be measured by the distance between their vectors, then we cluster log messages via clustering the vectors, and finally we extract log templates from the resultant clusters. By the end of the offline log parsing, each log template is assigned with an average vector, so that in the online log parsing, the similarity between an incoming log message and each log template can also be measured by the distance between their vectors. Extensive experiments have been conducted based on several public log datasets to evaluate LPV with three different vectorization methods. The results demonstrate that, with a proper vectorization method, LPV performs competitive with state-of-the-art log parsing methods, in both effectiveness and efficiency. Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Kaiqi Zhao 0001, Xiangke Liao, Yunfei Du 0001, Kenli Li 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | Loader: A Log Anomaly Detector Based on TransformerabstractDetecting anomalies in logs is crucial for service and system management, since logs are widely used to record the runtime status, and are often the only data available for postmortem analysis. Since anomalies are usually rare in real-world services and systems, a common and feasible practice is to mine or learn normal patterns from logs, and deem those violating the normal patterns as anomalies. As log sequences are a kind of time series data, RNN (Recurrent Neural Network) and its variants have been extensively employed to capture the normal patterns. Nevertheless, the sequential nature of RNN and its variants makes them hard to parallelize and capture long-term dependencies, which may hinder their performance. To address this issue, in this paper we propose Loader, a novel semi-supervisedloganomalydetector based on Transformer, because the Transformer architecture eschews recurrence and is able to draw global dependencies. Loader leverages the Transformer encoder to capture normal patterns from normal log sequences. When detecting, it gives a set of candidate log templates, that may appear after the input log substring under normal conditions. If the template of the actual next log message is not within the candidate set, this implies an anomaly. Previous similar methods select the most possible$k$log templates as candidates in any case, so the performance is sensitive to$k$, and it is nontrivial to pick a proper$k$. To alleviate this, we design a more flexible and robust ‘top-$p$’ algorithm, which determines the candidate set based on the cumulative probability of the most possible log templates. Extensive experiments are conducted based on three public log datasets, the experimental results validate the effectiveness and competitiveness of our approach. Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Yuquan Le, Yunfei Du 0001, Xiangke Liao, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Salient Object Detection Based on Multiscale Segmentation and Fuzzy Broad LearningabstractAbstract Saliency detection has been a hot topic in the field of computer vision. In this paper, we propose a novel approach that is based on multiscale segmentation and fuzzy broad learning. The core idea of our method is to segment the image into different scales, and then the extracted features are fed to the fuzzy broad learning system (FBLS) for training. More specifically, it first segments the image into superpixel blocks at different scales based on the simple linear iterative clustering algorithm. Then, it uses the local binary pattern algorithm to extract texture features and computes the average color information for each superpixel of these segmentation images. These extracted features are then fed to the FBLS to obtain multiscale saliency maps. After that, it fuses these saliency maps into an initial saliency map and uses the label propagation algorithm to further optimize it, obtaining the final saliency map. We have conducted experiments based on several benchmark datasets. The results show that our solution can outperform several existing algorithms. Particularly, our method is significantly faster than most of deep learning-based saliency detection algorithms, in terms of training and inferring time. Xiao Lin 0012, Zhi-Jie Wang 0009, Lizhuang Ma, Meie Fang |
Comput. J. | 2 |
| 2021 | Flexible Aggregate Nearest Neighbor Queries and its Keyword-Aware Variant on Road NetworksabstractAggregate nearest neighbor (Ann) query in both the euclidean space and road networks has been extensively studied, and the flexible aggregate nearest neighbor (Fann) problem further generalizesAnnby introducing an extra flexibility parameter$\phi$that ranges in$(0, 1]$. In this article, we focus onFannon road networks, denoted asFann$_\mathcal {R}$, and its keyword-aware variant, denoted asKFann$_\mathcal {R}$. To solve these problems, we propose a series of universal (i.e., suitable for bothmaxandsum) algorithms, including a Dijkstra-based algorithm that enumerates$P$instead of$\phi |Q|$-combinations of$Q$, a queue-based approach that processes data points from-near-to-far, and a framework that combinesincremental euclidean restriction(IER) and$k$NN. We also propose a specific exact solution tomax-Fann$_\mathcal {R}$and a constant-factor ratio approximate solution tosum-Fann$_\mathcal {R}$. These specific algorithms are easy to implement and can achieve excellent performance in some scenarios. Besides, we further extend this problem to top-$k$and multipleFann$_\mathcal {R}$(resp.,KFann$_\mathcal {R}$) queries. We conduct a comprehensive experimental evaluation for the proposed algorithms on real datasets to demonstrate their superior efficiency and high quality. Zhongpu Chen, Bin Yao 0002, Zhi-Jie Wang 0009, Xiaofeng Gao 0001, Shuo Shang, Shuai Ma 0001, Minyi Guo |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Utilizing Two-Phase Processing With FBLS for Single Image DerainingabstractRain removal from a single image is a challenging problem and has attracted much attention in recent years. In this paper, we revisit the single image deraining problem, and present a novel solution. The central idea of our solution is to merge the merits of two-phase processing methods and the Fuzzy Broad Learning System (FBLS). Specifically, our solution first uses the dehazing algorithm to preprocess the input rainy image and separates it into the detail layer and the base layer. After that, it puts the Y-channel image of the detail layer into the FBLS to obtain the derained Y channel image, which is then combined with the Cb and Cr channel images to obtain the derained detail layer. Later, it fuses the derained detail layer and the base layer to get a preliminary derained image. Finally, it superimposes the details extracted from the dehazed image with some transparency on the preliminary result, obtaining the final result. Experimental results based on both real and synthetic rainy images demonstrate that our proposed solution can outperform several state-of-the-art algorithms, while it consumes much less running time and training time, compared against the competitors. Xiao Lin 0012, Lizhuang Ma, Bin Sheng 0001, Zhi-Jie Wang 0009, Wansheng Chen |
IEEE Trans. Multim. | 4 |
| 2020 | A Novel Model for Imbalanced Data ClassificationabstractRecently, imbalanced data classification has received much attention due to its wide applications. In the literature, existing researches have attempted to improve the classification performance by considering various factors such as the imbalanced distribution, cost-sensitive learning, data space improvement, and ensemble learning. Nevertheless, most of the existing methods focus on only part of these main aspects/factors. In this work, we propose a novel imbalanced data classification model that considers all these main aspects. To evaluate the performance of our proposed model, we have conducted experiments based on 14 public datasets. The results show that our model outperforms the state-of-the-art methods in terms of recall, G-mean, F-measure and AUC. Jian Yin 0001, Chunjing Gan, Kaiqi Zhao 0001, Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009 |
AAAI | 6 |
| 2020 | DeepGS: Deep Representation Learning of Graphs and Sequences for Drug-Target Binding Affinity PredictionabstractAccurately predicting drug-target binding affinity (DTA) in silico is a key task in drug discovery. Most of the conventional DTA prediction methods are simulation-based, which rely heavily on domain knowledge or the assumption of having the 3D structure of the targets, which are often difficult to obtain. Meanwhile, traditional machine learning-based methods apply various features and descriptors, and simply depend on the similarities between drug-target pairs. Recently, with the increasing amount of affinity data available and the success of deep representation learning models on various domains, deep learning techniques have been applied to DTA prediction. However, these methods consider either label/one-hot encodings or the topological structure of molecules, without considering the local chemical context of amino acids and SMILES sequences. Motivated by this, we propose a novel end-to-end learning framework, called DeepGS, which uses deep neural networks to extract the local chemical context from amino acids and SMILES sequences, as well as the molecular structure from the drugs. To assist the operations on the symbolic data, we propose to use advanced embedding techniques (i.e., Smi2Vec and Prot2Vec) to encode the amino acids and SMILES sequences to a distributed representation. Meanwhile, we suggest a new molecular structure modeling approach that works well under our framework. We have conducted extensive experiments to compare our proposed method with state-of-the-art models including KronRLS, SimBoost, DeepDTA and DeepCPI. Extensive experimental results demonstrate the superiorities and competitiveness of DeepGS. Xuan Lin, Kaiqi Zhao 0001, Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Philip S. Yu |
ECAI | 5 |
| 2020 | LPV: A Log Parser Based on Vectorization for Offline and Online Log ParsingabstractAs the first and foremost step of typical automatic log analysis, log parsing has attracted a lot of interest. Most of existing studies treat log messages as pure strings and rely on string matching or string distance. In NLP, word2vec has shown very efficient and effective in representing words with low dimensional vectors. Inspired by this, in this paper we propose a novel method, called LPV (Log Parser based on Vectorization), for both offline and online log parsing. The central idea of our method in offline log parsing is to first convert log messages into vectors, and measure the similarity between two log messages by the distance between two vectors, then log messages can be clustered via clustering the vectors, and log templates can be extracted from the resulting clusters. For online log parsing, we also assign log templates with some kind of average vectors, so that the similarity between an incoming log message and each log template can also be measured by the distance between two vectors. We have conducted extensive experiments based on three widely used log datasets, and the results demonstrate that our proposed method LPV can achieve a competitive performance, compared against state-of-the-art log parsing methods. Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Kaiqi Zhao 0001, Xiangke Liao |
ICDM | 3 |
| 2020 | Text to Image Synthesis With Bidirectional Generative Adversarial NetworkabstractGenerating realistic images from text descriptions is a challenging problem in computer vision. Although previous works have shown remarkable progress, guaranteeing semantic consistency between text descriptions and images remains challenging. To generate semantically consistent images, we propose two semantics-enhanced modules and a novel Textual-Visual Bidirectional Generative Adversarial Network (TVBi-GAN). Specifically, this paper proposes a semanticsenhanced attention module and a semantics-enhanced batch normalization module. These modules improve consistency of synthesized images by involving precisely semantic features. What's more, an encoder network is proposed to extract semantic features from images. During the adversarial process, the encoder could guide our generator to explore corresponding features behind descriptions. With extensive experiments on CUB and COCO datasets, we demonstrate that our TVBi-GAN outperforms state-of-the-art methods. Zhe Quan, Zhi-Jie Wang 0009, Xinjian Hu, Yangyang Chen 0006 |
ICME | 3 |
| 2020 | KGNN: Knowledge Graph Neural Network for Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) prediction is a challenging problem in pharmacology and clinical application, and effectively identifying potential DDIs during clinical trials is critical for patients and society. Most of existing computational models with AI techniques often concentrate on integrating multiple data sources and combining popular embedding methods together. Yet, researchers pay less attention to the potential correlations between drug and other entities such as targets and genes. Moreover, recent studies also adopted knowledge graph (KG) for DDI prediction. Yet, this line of methods learn node latent embedding directly, but they are limited in obtaining the rich neighborhood information of each entity in the KG. To address the above limitations, we propose an end-to-end framework, called Knowledge Graph Neural Network (KGNN), to resolve the DDI prediction. Our framework can effectively capture drug and its potential neighborhoods by mining their associated relations in KG. To extract both high-order structures and semantic relations of the KG, we learn from the neighborhoods for each entity in the KG as their local receptive, and then integrate neighborhood information with bias from representation of the current entity. This way, the receptive field can be naturally extended to multiple hops away to model high-order topological information and to obtain drugs potential long-distance correlations. We have implemented our method and conducted experiments based on several widely-used datasets. Empirical results show that KGNN outperforms the classic and state-of-the-art models. Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Tengfei Ma 0002, Xiangxiang Zeng |
IJCAI | 3 |
| 2020 | A novel molecular representation with BiGRU neural networks for learning atomabstractMolecular representations play critical roles in researching drug design and properties, and effective methods are beneficial to assisting in the calculation of molecules and solving related problem in drug discovery. In previous years, most of the traditional molecular representations are based on hand-crafted features and rely heavily on biological experimentations, which are often costly and time consuming. However, recent researches achieve promising results using machine learning on various domains. In this article, we present a novel method named Smi2Vec-BiGRU that is designed for learning atoms and solving the single- and multitask binary classification problems in the field of drug discovery, which are the basic and also key problems in this field. Specifically, our approach transforms the molecule data in the SMILES format into a set of sample vectors and then feeds them into the bidirectional gated recurrent unit neural networks for training, which learns low-dimensional vector representations for molecular drug. We conduct extensive experiments on several widely used benchmarks including Tox21, SIDER and ClinTox. The experimental results show that our approach can achieve state-of-the-art performance on these benchmarking datasets, demonstrating the feasibility and competitiveness of our proposed approach. Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Xiangxiang Zeng |
Briefings Bioinform. | 3 |
| 2020 | ITISS: an efficient framework for querying big temporal data
Zhongpu Chen, Bin Yao 0002, Zhi-Jie Wang 0009, Wei Zhang 0398, Kai Zheng 0001, Panos Kalnis, Feilong Tang 0001 |
GeoInformatica | 3 |
| 2020 | Scalable Tensor-Train-Based Tensor Computations for Cyber-Physical-Social Big DataabstractTensor-based big data analysis approaches are effectively exploited to handle multisource and heterogeneous cyber-physical-social big data generated from diverse spaces. However, the curse of dimensionality seriously restricts their widespread exploitation, especially under edge/fog computing environments. To alleviate the dilemma, we attempt to present a set of tensor-train (TT)-based tensor operations with their scalable computations and then propose a novel TT-based big data processing framework under edge/fog computing environments. Specifically, in this article, we first summarize and present a set of TT-based tensor operations by converting the original high-order tensor operation to a series of low-order (second- or third-order) TTcore-based operations. Then, we propose a two-layer scalable TT-based computation architecture, including inter-TTcore and intra-TTcore scalable models. Afterward, according to various scalable models, a series of scalable TT-based tensor computations (STT-TCs) with their complexity analysis are proposed in detail. Finally, we propose a novel TT-based big data processing framework to adapt to edge/fog computing environments. We conduct extensive experiments based on both random data sets and real-world ubiquitous bus traffic data sets. Experimental results demonstrate that the proposed STT-TCs can significantly improve computation efficiency and are suitable for edge/fog computing environments. Huazhong Liu, Laurence T. Yang, Jihong Ding, Yimu Guo, Zhi-Jie Wang 0009 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2020 | Skia: Scalable and Efficient In-Memory Analytics for Big Spatial-Textual DataabstractIn recent years, spatial-keyword queries have attracted much attention with the fast development of location-based services. However, current spatial-keyword techniques are disk-based, which cannot fulfill the requirements of high throughput and low response time. With the surging data size, people tend to process data in distributed in-memory environments to achieve low latency. In this paper, we present the distributed solution, i.e., Skia (Spatial-Keyword In-memory Analytics), to provide a scalable backend for spatial-textual analytics. Skia introduces a two-level index framework for big spatial-textual data including: (1) efficient and scalable global index, which prunes the candidate partitions a lot while achieving small space budget; and (2) four novel local indexes, that further support low latency services for exact and approximate spatial-keyword queries. Skia can support common spatial-keyword queries via traditional SQL programming interfaces. The experiments conducted on large-scale real datasets have demonstrated the promising performance of the proposed indexes and our distributed solution. Yang Xu 0031, Bin Yao 0002, Zhi-Jie Wang 0009, Xiaofeng Gao 0001, Jiong Xie, Minyi Guo |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Task Scheduling for Energy Consumption Constrained Parallel Applications on Heterogeneous Computing SystemsabstractPower-aware task scheduling on processors has been a research hotspot in computing systems. Given an application G containing a set N of tasks {n1,...,n|N|}, and a system containing a set U of processors {u1,..., u|U|}, the power-aware task scheduling generally refers to finding the appropriate processor and frequency for each task ni, so as to make sure that all the tasks can be finished efficiently and the overall energy consumption is guaranteed. In this article, we study the problem of minimizing the schedule length for energy consumption constrained parallel applications on heterogeneous computing systems, where the schedule length refers to the time interval between starting the first task and finishing the last task. For this problem, existing work adopts a policy that preassigns the minimum energy consumption for each unassigned task. Nevertheless, our analysis reveals that, such a pre-assignment policy could be unfair for the low priority tasks, and it may not achieve an optimistic schedule length. Thereby, we propose a new task scheduling algorithm that suggests a weight-based mechanism to preassign energy consumption for unassigned tasks, and we provide the rigorous proof to show its feasibility. Further, we show that this idea can be extended to solve reliability maximization problems with energy consumption constraint or with both deadline and energy consumption constraints, where the reliability refers to the probability of executing application G without failures, and the deadline constraint refers to the “allowable” maximum schedule length. We have conducted extensive experiments based on real parallel applications. The experimental results consistently demonstrate that our proposed algorithms can achieve favourable performance, compared to state-of-the-art algorithms. Zhe Quan, Zhi-Jie Wang 0009, Ting Ye, Song Guo 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | An Improved Hierarchical Datastructure for Nearest Neighbor Search
Mengdie Nie, Zhi-Jie Wang 0009, Chunjing Gan, Zhe Quan, Bin Yao 0002, Jian Yin 0001 |
AAAI | 2 |
| 2019 | GraphCPI: Graph Neural Representation Learning for Compound-Protein InteractionabstractAccurately predicting compound-protein interactions (CPIs) is of great help to increase the efficiency and reduce costs in drug development. Most of existing machine learning models for CPI prediction often represent compounds and proteins in one-dimensional strings, or use the descriptor-based methods. These models might ignore the fact that molecules are essentially structured by the chemical bond of atoms. However, in real-world scenarios, the topological structure information usually provides an overview of how the atoms are connected, and the local chemical context reveals the functionality of the protein sequence in CPI. These two types of information are complementary to each other and they are both important for modeling compounds and proteins. Motivated by this, this paper suggests an end-to-end deep learning framework called GraphCPI, which captures the structural information of compounds and leverages the chemical context of protein sequences for solving the CPI prediction task. Our framework can integrate any popular graph nerual networks for learning compounds, and it combines with a convolutional neural network for embedding sequences. We conduct extensive experiments based on two benchmark CPI datasets. The experimental results demonstrate that our proposed framework is feasible and also competitive, comparing against classic and state-of-the-art methods. Zhe Quan, Xuan Lin, Zhi-Jie Wang 0009, Xiangxiang Zeng |
BIBM | 4 |
| 2019 | A Solution for High Availability Memory Access
Chunjing Gan, Bin Wang 0015, Zhi-Jie Wang 0009, Huazhong Liu, Dingyu Yang, Jian Yin 0001, Shiyou Qian, Song Guo 0001 |
ICA3PP (1) | 3 |
| 2019 | Optimizing the Waiting Time of Sensors in a MANET to Strike a Balance between Energy Consumption and Data TimelinessabstractOceans are important for scientific research and also for global economic and military security. Usually, wireless ad hoc networks are chosen to transform real-time data collected by ocean monitoring sensors (nodes). Due to the random motion of waves or the random direction of the wind, nodes in the network might become detached from the coverage of the network. In this case, the detached nodes can either send the collected data directly to the base station at the cost of consuming more energy or wait for a period of time to rejoin the network with the price of sacrificing the real time of the collected data. In this paper, we model the optimal waiting time for detached nodes before directly sending the data in the dynamic environment of ocean monitoring. For this purpose, we need to address two problems. The first is how to calculate the rate of coverage with a different number and different broadcast radii of nodes. The second is when a node detaches from the coverage of the network, how much time will it need to wait before it rejoins the network. We first establish the motion model of nodes, which is the basis to deduce the probability distribution of a certain time when the detached node rejoins the network. Based on the probability distribution, the waiting time of the detached nodes can be optimally determined, aiming to achieve a good balance between energy consumption and data timeliness. Finally, a series of simulations is conducted to validate the effectiveness of our proposed method. Hanwen Hu, Shiyou Qian, Jian Cao 0001, Jiadi Yu, Guangtao Xue, Yanmin Zhu 0006, Minglu Li 0001, Zhi-Jie Wang 0009 |
ICPADS | 8 |
| 2019 | Geo-ALM: POI Recommendation by Fusing Geographical Information and Adversarial Learning MechanismabstractLearning user’s preference from check-in data is important for POI recommendation. Yet, a user usually has visited some POIs while most of POIs are unvisited (i.e., negative samples). To leverage these “no-behavior” POIs, a typical approach is pairwise ranking, which constructs ranking pairs for the user and POIs. Although this approach is generally effective, the negative samples in ranking pairs are obtained randomly, which may fail to leverage “critical” negative samples in the model training. On the other hand, previous studies also utilized geographical feature to improve the recommendation quality. Nevertheless, most of previous works did not exploit geographical information comprehensively, which may also affect the performance. To alleviate these issues, we propose a geographical information based adversarial learning model (Geo-ALM), which can be viewed as a fusion of geographic features and generative adversarial networks. Its core idea is to learn the discriminator and generator interactively, by exploiting two granularity of geographic features (i.e., region and POI features). Experimental results show that Geo- ALM can achieve competitive performance, compared to several state-of-the-arts. Wei Liu 0061, Zhi-Jie Wang 0009, Bin Yao 0002, Jian Yin 0001 |
IJCAI | 2 |
| 2019 | MCCH: A novel convex hull prior based solution for saliency detection
Xiao Lin 0012, Zhi-Jie Wang 0009, Xin Tan 0002, Meie Fang, Naixue Xiong, Lizhuang Ma |
Inf. Sci. | 2 |
| 2019 | Reachable region query and its applications
Mengdie Nie, Zhi-Jie Wang 0009, Jian Yin 0001, Bin Yao 0002 |
Inf. Sci. | 2 |
| 2019 | CATIRI: An Efficient Method for Content-and-Text Based Image Retrieval
Mengqi Zeng, Bin Yao 0002, Zhi-Jie Wang 0009, Yanyan Shen, Feifei Li 0001, Minyi Guo |
J. Comput. Sci. Technol. | 3 |
| 2019 | An Efficient Framework for Sentence Similarity ModelingabstractSentence similarity modeling lies at the core of many natural language processing applications, and thus has received much attention. Owing to the success of word embeddings, recently, popular neural network methods achieved sentence embedding. Most of them focused on learning semantic information and modeling it as a continuous vector, yet the syntactic information of sentences has not been fully exploited. On the other hand, prior works have shown the benefits of structured trees that include syntactic information, while few methods in this branch utilized the advantages of word embeddings and another powerful technique-attention weight mechanism. This paper suggests to absorb their advantages by merging these techniques in a unified structure, dubbed as attention constituency vector tree (ACVT). Meanwhile, this paper develops a new tree kernel, known as ACVT kernel, which is tailored for sentence similarity measure based on the proposed structure. The experimental results, based on 19 widely used semantic textual similarity datasets, demonstrate that our model is effective and competitive, when compared against state-of-the-art models. Additionally, the experimental results validate that many attention weight mechanisms and word embedding techniques can be seamlessly integrated into our model, demonstrating the robustness and universality of our model. Zhe Quan, Zhi-Jie Wang 0009, Yuquan Le, Bin Yao 0002, Kenli Li 0001, Jian Yin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Reduced Prediction Saturation and View Effects for Estimating the Leaf Area Index of Winter WheatabstractRelationships between vegetation indices (VIs) and leaf area index (LAI) tend to saturate in the nadir direction, and vary with crop canopy structure and view zenith angles (VZAs). The objective of this paper was to improve the monitoring accuracy and angular stability of VIs for estimating LAI using multiangular remote sensing data. The relationship between LAI and ground-based hyperspectral spectral reflectance was quantified in winter wheat (Triticum aestivum L.) exhibiting erectophile and planophile growth habits. To further reduce the saturation, species specificity, and angular sensitivity, we developed a saturation factor (SF), based on near-infrared and green bands. Multiplying all VIs by SF greatly improved the association with LAI across all VZAs (R2= 0.73-0.82). Most VI × SF values, particularly optimized soil-adjusted VI × SF, were able to construct universal algorithms across VZAs for accurate estimation of LAI due to the sensitivity of SF to LAI in a dense canopy, and the insensitivity of SF to view effects with larger VZAs. This approach is also promising for the exploitation of multiangular satellite data for the design and calibration of nonview-angle-corrected spectral reflectance, for which the sensor is only deployed at fixed observation direction. Craig A. Coburn, Zhi-Jie Wang 0009, Wei Feng 0013, Tian-Cai Guo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Saliency Detection via Multi-Scale Global CuesabstractThe saliency detection technologies are very useful to analyze and extract important information from given multimedia data, and have already been extensively used in many multimedia applications. Past studies have revealed that utilizing the global cues is effective in saliency detection. Nevertheless, most of prior works mainly considered the single-scale segmentation when the global cues are employed. In this paper, we attempt to incorporate the multi-scale global cues for saliency detection problem. Achieving this proposal is interesting and also challenging (e.g., How to obtain appropriate foreground and background seeds effectively? How to merge rough saliency results into the final saliency map efficiently?). To alleviate the challenges, we present a three-phase solution that integrates several targeted strategies, first, a self-adaptive strategy for obtaining appropriate filter parameters; second, a cross-validation scheme for selecting appropriate background and foreground seeds; and third, a weight-based approach for merging the rough saliency maps. Our solution is easy to understand and implement, but without loss of effectiveness. Extensive experimental results based on benchmark datasets demonstrate the feasibility and competitiveness of our proposed solution. Xiao Lin 0012, Zhi-Jie Wang 0009, Lizhuang Ma, Xiabao Wu |
IEEE Trans. Multim. | 2 |
| 2019 | Improving Power Efficiency for Online Video Streaming Service: A Self-Adaptive ApproachabstractThe video streaming technique and 4G LTE networks have been developing at a high speed, while the advancement of the battery technology is relatively slow. A lot of efforts have been made on understanding the power consumption in general 4G LTE networks, while little attention has been focused on reducing the power consumption of mobile devices for online video streaming in 4G LTE networks. In this paper, we attempt to develop efficient optimization techniques to address this problem. It is an important and also interesting problem, while there are many challenges needing to be alleviated (e.g., the uncertainty of users' behavior modes). To attack the challenges, we suggest a self-adaptive method that can allow us to adjust various parameters dynamically and efficiently, achieving a relatively small energy consumption. We give the rigorous theoretical analysis for the proposed method, and conduct the empirical study to validate its effectiveness. Zhi-Jie Wang 0009, Kun Wang 0005, Song Guo 0001, Bin Wang 0015, Minyi Guo |
IEEE Trans. Sustain. Comput. | 2 |
| 2018 | A System for Learning Atoms Based on Long Short-Term Memory Recurrent Neural Networks
Zhe Quan, Xuan Lin, Zhi-Jie Wang 0009, Yan Liu 0032, Kenli Li 0001 |
BIBM | 3 |
| 2018 | Geographical Relevance Model for Long Tail Point-of-Interest Recommendation
Wei Liu 0061, Zhi-Jie Wang 0009, Bin Yao 0002, Mengdie Nie, Jing Wang 0030, Rui Mao 0001, Jian Yin 0001 |
DASFAA (1) | 2 |
| 2018 | Distributed In-Memory Analytics for Big Temporal Data
Bin Yao 0002, Wei Zhang 0398, Zhi-Jie Wang 0009, Zhongpu Chen, Shuo Shang, Kai Zheng 0001, Minyi Guo |
DASFAA (1) | 3 |
| 2018 | Saliency Detection via Multi-Center Convex Hull PriorabstractSaliency detection has been a hot topic in computer vision. Among existing approaches, a representative one is to use the convex hull prior to find the salient object in the image; and there are many variants that are based on the convex hull prior. Most of these works used a single center to construct the convex hull center prior map, while few attention has been made on the use of multiple centers. In this paper, we propose a multi-center convex hull prior based solution for saliency detection. Particularly, our solution also integrates two non-trivial optimizations: one is for obtaining an enhanced global color distinction prior map, and another is for refining the preliminary saliency map. We experimentally evaluate our solution through comparing against state-of-the-art algorithms. The results demonstrate the effectiveness and superiorities of the proposed solution. Zhi-Jie Wang 0009, Lizhuang Ma, Xiao Lin 0012 |
ICASSP | 1 |
| 2018 | MSGC: A New Bottom-Up Model for Salient Object DetectionabstractSaliency detection has been a hot topic in computer vision and image processing communities. Utilizing the global cues has been shown effective in saliency detection, whereas most of prior works mainly considered the single-scale segmentation when the global cues are employed. In this paper, we attempt to incorporate the multi-scale global cues (MSGC) for saliency detection. Achieving this proposal is interesting and also challenging (e.g., how to obtain appropriate foreground and background seeds; how to merge rough saliency results into the final saliency map efficiently). To alleviate various challenges, we present a solution that integrates three targeted techniques: (i) a self-adaptive approach for obtaining appropriate filter parameters; (ii) a cross-validation approach for selecting appropriate background and foreground seeds; and (iii) a weight-based approach for merging the rough saliency maps. Our solution is easy-to-understand and implement, but without loss of effectiveness. We have validated its competitiveness through widely used benchmark datasets. Zhi-Jie Wang 0009, Lizhuang Ma, Xiao Lin 0012, Xiabao Wu |
ICME | 1 |
| 2018 | ISAECC: An Improved Scheduling Approach for Energy Consumption Constrained Parallel Applications on Heterogeneous Distributed SystemsabstractPower-aware task scheduling on processors has been a hot topic. In this paper, we study the problem of minimizing the schedule length for energy consumption constrained parallel applications on heterogeneous distributed systems. Previous work (solving this problem) adopts a policy that preassigns the minimum energy consumption for each unassigned task. Nevertheless, our analysis reveals that such a preassignment policy could be unfair, and it may not achieve an optimistic schedule length. Motivated by this, we propose a new task scheduling algorithm that suggests a weight-based mechanism to preassign energy consumption for unassigned tasks. We theoretically prove that our preassignment mechanism can guarantee the energy consumption constraint. Also, we have conducted extensive experiments based on two real parallel applications. The results consistently demonstrate that, compared to state-of-the-art algorithms, our approach can achieve smaller schedule length while satisfying the energy consumption constraint. Ting Ye, Zhi-Jie Wang 0009, Zhe Quan, Song Guo 0001, Kenli Li 0001, Keqin Li 0001 |
ICPADS | 2 |
| 2018 | ACV-tree: A New Method for Sentence Similarity ModelingabstractSentence similarity modeling lies at the core of many natural language processing applications, and thus has received much attention. Owing to the success of word embeddings, recently, popular neural network methods have achieved sentence embedding, obtaining attractive performance. Nevertheless, most of them focused on learning semantic information and modeling it as a continuous vector, while the syntactic information of sentences has not been fully exploited. On the other hand, prior works have shown the benefits of structured trees that include syntactic information, while few methods in this branch utilized the advantages of word embeddings and another powerful technique ? attention weight mechanism. This paper makes the first attempt to absorb their advantages by merging these techniques in a unified structure, dubbed as ACV-tree. Meanwhile, this paper develops a new tree kernel, known as ACVT kernel, that is tailored for sentence similarity measure based on the proposed structure. The experimental results, based on 19 widely-used datasets, demonstrate that our model is effective and competitive, compared against state-of-the-art models. Yuquan Le, Zhi-Jie Wang 0009, Zhe Quan, Jiawei He 0003, Bin Yao 0002 |
IJCAI | 2 |
| 2018 | FastPM: An approach to pattern matching via distributed stream processing
Dingyu Yang, Jianmei Guo, Zhi-Jie Wang 0009, Yuan Wang 0003, Jingsong Zhang, Liang Hu 0004, Jian Yin 0001, Jian Cao 0001 |
Inf. Sci. | 3 |
| 2018 | Optimizing power consumption of mobile devices for video streaming over 4G LTE networks
Zhi-Jie Wang 0009, Zhe Quan, Jian Yin 0001, Minyi Guo |
Peer-to-Peer Netw. Appl. | 2 |
| 2018 | Power consumption analysis of video streaming in 4G LTE networks
Zhi-Jie Wang 0009, Song Guo 0001, Dingyu Yang, Gan Fang, Chunyi Peng 0001, Minyi Guo |
Wirel. Networks | 2 |
| 2017 | Re2l: An efficient output-sensitive algorithm for computing Boolean operations on circular-arc polygons and its applications
Zhi-Jie Wang 0009, Xiao Lin 0012, Meie Fang, Bin Yao 0002, Yong Peng 0001, Haibing Guan, Minyi Guo |
Comput. Aided Des. | 1 |
| 2016 | SMe: explicit & implicit constrained-space probabilistic threshold range queries for moving objects
Zhi-Jie Wang 0009, Bin Yao 0002, Reynold Cheng, Xiaofeng Gao 0001, Lei Zou 0001, Haibing Guan, Minyi Guo |
GeoInformatica | 1 |
| 2015 | Probabilistic Range Query over Uncertain Moving Objects in Constrained Two-Dimensional SpaceabstractProbabilistic range query (PRQ) over uncertain moving objects has attracted much attentions in recent years. Most of existing works focus on the PRQ for objects moving freely in two-dimensional (2D) space. In contrast, this paper studies the PRQ over objects moving in a constrained 2D space where objects are forbidden to be located in some specific areas. We dub it the constrained space probabilistic range query (CSPRQ). We analyze its unique properties and show that to process the CSPRQ using a straightforward solution is infeasible. The key idea of our solution is to use a strategy calledpre-approximationthat can reduce the initial problem to a highly simplified version, implying that it makes the rest of steps easy to tackle. In particular, this strategy itself is pretty simple and easy to implement. Furthermore, motivated by the cost analysis, we further optimize our solution. The optimizations are mainly based on two insights: (i) the number ofeffective subdivisions is no more than 1; and (ii) an entity with the largerspanis more likely to subdivide a single region. We demonstrate the effectiveness and efficiency of our proposed approaches through extensive experiments under various experimental settings, and highlight an extra finding—the precomputation based method suffers a non-trivial preprocessing time, which offers an important indication sign for the future research. Zhi-Jie Wang 0009, Dong-Hua Wang, Bin Yao 0002, Minyi Guo |
IEEE Trans. Knowl. Data Eng. | 1 |