VLDB 2026 Research / reviewers in the wild / expert
Jianwu Xu
dblp:81/8437 · also Jian-Wu Xu
· DBLP profile ↗
23ranked-venue papers
10as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3Security and privacy · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 77% Information retrieval · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 50% Distributed systems · 50% | |
| Artificial intelligence
1 paper |
Question answering and dialogue systems · 100% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › community question answering
question routing |
0.4 | 1 | 2020 | Temporal Context-Aware Representation Learning for Question Routing · WSDM 2020 |
Recommender systems › user recommendation
expert recommendation |
0.4 | 1 | 2020 | Temporal Context-Aware Representation Learning for Question Routing · WSDM 2020 |
Cloud and datacenter computing › cloud service management
cloud infrastructure management |
0.2 | 1 | 2016 | CloudSeer: Workflow Monitoring of Cloud Infrastructures via Interleaved Logs · ASPLOS 2016 |
Distributed systems
fault tolerance |
0.2 | 1 | 2016 | CloudSeer: Workflow Monitoring of Cloud Infrastructures via Interleaved Logs · ASPLOS 2016 |
Audio and music processing › speech processing
pitch estimation |
0.1 | 1 | 2008 | A Pitch Detector Based on a Generalized Correlation Function · IEEE Trans. Speech Audio Process. 2008 |
Methods — techniques the papers use, named apart from their topics
multi-resolution temporal modeling · 0.9attention · 0.9streaming log checking · 0.2automaton-based log analysis · 0.2reproducing kernel hilbert space · 0.1generalized correlation function · 0.1correntropy · 0.1ERB filter bank · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PPDO: a privacy-preservation-aware delay optimization task-offloading algorithm for collaborative edge computingabstractAlthough collaborative edge computing (CEC) systems are beneficial in enhancing the performance of mobile edge computing (MEC), the issue of user privacy leakage becomes prominent during task offloading. To address this issue, we design a privacy-preservation-aware delay optimization task-offloading algorithm (PPDO) in a CEC system. By considering location and usage pattern privacy protection, we establish a privacy task model to interfere with the edge server and ensure user privacy. To address the extra delay arising from privacy protection, we subsequently leverage a Markov decision processing (MDP) policy-iteration-based algorithm to minimize delays without compromising privacy. To simultaneously accelerate the MDP operation, we develop an extension that improves the PPDO by optimizing the action set. Finally, a comprehensive simulation was conducted using the edge user allocation (EUA) dataset. The results demonstrated that PPDO achieves an optimal trade-off between privacy protection and delay with a minimum delay compared with existing algorithms. Moreover, we examined the advantages and disadvantages of improving PPDO. Chao Jing, Jianwu Xu |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2020 | Mining Multivariate Discrete Event Sequences for Knowledge Discovery and Anomaly DetectionabstractModern physical systems deploy large numbers of sensors to record at different time-stamps the status of different systems components via measurements such as temperature, pressure, speed, but also the component's categorical state. Depending on the measurement values, there are two kinds of sequences: continuous and discrete. For continuous sequences, there is a host of state-of-the-art algorithms for anomaly detection based on time-series analysis, but there is a lack of effective methodologies that are tailored specifically to discrete event sequences. This paper proposes an analytics framework for discrete event sequences for knowledge discovery and anomaly detection. During the training phase, the framework extracts pairwise relationships among discrete event sequences using a neural machine translation model by viewing each discrete event sequence as a "natural language". The relationship between sequences is quantified by how well one discrete event sequence is "translated" into another sequence. These pairwise relationships among sequences are aggregated into a multivariate relationship graph that clusters the structural knowledge of the underlying system and essentially discovers the hidden relationships among discrete sequences. This graph quantifies system behavior during normal operation. During testing, if one or more pairwise relationships are violated, an anomaly is detected. The proposed framework is evaluated on two real-world datasets: a proprietary dataset collected from a physical plant where it is shown to be effective in extracting sensor pairwise relationships for knowledge discovery and anomaly detection, and a public hard disk drive dataset where its ability to effectively predict upcoming disk failures is illustrated. Bin Nie, Jianwu Xu, Jacob Alter, Evgenia Smirni |
DSN | 2 |
| 2020 | Temporal Context-Aware Representation Learning for Question RoutingabstractQuestion routing (QR) aims at recommending newly posted questions to the potential answerers who are most likely to answer the questions. The existing approaches that learn users' expertise from their past question-answering activities usually suffer from challenges in two aspects: 1) multi-faceted expertise and 2) temporal dynamics in the answering behavior. This paper proposes a novel temporal context-aware model in multiple granularities of temporal dynamics that concurrently address the above challenges. Specifically, the temporal context-aware attention characterizes the answerer's multi-faceted expertise in terms of the questions' semantic and temporal information simultaneously. Moreover, the design of the multi-shift and multi-resolution module enables our model to handle temporal impact on different time granularities. Extensive experiments on six datasets from different domains demonstrate that the proposed model significantly outperforms competitive baseline models. Xuchao Zhang, Wei Cheng 0002, Bo Zong, Yuncong Chen, Jianwu Xu, Ding Li 0001 |
WSDM | 5 |
| 2018 | LogLens: A Real-Time Log Analysis SystemabstractAdministrators of most user-facing systems depend on periodic log data to get an idea of the health and status of production applications. Logs report information, which is crucial to diagnose the root cause of complex problems. In this paper, we present a real-time log analysis system called LogLens that automates the process of anomaly detection from logs with no (or minimal) target system knowledge and user specification. In LogLens, we employ unsupervised machine learning based techniques to discover patterns in application logs, and then leverage these patterns along with the real-time log parsing for designing advanced log analytics applications. Compared to the existing systems which are primarily limited to log indexing and search capabilities, LogLens presents an extensible system for supporting both stateless and stateful log analysis applications. Currently, LogLens is running at the core of a commercial log analysis solution handling millions of logs generated from the large-scale industrial environments and reported up to 12096x man-hours reduction in troubleshooting operational problems compared to the manual approach. Biplob Debnath, Mohiuddin Solaimani, Muhammad Ali Gulzar, Nipun Arora, Cristian Lumezanu, Jianwu Xu, Bo Zong, Hui Zhang 0002, Guofei Jiang, Latifur Khan |
ICDCS | 6 |
| 2016 | CloudSeer: Workflow Monitoring of Cloud Infrastructures via Interleaved LogsabstractCloud infrastructures provide a rich set of management tasks that operate computing, storage, and networking resources in the cloud. Monitoring the executions of these tasks is crucial for cloud providers to promptly find and understand problems that compromise cloud availability. However, such monitoring is challenging because there are multiple distributed service components involved in the executions. CloudSeer enables effective workflow monitoring. It takes a lightweight non-intrusive approach that purely works on interleaved logs widely existing in cloud infrastructures. CloudSeer first builds an automaton for the workflow of each management task based on normal executions, and then it checks log messages against a set of automata for workflow divergences in a streaming manner. Divergences found during the checking process indicate potential execution problems, which may or may not be accompanied by error log messages. For each potential problem, CloudSeer outputs necessary context information including the affected task automaton and related log messages hinting where the problem occurs to help further diagnosis. Our experiments on OpenStack, a popular open-source cloud infrastructure, show that CloudSeer's efficiency and problem-detection capability are suitable for online monitoring. Pallavi Joshi, Jianwu Xu, Guoliang Jin, Hui Zhang 0002, Guofei Jiang |
ASPLOS | 3 |
| 2016 | Automated IT system failure prediction: A deep learning approachabstractIn mission critical IT services, system failure prediction becomes increasingly important; it prevents unexpected system downtime, and assures service reliability for end users. While operational console logs record rich and descriptive information on the health status of those IT systems, existing system management technologies mostly use them in a labor-intensive forensics approach, i.e., identifying what went wrong after the fact. Recent efforts on log-based system management take an automation approach with text mining techniques, such as term frequency - inverse document frequency (TF-IDF). However, those techniques lead to a high-dimensional feature space, and are not easily generalizable to heterogeneous log formats. In this paper, we present a novel system that automatically parses streamed console logs and detects early warning signals for IT system failure prediction. In particular, our solution includes a log pattern extraction method by clustering together logs with similar format and content. We then resemble the TF-IDF idea by considering each pattern as a word and the set of patterns in each discretized epoch as a document. This leads to a feature space with significantly lower dimensionality that can provide robust signals for the status of the system. As system failures tend to occur very rare, we apply a recurrent neural network, namely, Long Short-Term Memory (LSTM), to deal with the “rarity” of labeled data in the training process. LSTM is able to capture the long-range dependency across sequences, therefore outperforms traditional supervised learning methods in our application domain. We evaluated and compared our proposed technology with state-of-the-art machine learning approaches using real log traces from two large enterprise systems. The results showed the advantage and potentials of our system in prediction of complex IT failures. To our knowledge, our work is the first that employs LSTM for log-based system failure prediction. Ke Zhang 0013, Jianwu Xu, Martin Renqiang Min, Guofei Jiang, Konstantinos Pelechrinis, Hui Zhang 0002 |
IEEE BigData | 2 |
| 2016 | LogMine: Fast Pattern Recognition for Log AnalyticsabstractModern engineering incorporates smart technologies in all aspects of our lives. Smart technologies are generating terabytes of log messages every day to report their status. It is crucial to analyze these log messages and present usable information (e.g. patterns) to administrators, so that they can manage and monitor these technologies. Patterns minimally represent large groups of log messages and enable the administrators to do further analysis, such as anomaly detection and event prediction. Although patterns exist commonly in automated log messages, recognizing them in massive set of log messages from heterogeneous sources without any prior information is a significant undertaking. We propose a method, named LogMine, that extracts high quality patterns for a given set of log messages. Our method is fast, memory efficient, accurate, and scalable. LogMine is implemented in map-reduce framework for distributed platforms to process millions of log messages in seconds. LogMine is a robust method that works for heterogeneous log messages generated in a wide variety of systems. Our method exploits algorithmic techniques to minimize the computational overhead based on the fact that log messages are always automatically generated. We evaluate the performance of LogMine on massive sets of log messages generated in industrial applications. LogMine has successfully generated patterns which are as good as the patterns generated by exact and unscalable method, while achieving a 500× speedup. Finally, we describe three applications of the patterns generated by LogMine in monitoring large scale industrial systems. Hossein Hamooni, Biplob Debnath, Jianwu Xu, Hui Zhang 0002, Guofei Jiang, Abdullah Mueen |
CIKM | 3 |
| 2014 | Max-AUC Feature Selection in Computer-Aided Detection of Polyps in CT ColonographyabstractWe propose a feature selection method based on a sequential forward floating selection (SFFS) procedure to improve the performance of a classifier in computerized detection of polyps in CT colonography (CTC). The feature selection method is coupled with a nonlinear support vector machine (SVM) classifier. Unlike the conventional linear method based on Wilks' lambda, the proposed method selected the most relevant features that would maximize the area under the receiver operating characteristic curve (AUC), which directly maximizes classification performance, evaluated based on AUC value, in the computer-aided detection (CADe) scheme. We presented two variants of the proposed method with different stopping criteria used in the SFFS procedure. The first variant searched all feature combinations allowed in the SFFS procedure and selected the subsets that maximize the AUC values. The second variant performed a statistical test at each step during the SFFS procedure, and it was terminated if the increase in the AUC value was not statistically significant. The advantage of the second variant is its lower computational cost. To test the performance of the proposed method, we compared it against the popular stepwise feature selection method based on Wilks' lambda for a colonic-polyp database (25 polyps and 2624 nonpolyps). We extracted 75 morphologic, gray-level-based, and texture features from the segmented lesion candidate regions. The two variants of the proposed feature selection method chose 29 and 7 features, respectively. Two SVM classifiers trained with these selected features yielded a 96% by-polyp sensitivity at false-positive (FP) rates of 4.1 and 6.5 per patient, respectively. Experiments showed a significant improvement in the performance of the classifier with the proposed feature selection method over that with the popular stepwise feature selection based on Wilks' lambda that yielded 18.0 FPs per patient at the same sensitivity level. Jianwu Xu, Kenji Suzuki 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2011 | A test of independence based on a generalized correlation function
Murali Rao, Sohan Seth, Jianwu Xu, Yunmei Chen, Hemant D. Tagare, José C. Príncipe |
Signal Process. | 3 |
| 2010 | Massive-Training Artificial Neural Network Coupled With Laplacian-Eigenfunction-Based Dimensionality Reduction for Computer-Aided Detection of Polyps in CT ColonographyabstractA major challenge in the current computer-aided detection (CAD) of polyps in CT colonography (CTC) is to reduce the number of false-positive (FP) detections while maintaining a high sensitivity level. A pattern-recognition technique based on the use of an artificial neural network (ANN) as a filter, which is called a massive-training ANN (MTANN), has been developed recently for this purpose. The MTANN is trained with a massive number of subvolumes extracted from input volumes together with the teaching volumes containing the distribution for the "likelihood of being a polyp;" hence the term "massive training." Because of the large number of subvolumes and the high dimensionality of voxels in each input subvolume, the training of an MTANN is time-consuming. In order to solve this time issue and make an MTANN work more efficiently, we propose here a dimension reduction method for an MTANN by using Laplacian eigenfunctions (LAPs), denoted as LAP-MTANN. Instead of input voxels, the LAP-MTANN uses the dependence structures of input voxels to compute the selected LAPs of the input voxels from each input subvolume and thus reduces the dimensions of the input vector to the MTANN. Our database consisted of 246 CTC datasets obtained from 123 patients, each of whom was scanned in both supine and prone positions. Seventeen patients had 29 polyps, 15 of which were 5-9 mm and 14 were 10-25 mm in size. We divided our database into a training set and a test set. The training set included 10 polyps in 10 patients and 20 negative patients. The test set had 93 patients including 19 polyps in seven patients and 86 negative patients. To investigate the basic properties of a LAP-MTANN, we trained the LAP-MTANN with actual polyps and a single source of FPs, which were rectal tubes. We applied the trained LAP-MTANN to simulated polyps and rectal tubes. The results showed that the performance of LAP-MTANNs with 20 LAPs was advantageous over that of the original MTANN with 171 inputs. To test the feasibility of the LAP-MTANN, we compared the LAP-MTANN with the original MTANN in the distinction between actual polyps and various types of FPs. The original MTANN yielded a 95% (18/19) by-polyp sensitivity at an FP rate of 3.6 (338/93) per patient, whereas the LAP-MTANN achieved a comparable performance, i.e., an FP rate of 3.9 (367/93) per patient at the same sensitivity level. With the use of the dimension reduction architecture, the time required for training was reduced from 38 h to 4 h. The classification performance in terms of the area under the receiver-operating-characteristic curve of the LAP-MTANN (0.84) was slightly higher than that of the original MTANN (0.82) with no statistically significant difference (p-value =0.48). Kenji Suzuki 0001, Jun Zhang 0091, Jianwu Xu |
IEEE Trans. Medical Imaging | 3 |
| 2009 | A Hierarchical Classification Model for Document CategorizationabstractWe propose a novel hierarchical classification method for documents categorization in this paper. The approach consists of multiple levels of classification for different hierarchies. Regularized Least Square (RLS)binary classifiers are applied in the middle levels of the hierarchy to classify documents into smaller set of categories and K-nearest-neighbor (KNN) multi-class classifiers are used at the bottom to classify documents into final classes. Experiments on large-scale real world tax documents show that the proposed hierarchical approach outperforms traditional flat classification method. Jianwu Xu, Vartika Singh, Venu Govindaraju, Depankar Neogi |
ICDAR | 1 |
| 2008 | A new nonlinear similarity measure for multichannel signals
Jianwu Xu, Hovagim Bakardjian, Andrzej Cichocki, José C. Príncipe |
Neural Networks | 1 |
| 2008 | A Pitch Detector Based on a Generalized Correlation FunctionabstractThis paper proposes a novel pitch determination algorithm (PDA) based on the newly introduced concept of a generalized correlation function called correntropy. Correntropy is a positive definite kernel function which implicitly transforms the original signal into a high-dimensional reproducing kernel Hilbert space (RKHS) in a nonlinear way, and calculates very efficiently the generalized correlation in that RKHS. By incorporating the kernel function, correntropy is able to utilize higher order statistics to enhance the resolution of pitch estimation. The proposed PDA computes the summary of correntropy functions from the outputs of an equivalent rectangular bandwidth (ERB) filter bank. We present simulations on pitch determination for a single vowel, double vowels, and a benchmark database test. Simulations show that correntropy exhibits much better resolution than conventional autocorrelation in pitch determination and outperforms other PDAs in the benchmark database test. Jianwu Xu, José C. Príncipe |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Automatic medical coding of patient records via weighted ridge regressionabstractIn this paper, we apply weighted ridge regression to tackle the highly unbalanced data issue in automatic large-scale ICD-9 coding of medical patient records. Since most of the ICD-9 codes are unevenly represented in the medical records, a weighted scheme is employed to balance positive and negative examples. The weights turn out to be associated with the instance priors from a probabilistic interpretation, and an efficient EM algorithm is developed to automatically update both the weights and the regularization parameter. Experiments on a large-scale real patient database suggest that the weighted ridge regression outperforms the conventional ridge regression and linear support vector machines (SVM). Jianwu Xu, Shipeng Yu, Jinbo Bi, Lucian Vlad Lita, Radu Stefan Niculescu, R. Bharat Rao |
ICMLA | 1 |
| 2007 | A New Nonlinear Similarity Measure for Multichannel Biological SignalsabstractWe propose a novel similarity measure, called the correntropy coefficient, sensitive to higher order moments of the signal statistics based on a similarity function called crosscorrentopy. Crossorrentropy nonlinearly maps the original time series into a high-dimensional reproducing kernel Hilbert space (RKHS). The correntropy coefficient computes the cosine of the angle between the transformed vectors. Preliminary experiments with simulated data and multichannel electroencephalogram (EEG) signals during behavior studies elucidate the performance of the new measure versus the well established correlation coefficient. Jianwu Xu, Hovagim Bakardjian, Andrzej Cichocki, José C. Príncipe |
IJCNN | 1 |
| 2006 | Kernel Based Synthetic Discriminant Function for Object RecognitionabstractIn this paper a non-linear extension to the synthetic discriminant function (SDF) is proposed. The SDF is a well known 2-D correlation filter for object recognition. The proposed non-linear version of the SDF is derived from kernel-based learning. The kernel SDF is implemented in a nonlinear high dimensional space by using the kernel trick and it can improve the performance of the linear SDF by incorporating the image's class higher order moments. We show that this kernelized composite correlation filter has an intrinsic connection with the recently proposed correntropy function. We apply this kernel SDF to face recognition and simulations show that the kernel SDF significantly outperforms the traditional SDF as well as is robust in noisy data environments. Kyu-Hwa Jeong, Puskal P. Pokharel, Jianwu Xu, Seungju Han 0001, José C. Príncipe |
ICASSP (5) | 3 |
| 2006 | A Closed Form Solution for a Nonlinear Wiener FilterabstractIn this paper a nonlinear extension to the Wiener filter is presented. A direct approach has been devised of replacing the autocorrelation function with a novel function called correntropy, derived from ideas on kernel-based learning theory and information theoretic learning. The linear Wiener filter, widely used because of its simplicity and optimality for linear systems and Gaussian distribution, is no longer effective when dealing with nonlinear time series data. The proposed method incorporates higher order moments in the general form of autocorrelation and improves upon the linear filter. Moreover, the computation cost is still lower than some kernel based methods and has a closed form solution to the problem unlike neural network based methods. Puskal P. Pokharel, Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
ICASSP (3) | 2 |
| 2006 | An Explicit Construction Of A Reproducing Gaussian Kernel Hilbert SpaceabstractIn this paper, we propose a method to explicitly construct a reproducing kernel Hilbert space (RKHS) associated with a Gaussian kernel by means of polynomial spaces. In contrast to the conventional Mercer's theorem approach that implicitly defines kernels by an eigendecomposition, the functionals in this reproducing kernel Hilbert space are explicitly constructed and are not necessary orthonormal. We also point out an intriguing connection between this reproducing kernel Hilbert space and a generalized Fock space. We give an experimental result on approximation of the constructed kernel to a Gaussian kernel. Jianwu Xu, Puskal P. Pokharel, Kyu-Hwa Jeong, José C. Príncipe |
ICASSP (5) | 1 |
| 2006 | Nonlinear Component Analysis Based on CorrentropyabstractIn this paper, we propose a new nonlinear prin- cipal component analysis based on a generalized correlation function which we call correntropy. The data is nonlinearly transformed to a feature space, and the principal directions are found by eigen-decomposition of the correntropy matrix, which has the same dimension as the standard covariance matrix for the original input data. The correntropy matrix characterizes the nonlinear correlations between the data. With the correntropy function, one can efficiently compute the principal components in the feature space by projecting the transformed data onto those principal directions. We give the derivation of the new method and present simulation results. Jianwu Xu, Puskal P. Pokharel, António R. C. Paiva, José C. Príncipe |
IJCNN | 1 |
| 2005 | An information-theoretic perspective to kernel independent components analysisabstractIn this paper, we investigate the intriguing relationship between information-theoretic learning (ITL), based on weighted Parzen window density estimator, and kernel-based learning algorithms. We prove the equivalence between kernel independent component analysis (kernel ICA) and the Cauchy-Schwartz (C-S) independence measure. This link gives a theoretical motivation for the selection of the Mercer kernel, based on density estimation. Demonstrating this equivalence requires introducing a weighted kernel density estimator, a modification of Parzen windowing. We also discuss the role of the weights in the weighted Parzen windowing and kernel ICA. Jianwu Xu, Deniz Erdogmus, Robert Jenssen, José C. Príncipe |
ICASSP (5) | 1 |
| 2005 | An information theoretic approach to adaptive system training using unlabeled dataabstractTraditionally, supervised learning is performed with pairwise input-output labelled data. After the training procedure, the adaptive system weights are fixed and the system is tested with unlabelled data. Recently, exploiting the unlabeled data to improve classification performance has been proposed in the machine learning community. In this paper, we present an information theoretic approach based on density divergence minimization to obtain an extended training algorithm using unlabeled data during testing. The simulations for classification problems suggest that our method can improve the performance of adaptive system in the application phase. Kyu-Hwa Jeong, Jianwu Xu, José C. Príncipe |
IJCNN | 2 |
| 2005 | A new classifier based on information theoretic learning with unlabeled data
Kyu-Hwa Jeong, Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
Neural Networks | 2 |
| 2004 | Minimizing Fisher information of the error in supervised adaptive filter trainingabstractIn this paper, we propose minimizing the Fisher information of the error in supervised training of linear and nonlinear adaptive filters. Fisher information considers the local structure of the error probability distribution and therefore it is a criterion that deserves to be investigated as an alternative to more common statistics such as minimum mean-square-error or minimum-error-entropy. A gradient-based training algorithm, based on a nonparametric estimator of Fisher information is presented and the performances of the three mentioned optimization criteria are compared using Monte Carlo simulations. Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
ICASSP (5) | 1 |