VLDB 2026 Research / reviewers in the wild / expert
James Zhang
dblp:58/2249
· DBLP profile ↗
20ranked-venue papers
7as first author
12since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LLMRG: Improving Recommendations through Large Language Model Reasoning GraphsabstractRecommendation systems aim to provide users with relevant suggestions, but often lack interpretability and fail to capture higher-level semantic relationships between user behaviors and profiles. In this paper, we propose a novel approach that leverages large language models (LLMs) to construct personalized reasoning graphs. These graphs link a user's profile and behavioral sequences through causal and logical inferences, representing the user's interests in an interpretable way. Our approach, LLM reasoning graphs (LLMRG), has four components: chained graph reasoning, divergent extension, self-verification and scoring, and knowledge base self-improvement. The resulting reasoning graph is encoded using graph neural networks, which serves as additional input to improve conventional recommender systems, without requiring extra user or item information. Our approach demonstrates how LLMs can enable more logical and interpretable recommender systems through personalized reasoning graphs. LLMRG allows recommendations to benefit from both engineered recommendation systems and LLM-derived reasoning graphs. We demonstrate the effectiveness of LLMRG on benchmarks and real-world scenarios in enhancing base recommendation models. Yan Wang 0002, Zhixuan Chu, Xin Ouyang, Simeng Wang, Hongyan Hao, Jinjie Gu, Siqiao Xue, James Zhang, Qing Cui, Jun Zhou 0011, Sheng Li 0001 |
AAAI | 9 |
| 2024 | Post-OCR Correction with OpenAI's GPT Models on Challenging English Prosody TextsabstractThe digitization of historical documents faces challenges with the accuracy of Optical Character Recognition (OCR). Noting the success of large language models (LLMs) on many text-based tasks, this paper explores the potential of OpenAI's GPT models (3.5-turbo, 4, 4-turbo) on the post-OCR correction task using works from the Princeton Prosody Archive (PPA), a full-text searchable database containing English texts published between 1559 and 1928 on versification and pronunciation. We conduct a comparative analysis across different model configurations and prompt strategies. Our results indicate that tailoring prompts with work metadata is less effective than anticipated, though adjusting the temperature parameter can be beneficial. The models tend to overcorrect works with already good OCR quality but perform well overall, with the best model setup improving the Character Error Rate (CER) by a mean of 18.92%. Additionally, after introducing a preliminary quality estimation step to process texts differently based on their original OCR quality, the best mean improvement increases to 38.83%. James Zhang, Wouter Haverals, Mary Naydan, Brian W. Kernighan |
DocEng | 1 |
| 2023 | Bellman Meets Hawkes: Model-Based Reinforcement Learning via Temporal Point ProcessesabstractWe consider a sequential decision making problem where the agent faces the environment characterized by the stochastic discrete events and seeks an optimal intervention policy such that its long-term reward is maximized. This problem exists ubiquitously in social media, finance and health informatics but is rarely investigated by the conventional research in reinforcement learning. To this end, we present a novel framework of the model-based reinforcement learning where the agent's actions and observations are asynchronous stochastic discrete events occurring in continuous-time. We model the dynamics of the environment by Hawkes process with external intervention control term and develop an algorithm to embed such process in the Bellman equation which guides the direction of the value gradient. We demonstrate the superiority of our method in both synthetic simulator and real-data experiments. Chao Qu, Xiaoyu Tan, Siqiao Xue, Xiaoming Shi 0001, James Zhang, Hongyuan Mei |
AAAI | 5 |
| 2023 | SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on HierarchiesabstractMultivariate time series forecasting with hierarchical structure is widely used in real-world applications, e.g., sales predictions for the geographical hierarchy formed by cities, states, and countries. The hierarchical time series (HTS) forecasting includes two sub-tasks, i.e., forecasting and reconciliation. In the previous works, hierarchical information is only integrated in the reconciliation step to maintain coherency, but not in forecasting step for accuracy improvement. In this paper, we propose two novel tree-based feature integration mechanisms, i.e., top-down convolution and bottom-up attention to leverage the information of the hierarchical structure to improve the forecasting performance. Moreover, unlike most previous reconciliation methods which either rely on strong assumptions or focus on coherent constraints only, we utilize deep neural optimization networks, which not only achieve coherency without any assumptions, but also allow more flexible and realistic constraints to achieve task-based targets, e.g., lower under-estimation penalty and meaningful decision-making loss to facilitate the subsequent downstream tasks. Experiments on real-world datasets demonstrate that our tree-based feature integration mechanism achieves superior performances on hierarchical forecasting tasks compared to the state-of-the-art methods, and our neural optimization networks can be applied to real-world tasks effectively without any additional effort under coherence and task-based constraints. Fan Zhou 0012, Lintao Ma, Yu Liu 0071, Shiyu Wang 0001, James Zhang, Xuanwei Hu, Yunhua Hu, Yangfei Zheng, Lei Lei 0001, Hu Yun |
AAAI | 6 |
| 2023 | Flow-Based End-to-End Model for Hierarchical Time Series Forecasting via Trainable Attentive-Reconciliation
Shiyu Wang 0001, Yinbo Sun, Yan Wang 0002, Fan Zhou 0012, Lintao Ma, James Zhang, Yangfei Zheng |
DASFAA (1) | 6 |
| 2023 | Full Scaling Automation for Sustainable Development of Green Data CentersabstractThe rapid rise in cloud computing has resulted in an alarming increase in data centers' carbon emissions, which now accounts for >3% of global greenhouse gas emissions, necessitating immediate steps to combat their mounting strain on the global climate. An important focus of this effort is to improve resource utilization in order to save electricity usage. Our proposed Full Scaling Automation (FSA) mechanism is an effective method of dynamically adapting resources to accommodate changing workloads in large-scale cloud computing clusters, enabling the clusters in data centers to maintain their desired CPU utilization target and thus improve energy efficiency. FSA harnesses the power of deep representation learning to accurately predict the future workload of each service and automatically stabilize the corresponding target CPU usage level, unlike the previous autoscaling methods, such as Autopilot or FIRM, that need to adjust computing resources with statistical models and expert knowledge. Our approach achieves significant performance improvement compared to the existing work in real-world datasets. We also deployed FSA on large-scale cloud computing clusters in industrial data centers, and according to the certification of the China Environmental United Certification Center (CEC), a reduction of 947 tons of carbon dioxide, equivalent to a saving of 1538,000 kWh of electricity, was achieved during the Double 11 shopping festival of 2022, marking a critical step for our company’s strategic goal towards carbon neutrality by 2030. Shiyu Wang 0001, Yinbo Sun, Xiaoming Shi 0001, Shiyi Zhu, Lintao Ma, James Zhang, Yangfei Zheng, Liu Jian |
IJCAI | 6 |
| 2023 | Prompt-augmented Temporal Point Process for Streaming Event SequenceabstractNeural Temporal Point Processes (TPPs) are the prevalent paradigm for modeling continuous-time event sequences, such as user activities on the web and financial transactions. In real world applications, the event data typically comes in a streaming manner, where the distribution of the patterns may shift over time. Under the privacy and memory constraints commonly seen in real scenarios, how to continuously monitor a TPP to learn the streaming event sequence is an important yet under-investigated problem. In this work, we approach this problem by adopting Continual Learning (CL), which aims to enable a model to continuously learn a sequence of tasks without catastrophic forgetting. While CL for event sequence is less well studied, we present a simple yet effective framework, PromptTPP, by integrating the base TPP with a continuous-time retrieval prompt pool. In our proposed framework, prompts are small learnable parameters, maintained in a memory space and jointly optimized with the base TPP so that the model is properly instructed to learn event streams arriving sequentially without buffering past examples or task-specific attributes. We formalize a novel and realistic experimental setup for modeling event streams, where PromptTPP consistently sets state-of-the-art performance across two real user behavior datasets. Siqiao Xue, Yan Wang 0002, Zhixuan Chu, Xiaoming Shi 0001, Caigao Jiang, Hongyan Hao, Gangwei Jiang, Xiaoyun Feng, James Zhang, Jun Zhou 0011 |
NeurIPS | 9 |
| 2022 | Memory Augmented State Space Model for Time Series ForecastingabstractState space model (SSM) provides a general and flexible forecasting framework for time series. Conventional SSM with fixed-order Markovian assumption often falls short in handling the long-range temporal dependencies and/or highly non-linear correlation in time-series data, which is crucial for accurate forecasting. To this extend, we present External Memory Augmented State Space Model (EMSSM) within the sequential Monte Carlo (SMC) framework. Unlike the common fixed-order Markovian SSM, our model features an external memory system, in which we store informative latent state experience, whereby to create ``memoryful" latent dynamics modeling complex long-term dependencies. Moreover, conditional normalizing flows are incorporated in our emission model, enabling the adaptation to a broad class of underlying data distributions. We further propose a Monte Carlo Objective that employs an efficient variational proposal distribution, which fuses the filtering and the dynamic prior information, to approximate the posterior state with proper particles. Our results demonstrate the competitiveness of forecasting performance of our proposed model comparing with other state-of-the-art SSMs. Yinbo Sun, Lintao Ma, Yu Liu 0071, James Zhang, Yangfei Zheng, Hu Yun, Lei Lei 0001, Yulin Kang, Llinbao Ye |
IJCAI | 5 |
| 2022 | A Meta Reinforcement Learning Approach for Predictive Autoscaling in the CloudabstractPredictive autoscaling (autoscaling with workload forecasting) is an important mechanism that supports autonomous adjustment of computing resources in accordance with fluctuating workload demands in the Cloud. In recent works, Reinforcement Learning (RL) has been introduced as a promising approach to learn the resource management policies to guide the scaling actions under the dynamic and uncertain cloud environment. However, RL methods face the following challenges in steering predictive autoscaling, such as lack of accuracy in decision-making, inefficient sampling and significant variability in workload patterns that may cause policies to fail at test time. To this end, we propose an end-to-end predictive meta model-based RL algorithm, aiming to optimally allocate resource to maintain a stable CPU utilization level, which incorporates a specially-designed deep periodic workload prediction model as the input and embeds the Neural Process [11, 16] to guide the learning of the optimal scaling actions over numerous application services in the Cloud. Our algorithm not only ensures the predictability and accuracy of the scaling strategy, but also enables the scaling decisions to adapt to the changing workloads with high sample efficiency. Our method has achieved significant performance improvement compared to the existing algorithms and has been deployed online at Alipay, supporting the autoscaling of applications for the world-leading payment platform. Siqiao Xue, Chao Qu, Xiaoming Shi 0001, Cong Liao, Shiyi Zhu, Xiaoyu Tan, Lintao Ma, Shiyu Wang 0001, Yun Hu 0001, Lei Lei 0001, Yangfei Zheng, James Zhang |
KDD | 14 |
| 2021 | A Graph Regularized Point Process Model For Event Propagation SequenceabstractPoint process is the dominant paradigm for modeling event sequences occurring at irregular intervals. In this paper we aim at modeling latent dynamics of event propagation in graph, where the event sequence propagates in a directed weighted graph whose nodes represent event marks (e.g., event types). Most existing works have only considered encoding sequential event history into event representation and ignored the information from the latent graph structure. Besides they also suffer from poor model explainability, i.e., failing to uncover causal influence across a wide variety of nodes. To address these problems, we propose a Graph Regularized Point Process (GRPP) that can be decomposed into: 1) a graph propagation model that characterizes the event interactions across nodes with neighbors and inductively learns node representations; 2) a temporal attentive intensity model, whose excitation and time decay factors of past events on the current event are constructed via the contextualization of the node embedding. Moreover, by applying a graph regularization method, GRPP provides model interpretability by uncovering influence strengths between nodes. Numerical experiments on various datasets show that GRPP outperforms existing models on both the propagation time and node prediction by notable margins. Siqiao Xue, Xiaoming Shi 0001, Hongyan Hao, Lintao Ma, James Zhang, Shiyu Wang 0001 |
IJCNN | 5 |
| 2021 | Scientific Collaboration Network Analysis for Computing Education ConferencesabstractThe computing education community is growing, but there is little information about the geographic distribution of the community or collaboration between members. Our research investigates three computer science education conferences (SIGCSE Technical Symposium, ITiCSE and ICER) by analysing authorship and affiliation details for publications in the proceedings over the lifetime of the respective conferences, totalling over 4500 publications. We examine the geographic location of authors and model the scientific collaboration network of each conference. We conclude that the community is open to newcomers, and both the number of authors, and the overall level of collaboration is growing. James Zhang, Andrew Luxton-Reilly, Paul Denny 0001, Jacqueline L. Whalley |
ITiCSE (1) | 1 |
| 2021 | Dynamic time warp-based clustering: Application of machine learning algorithms to simulation input modelling
James Zhang, Michael Johnstone, Vu Le 0001, Burhan Khan, Mohammad Anwar Hosen, Douglas C. Creighton, Jessica Carney, Andy Wilson, Michael Lynch |
Expert Syst. Appl. | 1 |
| 2020 | Robust Optimal Parameter Estimation (OPE) for Unsupervised Clustering of Spikes Using Neural NetworksabstractSpike sorting of electrophysiological data plays an important role in deciphering useful information from the brain. Unsupervised clustering of brain data relative to respective neurons is important to understand single cell and networks dynamics. A large number of clustering techniques exist in the literature; however, the dependency of these clustering algorithms on the selection of appropriate parameters, such as, bandwidth or threshold window size is critical. Iterative methods are generally employed to estimate optimal parameters, however, significant computational time and associated large number of iterations make the clustering inefficient to implement. To address this issue, we introduce a robust Optimal Parameter Estimation (OPE) Algorithm that can estimate the optimized parameters in a fast and efficient way. The performance of the OPE algorithm is tested on MeanShift and DBSCAN clustering algorithms. Three different extracellular recorded datasets including two simulated and one single human cell, as well as two feature sets including PCA and Haar Wavelets are used for validation purposes. Masood Ul Hassan, Rakesh Veerabhadrappa, James Zhang, Asim Bhatti |
SMC | 3 |
| 2019 | Anomaly detection in wide area network meshes using two machine learning algorithms
James Zhang, Robert W. Gardner, Ilija Vukotic |
Future Gener. Comput. Syst. | 1 |
| 2019 | A unified framework for interactive image segmentation via Fisher rules
Lingkun Luo, Shiqiang Hu, Xing Hu 0007, Huanlong Zhang, James Zhang |
Vis. Comput. | 7 |
| 2017 | Cohort analysis of simulation-based medical training for decision supportabstractDebriefing is the practice of after session review of training performance to enhance self-reflection through feedback. It has been considered as a vital and crucial part of simulation-based medical training. However an accurate, objective and in-depth evaluation of trainee performance has been a significant challenge. To address this we developed a knowledge-based framework in which the criteria of performance for clinical training are distilled into expert rules. These rules are then matched to data streams from a training session to evaluate the strengths and weaknesses of a trainee. We applied the evaluation technology to a dataset collected from two medical cohorts. The cohort characteristics are calculated, visualised, and validated by the medical experts. The cohort analysis results inform decision making at the levels of both the trainers and the enterprise. Trainers can compare and characterise the performance of different student cohorts or the same cohort over a period of time. The course coordinators can use the cohort analysis result to adjust the course design to target the identified common problems in trainee cohorts. James Zhang, Samer Hanoun, Burhan Khan, Douglas C. Creighton, Saeid Nahavandi, Kellie Britt, Karen D'Souza, Jon Watson, Richard Yanieri |
SMC | 1 |
| 2015 | A dynamic time warped clustering technique for discrete event simulation-based system analysis
Michael Johnstone, Vu Le 0001, James Zhang, Bruce Gunn, Saeid Nahavandi, Douglas C. Creighton |
Expert Syst. Appl. | 3 |
| 2014 | Knowledge-based automatic performance evaluation for medical training debriefingabstractManikin-based medical simulation has been shown to benefit the knowledge, skills and attitudes of the learner, and to impart favourable patient effects. A vital component of any training simulation is the after-session discussion with trainees to debrief their performance. In this study we develop a rule-based debriefing tool for improving the efficacy of medical training sessions. Unlike most existing de-briefing tools, the tool presented here has been designed to reduce medical trainer assessment time and to improve evaluation accuracy through a largely automated evaluation of trainee performance. The developed tool is acknowledged by the School of Medicine of Deakin University as an important advancement in assisting medical trainers carry out the debriefing process effectively and efficiently. James Zhang, Samer Hanoun, Douglas C. Creighton, Saeid Nahavandi, Karen D'Souza, Kellie Britt, Richard Yanieri |
SMC | 1 |
| 2012 | A generalised data analysis approach for baggage handling systems simulationabstractAirport baggage handling systems are a critical infrastructure component within major airports, and essential to ensure smooth luggage transfer while preventing dangerous material being loaded onto aircraft. This paper proposes a standard set of measures to assess the expected performance of a baggage handling system through discrete event simulation. These evaluation methods also have application in the study of general network systems. Results from the application of these methods reveal operational characteristics of the studied BHS, in terms of metrics such as peak throughput, in-system time and system recovery time. Vu Le 0001, James Zhang, Michael Johnstone, Saeid Nahavandi, Douglas C. Creighton |
SMC | 2 |
| 2008 | Toward a Synergy between Simulation and Knowledge Management for Business IntelligenceabstractThe adoption of simulation as a powerful enabling method for knowledge management is hampered by the relatively high cost of model construction and maintenance. A two-step procedure, based on a divide and conquer strategy, is proposed in this paper. First, a simulation program is partitioned based on a reinterpretation of the model-view-controller architecture. Individual parts are then connected, in terms of abstraction, to guard against possible changes that resulted from shifting user requirements. We explore the applicability of these design principles through a detailed discussion of an industry case study. The knowledge-based perspective guides the design of architecture to accommodate the need of emulation without compromising the integrity of the simulation program. The synergy between simulation and a knowledge management perspective, as shown in the case study, has the potential to achieve the objectives of rapid development of models, with low maintenance cost. This could, in turn, facilitate an extension of the use of simulation in the knowledge management domain. James Zhang, Douglas C. Creighton, Saeid Nahavandi |
Cybern. Syst. | 1 |