Jiajia Zhang 0001

dblp:08/267-1 · DBLP profile ↗
← Back
32ranked-venue papers
1as first author
26since 2021 · last 2026
0000-0001-6611-2046ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 PT-DCFR: Accelerating and Improving Deep CFR Using Population Based Training (Student Abstract)
abstract
Deep CFR enables end-to-end approximation of Nash equilibria in imperfect-information games(IIGs) but is sensitive to hyperparameters, making manual tuning inefficient. To address this, we propose PT-DCFR, which integrates Population-Based Training(PBT) with Deep CFR to dynamically optimize hyperparameters during training. Building upon this, we further introduce P2T-DCFR, which decouples parameter selection from model performance.
Dingzhong Cai, Huale Li, Shuhan Qi, Jiajia Zhang 0001
AAAI5
2026 NegLLM: Enhancing Strategic LLM Negotiation Agents with Case-Based Reasoning
Ruoke Wang, Tianzi Ma, Yulin Wu 0001, Jiajia Zhang 0001, Xuan Wang 0002
ICCBR4
2026 Solving equilibrium for adversarial team games utilizing fictitious team play with refined team plans
Jinheng Xiao, Chen Qiu 0003, Jiajia Zhang 0001, Shuhan Qi, Xuan Wang 0002
Expert Syst. Appl.4
2026 Distributed scalable multi-agent reinforcement learning with intrinsic-episodic dual exploration
Shuhan Qi, Shuhao Zhang 0009, Qiang Wang 0022, Jiajia Zhang 0001, Xuan Wang 0002
Future Gener. Comput. Syst.4
2026 Castor: Optimizing Deep Learning Job Scheduling in Multi-Tenant GPU Clusters via Intelligent Colocation
abstract
Deep learning (DL) has achieved significant success across a wide range of domains, prompting the widespread deployment of GPU clusters equipped with specialized accelerators to support high-performance training workloads. To minimize operational costs while maximizing resource utilization, efficient job scheduling in these clusters is essential. Although recent schedulers have improved cluster efficiency through periodic reallocation or selection of GPU resources, they still face challenges such as preemption and migration overheads, along with the risk of degrading model accuracy. Despite these limitations, the potential of GPUsharing remains largely underexplored. Few existing studies have systematically examined GPU sharing as a strategy to enhance resource utilization and reduce job queuing delays in multi-tenant DL clusters. Motivated by these insights, we propose a job scheduling model that enables multiple jobs to share the same set of GPUs without modifying their original training configurations. We introduce Castor, a simple yet efficient scheduling system, to achieve intelligent GPU colocation for multiple DL jobs. Castor intelligently selects job pairs for GPU sharing and determines runtime parameters (sub-batch size and scheduling time point) to optimize overall system performance while preserving the accuracy of DL convergence through gradient accumulation. Through a combination of physical DL workloads and trace-driven simulations across various configurations, we demonstrate that Castor reduces average job completion time by 26–52% compared to state-of-the-art preemptive DL schedulers, despite operating under a preemption-free policy. Furthermore, Castor effectively identifies optimal resource-sharing configurations, outperforming the baseline first-fit sharing policy (SJF-FFS) by up to 20% on large-scale workload traces.
Yizhou Luo, Jiaxin Lai, Shaohuai Shi, Chen Chen 0067, Shuhan Qi, Jiajia Zhang 0001, Qiang Wang 0022
IEEE Trans. Cloud Comput.6
2025 Towards Building Human-like Smart Agents in Modern 3D Video Games (Student Abstract)
abstract
In recent years, reinforcement learning has been widely applied in the field of games. However, most studies focus on assisting agents to achieve victory, with less attention paid to whether the agents exhibit human-like characteristics. In order to build human-like agents with high performance, we propose a method for learning the strategies of human players in modern three-dimensional video games. Our method utilizes a hierarchical framework, learning basic behaviors and intentions of human players at the lower level through imitation learning, and generalized policies at the high level through reinforcement learning. Compared with other existing methods, our method demonstrates significant advantages in learning human-like strategies in complex environments.
Zhihang Sun, Shuhan Qi, Xinhao Huang, Xinyu Xiao, Jiajia Zhang 0001, Xuan Wang 0002, Peixi Peng
AAAI5
2025 Last-Iterate Convergence in Adaptive Regret Minimization for Approximate Extensive-Form Perfect Equilibrium
abstract
The Nash Equilibrium (NE) assumes rational play in imperfect-information Extensive-Form Games (EFGs) but fails to ensure optimal strategies for off-equilibrium branches of the game tree, potentially leading to suboptimal outcomes in practical settings. To address this, the Extensive-Form Perfect Equilibrium (EFPE), a refinement of NE, introduces controlled perturbations to model potential player errors. However, existing EFPE-finding algorithms, which typically rely on average strategy convergence and fixed perturbations, face significant limitations: computing average strategies incurs high computational costs and approximation errors, while fixed perturbations create a trade-off between NE approximation accuracy and the convergence rate of NE refinements. To tackle these challenges, we propose an efficient adaptive regret minimization algorithm for computing approximate EFPE, achieving last-iterate convergence in two-player zero-sum EFGs. Our approach introduces Reward Transformation Counterfactual Regret Minimization (RTCFR) to solve perturbed games and defines a novel metric, the Information Set Nash Equilibrium (ISNE), to dynamically adjust perturbations. Theoretical analysis confirms convergence to EFPE, and experimental results demonstrate that our method significantly outperforms state-of-the-art algorithms in both NE and EFPE-finding tasks.
Hang Ren 0002, Xiaozhen Sun, Tianzi Ma, Jiajia Zhang 0001, Xuan Wang 0002
ECAI4
2025 FGLight: Learning Neighbor-level Information for Traffic Signal Control
Huale Li, Shuhan Qi, Jiajia Zhang 0001, Dingzhong Cai
AAMAS4
2025 Optimizing strategy selection in hidden role games
Chen Qiu 0003, Jinheng Xiao, Jiajia Zhang 0001, Shuhan Qi, Xuan Wang 0002
Eng. Appl. Artif. Intell.4
2024 Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
abstract
Deep learning (DL) has demonstrated significant success across diverse fields, leading to the construction of dedicated GPU accelerators within GPU clusters for high-quality training services. Efficient scheduler designs for such clusters are vital to reduce operational costs and enhance resource utilization. While recent schedulers have shown impressive performance in optimizing DL job performance and cluster utilization through periodic reallocation or selection of GPU resources, they also encounter challenges such as preemption and migration overhead, along with potential DL accuracy degradation. Nonetheless, few explore the potential benefits of GPU sharing to improve resource utilization and reduce job queuing times.Motivated by these insights, we present a job scheduling model allowing multiple jobs to share the same set of GPUs without altering job training settings. We introduce SJF-BSBF (shortest job first with best sharing benefit first), a straightforward yet effective heuristic scheduling algorithm. SJF-BSBF intelligently selects job pairs for GPU resource sharing and runtime settings (sub-batch size and scheduling time point) to optimize overall performance while ensuring DL convergence accuracy through gradient accumulation. In experiments with both physical DL workloads and trace-driven simulations, even as a preemptionfree policy, SJF-BSBF reduces the average job completion time by 27-33% relative to the state-of-the-art preemptive DL schedulers. Moreover, SJF-BSBF can wisely determine the optimal resource sharing settings, such as the sharing time point and sub-batch size for gradient accumulation, outperforming the aggressive GPU sharing approach (baseline SJF-FFS policy) by up to 17% in large-scale traces.
Yizhou Luo, Qiang Wang 0022, Shaohuai Shi, Jiaxin Lai, Shuhan Qi, Jiajia Zhang 0001, Xuan Wang 0002
IWQoS6
2024 Improving Real-Time Service Quality Through Parallel Monte Carlo Tree Search
abstract
In this paper, we present a new algorithm named Distributed Multi-root Pipeline MCTS (DMP-MCTS), to improve real-time search efficiency in multi-machine scenarios. By utilizing root parallel technology with significance detection and pipeline pattern for parallel MCTS (3PMCTS), we not only reduce the repetitive computing tasks among subtrees but also allow for flexible operation time based resource allocation. The experiment shows this work achieves better computational performance under linear acceleration conditions, compared with other existing works.
Jiajia Zhang 0001, Shuman Zhuang, Yulin Wu 0001
IWQoS1
2024 Combining Counterfactual Regret Minimization With Information Gain to Solve Extensive Games With Unknown Environments
abstract
Counterfactual regret minimization (CFR) is an effective algorithm for solving extensive‐form games with imperfect information (IIEGs). However, CFR is only allowed to be applied in known environments, where the transition function of the chance player and the reward function of the terminal node in IIEGs are known. In uncertain situations, such as reinforcement learning (RL) problems, CFR is not applicable. Thus, applying CFR in unknown environments is a significant challenge that can also address some difficulties in the real world. Currently, advanced solutions require more interactions with the environment and are limited by large single‐sampling variances to narrow the gap with the real environment. In this paper, we propose a method that combines CFR with information gain to compute the Nash equilibrium (NE) of IIEGs with unknown environments. We use a curiosity‐driven approach to explore unknown environments and minimize the discrepancy between uncertain and real environments. In addition, by incorporating information into the reward, the average strategy calculated by CFR can be directly implemented as the interaction policy with the environment, thereby improving the exploration efficiency of our method in uncertain environments. Through experiments on standard testbeds such as Kuhn poker and Leduc poker, our method significantly reduces the number of interactions with the environment compared to the different baselines and computes a more accurate approximate NE within the same number of interaction rounds.
Chen Qiu 0003, Xuan Wang 0002, Tianzi Ma, Yaojun Wen, Jiajia Zhang 0001
Int. J. Intell. Syst.5
2024 A fast strategy-solving method for adversarial team games utilizing warm starting
Chen Qiu 0003, Jiajia Zhang 0001, Xuan Wang 0002
Neurocomputing4
2024 A Novel Tree-Based Method for Interpretable Reinforcement Learning
abstract
Deep reinforcement learning (DRL) has garnered remarkable success across various domains, propelled by advancements in deep learning (DL) technologies. However, the opacity of DL presents significant challenges, limiting the application of DRL in critical systems. In response, decision tree (DT)-based methods, known for their transparent decision-making mechanisms, have shown promise in making interpretable policies for decision-making problems. Existing methods often employ differential DTs to model RL policies and discretize them to conventional DTs for higher interpretability. Yet, this method leads to discrepancies between the trained differential DTs and the discretized DTs. To address this issue, we introduce Generative Consistent Trees (GCTs), a novel solution that circumvents the information loss typically associated with the argmax operation in prior research. By implementing a reparameterization technique to approximate the categorical distribution, GCTs ensure the consistencies between trained GCTs and discretized counterparts. Moreover, we have developed an imitation-learning-based framework for interpretable reinforcement learning. This framework is designed to train GCTs by efficiently mimicking expert policies. Our extensive experiments across multiple environments have validated the effectiveness of this approach, highlighting the potential of GCTs in enhancing the interpretability and applicability of DRL.
Shuhan Qi, Xuan Wang 0002, Jiajia Zhang 0001
ACM Trans. Knowl. Discov. Data4
2024 D2CFR: Minimize Counterfactual Regret With Deep Dueling Neural Network
abstract
Counterfactual regret minimization (CFR) is a popular method for finding approximate Nash equilibrium in two-player zero-sum games with imperfect information. Solving large-scale games with CFR needs a combination of abstraction techniques and certain expert knowledge, which constrains its scalability. Recent neural-based CFR methods mitigate the need for abstraction and expert knowledge by training an efficient network to directly obtain counterfactual regret without abstraction. However, these methods only consider estimating regret values for individual actions, neglecting the evaluation of state values, which are significant for decision-making. In this article, we introduce deep dueling CFR (D2CFR), which emphasizes the state value estimation by employing a novel value network with a dueling structure. Moreover, a rectification module based on a time-shifted Monte Carlo simulation is designed to rectify the inaccurate state value estimation. Extensive experimental results are conducted to show that D2CFR converges faster and outperforms comparison methods on test games.
Huale Li, Xuan Wang 0002, Zengyue Guo, Jiajia Zhang 0001, Shuhan Qi
IEEE Trans. Neural Networks Learn. Syst.4
2024 Cascaded Attention: Adaptive and Gated Graph Attention Network for Multiagent Reinforcement Learning
abstract
Modeling the interactive relationships of agents is critical to improving the collaborative capability of a multiagent system. Some methods model these by predefined rules. However, due to the nonstationary problem, the interactive relationship changes over time and cannot be well captured by rules. Other methods adopt a simple mechanism such as an attention network to select the neighbors the current agent should collaborate with. However, in large-scale multiagent systems, collaborative relationships are too complicated to be described by a simple attention network. We propose an adaptive and gated graph attention network (AGGAT), which models the interactive relationships between agents in a cascaded manner. In the AGGAT, we first propose a graph-based hard attention network that roughly filters irrelevant agents. Then, normal soft attention is adopted to decide the importance of each neighbor. Finally, gated attention further refines the collaborative relationship of agents. By using cascaded attention, the collaborative relationship of agents is precisely learned in a coarse-to-fine style. Extensive experiments are conducted on a variety of cooperative tasks. The results indicate that our proposed method outperforms state-of-the-art baselines.
Shuhan Qi, Xinhao Huang, Peixi Peng, Xuzhong Huang, Jiajia Zhang 0001, Xuan Wang 0002
IEEE Trans. Neural Networks Learn. Syst.5
2023 Kdb-D2CFR: Solving Multiplayer imperfect-information games with knowledge distillation-based DeepCFR
Huale Li, Zengyue Guo, Yang Liu 0039, Xuan Wang 0002, Shuhan Qi, Jiajia Zhang 0001, Jing Xiao 0006
Knowl. Based Syst.6
2022 RLCFR: Minimize counterfactual regret by deep reinforcement learning
Huale Li, Xuan Wang 0002, Fengwei Jia, Yulin Wu 0001, Jiajia Zhang 0001, Shuhan Qi
Expert Syst. Appl.5
2022 Finding top-K solutions for the decision-maker in multiobjective optimization
Wenjian Luo, Luming Shi, Xin Lin 0004, Jiajia Zhang 0001, Miqing Li, Xin Yao 0001
Inf. Sci.4
2022 Parallel learner: A practical deep reinforcement learning framework for multi-scenario games
Xiaohan Hou, Zhenyang Guo, Xuan Wang 0002, Tao Qian 0003, Jiajia Zhang 0001, Shuhan Qi, Jing Xiao 0006
Knowl. Based Syst.5
2022 3D face recognition algorithm based on nose tip contour and radial curve
Linlin Tang, Zhangyan Li, Yang Liu 0039, Shuhan Qi, Jiajia Zhang 0001, Jiancheng Pan, Shuaijie Shi
Multim. Tools Appl.5
2021 A Survey of Nearest-Better Clustering in Swarm and Evolutionary Computation
abstract
Nearest-Better Clustering (NBC) is an emergent niching technique in Swarm and Evolutionary Computation for optimization, which does not need to fix the number or radius of clusters in advance. The key idea of NBC is to first link each individual to its nearest better neighbor to form a spanning tree of all individuals in the population, and then partition all individuals into clusters by deleting the longer edges in the spanning tree. In this paper, a survey on the Nearest-Better Clustering algorithms and applications in multimodal and dynamic optimization is provided. First, the basic NBC algorithm is introduced. Second, the improvements of the basic NBC are detailed. Third, multimodal and dynamic optimization algorithms powered by NBC are enlisted and discussed.
Wenjian Luo, Xin Lin 0004, Jiajia Zhang 0001, Mike Preuss
CEC3
2021 Hiding All Labels for Multi-label Images: An Empirical Study of Adversarial Examples
abstract
Adversarial examples about deep learning models have been paid much attention in recent years, including single-label adversarial examples and multi-label adversarial examples. In this paper, for the first time, an empirical study of generating a multi-label adversarial example to hide all labels in a multi-label example is presented. The objective of hiding all labels in a multi-label example is to make deep learning models know nothing about the environments. That is very worthy of studying because deep learning models will say there is nothing, although the real input has more than one label. In the empirical study, we use five state-of-the-art multi-label attack algorithms, i.e., ML-CW, ML-DP, FGSM, MI-FGSM, MLA-LP, four popular datasets, i.e., VOC2007, VOC2012, NUS-WIDE and COCO, and two typical models ML-GCN and ASL for evaluation. We conduct extensive experiments and report the attack success rates, the amount of perturbations of the adversarial examples generated by state-of-the-art multi-label attack algorithms. We also report the attack performance when a typical defending algorithm based on JPEG compression is used. The work in this paper is beneficial to the future study of generating multi-label adversarial examples as well as defending them.
Wenjian Luo, Jiajia Zhang 0001, Linghao Kong
IJCNN3
2021 Scalable sub-game solving for imperfect-information games
Huale Li, Xuan Wang 0002, Kunchi Li, Fengwei Jia, Yulin Wu 0001, Jiajia Zhang 0001, Shuhan Qi
Knowl. Based Syst.6
2021 Autoencoder-based self-supervised hashing for cross-modal retrieval
Xuan Wang 0002, Jiajia Zhang 0001, Chengkai Huang, Shuhan Qi
Multim. Tools Appl.4
2021 Constraint-Objective Cooperative Coevolution for Large-scale Constrained Optimization
abstract
Large-scale optimization problems and constrained optimization problems have attracted considerable attention in the swarm and evolutionary intelligence communities and exemplify two common features of real problems, i.e., a large scale and constraint limitations. However, only a little work on solving large-scale continuous constrained optimization problems exists. Moreover, the types of benchmarks proposed for large-scale continuous constrained optimization algorithms are not comprehensive at present. In this article, first, a constraint-objective cooperative coevolution (COCC) framework is proposed for large-scale continuous constrained optimization problems, which is based on the dual nature of the objective and constraint functions: modular and imbalanced components. The COCC framework allocates the computing resources to different components according to the impact of objective values and constraint violations. Second, a benchmark for large-scale continuous constrained optimization is presented, which takes into account the modular nature, as well as both imbalanced and overlapping characteristics of components. Finally, three different evolutionary algorithms are embedded into the COCC framework for experiments, and the experimental results show that COCC performs competitively.
Peilan Xu, Wenjian Luo, Xin Lin 0004, Jiajia Zhang 0001, Yingying Qiao, Xuan Wang 0002
ACM Trans. Evol. Learn. Optim.4
2020 Ransomware classification using patch-based CNN and self-attention network on embedded N-grams of opcodes
Bin Zhang 0048, Wentao Xiao, Xi Xiao 0001, Arun Kumar Sangaiah, Weizhe Zhang, Jiajia Zhang 0001
Future Gener. Comput. Syst.6
2020 Bi-Connect Net for salient object detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Qing Liao 0001, Jiajia Zhang 0001, Huale Li, Shuhan Qi
Neurocomputing5
2020 Explore instance similarity: An instance correlation based hashing method for multi-label cross-model retrieval
Chengkai Huang, Jiajia Zhang 0001, Qing Liao 0001, Xuan Wang 0002, Zoe Lin Jiang, Shuhan Qi
Inf. Process. Manag.3
2019 Solving Six-Player Games via Online Situation Estimation
abstract
While the artificial intelligence theory for solving the perfect-information games has been well developed in recent years, great challenges are still posed in dealing with the imperfect-information game due to the huge state space and hidden information involved in it. In this paper, we design an online strategy solving framework for six-player no-limit Texas hold'em poker. Based on hand isomorphism and hand strength evalution, the framework provides an efficient situation estimation method for six-player poker. Such method could greatly reduce the the state space in six-player poker as well as effectively evaluate the current hands. The poker agent based on our method won the third place in the 2018 AAAI-ACPC.
Huale Li, Xuan Wang 0002, Shuhan Qi, Yang Liu 0039, Fengwei Jia, Jiajia Zhang 0001
ICTAI7
2019 Graph-based supervised discrete image hashing
Jian Guan 0001, Xuan Wang 0002, Hainan Zhao, Jiajia Zhang 0001, Zechao Liu, Shuhan Qi
J. Vis. Commun. Image Represent.6
2009 The Improvement of Q-learning Applied to Imperfect Information Game
abstract
There exist problems of slow convergence and local optimum in standard Q-learning algorithm. Truncated TD estimate returns efficiency and simulated annealing algorithm increase the chance of exploration. To accelerate the algorithm convergence speed and to avoid results in local optimum, this paper combines Q-learning algorithm, truncated TD estimation and simulated annealing algorithm. We apply improved Q-learning algorithm using into the imperfect information game (SiGuo military chess game), and realize a self-learning of imperfect information game system. Experimental outcomes show that this system can dynamically adjust each weight which describes game state according to the results. Further, it speeds up the process of learning, effectively simulates human intelligence and makes reasonable step, and significantly improves system performance.
Xuan Wang 0002, Lijiao Han, Jiajia Zhang 0001
SMC4