Ziwei Deng

dblp:203/3237 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-3989-1424ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Planning, search and constraint satisfaction · 30% Reinforcement learning · 28% Video understanding and tracking · 18%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.622025
DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning · ICLR 2025
Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning · ICLR 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
offline multi-agent reinforcement learning
0.912025
DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning · ICLR 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
planning under uncertainty
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning › markov games
mixed-motive games
0.812024
Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence · NeurIPS 2024
Computer vision › Video understanding and tracking
action recognition
0.412020
Cycle-Contrast for Self-Supervised Video Representation Learning · NeurIPS 2020
Machine learning › Representation and self-supervised learning
contrastive learning
0.412020
Cycle-Contrast for Self-Supervised Video Representation Learning · NeurIPS 2020
Computer vision › Video understanding and tracking › video representation learning
self-supervised video representation learning
0.412020
Cycle-Contrast for Self-Supervised Video Representation Learning · NeurIPS 2020
Computer vision › Video understanding and tracking
video representation learning
0.412020
Cycle-Contrast for Self-Supervised Video Representation Learning · NeurIPS 2020
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
cross-modal distillation
0.412019
MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding · ICCV 2019
Computer vision › Video understanding and tracking › action recognition
human action recognition
0.412019
MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding · ICCV 2019
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.412019
MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding · ICCV 2019
Computer vision › Video understanding and tracking › action recognition
multimodal action recognition
0.412019
MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding · ICCV 2019
Information retrieval
similarity search
0.112020
Cycle-Contrast for Self-Supervised Video Representation Learning · NeurIPS 2020
Information retrieval › multimedia analysis and retrieval
video retrieval
0.112020
Cycle-Contrast for Self-Supervised Video Representation Learning · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

upper confidence bounds with curiosity · 0.9quantile distribution · 0.9noise factorization · 0.9monte carlo tree search · 0.9diffusion model · 0.9r3d · 0.9cycle-contrastive loss · 0.9contrastive learning · 0.9multi-agent reinforcement learning · 0.8attention mechanism · 0.8knowledge distillation · 0.4
YearPublicationVenuePosition
2025 DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning
abstract
Diffusion models have been widely adopted in image and language generation and are now being applied to reinforcement learning. However, the application of diffusion models in offline cooperative Multi-Agent Reinforcement Learning (MARL) remains limited. Although existing studies explore this direction, they suffer from scalability or poor cooperation issues due to the lack of design principles for diffusion-based MARL. The Individual-Global-Max (IGM) principle is a popular design principle for cooperative MARL. By satisfying this principle, MARL algorithms achieve remarkable performance with good scalability. In this work, we extend the IGM principle to the Individual-Global-identically-Distributed (IGD) principle. This principle stipulates that the generated outcome of a multi-agent diffusion model should be identically distributed as the collective outcomes from multiple individual-agent diffusion models. We propose DoF, a diffusion factorization framework for Offline MARL. It uses noise factorization function to factorize a centralized diffusion model into multiple diffusion models. We theoretically show that the noise factorization functions satisfy the IGD principle. Furthermore, DoF uses data factorization function to model the complex relationship among data generated by multiple diffusion models. Through extensive experiments, we demonstrate the effectiveness of DoF. The source code is available at [https://github.com/xmu-rl-3dv/DoF](https://github.com/xmu-rl-3dv/DoF).
Ziwei Deng, Chenxing Lin, Yongquan Fu, Weiquan Liu, Chenglu Wen, Cheng Wang 0003
ICLR2
2025 PlanU: Large Language Model Reasoning through Planning under Uncertainty
abstract
Large Language Models (LLMs) are increasingly being explored across a range of reasoning tasks. However, LLMs sometimes struggle with reasoning tasks under uncertainty that are relatively easy for humans, such as planning actions in stochastic environments. The adoption of LLMs for reasoning is impeded by uncertainty challenges, such as LLM uncertainty and environmental uncertainty. LLM uncertainty arises from the stochastic sampling process inherent to LLMs. Most LLM-based Decision-Making (LDM) approaches address LLM uncertainty through multiple reasoning chains or search trees. However, these approaches overlook environmental uncertainty, which leads to poor performance in environments with stochastic state transitions. Some recent LDM approaches deal with uncertainty by forecasting the probability of unknown variables. However, they are not designed for multi-step reasoning tasks that require interaction with the environment. To address uncertainty in LLM decision-making, we introduce PlanU, an LLM-based planning method that captures uncertainty within Monte Carlo Tree Search (MCTS). PlanU models the return of each node in the MCTS as a quantile distribution, which uses a set of quantiles to represent the return distribution. To balance exploration and exploitation during tree search, PlanU introduces an Upper Confidence Bounds with Curiosity (UCC) score which estimates the uncertainty of MCTS nodes. Through extensive experiments, we demonstrate the effectiveness of PlanU in LLM-based reasoning tasks under uncertainty.
Ziwei Deng, Mian Deng, Chenjing Liang, Zeming Gao, Chennan Ma, Chenxing Lin, Songzhu Mei, Cheng Wang 0003
NeurIPS1
2025 Pure profit-oriented continuous influence maximization considering cost budget: A gradient descent-based approach
Ziwei Deng, Ling Chen 0005
Inf. Process. Manag.2
2025 An Aeromagnetic Compensation Method Based on Differentiable Architecture Search-Guided Physics-Informed Neural Network
abstract
The aeromagnetic compensation method is critical for mitigating magnetic interference in airborne geophysical surveys. In the case of complex magnetic interference, the model-driven linear regression method based on the Toles-Lawson (T-L) model cannot guarantee compensation accuracy. Additionally, purely data-driven methods often require large datasets and lack interpretability. Recent hybrid solutions, particularly Physics-Informed Neural Network (PINN), effectively combine the advantages of model-driven and data-driven methods. However, PINN is heavily reliant on manual architecture tuning, which severely limits optimization efficiency and compensation accuracy. To address this issue, this letter proposes a Differentiable Architecture Search-Guided Physics-Informed Neural Network (DARTS-PINN) method. Experimental results demonstrate that DARTS-PINN can effectively mitigate magnetic interference, improve optimization efficiency, and substantially reduce data dependence. For aeromagnetic compensation of magnetic sensors fixed inside the cabin, the root mean square error (RMSE) of DARTS-PINN is 1.089 nT. In contrast, the optimal RMSEs of the model-driven and data-driven methods are 6.195 nT and 2.026 nT, respectively.
Zhijian Jiang, Taoran Zhao, Menglei Wang, Junfeng Zhou, Ziwei Deng, Xinhua Lin
IEEE Geosci. Remote. Sens. Lett.5
2025 A Greedy Descent Method for Budget Constrained Continuous Influence Maximization in Online Social Network
abstract
Continuous influence maximization (CIM) in social networks aims to maximize the expected influence spreading by assigning each user a continuous weight reflecting the likelihood and cost for him becoming a seed. Traditional CIM assumes a budget constrain to limit the total cost of all users. However, this assumption does not tenable in practical applications. In practice, it is not necessary to incur costs for all the users. Instead, the budget should be set only for the cost associated with the seed set. In this article, an extended CIM problem of cost distribution under budget (CDB) is defined, which aims to assign different costs to the customers according to their ability to spread influence, ensuring that the cost of each potential seed set does not exceed the budget, while the expected spreading of the product’s influence is maximized. The NP-hardness of CDB and the monotonicity and submodularity of its objective function are investigated. We formulate the CDB problem into a constrained optimization, and present a greedy descent-based algorithm for the problem. In each iteration of the greedy descent method, the influence increment of each node is calculated according to its estimated influence spreading range. The cost distribution is updated along the direction with the maximum increment. The optimal cost distribution can be obtained after several iterations. Precision of the results obtained by the proposed algorithm is analyzed. To avoid the time-consuming simulations, we design an effective algorithm for estimating the seed influence spreading range. Experiment results on real and synthetic networks show that the proposed algorithm can significantly improve the expected influence spreading.
Wei Liu 0010, Ziwei Deng, Yixin Chen 0001, Ling Chen 0005
IEEE Trans. Comput. Soc. Syst.2
2024 Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence
abstract
Multi-agent AI research promises a path to develop human-like and human-compatible intelligent technologies that complement the solipsistic view of other approaches, which mostly do not consider interactions between agents. Aiming to make progress in this direction, the Melting Pot contest 2023 focused on the problem of cooperation among interacting agents and challenged researchers to push the boundaries of multi-agent reinforcement learning (MARL) for mixed-motive games. The contest leveraged the Melting Pot environment suite to rigorously evaluate how well agents can adapt their cooperative skills to interact with novel partners in unforeseen situations. Unlike other reinforcement learning challenges, this challenge focused on social rather than environmental generalization. In particular, a population of agents performs well in Melting Pot when its component individuals are adept at finding ways to cooperate both with others in their population and with strangers. Thus Melting Pot measures cooperative intelligence.The contest attracted over 600 participants across 100+ teams globally and was a success on multiple fronts: (i) it contributed to our goal of pushing the frontiers of MARL towards building more cooperatively intelligent agents, evidenced by several submissions that outperformed established baselines; (ii) it attracted a diverse range of participants, from independent researchers to industry affiliates and academic labs, both with strong background and new interest in the area alike, broadening the field’s demographic and intellectual diversity; and (iii) analyzing the submitted agents provided important insights, highlighting areas for improvement in evaluating agents' cooperative intelligence. This paper summarizes the design aspects and results of the contest and explores the potential of Melting Pot as a benchmark for studying Cooperative AI. We further analyze the top solutions and conclude with a discussion on promising directions for future research.
Rakshit S. Trivedi, Akbir Khan, Jesse Clifton, Lewis Hammond, Edgar A. Duéñez-Guzmán, Dipam Chakraborty, John P. Agapiou, Jayd Matyas, Alexander Vezhnevets, Barna Pásztor, Yunke Ao, Omar G. Younis, Benjamin Swain, Haoyuan Qin, Mian Deng, Ziwei Deng, Utku Erdoganaras, Yue Zhao 0023, Marko Tesic, Natasha Jaques, Jakob N. Foerster, Vincent Conitzer, José Hernández-Orallo, Dylan Hadfield-Menell, Joel Z. Leibo
NeurIPS17
2022 Hierarchical contrastive adaptation for cross-domain object detection
Ziwei Deng, Quan Kong, Naoto Akira, Tomoaki Yoshinaga
Mach. Vis. Appl.1
2020 Cycle-Contrast for Self-Supervised Video Representation Learning
abstract
We present Cycle-Contrastive Learning (CCL), a novel self-supervised method for learning video representation. Following a nature that there is a belong and inclusion relation of video and its frames, CCL is designed to find correspondences across frames and videos considering the contrastive representation in their domains respectively. It is different from recent approaches that merely learn correspondences across frames or clips. In our method, the frame and video representations are learned from a single network based on an R3D network, with a shared non-linear transformation for embedding both frame and video features before the cycle-contrastive loss. We demonstrate that the video representation learned by CCL can be transferred well to downstream tasks of video understanding, outperforming previous methods in nearest neighbour retrieval and action recognition tasks on UCF101, HMDB51 and MMAct.
Quan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga, Tomokazu Murakami
NeurIPS3
2019 MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding
abstract
Unlike vision modalities, body-worn sensors or passive sensing can avoid the failure of action understanding in vision related challenges, e.g. occlusion and appearance variation. However, a standard large-scale dataset does not exist, in which different types of modalities across vision and sensors are integrated. To address the disadvantage of vision-based modalities and push towards multi/cross modal action understanding, this paper introduces a new large-scale dataset recorded from 20 distinct subjects with seven different types of modalities: RGB videos, keypoints, acceleration, gyroscope, orientation, Wi-Fi and pressure signal. The dataset consists of more than 36k video clips for 37 action classes covering a wide range of daily life activities such as desktop-related and check-in-based ones in four different distinct scenarios. On the basis of our dataset, we propose a novel multi modality distillation model with attention mechanism to realize an adaptive knowledge transfer from sensor-based modalities to vision-based modalities. The proposed model significantly improves performance of action recognition compared to models trained with only RGB information. The experimental results confirm the effectiveness of our model on cross-subject, -view, -scene and -session evaluation criteria. We believe that this new large-scale multimodal dataset will contribute the community of multimodal based action understanding.
Quan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt, Bin Tong, Tomokazu Murakami
ICCV3