VLDB 2026 Research / reviewers in the wild / expert
Zhuoran Chen
dblp:05/6278
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 41% 3D vision · 38% Generative modeling · 17% | |
| Computer graphics and multimedia
1 paper |
Computer animation and physical simulation · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.4 | 2 | 2024 | Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning · AAAI 2024 Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning · IJCAI 2023 |
Computer vision › 3D vision
3d scene understanding |
1.0 | 1 | 2026 | Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy · AAAI 2026 |
Computer vision › 3D vision › 3d scene modeling › scene representation
semantic scene representation |
1.0 | 1 | 2026 | Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy · AAAI 2026 |
Computer animation and physical simulation › motion synthesis
human motion synthesis |
1.0 | 1 | 2026 | Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy · AAAI 2026 |
Computer animation and physical simulation › motion synthesis › human motion synthesis
scene-aware motion synthesis |
1.0 | 1 | 2026 | Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy · AAAI 2026 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization · ICCV 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent exploration |
0.8 | 1 | 2024 | Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning · AAAI 2024 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration |
0.2 | 1 | 2023 | Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning · IJCAI 2023 |
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search |
0.0 | 1 | 2003 | Assessing sequence comparison methods with the average precision criterion · Bioinform. 2003 |
Information retrieval › web search
mobile search |
0.0 | 1 | 2002 | A system for spoken query information retrieval on mobile devices · IEEE Trans. Speech Audio Process. 2002 |
Information retrieval › document retrieval › spoken document retrieval
spoken query retrieval |
0.0 | 1 | 2002 | A system for spoken query information retrieval on mobile devices · IEEE Trans. Speech Audio Process. 2002 |
Bioinformatics and computational biology › sequence analysis
sequence similarity search |
0.0 | 1 | 2003 | Assessing sequence comparison methods with the average precision criterion · Bioinform. 2003 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.0 | 1 | 2002 | A system for spoken query information retrieval on mobile devices · IEEE Trans. Speech Audio Process. 2002 |
Methods — techniques the papers use, named apart from their topics
tri-plane decomposition · 2.0scene semantic occupancy · 2.0CLIP encoding · 2.0vision-language model · 1.7preference alignment · 1.7LoRA · 1.7transformer sequence modeling · 0.8intrinsic reward · 0.8curriculum learning · 0.8acyclicity constraint · 0.7time-AP plot · 0.0average precision · 0.0speech recognition · 0.0information retrieval · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human Motion Synthesis in 3D Scenes via Unified Scene Semantic OccupancyabstractHuman motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take an unified Scene Semantic Occupancy (SSO) for scene representation, termed SSOMotion. We design a bi-directional tri-plane decomposition to derive a compact version of the SSO, and scene semantics are mapped to an unified feature space via CLIP encoding and shared linear dimensionality reduction. Such strategy can derive the fine-grained scene semantic structures while significantly reduce redundant computations. We further take these scene hints and movement direction derived from instructions for motion control via frame-wise scene query. Extensive experiments and ablation studies conducted on cluttered scenes using ShapeNet furniture, as well as scanned scenes from PROX and Replica datasets, demonstrate its cutting-edge performance while validating its effectiveness and generalization ability. Jingyu Gong, Kunkun Tong, Zhuoran Chen, Chuanhan Yuan, Mingang Chen, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006 |
AAAI | 3 |
| 2025 | ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model PersonalizationabstractWe introduce ImageGem, a dataset for studying generative models that understand fine-grained individual preferences. We posit that a key challenge hindering the development of such a generative model is the lack of in-the-wild and fine-grained user preference annotations. Our dataset features real-world interaction data from 57K users, who collectively have built 242K customized LoRAs, written 3M text prompts, and created 5M generated images. With user preference annotations from our dataset, we were able to train better preference alignment models. In addition, leveraging individual user preference, we investigated the performance of retrieval models and a vision-language model on personalized image retrieval and generative model recommendation. Finally, we propose an end-to-end framework for editing customized diffusion models in a latent weight space to align with individual user preferences. Our results demonstrate that the ImageGem dataset enables, for the first time, a new paradigm for generative model personalization. Yuanhe Guo, Linxi Xie, Zhuoran Chen, Kangrui Yu, Ryan Po, Guandao Yang, Gordon Wetztein, Hongyi Wen |
ICCV | 3 |
| 2025 | Offline Multi-Agent Preference-based Reinforcement Learning with Agent-aware Direct Preference Optimization
Qian Kou, Zeyang Liu 0001, Zhuoran Chen, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan |
AAMAS | 5 |
| 2025 | AlignCAPE: Support and Query Feature Aligning for Category-Agnostic Pose EstimationabstractRecent advancements in category-agnostic pose estimation have focused on developing a unified model capable of localizing keypoint coordinates across arbitrary categories, which enables robots to accurately interact with diverse objects by understanding their poses. While existing methods predominantly concentrate on local features surrounding the keypoints of the support image, they often overlook the importance of global features, leading to potential misalignment between the support and query image. To address the inherent conflicts between the two images, we propose AlignCAPE, a novel approach designed to mitigate such misalignment and enhance the model performance. Our method formulates a two-stage pipeline, generating initial proposals in the first stage, followed by another stage to refine iteratively. Specifically, we introduce two modules, Feature Alignment Module(FAM) and Keypoint Perception Module(KPM). FAM utilizes bidirectional cross-attention operation to align the support image feature and query image feature, thereby compensating for the limitations of previous methods. KPM employs self-attention mechanism to capture the interactions among keypoints, facilitating to localize keypoints in the query image. Experiments on MP-100 benchmark demonstrate that our method outperforms the widely-used baseline model in CAPE by 0.68% in [email protected] metric under 1-shot setting. Zhuoran Chen, Shaojie Zhang 0004, Jianqin Yin |
IROS | 1 |
| 2025 | RoboCam: Model-Based Robotic Visual Sensing for Precise Inspection of Mesh ScreensabstractThe 3D-printed mesh screen with dense penetrating pores is a new structure for massive manufacturing of molded pulp package products. However, some of the pores may be clogged by the printing material powder during the printing process. Such defects negatively affect the quality of the pulp packages produced using the mesh screen mold. To pinpoint the defects, we design a model-based robotic visual sensing system, called RoboCam, which uses a robotic arm to carry a high-resolution camera for full inspection of a mold consisting of joined mesh screens. To inspect the entire mold, RoboCam plans the camera poses to capture multiple images of the mold and render synthesized images as references for identifying the clogged pores. In particular, we propose novel designs to rectify the inherent run-time pose errors of the robotic system for ensuring the reference quality and to accelerate the reference rendering for reducing inspection latency. Extensive evaluation shows that RoboCam’s design outperforms various baselines, including three existing computer vision and convolution neural network-based inspection systems. RoboCam achieves a recall rate of 94.95% within 528 seconds latency for inspecting an entire mold with 13,000 designed pores. Duc Van Le, Linshan Jiang, Zhuoran Chen, Xiaohua Peng, Daren Ho, Jianmin Zheng, Rui Tan 0001 |
ACM Trans. Sens. Networks | 4 |
| 2024 | Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement LearningabstractEffective exploration is crucial to discovering optimal strategies for multi-agent reinforcement learning (MARL) in complex coordination tasks. Existing methods mainly utilize intrinsic rewards to enable committed exploration or use role-based learning for decomposing joint action spaces instead of directly conducting a collective search in the entire action-observation space. However, they often face challenges obtaining specific joint action sequences to reach successful states in long-horizon tasks. To address this limitation, we propose Imagine, Initialize, and Explore (IIE), a novel method that offers a promising solution for efficient multi-agent exploration in complex scenarios. IIE employs a transformer model to imagine how the agents reach a critical state that can influence each other's transition functions. Then, we initialize the environment at this state using a simulator before the exploration phase. We formulate the imagination as a sequence modeling problem, where the states, observations, prompts, actions, and rewards are predicted autoregressively. The prompt consists of timestep-to-go, return-to-go, influence value, and one-shot demonstration, specifying the desired state and trajectory as well as guiding the action generation. By initializing agents at the critical states, IIE significantly increases the likelihood of discovering potentially important under-explored regions. Despite its simplicity, empirical results demonstrate that our method outperforms multi-agent exploration baselines on the StarCraft Multi-Agent Challenge (SMAC) and SMACv2 environments. Particularly, IIE shows improved performance in the sparse-reward SMAC tasks and produces more effective curricula over the initialized states than other generative methods, such as CVAE-GAN and diffusion models. Zeyang Liu 0001, Lipeng Wan 0003, Zhuoran Chen, Xingyu Chen 0001, Xuguang Lan |
AAAI | 4 |
| 2023 | Deep Hierarchical Communication Graph in Multi-Agent Reinforcement LearningabstractSharing intentions is crucial for efficient cooperation in communication-enabled multi-agent reinforcement learning. Recent work applies static or undirected graphs to determine the order of interaction. However, the static graph is not general for complex cooperative tasks, and the parallel message-passing update in the undirected graph with cycles cannot guarantee convergence. To solve this problem, we propose Deep Hierarchical Communication Graph (DHCG) to learn the dependency relationships between agents based on their messages. The relationships are formulated as directed acyclic graphs (DAGs), where the selection of the proper topology is viewed as an action and trained in an end-to-end fashion. To eliminate the cycles in the graph, we apply an acyclicity constraint as intrinsic rewards and then project the graph in the admissible solution set of DAGs. As a result, DHCG removes redundant communication edges for cost improvement and guarantees convergence. To show the effectiveness of the learned graphs, we propose policy-based and value-based DHCG. Policy-based DHCG factorizes the joint policy in an auto-regressive manner, and value-based DHCG factorizes the joint value function to individual value functions and pairwise payoff functions. Empirical results show that our method improves performance across various cooperative multi-agent tasks, including Predator-Prey, Multi-Agent Coordination Challenge, and StarCraft Multi-Agent Challenge. Zeyang Liu 0001, Lipeng Wan 0003, Xue Sui, Zhuoran Chen, Kewu Sun, Xuguang Lan |
IJCAI | 4 |
| 2003 | Assessing sequence comparison methods with the average precision criterionabstractMOTIVATION: Comprehensive performance assessment is important for improving sequence database search methods. Sensitivity, selectivity and speed are three major yet usually conflicting evaluation criteria. The average precision (AP) measure aims to combine the sensitivity and selectivity features of a search algorithm. It can be easily visualized and extended to analyze results from a set of queries. Finally, the time-AP plot can clearly show the overall performance of different search methods. RESULTS: Experiments are performed based on the SCOP database. Popular sequence comparison algorithms, namely Smith-Waterman (SSEARCH), FASTA, BLAST and PSI-BLAST are evaluated. We find that (1) the low-complexity segment filtration procedure in BLAST actually harms its overall search quality; (2) AP scores of different search methods are approximately in proportion of the logarithm of search time; and (3) homologs in protein families with many members tend to be more obscure than those in small families. This measure may be helpful for developing new search algorithms and can guide researchers in selecting most suitable search methods. AVAILABILITY: Test sets and source code of this evaluation tool are available upon request. Zhuoran Chen |
Bioinform. | 1 |
| 2002 | A system for spoken query information retrieval on mobile devicesabstractWith the proliferation of handheld devices, information access on mobile devices is a topic of growing relevance. This paper presents a system that allows the user to search for information on mobile devices using spoken natural-language queries. We explore several issues related to the creation of this system, which combines state-of-the-art speech-recognition and information-retrieval technologies. This is the first work that we are aware of which evaluates spoken query based information retrieval on a commonly available and well researched text database, the Chinese news corpus used in the National Institute of Standards and Technology (NIST)s TREC-5 and TREC-6 benchmarks. To compare spoken-query retrieval performance for different relevant scenarios and recognition accuracies, the benchmark queries-read verbatim by 20 speakers-were recorded simultaneously through three channels: headset microphone, PDA microphone, and cellular phone. Our results show that for mobile devices with high-quality microphones, spoken-query retrieval based on existing technologies yields retrieval precisions that come close to that for perfect text input (mean average precision 0.459 and 0.489, respectively, on TREC-6). Eric Chang, Frank Seide, Helen M. Meng, Zhuoran Chen, Yu Shi 0001, Yuk-Chi Li |
IEEE Trans. Speech Audio Process. | 4 |