VLDB 2026 Research / reviewers in the wild / expert
Zerui Chen
dblp:254/8116
· DBLP profile ↗
25ranked-venue papers
14as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 10 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal ReasoningabstractZerui Chen, Qinggang Zhang, Zhishang Xiang, Zhimin Wei, Linfeng Gao, Xiao Huang, Zhihong Zhang, Jinsong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zerui Chen, Qinggang Zhang, Zhishang Xiang, Zhimin Wei, Linfeng Gao, Jinsong Su |
ACL (1) | 1 |
| 2026 | Beyond the Flat Sequence: Hierarchical and Preference-Aware Generative RecommendationsabstractGenerative Recommenders (GRs), exemplified by the Hierarchical Sequential Transduction Unit (HSTU), have emerged as a powerful paradigm for modeling long user interaction sequences. However, we observe that their ''flat-sequence'' assumption overlooks the rich, intrinsic structure of user behavior. This leads to two key limitations: a failure to capture the temporal hierarchy of session-based engagement, and computational inefficiency, as dense attention introduces significant noise that obscures true preference signals within semantically sparse histories, which deteriorates the quality of the learned representations. To this end, we propose a novel framework named HPGR (Hierarchical and Preference-aware Generative Recommender), built upon a two-stage paradigm that injects these crucial structural priors into the model to handle the drawback. Specifically, HPGR comprises two synergistic stages. First, a structure-aware pre-training stage employs a session-based Masked Item Modeling (MIM) objective to learn a hierarchically-informed and semantically rich item representation space. Second, a preference-aware fine-tuning stage leverages these powerful representations to implement a Preference-Guided Sparse Attention mechanism, which dynamically constrains computation to only the most relevant historical items, enhancing both efficiency and signal-to-noise ratio. Empirical experiments on a large-scale proprietary industrial dataset from APPGallery and an online A/B test verify that HPGR achieves state-of-the-art performance over multiple strong baselines, including HSTU and MTGR. Zerui Chen, Heng Chang, Tianying Liu, Chuantian Zhou, Yi Cao 0003, Jiandong Ding, Ming Liu 0004, Bing Qin 0001 |
WWW | 1 |
| 2026 | Subgraph-Centric Multi-Agent Reinforcement Learning for Multi-Hop Knowledge Graph ReasoningabstractMulti-hop Knowledge Graph Reasoning (KGR) seeks to identify accurate answers within Knowledge Graphs (KGs) via multi-step reasoning, predominantly utilizing reinforcement learning (RL) to enhance the efficiency of the reasoning process. Unlike traditional Knowledge Graph Embedding (KGE) methods, RL-based approaches offer superior interpretability. However, these methods often underperform due to two critical limitations: (1) their over-reliance on Horn rules for reasoning paths, which restricts their expressive power; and (2) inadequate utilization of reasoning states during the process. To address these issues, we propose a novel RL-based framework, RAR, which shifts focus from individual paths to subgraph structures for more robust predictions. RAR frames the retrieval of reasoning subgraphs from the KG as a Markov Decision Process (MDP) and incorporates a subgraph retriever. To efficiently explore the extensive subgraph space, we integrate multi-agent RL to enhance the retriever's capabilities. Additionally, RAR features an advanced analyst module that meticulously examines reasoning states. These modules function iteratively: the retriever expands the subgraph, followed by the analyst module's in-depth analysis. The insights gained are then used to inform subsequent retrieval steps. Ultimately, the predicted scores from both modules are synthesized to produce more precise posterior scores. Experimental results across multiple datasets demonstrate RAR's efficacy, showcasing a notable improvement over existing state-of-the-art RL-based KGR methods. Tao He 0014, Zerui Chen, Lizi Liao, Yixin Cao 0002, Yuanxing Liu 0001, Wei Tang 0015, Xun Mao, Ming Liu 0004, Bing Qin 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Simulation-Free Hierarchical Latent Policy Planning for Proactive DialoguesabstractRecent advancements in proactive dialogues have garnered significant attention, particularly for more complex objectives (e.g. emotion support and persuasion). Unlike traditional task-oriented dialogues, proactive dialogues demand advanced policy planning and adaptability, requiring rich scenarios and comprehensive policy repositories to develop such systems. However, existing approaches tend to rely on Large Language Models (LLMs) for user simulation and online learning, leading to biases that diverge from realistic scenarios and result in suboptimal efficiency. Moreover, these methods depend on manually defined, context-independent, coarse-grained policies, which not only incur high expert costs but also raise concerns regarding their completeness. In our work, we highlight the potential for automatically discovering policies directly from raw, real-world dialogue records. To this end, we introduce a novel dialogue policy planning framework, LDPP. It fully automates the process from mining policies in dialogue records to learning policy planning. Specifically, we employ a variant of the Variational Autoencoder to discover fine-grained policies represented as latent vectors. After automatically annotating the data with these latent policy labels, we propose an Offline Hierarchical Reinforcement Learning (RL) algorithm in the latent space to develop effective policy planning capabilities. Our experiments demonstrate that LDPP outperforms existing methods on two proactive scenarios, even surpassing ChatGPT with only a 1.8-billion-parameter LLM. Tao He 0014, Lizi Liao, Yixin Cao 0002, Yuanxing Liu 0001, Yiheng Sun, Zerui Chen, Ming Liu 0004, Bing Qin 0001 |
AAAI | 6 |
| 2025 | How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and FutureabstractEntity alignment (EA), critical for knowledge graph (KG) integration, identifies equivalent entities across different KGs.Traditional methods often face challenges in semantic understanding and scalability.The rise of language models (LMs), particularly large language models (LLMs), has provided powerful new strategies.This paper systematically reviews LM-driven EA methods, proposing a novel taxonomy that categorizes methods in three key stages: data preparation, feature embedding, and alignment.We further summarize key benchmarks, evaluation metrics, and discuss future directions.This paper aims to provide researchers and practitioners with a clear and comprehensive understanding of how language models reshape the field of entity alignment.* These authors contributed equally. Zerui Chen, Huiming Fan, Tao He 0014, Ming Liu 0004, Heng Chang, Weijiang Yu, Bing Qin 0001 |
EMNLP | 1 |
| 2025 | HORT: Monocular Hand-held Objects Reconstruction with TransformersabstractReconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which produce overly smooth reconstructions and are time-consuming to generate explicit 3D shapes. While more recent methods directly reconstruct point clouds with diffusion models, the multi-step denoising makes high-resolution reconstruction inefficient. To address these limitations, we propose a transformer-based model to efficiently reconstruct dense 3D point clouds of hand-held objects. Our method follows a coarse-to-fine strategy, first generating a sparse point cloud from the image and progressively refining it into a dense representation using pixel-aligned image features. To enhance reconstruction accuracy, we integrate image features with 3D hand geometry to jointly predict the object point cloud and its pose relative to the hand. Our model is trained end-to-end for optimal performance. Experimental results on both synthetic and real datasets demonstrate that our method achieves state-of-the-art accuracy with much faster inference speed, while generalizing well to in-the-wild images. Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Cordelia Schmid |
ICCV | 1 |
| 2025 | ViViDex: Learning Vision-Based Dexterous Manipulation from Human VideosabstractIn this work, we aim to learn a unified vision-based policy for multi-fingered robot hands to manipulate a variety of objects in diverse poses. Though prior work has shown benefits of using human videos for policy learning, performance gains have been limited by the noise in estimated trajectories. Moreover, reliance on privileged object information such as ground-truth object states further limits the applicability in realistic scenarios. To address these limitations, we propose a new framework ViViDex to improve vision-based policy learning from human videos. It first uses reinforcement learning with trajectory guided rewards to train state-based policies for each video, obtaining both visually natural and physically plausible trajectories from the video. We then rollout successful episodes from state-based policies and train a unified visual policy without using any privileged information. We propose coordinate transformation to further enhance the visual point cloud representation, and compare behavior cloning and diffusion policy for the visual policy training. Experiments both in simulation and on the real robot demonstrate that ViViDex outperforms state-of-theart approaches on three dexterous manipulation tasks. Project website: zerchen.github.io/projects/vividex.html. Zerui Chen, Shizhe Chen, Etienne Arlaud, Ivan Laptev, Cordelia Schmid |
ICRA | 1 |
| 2024 | Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular ImagesabstractReconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent works have employed implicit functions to achieve impressive progress, they ignore formulating contacts in their frameworks, which results in producing less realistic object meshes. In this work, we explore how to model contacts in an explicit way to benefit the implicit reconstruction of hand-held objects. Our method consists of two components: explicit contact prediction and implicit shape reconstruction. In the first part, we propose a new subtask of directly estimating 3D hand-object contacts from a single image. The part-level and vertex-level graph-based transformers are cascaded and jointly learned in a coarse-to-fine manner for more accurate contact probabilities. In the second part, we introduce a novel method to diffuse estimated contact states from the hand mesh surface to nearby 3D space and leverage diffused contact probabilities to construct the implicit neural representation for the manipulated object. Benefiting from estimating the interaction patterns between the hand and the object, our method can reconstruct more realistic object meshes, especially for object parts that are in contact with hands. Extensive experiments on challenging benchmarks show that the proposed method outperforms the current state of the arts by a great margin. Our code is publicly available at https://junxinghu.github.io/projects/hoi.html. Junxing Hu, Hongwen Zhang 0001, Zerui Chen, Mengcheng Li, Yunlong Wang 0003, Yebin Liu, Zhenan Sun |
AAAI | 3 |
| 2024 | Planning Like Human: A Dual-process Framework for Dialogue PlanningabstractIn proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature.Traditional approaches to enhance dialogue planning in LLMs, ranging from elaborate prompt engineering to the integration of policy networks, either face efficiency issues or deliver suboptimal performance.Inspired by the dualprocess theory in psychology, which identifies two distinct modes of thinking-intuitive (fast) and analytical (slow), we propose the Dual-Process Dialogue Planning (DPDP) framework.DPDP embodies this theory through two complementary planning systems: an instinctive policy model for familiar contexts and a deliberative Monte Carlo Tree Search (MCTS) mechanism for complex, novel scenarios.This dual strategy is further coupled with a novel two-stage training regimen: offline Reinforcement Learning for robust initial policy model formation followed by MCTS-enhanced on-thefly learning, which ensures a dynamic balance between efficiency and strategic depth.Our empirical evaluations across diverse dialogue tasks affirm DPDP's superiority in achieving both high-quality dialogues and operational efficiency, outpacing existing methods. 1 Tao He 0014, Lizi Liao, Yixin Cao 0002, Yuanxing Liu 0001, Ming Liu 0004, Zerui Chen, Bing Qin 0001 |
ACL (1) | 6 |
| 2024 | Pseudo Labels Regularization for Imbalanced Partial-Label LearningabstractPartial-label learning (PLL) is an important branch of weakly supervised learning where the single ground truth resides in a set of candidate labels, while the research rarely considers the label imbalance. A recent study for imbalanced PLL propose that the combinatorial challenge of partial-label learning and long-tail learning lies in matching between a decent marginal prior distribution with drawing the pseudo labels. However, even if the pseudo label matches the prior distribution, the tail classes will still be difficult to learn because the total weight of tail classes is too small. Therefore, we propose a pseudo-label regularization technique specially designed for imbalanced PLL. By punishing the pseudo labels of head classes, our method implements state-of-art under the standardized benchmarks compared to the previous PLL methods. Zheng Lian 0004, Bin Liu 0041, Zerui Chen, Jianhua Tao 0001 |
ICASSP | 4 |
| 2024 | GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
Jiafeng Liang, Shixin Jiang, Zekun Wang 0001, Haojie Pan, Zerui Chen, Ming Liu 0004, Ruiji Fu, Zhongyuan Wang 0006, Bing Qin 0001 |
IJCAI | 5 |
| 2024 | CEC-DL: Cloud-Edge Collaborative Delegation Learning Against Covert AdversariesabstractDelegation learning is indeed a prevalent approach in privacy-preserving machine learning (PPML), especially when dealing with big data. It specifically involves data owners delegating their data to servers with computational capabilities for training and inference. These servers provide services on a pay-per-use basis. The essence of delegation learning lies in maintaining the integrity of the server’s training while ensuring the privacy of the delegator’s data. However, existing delegation learning schemes struggle to balance security and efficiency, and they cannot guarantee correctness. To tackle these challenges, we propose a cloud-edge collaborative delegation learning framework (CEC-DL) against covert adversaries, which is verifiable and satisfies the guaranteed output delivery (GOD) in security. This is the first time that the covert security assumption is used in a PPML scenario. Furthermore, we design probabilistic verifiable secure addition and subtraction computation protocol (PVS-AaS) and probabilistic verifiable secure multiplication computation protocol (PVS-MUL), which can be used to realize secure addition and multiplication computations in delegation learning without expensive message authentication code (MAC) verification. At the same time, we develop a malicious adversary detection protocol (MADP) that could prevent the malicious actions of potential covert adversaries while ensuring the correct output. Finally, we apply the CEC-DL to the linear regression model to construct privacy preserving linear regression protocol (PP-LRP). Through theoretical analysis and experiments, CEC-DL improves the security, and is more efficient than the verifiable computation of malicious adversaries. Youliang Tian, Zerui Chen, Xinhua Cui, Jinbo Xiong, Jianfeng Ma 0001 |
IEEE Internet Things J. | 3 |
| 2023 | gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object ReconstructionabstractSigned distance functions (SDFs) is an attractive frame-work that has recently shown promising results for 3D shape reconstruction from images. SDFs seamlessly generalize to different shape resolutions and topologies but lack explicit modelling of the underlying 3D geometry. In this work, we ex-ploit the hand structure and use it as guidance for SDF-based shape reconstruction. In particular, we address reconstruction of hands and manipulated objects from monocular RGB images. To this end, we estimate poses of hands and objects and use them to guide 3D reconstruction. More specifically, we predict kinematic chains of pose transformations and align SDFs with highly-articulated hand poses. We improve the visual features of 3D points with geometry alignment and further leverage temporal information to enhance the robustness to occlusion and motion blurs. We conduct extensive experiments on the challenging ObMan and DexYCB benchmarks and demonstrate significant improvements of the proposed method over the state of the art. Zerui Chen, Shizhe Chen, Cordelia Schmid, Ivan Laptev |
CVPR | 1 |
| 2023 | Electronic Sheepdog: A Novel Method in With UAV-Assisted Wearable Grazing MonitoringabstractThe application of Internet of Things (IoT) and unmanned aerial vehicle (UAV) technology to help herders monitor livestock makes intelligent and scientific grazing possible. In order to extend the limited functionality of the existing monitoring equipment and reduce the flight time of UAVs, a novel wearable grazing system with UAV-assisted monitoring is proposed in this article. First, the system can monitor physiological parameters, such as running, falling, and feeding, as well as geographic parameters. The acquired geographic location is used to implement electronic fences. When livestock leave a specified area, a voice playback module makes the sound of sheepdogs to drive them back. Then, narrowband IoT (NB-IoT) technology is used to transmit data. The monitoring values of physiological and geographic parameters can be displayed on a terminal. Finally, when an alarm goes off, a UAV instead of a herder arrives at the grazing area, where the UAV is used as an auxiliary device. Combining these technologies can realize the design of electronic sheepdogs. A standby algorithm is used to reduce the system’s power consumption. Experimental results show the reliability of livestock posture monitoring high accuracy of the alarm system, and safety of the UAV. Therefore, using a combination of UAV and IoT systems greatly improves the efficiency of grazing monitoring. Hao Wang 0173, Xihai Zhang, Xiangyu Meng 0001, Weixian Song, Zerui Chen |
IEEE Internet Things J. | 5 |
| 2022 | AlignSDF: Pose-Aligned Signed Distance Fields for Hand-Object Reconstruction
Zerui Chen, Yana Hasson, Cordelia Schmid, Ivan Laptev |
ECCV (1) | 1 |
| 2022 | Learning a Robust Part-Aware Monocular 3D Human Pose Estimator via Neural Architecture Search
Zerui Chen, Yan Huang 0008, Hongyuan Yu, Liang Wang 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | An incentive-compatible rational secret sharing scheme using blockchain and smart contract
Zerui Chen, Youliang Tian, Changgen Peng |
Sci. China Inf. Sci. | 1 |
| 2021 | Towards reducing delegation overhead in replication-based verification: An incentive-compatible rational delegation computing scheme
Zerui Chen, Youliang Tian, Jinbo Xiong, Changgen Peng, Jianfeng Ma 0001 |
Inf. Sci. | 1 |
| 2020 | Towards Part-Aware Monocular 3D Human Pose Estimation: An Architecture Search Approach
Zerui Chen, Yan Huang 0008, Hongyuan Yu, Yiru Guo, Liang Wang 0001 |
ECCV (3) | 1 |
| 2020 | Prediction and Recovery for Adaptive Low-Resolution Person Re-Identification
Yan Huang 0008, Zerui Chen, Liang Wang 0001, Tieniu Tan |
ECCV (26) | 3 |
| 2020 | On the Robustness of 3D Human Pose EstimationabstractIt is widely shown that Convolutional Neural Networks (CNNs) are vulnerable to adversarial examples on most recognition tasks, such as image classification and segmentation. However, few work studies the more complicated task - 3D human pose estimation. This task often requires large-scale datasets, specialized network architectures, and it can be solved either from single-view RGB images or from multi-view RGB images. In this paper, we make the first attempt to investigate the robustness of current state-of-the-art 3D human pose estimation methods. To this end, we build four representative baseline models, where most of the current methods can be generally classified as one of them. Furthermore, we design targeted adversarial attacks to detect whether 3D pose estimators are robust to different camera parameters. For different types of methods, we present a comprehensive study of their robustness on the large-scale Human3.6M benchmark. Our work shows that different methods vary significantly in their resistance to adversarial attacks. Through extensive experiments, we show that multi-view 3D pose estimators can be more vulnerable to adversarial examples. We believe that our efforts can shed light on future works to design more robust 3D human pose estimators. Zerui Chen, Yan Huang 0008, Liang Wang 0001 |
ICPR | 1 |
| 2020 | VSR++: Improving Visual Semantic Reasoning for Fine-Grained Image-Text MatchingabstractImage-text matching has made great progresses recently, but there still remains challenges in fine-grained matching. To deal with this problem, we propose an Improved Visual Semantic Reasoning model (VSR++), which jointly models 1) global alignment between images and texts and 2) local correspondence between regions and words in a unified framework. To exploit their complementary advantages, we also develop a suitable learning strategy to balance their relative importance. As a result, our model can distinguish image regions and text words in a fine-grained level, and thus achieves the current state-of-the-art performance on two benchmark datasets. Yan Huang 0008, Dongbo Zhang 0003, Zerui Chen, Liang Wang 0001 |
ICPR | 4 |
| 2020 | Absolute 3D Human Pose Estimation via Weakly-supervised LearningabstractIn this paper, we attempt to estimate absolute 3D human poses directly from monocular images. Not limited to estimating root-relative 3D human poses, our method can recover the absolute depth for each joint. Our method is trained with multi-view images in a weakly-supervised manner removing the need for 3D ground-truth annotations. We conduct extensive experiments on the Human3.6M benchmark to show the effectiveness of our method. Yiru Guo, Zerui Chen |
VCIP | 2 |
| 2019 | Learning Depth-aware Heatmaps for 3D Human Pose Estimation in the Wild
Zerui Chen, Yiru Guo, Yan Huang 0008, Liang Wang 0001 |
BMVC | 1 |
| 2019 | Augmented Visual-Semantic Embeddings for Image and Sentence MatchingabstractThe task of image and sentence matching has witnessed significant progress recently, but it is still challenging arising from the tremendous semantic gap between a pixel-level image and its matched sentences. Due to limited training data, it is rather challenging to optimize the visual-semantic embeddings. In this work, we propose to augment visual-semantic embeddings via enlarging the training dataset. With more data, models can learn discriminative features with high-quality semantic concepts. More specifically, we augment data by generating sentences for given images. Our method consists of two steps. At first, to enlarge the training dataset, given an image, we perform image captioning. Instead of introducing redundancy to our augmented dataset, we hope that our generated sentences are in diverse style and maintain its fidelity at the same time. Therefore, we consult to generative adversarial networks (GANs) which can produce more flexible expressions compared to methods based on the maximum likelihood principle. Then, we augment visual-semantic embeddings with the augmented training dataset and obtain the model for the task of image and sentence matching. Experiments on the popular benchmark demonstrate the effectiveness of our method by achieving superior results compared to our baseline. Zerui Chen, Yan Huang 0008, Liang Wang 0001 |
ICIP | 1 |