EDBT 2026 Demo / reviewers in the wild / expert
Jiadai Sun
dblp:271/7474
· DBLP profile ↗
17ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0001-9593-0873ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EC-MVSNet: Enhanced Cascaded Multi-View Stereo with Cross-Scale Relevance IntegrationabstractCascade-based multi-scale architectures are currently the mainstream in Multi-view Stereo (MVS), achieving a balance between computational efficiency and reconstruction accuracy. However, existing cascade MVS methods suffer from significant limitations in cross-scale information utilization, where depth estimation processes operate independently across scales without fully exploiting the rich relevance between adjacent scales. To address this fundamental limitation, we propose an Enhanced Cascade Multi-View Stereo framework (EC-MVSNet), which introduces a novel cross-scale relevance integration strategy. Specifically, we introduce a Cross-Scale Feature-based Joint Construction (CFC) module to synergistically combine features from adjacent scales to build more reliable cost volumes. Additionally, a Cross-Scale Probability-guided Enhancement (CPE) module is proposed to propagate depth probability distributions across scales to guide cost volume enhancement. Furthermore, we propose a Monocular Feature-based Refinement (MFR) module to further enhance depth prediction accuracy by leveraging monocular priors. Extensive experiments demonstrate that EC-MVSNet achieves state-of-the-art performance on multiple benchmarks, validating the effectiveness of the cross-scale integration in improving MVS reconstruction quality. Shaoqian Wang, Jiadai Sun, Bin Fan 0002, Qiang Wang 0023, Yuchao Dai |
AAAI | 2 |
| 2025 | VisualAgentBench: Towards Large Multimodal Models as Visual Foundation AgentsabstractLarge Multimodal Models (LMMs) have ushered in a new era in artificial intelligence, merging capabilities in both language and vision to form highly capable \textbf{Visual Foundation Agents} that are postulated to excel across a myriad of tasks. However, existing benchmarks fail to sufficiently challenge or showcase the full potential of LMMs as visual foundation agents in complex, real-world environments. To address this gap, we introduce VisualAgentBench (VAB), a comprehensive and unified benchmark specifically designed to train and evaluate LMMs as visual foundation agents across diverse scenarios in one standard setting, including Embodied, Graphical User Interface, and Visual Design, with tasks formulated to probe the depth of LMMs' understanding and interaction capabilities. Through rigorous testing across 9 proprietary LMM APIs and 9 open models (18 in total), we demonstrate the considerable yet still developing visual agent capabilities of these models. Additionally, VAB explores the synthesizing of visual agent trajectory data through hybrid methods including Program-based Solvers, LMM Agent Bootstrapping, and Human Demonstrations, offering insights into obstacles, solutions, and trade-offs one may meet in developing open LMM agents. Our work not only aims to benchmark existing models but also provides an instrumental playground for future development into visual foundation agents. Code, train, and test data are available at \url{https://github.com/THUDM/VisualAgentBench}. Xiao Liu 0036, Tianjie Zhang, Yu Gu 0016, Iat Long Iong, Xixuan Song, Yifan Xu 0014, Shudan Zhang, Hanyu Lai, Jiadai Sun, Zehan Qi, Shuntian Yao, Xueqiao Sun, Qinkai Zheng, Hao Yu 0030, Hanchen Zhang, Wenyi Hong, Ming Ding 0004, Lihang Pan, Xiaotao Gu, Aohan Zeng, Zhengxiao Du, Chan Hee Song, Yu Su 0001, Yuxiao Dong, Jie Tang 0001 |
ICLR | 9 |
| 2025 | WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement LearningabstractLarge language models (LLMs) have shown remarkable potential as autonomous agents, particularly in web-based tasks.
However, existing LLM web agents face significant limitations: high-performing agents rely on expensive proprietary LLM APIs, while open LLMs lack the necessary decision-making capabilities.
This paper introduces WebRL, a novel self-evolving online curriculum reinforcement learning framework designed to train high-performance web agents using open LLMs.
Our approach addresses key challenges in this domain, including the scarcity of training tasks, sparse feedback signals, and policy distribution drift in online learning.
WebRL incorporates a self-evolving curriculum that generates new tasks from unsuccessful attempts, a robust outcome-supervised reward model (ORM), and adaptive reinforcement learning strategies to ensure consistent improvement.
We apply WebRL to transform Llama-3.1 models into proficient web agents, achieving remarkable results on the WebArena-Lite benchmark.
Our Llama-3.1-8B agent improves from an initial 4.8\% success rate to 42.4\%, while the Llama-3.1-70B agent achieves a 47.3\% success rate across five diverse websites.
These results surpass the performance of GPT-4-Turbo (17.6\%) by over 160\% relatively and significantly outperform previous state-of-the-art web agents trained on open LLMs (AutoWebGLM, 18.2\%).
Our findings demonstrate WebRL's effectiveness in bridging the gap between open and proprietary LLM-based web agents, paving the way for more accessible and powerful autonomous web interaction systems. Zehan Qi, Xiao Liu 0036, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Jiadai Sun, Shuntian Yao, Wei Xu 0017, Jie Tang 0001, Yuxiao Dong |
ICLR | 6 |
| 2024 | Joint Scene Flow Estimation and Moving Object Segmentation on Rotational LiDAR DataabstractLiDAR-based scene flow estimation (SFE) and moving object segmentation (MOS) are important tasks with broad-ranging applications in autonomous driving, such as traffic surveillance, motion analysis, obstacle avoidance, etc. Most existing works address SFE and MOS separately, ignoring the underlying shared geometric constraints and their inherent correlation. This article rethinks LiDAR-based SFE and MOS tasks, providing our key insight that jointly addressing them can tackle challenges in both tasks, and their solutions can reinforce one another to improve the performance of both. Based on this insight, we introduce a novel framework that exploits shared geometric constraints by explicitly partitioning the scene into static and moving regions and subsequently estimating flow differently for these regions. A lightweight and interpretable neural network dubbed SFEMOS is proposed. It employs an encoder and two specially designed head modules for each task, achieving MOS without relying on prior poses and online point-wise flow estimation for 360-degree point clouds. Due to the absence of public datasets for concurrently evaluating both tasks, we generate ground truth flow data using MOS labels from SemanticKITTI. Additionally, we establish a new dataset using a rotational LiDAR mounted on our own autonomous vehicle. Evaluation results on both datasets validate the superior performance of our proposed SFEMOS. Our dataset and label generation method are released athttps://github.com/nubot-nudt/SFEMOS. Xieyuanli Chen, Jiafeng Cui, Xianjing Zhang, Jiadai Sun, Rui Ai 0001, Weihao Gu, Jintao Xu 0001, Huimin Lu 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Forward Flow for Novel View Synthesis of Dynamic ScenesabstractThis paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canonical space with the learned backward flow field. However, this backward flow field is non-smooth and discontinuous, which is difficult to be fitted by commonly used smooth motion models. To address this problem, we propose to estimate the forward flow field and directly warp the canonical radiance field to other time steps. Such forward flow field is smooth and continuous within the object region, which benefits the motion model learning. To achieve this goal, we represent the canonical radiance field with voxel grids to enable efficient forward warping, and propose a differentiable warping process, including an average splatting operation and an inpaint network, to resolve the many-to-one and one-to-many mapping issues. Thorough experiments show that our method outperforms existing methods in both novel view rendering and motion modeling, demonstrating the effectiveness of our forward flow motion modeling. Project page: https://npucvr.github.io/ForwardFlowDNeRF. Jiadai Sun, Yuchao Dai, Guanying Chen, Xiaoqing Ye, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001 |
ICCV | 2 |
| 2023 | MapNeRF: Incorporating Map Priors into Neural Radiance Fields for Driving View SimulationabstractSimulating camera sensors is a crucial task in autonomous driving. Although neural radiance fields are exceptional at synthesizing photorealistic views in driving simulations, they still fail to generate extrapolated views. This paper proposes to incorporate map priors into neural radiance fields to synthesize out-of-trajectory driving views with semantic road consistency. The key insight is that map information can be utilized as a prior to guiding the training of the radiance fields with uncertainty. Specifically, we utilize the coarse ground surface as uncertain information to supervise the density field and warp depth with uncertainty from unknown camera poses to ensure multi-view consistency. Experimental results demonstrate that our approach can produce semantic consistency in deviated views for vehicle camera simulation. The supplementary video can be viewed at https://youtu.be/jEQWr-Rfh3A. Chenming Wu, Jiadai Sun, Zhelun Shen, Liangjun Zhang |
IROS | 2 |
| 2023 | Digging into Depth Priors for Outdoor Neural Radiance FieldsabstractNeural Radiance Fields (NeRFs) have demonstrated impressive performance in vision and graphics tasks, such as novel view synthesis and immersive reality. However, the shape-radiance ambiguity of radiance fields remains a challenge, especially in the sparse viewpoints setting. Recent work resorts to integrating depth priors into outdoor NeRF training to alleviate the issue. However, the criteria for selecting depth priors and the relative merits of different priors have not been thoroughly investigated. Moreover, the relative merits of selecting different approaches to use the depth priors is also an unexplored problem. In this paper, we provide a comprehensive study and evaluation of employing depth priors to outdoor neural radiance fields, covering common depth sensing technologies and most application ways. Specifically, we conduct extensive experiments with two representative NeRF methods equipped with four commonly-used depth priors and different depth usages on two widely used outdoor datasets. Our experimental results reveal several interesting findings that can potentially benefit practitioners and researchers in training their NeRF models with depth priors. Project page: https://cwchenwang.github.io/outdoor-nerf-depth Chen Wang 0049, Jiadai Sun, Lina Liu 0010, Chenming Wu, Zhelun Shen, Dayan Wu, Yuchao Dai, Liangjun Zhang |
ACM Multimedia | 2 |
| 2023 | MUNet: Motion uncertainty-aware semi-supervised video object segmentation
Jiadai Sun, Yuxin Mao, Yuchao Dai, Yiran Zhong |
Pattern Recognit. | 1 |
| 2022 | End-to-End Learning the Partial Permutation Matrix for Robust 3D Point Cloud RegistrationabstractEven though considerable progress has been made in deep learning-based 3D point cloud processing, how to obtain accurate correspondences for robust registration remains a major challenge because existing hard assignment methods cannot deal with outliers naturally. Alternatively, the soft matching-based methods have been proposed to learn the matching probability rather than hard assignment. However, in this paper, we prove that these methods have an inherent ambiguity causing many deceptive correspondences. To address the above challenges, we propose to learn a partial permutation matching matrix, which does not assign corresponding points to outliers, and implements hard assignment to prevent ambiguity. However, this proposal poses two new problems, i.e. existing hard assignment algorithms can only solve a full rank permutation matrix rather than a partial permutation matrix, and this desired matrix is defined in the discrete space, which is non-differentiable. In response, we design a dedicated soft-to-hard (S2H) matching procedure within the registration pipeline consisting of two steps: solving the soft matching matrix (S-step) and projecting this soft matrix to the partial permutation matrix (H-step). Specifically, we augment the profit matrix before the hard assignment to solve an augmented permutation matrix, which is cropped to achieve the final partial permutation matrix. Moreover, to guarantee end-to-end learning, we supervise the learned partial permutation matrix but propagate the gradient to the soft matrix instead. Our S2H matching procedure can be easily integrated with existing registration frameworks, which has been verified in representative frameworks including DCP, RPMNet, and DGR. Extensive experiments have validated our method, which creates a new state-of-the-art performance. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He |
AAAI | 2 |
| 2022 | Neural Deformable Voxel Grid for Fast Optimization of Dynamic View Synthesis
Guanying Chen, Yuchao Dai, Xiaoqing Ye, Jiadai Sun, Xiao Tan 0001, Errui Ding |
ACCV (1) | 5 |
| 2022 | Efficient Spatial-Temporal Information Fusion for LiDAR-Based 3D Moving Object SegmentationabstractAccurate moving object segmentation is an es-sential task for autonomous driving. It can provide effective information for many downstream tasks, such as collision avoidance, path planning, and static map construction. How to effectively exploit the spatial-temporal information is a critical question for 3D LiDAR moving object segmentation (LiDAR-MOS). In this work, we propose a novel deep neural network exploiting both spatial-temporal information and different representation modalities of LiDAR scans to improve LiDAR-MOS performance. Specifically, we first use a range image-based dual-branch structure to separately deal with spatial and temporal information that can be obtained from sequential LiDAR scans, and later combine them using motion-guided attention modules. We also use a point refinement module via 3D sparse convolution to fuse the information from both LiDAR range image and point cloud representations and reduce the artifacts on the borders of the objects. We verify the effectiveness of our proposed approach on the LiDAR-MOS benchmark of SemanticKITTI. Our method outperforms the state-of-the-art methods significantly in terms of LiDAR-MOS IoU. Benefiting from the devised coarse-to-fine architecture, our method operates online at sensor frame rate. Code is available at: https://github.com/haomo-ai/MotionSeg3D. Jiadai Sun, Yuchao Dai, Xianjing Zhang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Xieyuanli Chen |
IROS | 1 |
| 2022 | A Representation Separation Perspective to Correspondence-Free Unsupervised 3-D Point Cloud Registrationabstract3-D point cloud registration in remote sensing field has been greatly advanced by deep learning-based methods, where the rigid transformation is either directly regressed from the two point clouds (correspondences-free approaches) or computed from the learned correspondences (correspondences-based approaches). Existing correspondence-free methods generally learn the holistic representation of the entire point cloud, which is fragile for partial and noisy point clouds. In this letter, we propose a correspondence-free unsupervised point cloud registration (UPCR) method from the representation separation perspective. First, we model the input point cloud as a combination of pose-invariant representation and pose-related representation. Second, the pose-related representation is used to learn the relative pose w.r.t. a “latent canonical shape” for thesourceandtargetpoint clouds, respectively. Third, the rigid transformation is obtained from the above two learned relative poses. Our method not only filters out the disturbance in pose-invariant representation but also is robust to partial-to-partial point clouds or noise. Experiments on benchmark datasets demonstrate that our unsupervised method achieves comparable if not better performance than state-of-the-art supervised registration methods.The source code will be made public. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Self-supervised rigid transformation equivariance for accurate 3D point cloud registration
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He |
Pattern Recognit. | 2 |
| 2022 | Searching Dense Point Correspondences via Permutation Matrix LearningabstractAlthough 3D point cloud data has received widespread attentions as a general form of 3D signal expression, applying point clouds to the task of dense correspondence estimation between 3D shapes has not been investigated widely. Furthermore, even in the few existing 3D point cloud-based methods, an important and widely acknowledged principle,i.e. one-to-one matching, is usually ignored. In response, this paper presents a novel end-to-end learning-based method to estimate the dense correspondence of 3D point clouds, in which the problem of point matching is formulated as a zero-one assignment problem to achieve a permutation matching matrix to implement the one-to-one principle fundamentally. Note that the classical solutions of this assignment problem are always non-differentiable, which is fatal for deep learning frameworks. Thus we design a special matching module, which solves a doubly stochastic matrix at first and then projects this obtained approximate solution to the desired permutation matrix. Moreover, to guarantee end-to-end learning and the accuracy of the calculated loss, we calculate the loss from the learned permutation matrix but propagate the gradient to the doubly stochastic matrix directly which bypasses the permutation matrix during the backward propagation. Our method can be applied to both non-rigid and rigid 3D point cloud data and extensive experiments show that our method achieves state-of-the-art performance for dense correspondence learning.The code will be released. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Qi Liu 0054 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Learning a Task-Specific Descriptor for Robust Matching of 3D Point CloudsabstractExisting learning-based point feature descriptors are usually task-agnostic, which pursue describing the individual 3D point clouds as accurate as possible. However, the matching task aims at describing the corresponding points consistently across different 3D point clouds. Therefore these too accurate features may play a counterproductive role due to the inconsistent point feature representations of correspondences caused by the unpredictable noise, partiality, deformation, etc., in the local geometry. In this paper, we propose to learn a robust task-specific feature descriptor to consistently describe the correct point correspondence under interference. Born with anEncoder and aDynamicFusion module, our method EDFNet develops from two aspects. First, we augment the matchability of correspondences by utilizing their repetitive local structure. To this end, a special encoder is designed to exploit two input point clouds jointly for each point descriptor. It not only captures the local geometry of each point in the current point cloud by convolution, but also exploits the repetitive structure from paired point cloud by Transformer. Second, we propose a dynamical fusion module to jointly use different scale features. There is an inevitable struggle between robustness and discriminativeness of the single scale feature. Specifically, the small scale feature is robust since little interference exists in this small receptive field. But it is not sufficiently discriminative as there are many repetitive local structures within a point cloud. Thus the resultant descriptors will lead to many incorrect matches. In contrast, the large scale feature is more discriminative by integrating more neighborhood information. But it is easier to be disturbed since there is much more interference in the large receptive field. Compared with the conventional fusion strategy that handles multiple scale features equally, we analyze the consistency of them to judge the clean ones and perform larger aggregation weights on them during fusion. Then, a robust and discriminative feature descriptor is achieved by focusing on multiple clean scale features. Extensive evaluations validate that EDFNet learns a task-specific descriptor, which achieves state-of-the-art or comparable performance for robust matching of 3D point clouds. Zhiyuan Zhang 0002, Yuchao Dai, Bin Fan 0002, Jiadai Sun, Mingyi He |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | VRNet: Learning the Rectified Virtual Corresponding Points for 3D Point Cloud Registrationabstract3D point cloud registration is fragile to outliers, which are labeled as the points without corresponding points. To handle this problem, a widely adopted strategy is to estimate the relative pose based only on some accurate correspondences, which is achieved by building correspondences on the identified inliers or by selecting reliable ones. However, these approaches are usually complicated and time-consuming. By contrast, the virtual point-based methods learn the virtual corresponding points (VCPs) for allsourcepoints uniformly without distinguishing the outliers and the inliers. Although this strategy is time-efficient, the learned VCPs usually exhibit serious collapse degeneration due to insufficient supervision and the inherent distribution limitation. In this paper, we propose to exploit the best of both worlds and present a novel robust 3D point cloud registration framework. We follow the idea of the virtual point-based methods but learn a new type of virtual points called rectified virtual corresponding points (RCPs), which are defined as the point set with the same shape as thesourceand with the same pose as thetarget. Hence, a pair of consistent point clouds,i.e.sourceand RCPs, is formed by rectifying VCPs to RCPs (VRNet), through which reliable correspondences betweensourceand RCPs can be accurately obtained. Since the relative pose betweensourceand RCPs is the same as the relative pose betweensourceandtarget, the input point clouds can be registered naturally. Specifically, we first construct the initial VCPs by using an estimated soft matching matrix to perform a weighted average on thetargetpoints. Then, we design a correction-walk module to learn an offset to rectify VCPs to RCPs, which effectively breaks the distribution limitation of VCPs. Finally, we develop a hybrid loss function to enforce the shape and geometry structure consistency of the learned RCPs and thesourceto provide sufficient supervision. Extensive experiments on several benchmark datasets demonstrate that our method achieves advanced registration performance and time-efficiency simultaneously.The code will be made public. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Mingyi He |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Deep learning based point cloud registration: an overviewabstractPoint cloud registration aims at finding a rigid transformation to align one point cloud to another one. It is a fundamental problem in computer vision and robotics, which has been widely used in various applications, such as 3D reconstruction, SLAM (simultaneous localization and mapping), and autonomous driving. Over the last decades, many researchers have devoted themselves to tackle this challenging problem. Recently, the success of deep learning in high-level vision tasks has been extended to different geometric vision tasks. Various kinds of deep learning based point cloud registration methods have been proposed to exploit different aspects of the problem. However, a comprehensive overview of these approaches is still missing. To this end, in this paper, we summarize recent progress and present a comprehensive overview for deep learning based point cloud registration. We classify the popular approaches into different categories such as, correspondences-based or correspondences-free, effective modules: feature extractor, matching, outlier rejection, and motion estimation. Furthermore, we discuss the merits and demerits in detail. We provide a systematic and compact framework towards currently proposed methods and discuss future research directions. Zhiyuan Zhang 0002, Yuchao Dai, Jiadai Sun |
Virtual Real. Intell. Hardw. | 3 |