EDBT 2026 Demo / reviewers in the wild / expert
Sidun Liu
dblp:312/8271
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-7715-4698ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transolver Is a Linear Transformer: Revisiting Physics-Attention Through the Lens of Linear AttentionabstractRecent advances in Transformer-based Neural Operators have enabled significant progress in data-driven solvers for Partial Differential Equations (PDEs). Most current research has focused on reducing the quadratic complexity of attention to address the resulting low training and inference efficiency. Among these works, Transolver stands out as a representative method that introduces Physics-Attention to reduce computational costs. Physics-Attention projects grid points into slices for slice attention, then maps them back through deslicing. However, we observe that Physics-Attention can be reformulated as a special case of linear attention, and that the slice attention may even hurt the model performance. Based on these observations, we argue that its effectiveness primarily arises from the slice and deslice operations rather than interactions between slices. Building on this insight, we propose a two-step transformation to redesign Physics-Attention into a canonical linear attention, which we call Linear Attention Neural Operator (LinearNO). Our method achieves state-of-the-art performance on six standard PDE benchmarks, while reducing the number of parameters by an average of 40.0% and computational cost by 36.2%. Additionally, it delivers superior performance on two challenging, industrial-level datasets: AirfRANS and Shape-Net Car. Sidun Liu, Peng Qiao, Zhenglun Sun, Yong Dou |
AAAI | 2 |
| 2025 | MonoIR: Inpainting and Reconstruction for Monocular Endoscope Deformation ScenesabstractMonocular endoscopic scene reconstruction is challenging due to limited viewpoints and interference from surgical instruments. While 3D Gaussian-based methods are popular for their strong reconstruction capabilities and efficiency, they often rely on sensors or stereo depth, resulting in blurred tissue areas when instruments obstruct the view. To overcome these issues, we propose MonoIR, an inpainting and reconstruction method that produces spatiotemporally consistent videos and high-quality reconstructions. Our approach employs a propagation-based method to inpaint video holes, enhanced by optical flow constraints for robustness and an additional transformer-based module to address detail loss. We also introduce the normal and depth regularization with confidence to improve reconstruction quality. Extensive experiments demonstrate that MonoIR efficiently reconstructs deformed tissues and outperforms both monocular and binocular methods in key metrics. Ziteng Zhang, Sidun Liu, Peng Qiao, Yong Dou |
ICASSP | 3 |
| 2025 | End-To-End Casual Video Reconstruction: Geometry, Pose and MotionabstractFrom casual videos in daily life, humans can effortlessly recognize object shapes, perceive variations in viewpoints, decompose dynamic objects and infer their motions. This suggests that reconstruction algorithms should also, like humans, simultaneously achieve these capabilities. However, existing methods tackle the aforementioned tasks in multiple stages. The phased processing approaches mean that each stage’s performance heavily depends on the preceding one, and the entire reconstruction process cannot be optimized in an end-to-end manner, which limits its overall potential. In response, we present an algorithm, designed to integrate these capabilities for casual videos in an end-to-end manner. Specifically, we represent the 4D scene in a video as the combination of local multi-view depth maps and a shared canonical space, where a continuous bijective mapping is used to model motions between local and canonical space. We also extend the depth estimation network to decompose scenes into static and dynamic parts, which helps to avoid degenerate cases and leads to accurate tracking. Experiments demonstrate that our method performs well on casual videos, and achieves performance comparable to state-of-the-art methods across all tasks. Our project page is https://fullre.github.io/FullRe. Peng Qiao, Sidun Liu, Zongxin Ye, Ziteng Zhang, Zhenglun Sun, Yong Dou |
ICME | 3 |
| 2025 | Only One Stage: A Chemical-Aware Model for Accurate Combustion Chemical Kinetics PredictionabstractThe combustion chemical kinetics simulation, which focuses on the change in species mass fractions during reactions, is vital for clean energy development. Due to the sample complexities introduced by chemical kinetics, current multi-stage methods employ data preprocessing stage to separate it into subspaces, aiming to ease training. However, the current approaches to sample space separation are not effective, which not only affects the training accuracy of the network but also introduces a complex preprocessing procedure. To solve this, we propose a one-stage, end-to-end model with chemical-aware capabilities, using an auto dividing mechanism and spatio-temporal convolution for feature extraction. In hydrogen simulation experiments, our model achieved an L1 error of 10−6level, which is nearly identical to the standard results of numerical calculations, outperforming the usually used Multilayer Perceptron (MLP) method by 41.9 times. Zhenglun Sun, Peng Qiao, Yong Dou, Rongchun Li, Sidun Liu |
ICME | 5 |
| 2025 | Mono3R: Exploiting Monocular Cues for Geometric 3D ReconstructionabstractRecent advances in data-driven geometric multi-view 3D reconstruction foundation models (e.g., DUSt3R) have shown remarkable performance across various 3D vision tasks, facilitated by the release of large-scale, high-quality 3D datasets. However, as we observed, constrained by their matching-based principles, the reconstruction quality of existing models suffers significant degradation in challenging regions with limited matching cues, particularly in weakly textured areas and low-light conditions. To mitigate these limitations, we propose to harness the inherent robustness of monocular geometry estimation to compensate for the shortcomings. Specifically, we introduce a monocular-guided refinement module that integrates monocular geometric priors into multi-view reconstruction frameworks. This integration substantially enhances the robustness of multi-view reconstruction systems, leading to high-quality feed-forward reconstructions. Comprehensive experiments across multiple benchmarks demonstrate that our method achieves substantial improvements in both multi-view camera pose estimation and point cloud accuracy. Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 2 |
| 2025 | Regist3R: Incremental Registration with Stereo Foundation ModelabstractMulti-view 3D reconstruction has remained an essential yet challenging problem in the field of computer vision. While DUSt3R and its successors have achieved breakthroughs in 3D reconstruction from unposed images, these methods exhibit significant limitations when scaling to multi-view scenarios, including high computational cost and cumulative error induced by global alignment. To address these challenges, we propose Regist3R, a novel stereo foundation model tailored for efficient and scalable incremental reconstruction. Regist3R leverages an incremental reconstruction paradigm, enabling large-scale 3D reconstructions from unordered and many-view image collections. We evaluate Regist3R on public datasets for camera pose estimation and 3D reconstruction. Our experiments demonstrate that Regist3R achieves comparable performance with optimization-based methods while significantly improving computational efficiency, and outperforms existing multi-view reconstruction models. Furthermore, to assess its performance in real-world applications, we introduce a challenging oblique aerial dataset which has long spatial spans and hundreds of views. The results highlight the effectiveness of Regist3R. We also demonstrate the first attempt to reconstruct large-scale scenes encompassing over thousands of views through pointmap-based foundation models, showcasing its potential for practical applications in large-scale 3D reconstruction tasks, including urban modeling, aerial mapping, and beyond. Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 1 |
| 2024 | SAM-NeRF: NeRF-Based 3D Instance Segmentation with Segment Anything Model
Linglin Xie, Peng Qiao, Yong Dou, Sidun Liu, Kaijun Yang |
ICANN (2) | 5 |
| 2024 | Dual Dreamer: Extending Single-View Dreamer with Few Shot of Complementary Views
Ziteng Zhang, Peng Qiao, Dou Yong, Sidun Liu |
ICANN (3) | 4 |
| 2024 | Improving Motion Deblur By Multi-Output LearningabstractImage deblurring is an ill-posed task, where exists infinite feasible solutions for blurry images. Modem deep learning approaches usually discard the learning of blur kernels and directly employ end-to-end supervised learning. However, supervised learning can’t handle ill-posed tasks appropriately. It regresses the average thus losing sharp details. Therefore, we propose an extension to the network to learn from stochastic supervisions, where a novel multi-output architecture and Min-Out loss function are designed. Our approach enables the model to output multiple feasible solutions to fit various non-uniform motions. We then propose a novel parameter multiplexing method that reduces computations while improving performance with fewer parameters. After training, the best-performed head is fine-tuned to be used for inference where the sharp label is absent. The proposed approach is evaluated with multiple image deblur models on the GoPro motion deblur dataset. On average, the multi-output extension improves the PSNR by 0.08 dB. When applied to the popular attention-based model Restormer, the multi-output helps it achieve 33.05 dB (+0.13 dB) PSNR without modification on network architecture. Sidun Liu, Peng Qiao, Yong Dou |
ICASSP | 1 |
| 2024 | ParaSurRe: Parallel Surface Reconstruction with No Pose PriorabstractSurface reconstruction from multi-view images without pose prior is challenging. Recent advances integrate incremental Structure from Motion (SfM) pipeline into surface optimization, enabling simultaneous surface reconstruction and pose estimation. However, due to the inherent incremental registration scheme of SfM, the efficiency of these methods is far from satisfactory, e.g., reconstruction of an object captured by 49 images costs over 9 hours using a high-end GPU. Inspired by divide-and-conquer strategy, we present a Parallel Surface Reconstruction method, coined as ParaSurRe, where image collections are divided into non-overlapped clusters and the incremental reconstructions are performed individually in each cluster. Owing to image partitioning, each cluster only accurately reconstructs a part of the surface. The major challenge is to merge multiple partial neural implicit surfaces into a complete one. We propose a confidence-aware surface fusion strategy and a geometry-guided refinement to tackle this issue. Experiments on real-world datasets demonstrate that ParaSurRe reconstructs delicate surfaces from unposed images, and achieves competitive pose estimation performance compared with state-of-the-art methods, with up to 6.5× speedup on a scene captured by 81 images. Zongxin Ye, Sidun Liu, Ziteng Zhang, Peng Qiao, Yong Dou |
ICME | 3 |
| 2024 | AbsGS: Recovering Fine Details in 3D Gaussian Splatting
Zongxin Ye, Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 3 |
| 2024 | ER-SFM: Efficient and Robust Cluster-Based Structure from Motion
Zongxin Ye, Sidun Liu, Peng Qiao, Yong Dou |
PRCV (6) | 3 |
| 2022 | Searching Latent Sub-Goals in Hierarchical Reinforcement Learning as Riemannian Manifold OptimizationabstractHierarchical Reinforcement Learning (HRL) is promising to tackle the long-term sparse reward problem. However, goal conditioned HRL, which decomposes the goal into a series of sub-goals, suffers from sub-goal search inefficiency problems when the observation space is too large. This problem is more severe in a visual observation space, since its high latent dimensions, where the complete dynamics information is preserved, exponentially increase the difficulty of sub-goal search. In view of this, we propose to treat the latent space as a manifold, i.e., a Riemannian manifold. Assisted by the Riemannian manifold optimization, sub-goals can be efficiently searched in the higher-dimensional latent space, with the help of preserving the dynamics information efficiently. Experiments on a series of MuJoCo tasks with visual observation show that the proposed Riemannian manifold optimization, compared with the baseline that directly searches for sub-goals in bounded latent space, improves the success rate by 1.5 times on average. In much higher dimensions where the baseline no longer converges, the success rate of the proposed method is maintained. Sidun Liu, Peng Qiao, Yong Dou, Ruochun Jin |
ICME | 1 |
| 2021 | Ddper: Decentralized Distributed Prioritized Experience ReplayabstractIn off-policy reinforcement learning, prioritized experience replay plays an important role. However, the centralized prioritized experience replay becomes the bottleneck for efficient training. We propose to approximate the centralized prioritized experience replay in a distributed and decentralized way under certain mild assumptions. To be specific, each actor stores samples in its local replay in the same way as prioritized experience replay, the learner fetches a batch of samples from these replays following a certain strategy. We implement a Deep Q-Learning off-policy algorithm upon the proposed framework. The comparison experiments are performed on a commonly used subset of the Atari-57 learning environment. The experimental results show that the proposed framework speeds up training as the number of actors increases. With the same algorithm and hyper-parameter settings, the proposed framework with 16 actors achieves superior performance that Ape-X with 32 and even more actors does. Sidun Liu, Peng Qiao, Yong Dou, Rongchun Li |
ICME | 1 |