EDBT 2026 Demo / reviewers in the wild / expert
Sili Chen
dblp:176/3791
· DBLP profile ↗
7ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0003-5553-2119ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 66% Optimization for machine learning · 17% Video understanding and tracking · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | Towards In-the-wild 3D Plane Reconstruction from a Single Image · CVPR 2025 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | Towards In-the-wild 3D Plane Reconstruction from a Single Image · CVPR 2025 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | Video Depth Anything: Consistent Depth Estimation for Super-Long Videos · CVPR 2025 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
planar surface reconstruction |
0.9 | 1 | 2025 | Towards In-the-wild 3D Plane Reconstruction from a Single Image · CVPR 2025 |
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency |
0.9 | 1 | 2025 | Video Depth Anything: Consistent Depth Estimation for Super-Long Videos · CVPR 2025 |
Computer vision › 3D vision › depth estimation
video depth estimation |
0.9 | 1 | 2025 | Video Depth Anything: Consistent Depth Estimation for Super-Long Videos · CVPR 2025 |
Machine learning › Optimization for machine learning › adaptive optimization
adaptive learning rate |
0.6 | 1 | 2022 | Position-Transitional Particle Swarm Optimization-Incorporated Latent Factor Analysis · IEEE Trans. Knowl. Data Eng. 2022 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.6 | 1 | 2022 | Position-Transitional Particle Swarm Optimization-Incorporated Latent Factor Analysis · IEEE Trans. Knowl. Data Eng. 2022 |
Recommender systems › collaborative filtering
matrix factorization |
0.6 | 1 | 2022 | Position-Transitional Particle Swarm Optimization-Incorporated Latent Factor Analysis · IEEE Trans. Knowl. Data Eng. 2022 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2025 | Towards In-the-wild 3D Plane Reconstruction from a Single Image · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
position-transitional PSO · 1.1particle swarm optimization · 1.1temporal consistency loss · 0.9spatial-temporal head · 0.9pixel-geometry-enhanced embedding · 0.9key-frame inference · 0.9exemplar-guided classification-then-regression · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Video Depth Anything: Consistent Depth Estimation for Super-Long VideosabstractDepth Anything has achieved remarkable success in monocular depth estimation with strong generalization ability. However, it suffers from temporal inconsistency in videos, hindering its practical applications. Various methods have been proposed to alleviate this issue by leveraging video generation models or introducing priors from optical flow and camera poses. Nonetheless, these methods are only applicable to short videos (< 10 seconds) and require a trade-off between quality and computational efficiency. We propose Video Depth Anything for high-quality, consistent depth estimation in super-long videos (over several minutes) without sacrificing efficiency. We base our model on Depth Anything V2 and replace its head with an efficient spatial-temporal head. We design a straightforward yet effective temporal consistency loss by constraining the temporal depth gradient, eliminating the need for additional geometric priors. The model is trained on a joint dataset of video depth and unlabeled images, similar to Depth Anything V2. Moreover, a novel key-frame-based strategy is developed for long video inference. Experiments show that our model can be applied to arbitrarily long videos without compromising quality, consistency, or generalization ability. Comprehensive evaluations on multiple video benchmarks demonstrate that our approach sets a new state-of-the-art in zero-shot video depth estimation. We offer models of different scales to support a range of scenarios, with our smallest model capable of real-time performance at 30 FPS. Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Jiashi Feng, Bingyi Kang |
CVPR | 1 |
| 2025 | Towards In-the-wild 3D Plane Reconstruction from a Single Imageabstract3D plane reconstruction from a single image is a crucial yet challenging topic in 3D computer vision. Previous state-of-the-art (SOTA) methods have focused on training their system on a single dataset from either indoor or outdoor domain, limiting their generalizability across diverse testing data. In this work, we introduce a novel framework dubbed ZeroPlane, a Transformer-based model targeting zero-shot 3D plane detection and reconstruction from a single image, over diverse domains and environments. To enable data-driven models across multiple domains, we have curated a large-scale planar benchmark, comprising over 14 datasets and 560,000 high-resolution, dense planar annotations for diverse indoor and outdoor scenes. To address the challenge of achieving desirable planar geometry on multi-dataset training, we propose to disentangle the representation of plane normal and offset, and employ an exemplar-guided, classification-then-regression paradigm to learn plane and offset respectively. Additionally, we employ advanced backbones as image encoder, and present an effective pixel-geometry-enhanced plane embedding module to further facilitate planar reconstruction. Extensive experiments across multiple zero-shot evaluation datasets have demonstrated that our approach significantly outperforms previous methods on both reconstruction accuracy and generalizability, especially over in-the-wild data. Our code and data are available at: https://github.com/jcliu0428/ZeroPlane. Rui Yu 0002, Sili Chen, Sharon X. Huang, Hengkai Guo |
CVPR | 3 |
| 2024 | MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane ReconstructionabstractThis paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and establishes a plane reconstruction pipeline based on monocular geometric cues, resulting in accurate, robust and scalable 3D plane detection and reconstruction in the wild. Specifically, we first leverage large-scale pre-trained neural networks to obtain the depth and surface normals from a single image. These monocular geometric cues are then incorporated into a proximity-guided RANSAC framework to sequentially fit each plane instance. We exploit effective 3D point proximity and model such proximity via a graph within RANSAC to guide the plane fitting from noisy monocular depths, followed by image-level multi-plane joint optimization to improve the consistency among all plane instances. We further design a simple but effective pipeline to extend this single-view solution to sparse-view 3D plane reconstruction. Extensive experiments on a list of datasets demonstrate our superior zero-shot generalizability over baselines, achieving state-of-the-art plane reconstruction performance in a transferring setting. Our code is available at https://github.com/thuzhaowang/MonoPlane. Wang Zhao 0001, Yishu Li, Sili Chen, Sharon X. Huang, Yong-Jin Liu 0001, Hengkai Guo |
IROS | 5 |
| 2022 | Position-Transitional Particle Swarm Optimization-Incorporated Latent Factor AnalysisabstractHigh-dimensional and sparse (HiDS) matrices are frequently found in various industrial applications. A latent factor analysis (LFA) model is commonly adopted to extract useful knowledge from an HiDS matrix, whose parameter training mostly relies on a stochastic gradient descent (SGD) algorithm. However, an SGD-based LFA model's learning rate is hard to tune in real applications, making it vital to implement its self-adaptation. To address this critical issue, this study firstly investigates the evolution process of a particle swarm optimization algorithm with care, and then proposes to incorporate more dynamic information into it for avoiding accuracy loss caused by premature convergence without extra computation burden, thereby innovatively achieving a novel position-transitional particle swarm optimization (P2SO) algorithm. It is subsequently adopted to implement a P2SO-based LFA (PLFA) model that builds a learning rate swarm applied to the same group of LFs. Thus, a PLFA model implements highly efficient learning rate adaptation as well as represents an HiDS matrix precisely. Experimental results on four HiDS matrices emerging from real applications demonstrate that compared with an SGD-based LFA model, a PLFA model no longer suffers from a tedious and expensive tuning process of its learning rate, and it can achieve even higher prediction accuracy for missing data of an HiDS matrix. On the other hand, compared with state-of-the-art adaptive LFA models, a PLFA model's prediction accuracy and computational efficiency are highly competitive. Hence, it has high potential in addressing real industrial issues. Xin Luo 0001, Ye Yuan 0014, Sili Chen, Nianyin Zeng, Zidong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Momentum-incorporated Latent Factorization of Tensors for Extracting Temporal Patterns from QoS DataabstractQuality-of-service (QoS) of Web services vary over time, making it a significant issue to discover temporal patterns from them for addressing various subsequent analyzing tasks like missing QoS prediction. A Latent factorization of tensors (LFT)-based approach proves to be highly efficient in addressing this issue, which can be built through a stochastic gradient descent (SGD) solver efficiently. However, an SGD-based LFT model frequently suffers low-tail convergence. For addressing this issue, we present a momentum-incorporated latent factorization of tensors (MLFT) model, which integrates a momentum method into an SGD-based LFT model, thereby improving its convergence rate as well as maintaining the prediction accuracy for missing QoS data. Empirical studies on two dynamic industrial QoS datasets show that compared with an SGD-based LFT model, an MLFT model achieves faster convergence rate and higher prediction accuracy. Minzhi Chen, Sili Chen |
SMC | 4 |
| 2019 | An Adaptive Latent Factor Model via Particle Swarm Optimization for High-Dimensional and Sparse MatricesabstractLatent factor (LF) models are greatly efficient in extracting valuable knowledge from High-Dimensional and Sparse (HiDS) matrices which are commonly seen in many industrial applications. Stochastic gradient descent (SGD) is an efficient scheme to build an LF model, yet its convergence rate depends vastly on the learning rate which should be tuned with care. Therefore, automatic selection of an optimal learning rate for an SGD-based LF model is a significant issue. To address it, this study incorporates the principle of particle swarm optimization (PSO) into an SGD-based LF model for searching an optimal learning rate automatically. With it, we further propose an adaptive Latent Factor (ALF) model. Empirical studies on two HiDS matrices from industrial applications indicate that an ALF model obviously outperforms an LF model in terms of convergence rate, and maintains competitive prediction accuracy for missing data. Sili Chen |
SMC | 1 |
| 2019 | An adaptive latent factor model via particle swarm optimization
Qing-Xian Wang 0001, Sili Chen, Xin Luo 0001 |
Neurocomputing | 2 |