VLDB 2026 Research / reviewers in the wild / expert
Songlin Fei
dblp:65/3328
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0003-2772-0166ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context LearningabstractMultimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs still struggle to make full use of the provided context. While existing work mainly focuses on prompt engineering or post-hoc logit calibration, we study the attention mechanisms inside LVLMs to address their inherent limitations. We identify two important weaknesses in their self-attention that hinder effective ICL. To address these weaknesses, we propose Context-Aware Modulated Attention (CAMA), a training-free and plug-and-play method that dynamically adjusts attention logits based on the input in-context sequence. CAMA uses a two-stage modulation process that strengthens attention to semantically important tokens, especially visual ones. Across four LVLMs and seven benchmarks, CAMA consistently outperforms vanilla models and baselines, showing clear effectiveness and generalization. It can also activate the intended benefits of prompt engineering methods and remains robust across different sequence configurations. Therefore, CAMA opens up new directions for improving multimodal reasoning through a deeper understanding of attention dynamics. Yanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li, Ligong Han, Hongyang He, Zhengtao Yao, Victor Y. Chen, Songlin Fei, Dongfang Liu, Ruixiang Tang |
AAAI | 9 |
| 2025 | TreeStructor: Forest Reconstruction With Neural RankingabstractWe introduceTreeStructor, a novel approach for isolating and reconstructing forest trees. The key novelty is a deep neural model that uses neural ranking to assign pre-generated connectable 3D geometries to a point cloud.TreeStructoris trained on a large set of synthetically generated point clouds. The input to our method is a forest point cloud (FPC) that we first decompose into point clouds that approximately represent trees (TPC) and then into point clouds that represent their parts (PPC). We use a point cloud encoder-decoder to compute embedding vectors that retrieve the best-fitting surface mesh for eachPPCfrom a large set of predefined branch parts. Finally, the retrieved meshes are connected and oriented to obtain individual surface meshes of all trees represented by theFPC. We qualitatively and quantitatively validate that our method can reconstruct forest trees with unprecedented accuracy and visual fidelity.TreeStructoroutperforms the state-of-the-art reconstruction method for around 6% on quantitative metrics and 12% less error compared with QSM on low-quality scanned data. The code and data are available at https://lewkesy.github.io/TreeStructor/. Xiaochen Zhou, Bosheng Li, Bedrich Benes, Ayman Habib 0001, Songlin Fei, Jinyuan Shao, Sören Pirk |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Errata to "TreeStructor: Forest Reconstruction With Neural Ranking"abstractPresents corrections to the paper, (Errata to “TreeStructor: Forest Reconstruction With Neural Ranking”). Xiaochen Zhou, Bosheng Li, Bedrich Benes, Ayman Habib 0001, Songlin Fei, Jinyuan Shao, Sören Pirk |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Unsupervised Machine Learning for Detecting and Locating Human-Made Objects in 3D Point Cloudabstract3D point clouds are unstructured, sparse, and irregular data collected by airborne LiDAR systems over a geological region. Laser pulses emitted from the systems reflect off objects both on and above the ground, resulting in data with the longitude, latitude, and elevation of the points, and the corresponding laser pulse strengths. Ground filtering is important. The aim is to partition the points into ground and non-ground subsets. In addition, this research introduces a novel task: detecting and identifying human-made objects amidst natural tree structures. The task is performed on the non-ground subset derived given by the ground filtering stage. Marked Point Fields (MPFs) are used to these tasks. The proposed methodology consists of three stages: ground filtering, local information extraction (LIE), and clustering. In the ground filtering stage, a statistical method called One-Sided Regression (OSR) is devised to overcome the limitations of prior ground filtering methods on uneven terrains. In the LIE stage, a kernel-based method for the Hessian matrix of the MPF is developed. In the clustering stage, the Gaussian Mixture Model (GMM) is applied to the results of the LIE for partitioning the non-ground points into trees and human-made objects. The underlying assumption is that LiDAR points from trees exhibit a three-dimensional distribution, while those from human-made objects follow a two-dimensional distribution. The Hessian matrix of the MPF effectively captures the difference. Experimental results demonstrate that the proposed ground filtering method outperforms previous techniques, and the LIE method successfully distinguishes between points representing trees and human-made objects. Huyunting Huang, Tonglin Zhang, Baijian Yang 0001, Jin Wei-Kocsis, Songlin Fei |
IEEE Big Data | 6 |
| 2024 | Tree-D Fusion: Simulation-Ready Tree Dataset from Single Images with Diffusion Priors
Jae Joong Lee, Bosheng Li, Sara Beery, Jonathan Huang, Songlin Fei, Raymond A. Yeh, Bedrich Benes |
ECCV (41) | 5 |
| 2024 | Morphological Approach for Forest Woody Debris Detection Using Multi-Platform, Multi-Resolution Lidar DataabstractWoody Debris (WD) plays an important role in forest ecosystems. It provides critical habitat for plants, animals, and insects, but it is also a source of fuel contributing to fire propagation and sometimes leads to catastrophic wildfire. Traditional field surveys for WD assessments are usually restricted to transects and sample plots. Light Detection and Ranging (LiDAR) point clouds emerge as a valuable source for the development of comprehensive WD detection strategies. Although results from previous studies on LiDAR-based WD detection approaches have been promising, there is still a lack of general strategy for handling point clouds acquired by different platforms with varying characteristics (e.g., point density) in different forest types. In this study, we propose a general morphological WD detection strategy which requires few intuitive thresholds, making it applicable to multi-platform LiDAR datasets in both plantation and natural forests. Renato César dos Santos, Sang-Yeop Shin, Raja Manish, Tian Zhou 0001, Songlin Fei, Ayman Habib 0001 |
IGARSS | 5 |
| 2024 | Label-Efficient Video Object Segmentation With Motion CluesabstractVideo object segmentation (VOS) plays an important role in video analysis and understanding, which in turn facilitates a number of diverse applications, including video editing, video rendering, and augmented reality / virtual reality. However, existing deep learning-based approaches rely heavily on a large number of pixel-wise annotated video frames to achieve promising results, which is notoriously laborious and costly. To address this, in this paper, we formulate unsupervised video object detection by exploring simulated dense labels and explicit motion clues. Specifically, we first propose an effective video label generator network based on the sparsely annotated frames and the flow motion between them. It can largely alleviate our dependence and limitation on the sparse labels. Furthermore, we propose a transformer-based architecture to model the appearance and motion clues simultaneously with the cross-attention module, in order to maximally overcome non-linear motion with potential occlusions. Extensive experiments show that the proposed method outperforms recent VOS methods on four popular benchmarks (i.e., DAVIS-16, FBMS, Youtube-VOS and SegTrack-v2). Moreover, the proposed method can be further applied to a wide range of wild scenes such as wild forests and animals. Because of its effectiveness and generalization, we believe that our method could serve as a useful basis for alleviating the dependence on dense annotation in video data. Yawen Lu, Jie Zhang 0066, Su Sun, Zhiwen Cao, Songlin Fei, Baijian Yang 0001, Victor Y. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | MetaSegNet: Metadata-Collaborative Vision-Language Representation Learning for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images plays a vital role in a wide range of Earth Observation applications, such as land-use land-cover (LULC) mapping, environment monitoring, and sustainable development. Driven by rapid developments in artificial intelligence, deep learning (DL) has emerged as the mainstream for semantic segmentation and has achieved many breakthroughs in the field of remote sensing. However, most DL-based methods focus on unimodal visual data while ignoring rich multimodal information involved in the real world. Nonvisual data, such as text, can gather extra knowledge from the real world, which can strengthen the interpretability, reliability, and generalization of visual models. Inspired by this, we propose a novel metadata-collaborative segmentation network (MetaSegNet) that applies vision-language representation learning for the semantic segmentation of remote sensing images. Unlike the common model structure that only uses unimodal visual data, we extract the key characteristic (e.g., the climate zone) from freely available remote sensing image metadata and transfer it into geographic text prompts via the generic ChatGPT. Then, we construct an image encoder, a text encoder, and a crossmodal attention fusion subnetwork to extract the image and text feature and apply image-text interaction. Benefiting from such a design, the proposed MetaSegNet not only demonstrates superior generalization in zero-shot testing but also achieves competitive accuracy with the state-of-the-art semantic segmentation methods on the large-scale OpenEarthMap dataset [70.4% mean intersection over union (mIoU)] and the Potsdam dataset (93.3% mean${F}1$score) as well as the LoveDA dataset (52.0% mIoU). Sijun Dong, Xiaoliang Meng, Shenghui Fang, Songlin Fei |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | DeepTree: Modeling Trees With Situated LatentsabstractIn this article, we propose DeepTree, a novel method for modeling trees based on learning developmental rules for branching structures instead of manually defining them. We call our deep neural model "situated latent" because its behavior is determined by the intrinsic state -encoded as a latent space of a deep neural model- and by the extrinsic (environmental) data that is "situated" as the location in the 3D space and on the tree structure. We use a neural network pipeline to train a situated latent space that allows us to locally predict branch growth only based on a single node in the branch graph of a tree model. We use this representation to progressively develop new branch nodes, thereby mimicking the growth process of trees. Starting from a root node, a tree is generated by iteratively querying the neural network on the newly added nodes resulting in the branching structure of the whole tree. Our method enables generating a wide variety of tree shapes without the need to define intricate parameters that control their growth and behavior. Furthermore, we show that the situated latents can also be used to encode the environmental response of tree models, e.g., when trees grow next to obstacles. We validate the effectiveness of our method by measuring the similarity of our tree models and by procedurally generated ones based on a number of established metrics for tree form. Xiaochen Zhou, Bosheng Li, Bedrich Benes, Songlin Fei, Sören Pirk |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Radiometric And Geometric Approach For Major Woody Parts Segmentation In Forest Lidar Point CloudsabstractSegmenting major woody parts is a critical prerequisite to derive structural and biophysical attributes of trees. Static Terrestrial laser scanning (TLS) has been widely used due to its accurate and non-destructive scanning capability; wood parts segmentation has been experimented using the raw radiometric feature. However, due to the challenges of fixed scanning positions and occlusion, using TLS to capture an entire tree is time-consuming. Additionally, the raw intensity of TLS data cannot accurately represent objects’ physical characteristics. Here, using LiDAR data acquired by an inhouse developed backpack Mobile Mapping System (MMS), we introduce a fast and fully unsupervised method that combines automatic thresholding of normalized radiometric and geometric features to extract major woody parts in the point clouds. We show that using MMS LiDAR data, our method can achieve higher performance than existing methods for major woody parts segmentation on 14 trees with different sizes and species in both leaf-on and leaf-off seasons. Jinyuan Shao, Yi-Ting Cheng, Yerassyl Koshan, Raja Manish, Ayman Habib 0001, Songlin Fei |
IGARSS | 6 |