VLDB 2026 Research / reviewers in the wild / expert
Song Dai
dblp:166/5865
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0003-5413-7635ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HFSM: A Hierarchical Feature Structure-Driven Method for Multisource Sonar Image Registration of Subsea PipelinesabstractSubsea pipelines are prone to exposure due to natural factors such as earthquakes and vortices, which necessitates regular condition monitoring. Multi-beam echo sounders (MBES) can provide high-precision seabed topographic information, while side-scan sonar (SSS) excels at capturing high-resolution seabed texture features. The integration of these two data sources can complement each other, thereby improving the detection accuracy of subsea pipelines. To achieve effective fusion, high-precision spatial registration is required. However, existing registration algorithms still face challenges such as uneven feature point distribution, dependence on prior knowledge, and unstable matching. This paper proposes a multi-source sonar image registration algorithm for subsea pipelines, named A Hierarchical Feature Structure-Driven Method for Multi-Source Sonar Image Registration of Subsea Pipelines (HFSM). First, the method designs a grid-based multi-scale corner detection (MS-CD), which effectively enhances the spatial distribution balance of feature points. Next, a multi-window geometric-texture joint feature descriptor (MW-GTD) is proposed, which combines direction-sensitive curvature and spatial shadow distribution features within different scale windows. Finally, a multi-layer coarse-to-fine guided matching strategy (ML-CFGM) is introduced to enhance the matching stability of images in feature-sparse regions and realize multi-layer feature matching. The superiority of the proposed method is validated with real-world data, providing technical support for the efficient registration of MBES and SSS images and subsea pipeline detection. Xue-rong Cui, Juan Li 0009, Song Dai, Bin Jiang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Physics-Guided Joint Multisource GNSS-R Soil Moisture Retrieval in the Yellow River DeltaabstractGlobal navigation satellite system reflectometry (GNSS-R) is widely used for soil moisture retrieval. To mitigate heterogeneity between delay–doppler map (DDM) and auxiliary features and to impose physical constraints in fusion modeling, we propose a physics-guided cross-feature fusion network (PCF-Net) for the Yellow River Delta, comprising a multi-feature input, cross-feature fusion, and a physics-constrained retrieval module. Specifically, the multi-feature input module employs convolutional neural networks (CNNs) to extract spatial–local features from DDM and deep semantic information from auxiliary data, enabling structural alignment and unified embedding of heterogeneous modalities. The cross-feature fusion module adopts mamba-based cross-feature fusion (MCF) block and transformer-based cross-feature fusion (TCF) block to achieve bidirectional interaction and deep coupling between one-dimensional auxiliary features and DDM image features. The physics-constrained retrieval module constructs a physics-consistency loss between measured and simulated GNSS-R surface reflectivity to regularize the training process. We further construct a multi-source GNSS-R dataset from cyclone global navigation satellite system (CYGNSS), Tianmu-1, and fengyun-3E (FY-3E) and validate against soil moisture active passive (SMAP). The joint use of multi-source GNSS-R data markedly improves retrieval accuracy, achieving R of 0.9805 and RMSE of 0.0176 cm³/cm³, demonstrating effectiveness at the regional scale. Song Dai, Dongmei Song, Sumiya Erdenesukh, Ronghan Xu, Mei Yong, Bin Wang 0010, Yuhai Bao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | CCEnd-Net: Cross-Modal Cascaded Encoder-Decoder Network for Multisource Data Fusion ClassificationabstractMulti-source data fusion offers great potential for land cover classification. However, the substantial differences in data structures and content representations across various remote sensing sources present significant challenges in heterogeneous feature extraction and information fusion, ultimately constraining the effectiveness of fusion-based classification. To address the aforementioned limitations, this article proposes a multi-source data fusion classification method based on a cross-modal cascaded encoder-decoder network (CCEnd-Net). The proposed algorithm comprises two primary components: a blur feature extraction module and a cascaded encoder-decoder feature fusion module. Specifically, the blur feature extraction module utilizes multi-dimensional convolution to extract deep features from multi-source data and incorporates a blur pooling module to enhance aliasing resistance. This approach mitigates the original spatial discrepancies among heterogeneous data while preventing feature distortion. Meanwhile, the cascaded encoder-decoder feature fusion module reconstructs multi-source data features by integrating a multi-scale channel-spatial interaction attention (MCIA)-enhanced convolutional neural networks (CNNs) with a transformer-based cross-modal fusion (TCMF) block. Additionally, a multi-level fusion strategy is employed to comprehensively exploit the complementary information from different remote sensing data sources, thereby improving classification performance. To validate the effectiveness of the proposed method, comprehensive experiments were conducted on four benchmark datasets covering light detection and ranging (LiDAR), hyperspectral imaging (HSI), synthetic aperture radar (SAR), and very high resolution (VHR) data. The proposed CCEnd-Net was systematically evaluated against state-of-the-art models, including transformer-based architectures, CNNs, and traditional classifiers. The experimental results demonstrate that CCEnd-Net achieves superior performance across all datasets. Song Dai, Dongmei Song, Bin Wang 0010, Weimin Chen 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Annotation-Free, High-Fidelity SAR Oil-Spill Image Synthesis via Classification-Guided Diffusion ModelabstractSynthetic Aperture Radar (SAR) imagery is indispensable for rapid, weather-independent marine oil spill monitoring. Yet the acute shortage of annotated SAR spill imagery severely limits deep learning detectors. While Generative Adversarial Networks (GANs) have been used to generate synthetic data, their inherent limitations—training instability and mode collapse—often result in blurred and semantically inconsistent outputs. To overcome these challenges, we introduce a Classification-Guided Diffusion Model (CG-DM), which integrates the expressive power of diffusion processes with task-specific guidance. CG-DM incorporates three key innovations: (i) Morphology-Aware Classification Guidance: SAR oil spill images are categorized into different morphological categories (blocky, elongated and patchy). This category information conditions every step of the reverse-diffusion process, enabling fine-grained control over the global geometry of generated spills while preserving intra-category diversity. (ii) Label-Synchronized Generation: The model simultaneously generates the SAR image and its corresponding pixel-level annotation mask, which eliminates the need for the time-consuming and error-prone manual labeling process. (iii) Spatial-Aware Attention Mechanism: A lightweight attention mechanism performs localized window self-attention with relative positional offsets, which significantly enhances the sharpness of spill edges and the fidelity of speckle-textured details. Evaluated on the M4D benchmark (ITI/EMSA’s semantic segmentation dataset with Sentinel-1 SAR imagery from 2015–2017), CG-DM achieves a Fréchet Inception Distance (FID) of 203.14 and a Kernel Inception Distance (KID) of 0.124, surpassing state-of-the-art GAN baselines by substantial margins. Besides, ablation studies confirm the critical contributions of each component: spatial-aware attention mechanism significantly enhancing generation quality, while classification guidance effectively preserving morphological feature diversity. Crucially, under extreme data scarcity, training augmentation with CG-DM synthetic samples improves oil-spill detection IoU by up to 81%, demonstrating strong practical utility. This study establishes the first annotation-free paradigm for SAR oil-spill data generation, paving the way for high-accuracy maritime disaster monitoring systems. Bin Wang 0010, Song Dai, Dongmei Song, Weimin Chen 0005 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Few-Shot Learning With Label Smoothing and Metric Space Optimization for Hyperspectral Image ClassificationabstractDue to the high operational complexity of manual sample labeling for hyperspectral images (HSIs), few-shot learning (FSL) has been introduced to cope with the lack of training samples in the field of HSI classification and achieved good results by virtue of its excellent performance. However, factors such as category labeling noise, category distinguishability in the metric space, and completeness of effective feature mining are still the primary considerations that significantly affect the stability and robustness of FSL. Therefore, this study proposes a framework of FSL with label smoothing and metric space optimization (LMFSL) for HSI classification. The framework first incorporates a category label smoothing strategy into FSL, which mitigates the effect of noise by constructing a novel category label smoothing module (CLSM), to reduce the confidence of the classifier. Meanwhile, the study also designs a metric space optimization module (MSOM), which prompts similar samples in the metric space to largely aggregate together by maximizing the intraclass similarity and minimizing the interclass similarity in a more flexible way, so as to improve the decision boundary of the model and enhance the recognition performance of the model. Furthermore, to achieve effective feature extraction in the context of a few labeled samples, a lightweight feature extractor, LFE-MAs, incorporating multiple attention mechanisms is designed to extract the HSI features with high efficiency and low computational cost. Experimental results on four public HSI datasets show that LMFSL outperforms other state-of-the-art methods in terms of classification accuracy with limited labeled samples and has lower computational complexity. Dongmei Song, Fuhou Qin, Bin Wang 0010, Song Dai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Extracting news content with visual unit of web pagesabstractThe Document Object Model (DOM) provides a tree structure called DOM tree for representing with objects in HTML. Many researchers have considered using leaf nodes of DOM tree as basic objects in extracting information from web pages. However, web pages are more of information blocks which each have a consistent visual format rather than individual DOM tree nodes. And those information blocks do not necessarily have a direct map to DOM tree nodes. In this paper, we propose a visual oriented extraction method that extracts news content by visual unit (vu, for short). Visual units are identified by a top-down approach based on visual features and text features. After that, page content is extracted according to domain characteristic. In experiments, the proposed approach achieves 94.86% accuracy over 700 news web pages from 7 different news sites. The result demonstrates that our method represents a promising approach for news content extraction with visual units and domain characteristic. Song Dai, Zhiguo Lu |
SNPD | 2 |