EDBT 2026 Demo / reviewers in the wild / expert
Songhua Wu
dblp:197/4723
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Background Signal Characterization Analysis for ACDL/DQ-1 at Multichannel of 532 nmabstractThe Aerosol and Carbon Detection Lidar (ACDL) on board the Atmospheric Environment Monitoring Satellite (DQ-1) has been successfully operationalized, and it possesses the capability of detecting carbon dioxide, aerosols, and clouds around the world. However, in contrast to other spaceborne lidars currently/previously in operation (e.g., the Cloud–Aerosol Lidar and Infrared Pathfinder Satellite Observation-CALIPSO, the Ice, Cloud, and land Elevation Satellite-2 mission-ICESat-2 and so on), the ACDL does not incorporate a background signal monitor. Consequently, it is crucial to develop background signal acquisition algorithm for ACDL and to analyze the background characterization. In this paper, a three-segmented background signal acquisition algorithm is designed for ACDL to ensure that the acquired background signals match the instrument characteristics. Additionally, the multi-channel background signals measured by ACDL during daytime and nighttime are characterized in detail, including the features in single profile, along the orbit, on a monthly, semi-annually basis, and particular regions. The quantitative analysis of the data reveals a decrease in background signal intensity along latitudinal lines during nocturnal periods (<0.0005 V). Concurrently, elevated signal values are observed within specific regions of South Atlantic Anomaly (SAA). Additionally, the daytime background signal exhibits a high degree of sensitivity to the characteristics of the feature, and it is influenced by the relative positions of the Sun and the Earth. These findings corroborate the stability of the background signal and its consistency with the behavior of the background signal collected by CALIPSO. And the background removed signals collected by ACDL showed good agreement (difference less than 0.026 V) in the region of consistent diurnal aerosol loads (18-22 km) under clear-air conditions. The background acquisition algorithm and characterization analysis proposed in this paper provide accurate data and theoretical support for subsequent calibration and product inversion. Additionally, it offers a solution for background signal extraction in the case of spaceborne lidars that are not equipped with a background signal monitor. Fanqian Meng, Junwu Tang, Guangyao Dai, Songhua Wu, Wenrui Long, Kangwen Sun, Xinru He, Xiaoquan Song, Jiqiao Liu, Wei-Biao Chen, Xiuqing Hu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | An Ocean Optical Parameters Inversion Algorithm Based on Pulse Broadening Match With Varying System Impulse Response Functions for ICESat-2abstractThe Ice, Cloud and land elevation Satellite-2 (ICESat-2) carries the new generation photon-counting lidar payload, the Advanced Topographic Laser Altimeter System (ATLAS). After after-pulse correction, the water column profile acquired by ATLAS can be used to derive vertical profiles of ocean optical parameters. However, the stability and robustness of the correction method across different sea areas could be improved. This study analyzes the broadening effect of sea surface signals caused by varying sea states. It employs land signals with broadening characteristics similar to those of the sea surface as the system impulse response function (SIRF) to perform after-pulse correction on water column profiles. By matching pulse widths, the proposed algorithm reduces the impact of sea surface broadening on after-pulse correction results. Internal consistency validation among ATLAS three strong beams is conducted and achieve the correlation coefficient exceeding 0.8 in the Black Sea, Atlantic Ocean, and Pacific Ocean. Meanwhile, the ocean optical parameter inversions are compared and validated with MODIS and BGC-Argo data. The mean relative errors of inversion results in the three sea areas are approximately 20%, representing a reduction of over 30% compared to traditional algorithm. Finally, a sensitivity analysis of the broadening matching algorithm and cumulative length confirms the algorithm’s robustness. The ocean optical parameters inversion algorithm based on pulse broadening matching lays the foundation for high-precision inversion of parameters such as chlorophyll concentration and provides a reference for future data processing in related systems. Zhiyu Zhang 0006, Mingyu Shi, Junwu Tang, Songhua Wu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | High-Precision Inversion of the 180° Volume Scattering Function for Oceanographic Lidar: Airborne Experimental ValidationabstractLidar can provide three-dimensional detection of the subsurface layer of the global ocean and represents the future direction of ocean remote sensing. One equation and two unknowns in elastic scattering lidar greatly affect the inversion accuracy. Based on airborne oceanographic lidar and in-situ measurements, an error analysis of widely used profile parameter inversion methods was performed, and a high-precision inversion method was developed. After careful data preprocessing, error analysis was performed on bio-optical parameters obtained from inversion methods including Collis Slope method, Klett backscatter iteration method, Churnside perturbation retrieval method and empirical models. In response to the inversion accuracy problems, a parameters optimization method has been proposed that does not require the assumption of a lidar ratio or a constant of lidar attenuation coefficient. Inversion results from the airborne lidar under the conditions of Case I waters in the South China Sea show that the method effectively improves the inversion accuracy of the β(π) profile to about 10% at depths over 50 m where the bio-optical parameters underwater change rapidly, and it can avoid misjudging the subsurface phytoplankton layer. Through detailed data preprocessing, profile parameter inversion and error analysis, this study provides a theoretical foundation for the development of future spaceborne oceanographic lidar data products. Peizhi Zhu, Junwu Tang, Xinke Hao, Mingyu Shi, Bingyi Liu, Songhua Wu, Xiaoquan Song |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | Future Spaceborne Oceanographic Lidar: Exploring the Effects of Large Off-Nadir Angles on Signal Dynamic Range and Depth AliasingabstractThe large signal dynamic range affecting the profile recognition of refined structures is one of the major challenges for future spaceborne oceanographic light detection and ranging (lidar) systems. Reduce the intensity of the sea surface signal and ensure that the detector operates in a linear response region, which helps reduce the subsurface signal error and improves the capability to detect weak signals in deep water. As a solution, the off-nadir pointing could reduce the photon counts from the sea surface but leads to depth aliasing. This reduces the vertical resolution and makes it difficult to determine the sea surface’s position and retrieve the thin chlorophyll layer. The lidar signal’s dynamic range is simulated to improve the detection accuracy. Based on the oceanographic lidar simulator, the laser transmission characteristics are analyzed, taking into account various different environmental parameters (including wind speed, sea surface roughness, concentration of whitecaps and bubbles) and lidar specifications (including laser off-nadir angle, divergence angle, and pulsewidth). The results show that increasing the off-nadir angle to 7°–15° can effectively reduce the dynamic range of the sea surface signal by about one order of magnitude, while increasing the aliasing depth by about 4–8 m. Reducing the beam divergence angle is beneficial for accurate inversion of profiles within the limits of engineering realization. Other parameters, such as pulsewidth, wind speed, and sea surface roughness, have little influence on depth aliasing and depth estimation errors. Peizhi Zhu, Junwu Tang, Xiaoquan Song, Huixin He, Mingyu Shi, Bingyi Liu, Songhua Wu, Jiqiao Liu, Keli Zhang |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | A Time-Consistency Curriculum for Learning From Instance-Dependent Noisy LabelsabstractMany machine learning algorithms are known to be fragile on simple instance-independent noisy labels. However, noisy labels in real-world data are more devastating since they are produced by more complicated mechanisms in an instance-dependent manner. In this paper, we target this practical challenge of Instance-Dependent Noisy Labels by jointly training (1) a model reversely engineering the noise generating mechanism, which produces an instance-dependent mapping between the clean label posterior and the observed noisy label and (2) a robust classifier that produces clean label posteriors. Compared to previous methods, the former model is novel and enables end-to-end learning of the latter directly from noisy labels. An extensive empirical study indicates that the time-consistency of data is critical to the success of training both models and motivates us to develop a curriculum selecting training data based on their dynamics on the two models' outputs over the course of training. We show that the curriculum-selected data provide both clean labels and high-quality input-output pairs for training the two models. Therefore, it leads to promising and robust classification performance even in notably challenging settings of instance-dependent noisy labels where many SoTA methods could easily fail. Extensive experimental comparisons and ablation studies further demonstrate the advantages and significance of the time-consistency curriculum in learning from instance-dependent noisy labels on multiple benchmark datasets. Songhua Wu, Tianyi Zhou 0001, Jun Yu 0001, Bo Han 0003, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Rain Measurement Using Nacelle-Mounted Doppler LidarabstractRain measurements based on the remote sensing technique play an important role in many fields. In this article, the DTU-developed continuous-wave SpinnerLidar, deployed on top of a wind turbine, is used to investigate the feasibility of rain measurement. Considering the wind-driven effect on raindrops, a retrieval method is proposed to obtain the rain vector velocity, rain drop size distribution (DSD), and rain intensity from the combined Doppler spectrum of precipitation and wind based on a nacelle-mounted Doppler lidar. To evaluate the performance of the rain identification method, the rain vector velocities and wind speeds identified from the SpinnerLidar are compared with those from the disdrometer and the sonic anemometers, respectively. The results indicate that a nacelle-mounted wind lidar has the ability to observe the rainfall velocity even with low elevation angles of about 20°–30° of the laser beams. The multipeak structure produced by raindrops in the Doppler lidar spectrum was reported for the first time, which may be attributed to the wide range of sizes of raindrops and snowflakes, and their different corresponding falling velocities. Additionally, the rain and wind spectra retrieved from the SpinnerLidar are in good agreement with those calculated from the rain drop size and speed distributions from the disdrometer combined with backscattering models, and the probability density function of wind speeds measured by the sonic anemometer, respectively, on both clear and rainy days. Finally, the errors of the DSD and rain intensity estimation related to the rain attenuation effect, rain echo signals beyond the Rayleigh length of continuous-wave lidar, and the two assumptions underlying the proposed method were analyzed and discussed. The results indicated that under typical atmospheric conditions, the absolute value of the error induced by the two assumptions is mainly less than 0.5 m/s, fluctuating between 0.1 and 0.3 m/s. Charlotte B. Hasager, Jakob Mann, Mikael Sjöholm, Songhua Wu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | A Parametrical Model for Instance-Dependent Label NoiseabstractIn label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited on class-dependent label-noise (wherein all samples in a clean class share the same label transition matrix). However, the CLTM cannot handle the more common instance-dependent label-noise well (wherein the clean-to-noisy label transition matrix needs to be estimated at the instance level by considering the input quality). Motivated by the fact that classifiers mostly output Bayes optimal labels for prediction, in this paper, we study to directly model the transition from Bayes optimal labels to noisy labels (i.e., Bayes-Label Transition Matrix (BLTM)) and learn a classifier to predict Bayes optimal labels. Note that given only noisy data, it is ill-posed to estimate either the CLTM or the BLTM. But favorably, Bayes optimal labels have no uncertainty compared with the clean labels, i.e., the class posteriors of Bayes optimal labels are one-hot vectors while those of clean labels are not. This enables two advantages to estimate the BLTM, i.e., (a) a set of examples with theoretically guaranteed Bayes optimal labels can be collected out of noisy data; (b) the feasible solution space is much smaller. By exploiting the advantages, this work proposes a parametrical model for estimating the instance-dependent label-noise transition matrix by employing a deep neural network, leading to better generalization and superior classification performance. Shuo Yang 0006, Songhua Wu, Erkun Yang, Bo Han 0003, Yang Liu 0018, Min Xu 0001, Gang Niu 0001, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Classification of Cloud Phase Using Combined Ground-Based Polarization Lidar and Millimeter Cloud Radar Observations Over the Tibetan PlateauabstractThe distributions of cloud phases play an important role in influencing the weather and climate system. The characteristics of clouds above the Tibetan Plateau (TP) can profoundly affect regional and global atmospheric circulation. To research the distributions of cloud phases in the TP region, a retrieval algorithm was developed based on the combination of polarization lidar and millimeter cloud radar measurements and applied to the data from a comprehensive field campaign on the central TP in the summer of 2014. The structure and phase of four different types of clouds were retrieved accordingly, which validates the reliability of the algorithm. The result shows that the occurrence frequency of low clouds remains around 50%, which is very high throughout the whole day in Nagqu, Tibetan in summer. The liquid and mixed cloud frequencies are higher in the morning and afternoon, while ice cloud mainly occurs from the afternoon to midnight. Liquid and ice phase distributions show an inverse relationship in the atmospheric layer from 2 to 8 km in height. Meanwhile, the proportion of the liquid phase to the cloud top is significantly higher than that to the cloud body, which indicates that the supercooled water is more likely to appear at the cloud top than in the cloud. The fractional probabilities of the ice phase and liquid phase in the total cloud top phase intersect at about$-26.7\,\,^{\circ }\text{C}$. Yuxuan Bian, Jiafeng Zheng, Songhua Wu, Guangyao Dai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Learning from Noisy Pairwise Similarity and Unlabeled DataabstractSU classification employs similar (S) data pairs (two examples belong to the same class) and unlabeled (U) data points to build a classifier, which can serve as an alternative to the standard supervised trained classifiers requiring data points with class labels. SU classification is advantageous because in the era of big data, more attention has been paid to data privacy. Datasets with specific class labels are often difficult to obtain in real-world classification applications regarding privacy-sensitive matters, such as politics and religion, which can be a bottleneck in supervised classification. Fortunately, similarity labels do not reveal the explicit information and inherently protect the privacy, e.g., collecting answers to “With whom do you share the same opinion on issue $\mathcal{I}$?" instead of “What is your opinion on issue $\mathcal{I}$?". Nevertheless, SU classification still has an obvious limitation: respondents might answer these questions in a manner that is viewed favorably by others instead of answering truthfully. Therefore, there exist some dissimilar data pairs labeled as similar, which significantly degenerates the performance of SU classification. In this paper, we study how to learn from noisy similar (nS) data pairs and unlabeled (U) data, which is called nSU classification. Specifically, we carefully model the similarity noise and estimate the noise rate by using the mixture proportion estimation technique. Then, a clean classifier can be learned by minimizing a denoised and unbiased classification risk estimator, which only involves the noisy data. Moreover, we further derive a theoretical generalization error bound for the proposed method. Experimental results demonstrate the effectiveness of the proposed algorithm on several benchmark datasets. Songhua Wu, Tongliang Liu, Bo Han 0003, Jun Yu 0001, Gang Niu 0001, Masashi Sugiyama |
J. Mach. Learn. Res. | 1 |
| 2022 | Bridging the Gap Between Few-Shot and Many-Shot Learning via Distribution CalibrationabstractA major gap between few-shot and many-shot learning is the data distribution empirically oserved by the model during training. In few-shot learning, the learned model can easily become over-fitted based on the biased distribution formed by only a few training examples, while the ground-truth data distribution is more accurately uncovered in many-shot learning to learn a well-generalized model. In this paper, we propose to calibrate the distribution of these few-sample classes to be more unbiased to alleviate such an over-fitting problem. The distribution calibration is achieved by transferring statistics from the classes with sufficient examples to those few-sample classes. After calibration, an adequate number of examples can be sampled from the calibrated distribution to expand the inputs to the classifier. Specifically, we assume every dimension in the feature representation from the same class follows a Gaussian distribution so that the mean and the variance of the distribution can borrow from that of similar classes whose statistics are better estimated with an adequate number of samples. Extensive experiments on three datasets,miniImageNet,tieredImageNet, and CUB, show that a simple linear classifier trained using the features sampled from our calibrated distribution can outperform the state-of-the-art accuracy by a large margin. Besides the favorable performance, the proposed method also exhibits high flexibility by showing consistent accuracy improvement when it is built on top of any off-the-shelf pretrained feature extractors and classification models without extra learnable parameters. The visualization of these generated features demonstrates that our calibrated distribution is an accurate estimation thus the generalization ability gain is convincing. We also establish a generalization error bound for the proposed distribution-calibration-based few-shot learning, which consists of thedistribution assumption error, thedistribution approximation error, and theestimation error. This generalization error bound theoretically justifies the effectiveness of the proposed method. Shuo Yang 0006, Songhua Wu, Tongliang Liu, Min Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Class2Simi: A Noise Reduction Perspective on Learning with Noisy LabelsabstractLearning with noisy labels has attracted a lot of attention in recent years, where the mainstream approaches are in \emph{pointwise} manners. Meanwhile, \emph{pairwise} manners have shown great potential in supervised metric learning and unsupervised contrastive learning. Thus, a natural question is raised: does learning in a pairwise manner \emph{mitigate} label noise? To give an affirmative answer, in this paper, we propose a framework called \emph{Class2Simi}: it transforms data points with noisy \emph{class labels} to data pairs with noisy \emph{similarity labels}, where a similarity label denotes whether a pair shares the class label or not. Through this transformation, the \emph{reduction of the noise rate} is theoretically guaranteed, and hence it is in principle easier to handle noisy similarity labels. Amazingly, DNNs that predict the \emph{clean} class labels can be trained from noisy data pairs if they are first pretrained from noisy data points. Class2Simi is \emph{computationally efficient} because not only this transformation is on-the-fly in mini-batches, but also it just changes loss computation on top of model prediction into a pairwise manner. Its effectiveness is verified by extensive experiments. Songhua Wu, Xiaobo Xia, Tongliang Liu, Bo Han 0003, Mingming Gong, Nannan Wang 0001, Gang Niu 0001 |
ICML | 1 |