Yizhen Wu

dblp:157/9277 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Correspondence Calibrating and Dynamic Consistency Learning for Noisy Cross-Modal Retrieval
abstract
Cross-modal retrieval has drawn an increasing amount of attention due to its effective ability for searching semantic relative data points with different modalities. In spite of some progress obtained, such methods often require the data pair maintaining the correct cross-modal correspondence in the training process, which is impractical in real application. To tackle this issue, we propose a Correspondence Calibrating and Dynamic Consistency Learning Network (CCDCL), aiming at optimizing the correspondence of positive samples and deeply investigating the consistency of negative samples. Specifically, to effectively alleviate the false positive issues, co-teaching paradigm is introduced to optimize the correspondence of positive samples by calibrating their confidence scores. To address the false negative sample problem, we propose a Vision-Language Semantic Collaborative Dynamic Margin Adaptation (VLC-DMA) method, which integrates unimodal similarities with calibrated confidence to produce the cross-modal semantic similarities, and consequently employs a dynamic margin function for accurately discriminating false and true negative samples. Experimental results demonstrate that the proposed method effectively improves cross-modal retrieval performance across multiple image-text matching datasets. The code will be available on GitHub upon acceptance of this paper.
Yizhen Wu, Linliang Zhang, Guorui Sheng, Ying Li 0016, Qi Tian 0001
IEEE Trans. Multim.2
2024 Exploring the Drivers of Variations in Daily Nighttime Light Time Series From the Perspective of Periodic Factors
abstract
Periodic factors including lunar phase cycle, sensor viewing angle, and seasonality synergistically affect daily nighttime light (NTL) observations, which brings difficulties to the correction and application of daily NTL data. Here, assuming Lhasa is an ideal observation area, we quantified the periodic characteristics of NTL data with the normalized difference between hotspot and darkspot NTL radiance (NDHD) and further utilized the econometric models to verify the individual and synergistic effects. Our results showed periodic factors differentially affected NTL observations. Sensor viewing angle and lunar phase cycle mainly dominated the variations in NTL time series at the pixel and urban scale, respectively. Then, the econometric model results indicated different periodic factors had a significant individual effect on NTL observations. Lunar phase cycles positively increased observed NTL radiance; as for the sensor viewing angle, the sensor zenith angle (VZA) had a nonlinear link with observed NTL radiance, which was regulated by the sensor azimuth angle (VAA). Spring and summer had a negative impact on NTL observations, whereas the opposite was true for autumn and winter. Moreover, a positive synergistic effect existed between the lunar phase cycle and sensor viewing angle, while the synergistic effects of seasonality and other periodic factors were heterogeneous. Our analysis provides a more comprehensive understanding of the temporal changes in the daily NTL time series.
Yizhen Wu, Xi Li 0016
IEEE Geosci. Remote. Sens. Lett.1
2023 Adaptive Stanley Control Method Based on Dynamic Window Approach
abstract
Stanley algorithm is a classical algorithm in automatic driving path tracking algorithm, which has the advantages of low computational complexity and good tracking effect. However, the tracking accuracy of the traditional Stanley algorithm is limited by the gain parameters and the gain coefficients cannot be dynamically adjusted according to the road conditions. To improve the tracking accuracy and driving stability, an adaptive Stanley control algorithm based on DWA (DWA-Stanley) is proposed, which is used to sampling the velocity vector space and selecting the evaluation function, the vehicle motion planning is based on the best set of velocity values. In this paper, the DWA-Stanley algorithm is validated based on three different driving speeds. Simulation results show that an racecar using the DWA-Stanley algorithm has higher tracking accuracy and smoothness than a conventional Stanley.
Jiaxin Zhuang, Guanrong Huang, Yizhen Wu, Qiang Hua, Bian Gong, Xiaolin Mou
IECON4
2023 Exploring the Correlations Between SNPP-VIIRS Nighttime Light Data and Population From a Multiple Scale Perspective
abstract
Remote sensing nighttime light (NTL) data are widely used for fine-grained estimation of population due to their effectiveness in monitoring anthropic lights at night. However, there is a lack of research exploring the correlations between population and NTL data at various administrative and grid scales, as well as their related impact factors. Thus, to analyze the spatial effect of NTL on NTL-based population estimation, five models (linear, quadratic, exponential, logarithmic, and power function) were used in this letter to examine the correlations between NTL and population at multiscale by using the Visible Infrared Imaging Radiometer Suite data from the National Polar-orbiting Partnership. Results show that the correlations between NTL and population were significantly related to geographic scales. The power function model had the highest fitting accuracy between NTL and population when the grid scales were less than 5 km. The quadratic model had the highest precision on grid scales larger than 5 km. It showed a sharp increase in R2values between 0.5 and 10 km as the scale increase. At different scales, the correlations between NTL and population was negatively impacted by temperature, Normalized Difference Vegetation Index (NDVI), and Road Network Density (RND), but they were positively impacted by the Relief Degree of Land Surface (RDLS). We also found that RDLS had the greatest impact on model accuracy, followed by NDVI and temperature. In addition, RND had a low impact on the accuracy of population estimation on rough grids.
Zhijian Chang, Yizhen Wu, Jingwei Shen, Kaifang Shi
IEEE Geosci. Remote. Sens. Lett.2
2022 Identifying and Quantifying Urban Polycentric Development in China From DMSP-OLS Data and Urban Land Data Sets
abstract
This study attempted to identify and quantify the morphology of intercity urban polycentric development (UPD) from the Defense Meteorological Satellite Program’s Operational Linescan System (DMSP-OLS) data and urban land (UL) data sets in China at the provincial level. The spatiotemporal change and impact factors of UPD from 2000 to 2012 were also evaluated. The accuracy verification results indicated that the UPD could effectively and accurately identify and evaluate from the DMSP-OLS data and UL data sets in China. China’s urban structure presented a UPD trend from 2000 to 2012 and showed a pattern of high values in the eastern region and low values in the western region. In addition, the gross domestic product was proven to be a significant factor and had an inverted U-shaped impact on the UPD. The study can provide accurate time-series UPD data sets for decision-makers to evaluate the spatiotemporal change and driving mechanism of the intercity urban structure in China at the provincial level.
Kaifang Shi, Jingwei Shen, Yizhen Wu, Xuguang Tang
IEEE Geosci. Remote. Sens. Lett.3
2022 Population, GDP, and Carbon Emissions as Revealed by SNPP-VIIRS Nighttime Light Data in China With Different Scales
abstract
Satellite-based artificial nighttime brightness observations are typically considered proxy measures of socioeconomic indicators at large scales, such as population, gross domestic product (GDP), and carbon emissions. However, few studies have explored and compared the correlations between SNPP-VIIRS nighttime light data and socioeconomic indicators from administrative scale to grid scale, and further analyzed the potential mechanisms for the dissimilar correlations at different grid scales. Using regression model, dissimilarity index, and relief amplitude, the quantitative relationship and potential influence mechanism across different scales was investigated in this letter. Results show that the finer the scale is, the lower the correlations between total nighttime lights (NTL) and socioeconomic indicators when comparing 1 km, town, and county scales. The R2values of the NTL-socioeconomic indicator correlations increase sharply with the increase of grid scale at 1–10 km scale. The R2values increase volatilely between 10–30 km but are relatively stable above 30 km. The differences in R2values may be attributed to the diversity and distribution balance of industrial types and relief amplitude at different scales. This letter provides new insights into estimating and predicting population, GDP, and carbon emissions by using SNPP-VIIRS data.
Kaifang Shi, Yizhen Wu, DeRen Li, Xi Li 0016
IEEE Geosci. Remote. Sens. Lett.2
2022 Developing Improved Time-Series DMSP-OLS-Like Data (1992-2019) in China by Integrating DMSP-OLS and SNPP-VIIRS
abstract
Defense Meteorological Satellite Program Operational Linescan System (DMSP-OLS) and Suomi National Polar-orbiting Partnership Visible Infrared Imaging Radiometer Suite (SNPP-VIIRS) data are valuable records of nighttime lights (NTLs) in analyzing socioeconomic development. However, inconsistencies between these data have severely restricted long time-series analyses. Published time-series NTL data sets are not widely available or accurate because the DMSP-OLS calibration is inadequate and some missing data in the SNPP-VIIRS data are seldom considered for patching. To address these issues, we calibrated DMSP-OLS data (1992–2013) by using a quadratic model based on a “pseudo-invariant pixel” method. Thereafter, an exponential smoothing model was used to predict and patch missing data in the monthly SNPP-VIIRS data (2013–2019). Outliers and noise were also removed from the annual data. In addition, a sigmoid model was employed to generate improved simulated DMSP-OLS (SDMSP-OLS) data (2013–2019), which were appended with the calibrated DMSP-OLS data (1992–2013) to develop improved DMSP-OLS-like data (1992–2019) in China. Finally, we qualitatively and quantitatively compared these data with published NTL data to examine data availability. Results showed that choosing invariant pixels to calibrate DMSP-OLS data can minimize discontinuity. The correlation between the SNPP-VIIRS data synthesized by the patched monthly SNPP-VIIRS data and the official annual SNPP-VIIRS data in 2015 ($R^{2} =0.931$) and 2016 ($R^{2} =0.930$) was higher than those of the two existing correction methods with$R^{2}$values below 0.90. Spatial patterns of pixels in the improved SDMSP-OLS data in 2013 were more similar with the DMSP-OLS data than those in the published data. Strong correlations likewise existed between the total (average) pixel values of the improved SDMSP-OLS data (2013–2019) and the DMSP-OLS data in 2012. We also found that the improved DMSP-OLS-like data held strong linear correlations with different statistics, the average$R^{2}$values of which were 0.931 and 0.654 at the national and provincial levels, respectively. Meanwhile, the average regression$R^{2}$values between the two published data sets and statistics were 0.858/0.506 and 0.911/0.611, respectively. Our study has proven that the improved DMSP-OLS-like data (1992–2019) have immense potential to effectively evaluate socioeconomic development and anthropic activities.
Yizhen Wu, Kaifang Shi, Zuoqi Chen, Shirao Liu, Zhijian Chang
IEEE Trans. Geosci. Remote. Sens.1
2021 Self-Supervised Patch Localization for Cross-Domain Facial Action Unit Detection
abstract
Automatic detection of Facial Action Units (AUs) is a fundamental block for objective facial expression analysis. Computer vision-based detection of facial action units is susceptible to variations across corpora. To address this problem, we propose a novel architecture that can be jointly trained for self-supervised optical flow estimation, patch localization, supervised action unit detection, and adversarial domain adaptation. Patch localization allows the encoder to learn local features that are critical to detecting subtle changes caused by the presence of AUs in face. Specifically, an encoder-decoder architecture is used to estimate optical flow from every image. The optical flow is simultaneously used for AU detection, patch localization and adversarial domain adaptation. Majority of the existing work on facial expression analysis is evaluated within corpora. In this paper, we develop and evaluate this novel architecture for facial action unit detection across corpora. The experimental results indicate that our framework improves cross-domain performance (5.5 % F1-score on average), suggesting that the proposed patch localization guides the network to learn a more generalizable representation.
Yufeng Yin 0002, Liupei Lu, Yizhen Wu, Mohammad Soleymani 0001
FG3
2020 Speaker-Invariant Adversarial Domain Adaptation for Emotion Recognition
abstract
Automatic emotion recognition methods are sensitive to the variations across different datasets and their performance drops when evaluated across corpora. We can apply domain adaptation techniques e.g., Domain-Adversarial Neural Network (DANN) to mitigate this problem. Though the DANN can detect and remove the bias between corpora, the bias between speakers still remains which results in reduced performance. In this paper, we propose Speaker-Invariant Domain-Adversarial Neural Network (SIDANN) to reduce both the domain bias and the speaker bias. Specifically, based on the DANN, we add a speaker discriminator to unlearn information representing speakers' individual characteristics with a gradient reversal layer (GRL). Our experiments with multimodal data (speech, vision, and text) and the cross-domain evaluation indicate that the proposed SIDANN outperforms (+5.6% and +2.8% on average for detecting arousal and valence) the DANN model, suggesting that the SIDANN has a better domain adaptation ability than the DANN. Besides, the modality contribution analysis shows that the acoustic features are the most informative for arousal detection while the lexical features perform the best for valence detection.
Yufeng Yin 0002, Baiyu Huang, Yizhen Wu, Mohammad Soleymani 0001
ICMI3