VLDB 2026 Research / reviewers in the wild / expert
Zhenjie Liu
dblp:86/1226
· DBLP profile ↗
16ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise SearchabstractRecent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily stem from entangled appearance-motion representations and unstable inference strategies. In this paper, we introduce ConsistTalk, a novel intensity-controllable and temporally consistent talking head generation framework with diffusion noise search inference. First, we propose an optical flow-guided temporal module (OFT) that decouples motion features from static appearance by leveraging facial optical flow, thereby reducing visual flicker and improving temporal consistency. Second, we present an Audio-to-Intensity (A2I) model obtained through multimodal teacher-student knowledge distillation. By transforming audio and facial velocity features into a frame-wise intensity sequence, the A2I model enables joint modeling of audio and visual motion, resulting in more natural dynamics. This further enables fine-grained, frame-wise control of motion dynamics while maintaining tight audio-visual synchronization. Third, we introduce a diffusion noise initialization strategy (IC-Init). By enforcing explicit constraints on background coherence and motion continuity during inference-time noise search, we achieve better identity preservation and refine motion dynamics compared to the current autoregressive strategy. Extensive experiments demonstrate that ConsistTalk significantly outperforms prior methods in reducing flicker, preserving identity, and delivering temporally stable, high-fidelity talking head videos. Zhenjie Liu, Jianzhang Lu, Cong Liang 0002, Shangfei Wang |
AAAI | 1 |
| 2024 | One-to-Many Appropriate Reaction Mapping Modeling with Discrete Latent VariableabstractIn dyadic interaction, listener reaction generation can be treated as a one-to-many mapping problem since multiple listener reactions can correspond to a given speaker action. The existing methods have not modeled the diversity of contextual factors well and fail to generate diverse appropriate listener reactions. In response, we introduce discrete latent variables to tackle this one-to-many mapping problem. We conducted experiments on the datasets provided by the REACT2024 Challenge, and the results demonstrated that our approach is capable of generating appropriate listening reactions with higher diversity. Our method achieved first place in the offline track and second in the online track. Zhenjie Liu, Cong Liang 0002, Haofan Zhang, Yadong Liu 0003, Caichao Zhang, Jialin Gui, Shangfei Wang |
FG | 1 |
| 2024 | Fusing SAR Images and Social Media Data Through Domain Adaptation and Land Cover Information: A Case of 2017 Houston Flood EventabstractGlobal climate change leads to increasing frequency and severity of urban flood disasters, which restricts human sustainable development. So far, many studies have explored the potential of fusing remote sensing images and social media data for monitoring urban flood disasters. However, the SAR flooded characteristics of different land cover in complex urban environment are generally ignored. In this paper, we develop a new method for the fusion of heterogeneous SAR images and social media data by using domain adaptation and land cover information. The 2017 Houston flood event is taken as a case study for evaluation. According to our experiments, compared with traditional methods of optimal transport and geographic optimal transport, the proposed method can align the geo-tagged tweets with location uncertainty to nearby flooded areas of low-, medium-, and high-intensity developed areas with more remarkable performance. Zhenjie Liu, Jun Li 0009, Javier Plaza, Antonio Plaza |
IGARSS | 1 |
| 2024 | Adaptive Environment Geographic Optimal Transport Based on Remote Sensing Feature AnalysisabstractThe fusion of remote sensing data and social media data effectively enhances the spatiotemporal resolution of available datasets, providing assistance in addressing specific issues. For the integration of heterogeneous data, domain adaptation and Geographic Optimal Transport (GOT) methods have offered significant support. However, within the process of data fusion, the climatic environment of the data collection areas is often less considered, resulting in the omission of certain climate characteristics under remote sensing imagery, which typically significantly impact algorithm precision. In this paper, building upon the GOT, we further explore the impact of variations in different remote sensing parameters on data fusion accuracy under specific climatic conditions, leading to modifications in the GOT formula. In experiments conducted for flood scenarios, we have enhanced the accuracy of heterogeneous data fusion in this context by emphasizing the MNDWI term and removing the NDVI term from the GOT equation. Qiwang Yuan, Zhenjie Liu, Jun Li 0009 |
IGARSS | 2 |
| 2024 | A Novel Multiplatform Spatiotempoal Data Fusion Approach for Remote Sensing Imagery Based on Parameter SelectionabstractSpatiotemporal fusion is an important means to reconstruct the medium spatial resolution remote sensing image series. Presently, many spatiotemporal fusion approaches have been developed and adopted in research on agriculture, ecology, environment, and so on. Although these approaches have achieved remarkable performance in experiments and applications, most of them are designed to fuse all involved bands using the same model with the same parameters, which ignores the band difference. The ignorance may limit the fusion quality for some bands. To address this problem, we propose a novel spatiotemporal data fusion approach based on parameter selection (PSDFA) in this article. The core idea of the newly proposed PSDFA is producing the synthetic image pairs using available data via three means first and then selecting a similar image pair for each band to provide the parameters that are needed for their fusion. The PSDFA can not only be applied in local computers, and its simplified version can also be implemented in Google Earth Engine (GEE), which is a powerful and widely used cloud platform for remote sensing data computing. To test the PSDFA, we conduct two experiments, one in local computers and another in GEE. In local computers, the PSDFA is compared with five state-of-the-art fusion methods on two public Landsat–Moderate Resolution Imaging Spectroradiometer (MODIS) datasets. In GEE, it is used to produce the monthly 30-m image series in two study sites in the USA and compared with another GEE-based fusion approach. The experimental results demonstrate the outstanding performance of the proposed PSDFA in both local computers and GEE. Yunfei Li 0006, Liangli Meng, Zhenjie Liu, Qian Shi 0001, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Optimal Transport Under Land Cover Information Constraints: Fusing Heterogeneous SAR Imagery and Social Media DataabstractThe integration of remote sensing and citizen science offers an unprecedented opportunity for observing the Earth and human activities. Recently, heterogeneous data fusion models have been proposed to align representations and geolocations of remote sensing imagery and social media data, such as optimal transport (OT) and geographic OT (GOT). However, these models generally ignore the differences in remote sensing features of the same geographical phenomenon for different land cover types, which not only affects the fusion accuracy but also leads to the loss of fusion information. In this study, we develop a general model for heterogeneous SAR imagery and social media data fusion based on OT and land cover information, namely, land cover information-constraint OT (LCIOT). Taking the 2017 Houston flood event as a case study, the experimental findings demonstrate that the proposed LCIOT can align 99% of geotagged Twitter data to flooded areas with an average transport distance of 710 m, outperforming the state-of-the-art models, i.e., OT and GOT. By combining the supervised information obtained by LCIOT with the SAR imagery, the overall accuracy (OA) and Kappa of urban flood mapping in the three study areas ranges from 0.77 to 0.82 and 0.51 to 0.61, respectively. Overall, the proposed LCIOT provides a new perspective to solve the issues of discrepancies in data distribution and geolocation uncertainty in the context of heterogeneous data fusion. Zhenjie Liu, Jun Li 0009, Lizhe Wang 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Dual-channel spatio-temporal wavelet transform graph neural network for traffic forecastingabstractTimely and accurate traffic prediction is crucial for public safety and rational allocation of resources such as roads. However, it still remains an open challenge for timely accurate traffic forecasting, due to the highly nonlinear temporal correlation and dynamical spatial dependence of traffic data. In order to fully capture the temporal and spatial dependences, we propose a dual-channel spatio-temporal wavelet transform graph neural network (DSTwave) for traffic forecasting. Specifically, the wavelet transform neural network is used to obtain the low- and high-frequency parts from the original traffic sequence signals, and in order to accurately capture the spatio-temporal dependence of the low- and high-frequency components in the long - and short-term patterns, the dual-channel ST-GCN with trend-seasonal feature decomposition is carefully designed. In addition, Dynamic-adaptive adjacency matrix is introduced, which can flexibly adapt to changing data. A large number of experiments on two real datasets show that the proposed model has high prediction accuracy. Baowen Xu, Chengbao Liu, Zhenjie Liu, Liwen Kang |
IJCNN | 4 |
| 2022 | Automatic and accurate segmentation of peripherally inserted central catheter (PICC) from chest X-rays using multi-stage attention-guided learning
Xiaoyan Wang 0007, Ye Sheng, Chenglu Zhu, Cong Bai, Ming Xia 0005, Zhanpeng Shao, Ruiyi Zhao, Zhenjie Liu |
Neurocomputing | 12 |
| 2022 | Pansharpening-Based Spatio-Temporal Fusion for Predicting Intense Surface ChangesabstractSpatio-temporal fusion is a feasible way to provide synthetic satellite images with high spatial and high temporal resolution simultaneously. Due to its practicability, spatio-temporal fusion has gotten increasing attention, for which many spatio-temporal fusion approaches have been developed. Most spatio-temporal fusion methods follow the “base fine image guided” (BFIG) fusion mode, resulting in the fact that their fusion results are similar to the base fine images. Therefore, these methods can perform well in areas with limited surface changes due to high similarity between the base and the predicted fine images. However, they might not be applicable in areas with intense surface changes. In this article, we develop a pansharpening-based spatio-temporal fusion model (PSTFM) by introducing the pansharpening fusion mode, which is “coarse image guided” (CIG), into spatio-temporal fusion. PSTFM first trains a pansharpening convolutional neural network (CNN), which then fuses the coarse images and reconstructed panchromatic (Pan) images of the predicted time to recover the missing fine images. The newly proposed PSTFM is compared with three representative BFIG spatio-temporal fusion methods on two Landsat–Moderate Resolution Imaging Spectroradiometer (MODIS) datasets, both of which contain intense surface changes. After that, the experimental results are analyzed and discussed in detail. The experiments and the analysis demonstrate that the newly proposed PSTFM has remarkably qualitative and quantitative performance in predicting the intense surface changes while it is mediocre in areas with low surface change intensity. Yunfei Li 0006, Runlin Cai, Jun Li 0009, Zhenjie Liu, Liangli Meng, Lin He 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Moving Ship Optimal Association for Maritime Surveillance: Fusing AIS and Sentinel-2 DataabstractNowadays, a variety of different sources can be combined together to measure and monitor maritime human activities. Reliable data fusion techniques are essential to associate the targets from different systems for maritime surveillance. In particular, the fusion of data from Sentinel-2 satellites and the Automatic Identification System (AIS) has attracted wide attention due to their public availability and complementarity. However, most traditional methods for target association are not suitable for this particular case, due to the time lag phenomenon of Sentinel-2 data. In this study, we first construct two new datasets for the detection of moving ships and their wakes based on Sentinel-2 images. Combined with the detection results obtained by the You Only Look Once (YOLOv5) model, the position and course information of the detected ships are first extracted. After carefully analyzing the time lag phenomenon of Sentinel-2 data, we develop a new domain adaptation-based method for target association based on the fusion of Sentinel-2 and AIS data, called Moving Ship Optimal Association (MSOA). Different from standard domain adaptation methods only for representation alignment, the proposed MSOA is able to align representation, time and position simultaneously. A case study is provided in which the newly proposed method is tested over the Port of Long Beach, USA. Experimental results demonstrate that both moving ships and wakes are well detected. Specifically, our newly proposed MSOA exhibits more accurate and robust performance when compared to traditional methods, and the detected ships without corresponding AIS tracks can also be detected by our MSOA. Moreover, the real sensing time and time lag of Sentinel-2 data are deduced with high accuracy. Overall, it can be concluded that our MSOA provides a new perspective for accurate target association based on heterogeneous data fusion. Zhenjie Liu, Jun Li 0009, Antonio Plaza, Shaoquan Zhang, Lizhe Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Subspace Optimal Transport for Spatial Bias Correction of Social Media Data: A Case Study of 2013 Boulder Flood EventabstractSocial media data generated from individuals provides a unique opportunity to gain valuable insight on information flow, especially for emergency response. However, the inherent limitations associated to these data (particularly, the spatial bias) restrict its precise application. Existing research on spatial bias correction of social media data mainly face two issues: 1) the geographic extent in target domains may be underestimated, and 2) source elements may be transported within inappropriate distance. In this paper, we take 2013 Boulder, Colorado flood event as a case study, and present a new method called subspace optimal transport (SOT). Our proposed SOT aims at transporting biased tweets from dry to real flooded areas with a relatively close distance. Specifically, a comparison between our newly developed SOT and the traditional optimal transport (OT) and geographic optimal transport (GOT) is performed. Experimental results demonstrate that our new SOT method is able to correct the spatially biased geo-referenced tweets, with high precision and excellent computing performance. Zhenjie Liu, Jun Li 0009, Javier Plaza, Antonio Plaza |
IGARSS | 1 |
| 2021 | Balancing topology structure and node attribute in evolutionary multi-objective community detection for attributed networks
Haiping Ma, Zhenjie Liu, Xingyi Zhang 0001, Lei Zhang 0060, Hao Jiang 0023 |
Knowl. Based Syst. | 2 |
| 2021 | Distributed Fusion of Heterogeneous Remote Sensing and Social Media Data: A Review and New DevelopmentsabstractDespite the wide availability of remote sensing big data from numerous different Earth Observation (EO) instruments, the limitations in the spatial and temporal resolution of such EO sensors (as well as atmospheric opacity and other kinds of interferers) have led to many situations in which using only remote sensing data cannot fully meet the requirements of applications in which a (near) real-time response is needed. Examples of these applications include floods, earthquakes, and other kinds of natural disasters, such as typhoons. To address this issue, social media data have gradually been adopted to fill possible gaps in the analysis when remote sensing data are lacking or incomplete. In this case, the fusion of heterogeneous big data streams from multiple data sources introduces significant demands from a computational viewpoint. In order to meet these challenges, distributed computing is increasingly viewed as a feasible solution to parallelize the analysis of massive data coming from different sources (e.g., remote sensing and social media data). In this article, we provide an overview of available and new distributed strategies to address the computational challenges brought by massive heterogeneous data processing and fusion for real-time environmental monitoring and decision-making. The 2013 Boulder (Colorado) flood event is taken as a case study to evaluate several new distributed data fusion frameworks. Experimental results demonstrate that the proposed distributed frameworks are suitable in terms of response time and computational requirements for fusing large-volume heterogeneous data sources. Jun Li 0009, Zhenjie Liu, Xinya Lei, Lizhe Wang 0001 |
Proc. IEEE | 2 |
| 2021 | Geographic Optimal Transport for Heterogeneous Data: Fusing Remote Sensing and Social MediaabstractThe fusion of heterogeneous remote sensing and social media data can fill the gaps in satellite image collections and improve the spatiotemporal resolution of the available data sets. As a result, it is being gradually adopted in multimodal data analytics. Generally, the fusion of heterogeneous geographic data faces the following issues: 1) the probability density functions may differ from different data sources and 2) the geolocations may not be well aligned. The former one can be generally solved by performing an alignment of representations in the source and target domains using, for instance, domain adaptation. The latter issue is seldom considered in the fusion of heterogeneous geographic data. In this article, we present a new method called geographic optimal transport (GOT), which aims at aligning representations and geolocations in a simultaneous fashion. A flood event that took place in 2013 in Boulder, CO, USA, is taken as a case study to evaluate our GOT method. Here, we consider two remote sensing features derived from water indicators, i.e., the normalized difference vegetation index (NDVI) and the normalized difference water index (NDWI), for the fusion of Landsat 8 imagery and Twitter data. A comparison between our newly developed GOT and the traditional optimal transport (OT) is performed. Experimental results demonstrate that the proposed GOT can accurately align spatially biased georeferenced tweets to the flood phenomena, leading to the conclusion that GOT can effectively fuse heterogeneous remote sensing and social media data. Zhenjie Liu, Jun Li 0009, Lizhe Wang 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Spatial Bias Correction of Social Media Data by Exploiting Remote Sensing Knowledge in Data-Deficient RegionsabstractSocial media data have shown great potential for disaster response. However, the inherent limitations associated to these data (particu-larly, the spatial bias) restrict its precise application. In this work, we present a new spatial bias correction method based on remote sensing knowledge and spatio-temporal fusion, named locally optimal transport (LOT). Our method is first tested using a case study (2013 Boulder, Colorado flood event). Then, we apply our method to a 2016 Wuhan flood event to test its accuracy in a data deficient region. Our results show that combining remote sensing features and spatio-temporal fusion can help to address problems with a lack of prior data and limited disaster period data. According to the random ground verification points collected from news, pictures and videos, our new LOT method is able to accurately relocate spatially biased social media data to inundated areas, which are dangerous for users. Zhenjie Liu, Jun Li 0009, Javier Plaza, Antonio Plaza |
IGARSS | 1 |
| 2020 | Community detection in complex networks with an ambiguous structure using central node based link prediction
Hao Jiang 0023, Zhenjie Liu, Chunlong Liu, Yansen Su, Xingyi Zhang 0001 |
Knowl. Based Syst. | 2 |