Wenjin Wu

dblp:00/10779 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 9 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Align³GR: Unified Multi-Level Alignment for LLM-based Generative Recommendation
abstract
Large Language Models (LLMs) demonstrate significant advantages in leveraging structured world knowledge and multi-step reasoning capabilities. However, fundamental challenges arise when transforming LLMs into real-world recommendation systems due to semantic and behavioral misalignment. To bridge this gap, we propose Align³GR, a novel framework that unifies token-level, behavior modeling-level, and preference-level alignment. Our approach introduces: Dual tokenization fusing user-item semantic and collaborative signals. Enhanced behavior modeling with bidirectional semantic alignment. Progressive DPO strategy combining self-play (SP-DPO) and real-world feedback (RF-DPO) for dynamic preference adaptation. Experiments show Align³GR outperforms the SOTA baseline by +17.8% in Recall@10 and +20.2% in NDCG@10 on the public dataset, with significant gains in online A/B tests and full-scale deployment on an industrial large-scale recommendation platform.
Wencai Ye, Mingjie Sun, Wenjin Wu, Peng Jiang 0002
AAAI4
2026 CS3: Efficient Online Capability Synergy for Two-Tower Recommendation
abstract
To balance effectiveness and efficiency in recommender systems, multi-stage pipelines commonly use lightweight two-tower models for large-scale candidate retrieval. However, the isolated two-tower architecture restricts representation capacity, embedding-space alignment, and cross-feature interactions. Existing solutions such as late interaction and knowledge distillation can mitigate these issues, but often increase latency or are difficult to deploy in online learning settings. We propose Capability Synergy (CS3), an efficient online framework that strengthens two-tower retrievers while preserving real-time constraints. CS3 introduces three mechanisms: (1) Cycle-Adaptive Structure for self-revision via adaptive feature denoising within each tower; (2) Cross-Tower Synchronization to improve alignment through lightweight mutual awareness between towers; and (3) Cascade-Model Sharing to enhance cross-stage consistency by reusing knowledge from downstream models. CS3 is plug-and-play with diverse two-tower backbones and compatible with online learning. Experiments on three public datasets show consistent gains over strong baselines, and deployment in a large-scale advertising system yields up to 8.36% revenue improvement across three scenarios while maintaining ms-level latency.
Lixiang Wang, Shaoyun Shi, Wenjin Wu
SIGIR4
2025 Learning Multiple User Distributions for Recommendation via Guided Conditional Diffusion
abstract
Recommender systems are increasingly prevalent to provide personalized suggestions and enhance user satisfaction. Typical recommendation models encode users and items as embeddings, and generate recommendations by assessing the similarity between these embeddings. Despite their effectiveness, these embedding-based models struggle with modeling user uncertainty and capturing diverse user interests using a single fixed user embedding. Recent studies have begun to explore a user-distribution paradigm to learn distributions for users. However, this approach employs a single distribution per user, which fails to effectively delineate semantic boundaries, resulting in sub-optimal recommendations. To this end, we propose GCDR, a Guided Conditional Diffusion Recommender model, to learn multiple distributions for each user in this paper. Specifically, GCDR addresses two major challenges: 1) learning disentangled distributions, and 2) learning personalized distributions. GCDR captures inter-user and intra-user distribution properties through conditional and guided diffusion, respectively. It maintains user-specific embeddings to encode long-term interests for conditional diffusion, while for guided diffusion, it incorporates short-term interests encoded from recent interactions with category preferences. To align the diffusion model with the recommendation task, we train GCDR with three loss functions, included the user loss, the recommendation loss and the diffusion loss. Extensive experiments on four real-world datasets show that GCDR is able to learn effective user distributions and is superior to thirteen state-of-the-art baseline methods.
Cheng Wu 0004, Chaokun Wang, Shaoyun Shi, Ziyang Liu 0004, Wang Peng, Wenjin Wu, Peng Jiang 0002
AAAI8
2025 DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System
abstract
Semantic IDs are discrete identifiers generated by quantizing the Multi-modal Large Language Models embeddings, enabling efficient multi-modal content integration in recommendation systems. However, their lack of collaborative signals results in a misalignment with downstream discriminative and generative recommendation objectives. Recent studies have introduced various alignment mechanisms to address this problem, but their two-stage framework design still leads to two main limitations: (1) inevitable information loss during alignment, and (2) inflexibility in applying adaptive alignment strategies, consequently constraining the mutual information maximization during the alignment process.
Wencai Ye, Mingjie Sun, Shaoyun Shi, Wenjin Wu, Peng Jiang 0002
CIKM5
2025 SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
Zhentao Tan, Ben Xue, Jian Jia, Wencai Ye, Shaoyun Shi, Mingjie Sun, Wenjin Wu, Quan Chen 0006, Peng Jiang 0002
ICCV8
2025 LCTEG: A Spatiotemporal Deep Learning Model for Tropical Forest Growth Prediction
abstract
Tropical forest plays a critical role in climate regulation, carbon storage, and biodiversity conservation. Their vulnerability to climate change and human disturbances necessitates accurate, scalable tools for monitoring and predicting forest growth. While recent deep learning models have improved in capturing nonlinear and lagged effects, current remote sensing applications often lack explicit modeling of factor-variant climatic delays and neighborhood interactions. Here we present a remote sensing–oriented deep learning framework named Lag-aware Convolutional Transformer with Error Feedback and Geographic Feature Fusion (LCTEG). LCTEG explicitly models the delayed effects of climatic factors through structured multi-scale convolution, incorporates neighborhood features to learn local conditions, and leverages error-guided attention to improve the stability and interpretability of multi-step predictions. Applied to global tropical forest at 0.1° spatial and 16-day temporal resolution, LCTEG achieved a test R² of 0.966 and reduced MAE, MSE, and RMSE by at least 65%, 86%, and 64% compared to baseline models. Lag analysis confirms delayed climate effects. Perturbation experiments identify precipitation as the strongest driver and show proximity to water bodies as a key spatial factor, highlighting the dominant role of water availability. Furthermore, future projections (2030-2034) show consistent LAI increases under all three SSP scenarios, but with smaller gains under high-emission pathways, suggesting that potential water stress may constrain vegetation growth. Overall, LCTEG serves as a robust, interpretable tool for tropical forest growth prediction and for advancing climate-resilient ecosystem and policy strategies.
Wenjin Wu, Xinwu Li, Jiankang Shi, Bob O'Hara, Linlu Mei
IEEE Trans. Geosci. Remote. Sens.2
2024 Enhancing Recommendation Accuracy and Diversity with Box Embedding: A Universal Framework
abstract
Recommender systems have emerged as an indispensable mean to meet personalized interests of users and alleviate information overload. Despite the great success, accuracy-oriented recommendation models are creating information cocoons, i.e., it is becoming increasingly difficult for users to see other items they might be interested in. Although recent studies start paying attention to enhancing recommendation diversity, models based on point embedding fail to describe the range of user preferences and item features well, which is essential for diversified matching. To this end, we propose LCD-UC , a novel List-Check-Decide framework with UnCertainty masking based on box embedding to improve recommendation diversity with recommendation accuracy maintained. Specifically, LCD-UC creates hypercubes to represent users and items using box embedding for high model flexibility and expressiveness. Then, a hypercube similarity scoring function is designed to measure the similarity between hypercubes representing users and items. To make a balance between the accuracy and diversity of recommendations and achieve personalized diversity needs, we further develop a user-item pairwise attention mechanism as well as a user uncertainty masking mechanism in LCD-UC. Besides, we present two new metrics for better evaluation on recommendation diversity, which address the issue that existing metrics only consider the coverage of categories while ignore the frequency of categories. The extensive experiments on three real-world datasets show that LCD-UC can improve both recommendation accuracy and diversity over three base models, and is superior to six state-of-the-art recommendation models. An online 10-day AB test also demonstrates that LCD-UC can improve the performance of a real-world advertising system.
Cheng Wu 0004, Shaoyun Shi, Chaokun Wang, Ziyang Liu 0004, Wang Peng, Wenjin Wu, Dongying Kong, Han Li 0005, Kun Gai
WWW6
2024 Fast CU patition based on image similarity using neural network
Yinglie Cao, Wenjin Wu, Zhiheng Zhou 0001, Haoqi Xu, Wanlin Yue, Shang Zhuge
Multim. Tools Appl.2
2019 PolSAR Image Semantic Segmentation Based on Deep Transfer Learning - Realizing Smooth Classification With Small Training Sets
abstract
Suffering from speckle noise and complex scattering phenomena, classification results of SAR images are usually noisy and shattered, which makes them difficult to use in practical applications. Deep-learning-based semantic segmentation realizes segmentation and categorization at the same time, and thus can obtain smooth and fine-grained classification maps. However, this kind of methods require large data sets with pixel-wise categorical annotations, which are time consuming and tedious to retrieve. Compared with photographs and optical remote sensing images, manually annotating SAR data is even harder, which results in a delay of using relevant techniques in this field. In this letter, a new data set is proposed to support semantic segmentation for high-resolution PolSAR images. Limited by the aforementioned problems, the data set is only a small one with 50 image patches. Therefore, two transfer learning strategies are proposed, which adopt the fully convolutional network (FCN) and U-net architecture, respectively, and use distinct pretraining data sets to adapt to different situations. The experiments demonstrate the good performance of both methods and a promising applicability of using small training sets. Moreover, although trained with small patches, both networks can perfectly apply on large images. The new data set and methods are hopeful to support various PolSAR applications as baselines.
Wenjin Wu, Hailei Li, Xinwu Li, Huadong Guo, Lu Zhang 0017
IEEE Geosci. Remote. Sens. Lett.1
2018 Obtain the Patterns of Global Forest NPP and its Influence Factors with Google Earth Engine
abstract
Nowadays, we are easy to access time series earth observation data of multiple kinds. However, since the data is so big, to effectively use them, the traditional way of processing has been challenged. Google earth engine is a powerful cloud platform that equipped with petabyte-scale archive of publicly available remote sensed data and high-performance online analysis abilities. In this paper, the patterns of global forest NPP is analyzed using this amazing tool in a data-driven perspective. Nine categories of forest are obtained with the SOM method based on their NPP levels, and their correlation with different climate zones and forest types are obtained. The correlations between their influence factors are also presented, and different patterns are found with interesting indications.
Wenjin Wu, Xuejing Zhao, Xinwu Li
IGARSS1
2018 Millimeter-Wave Ultrahigh Resolution SAR Image Classification Based on a New Feature Set
abstract
Aiming at the problems and prospects in millimeter-wave ultrahigh resolution synthetic aperture radar applications, we have developed a method with a new feature set for sophisticated classification of large images. It includes innovative parameters derived from different kinds of spectral and characteristic signatures, such as the correlation signature, radial spectrum, and angular spectrum. These features can mine repetitive information from the fragmented patterns and enhance the texture description in different aspects. In the experiment, the proposed feature set achieves 89% overall accuracy which is 25% higher compared with the gray-level co-occurrence matrix feature set. The four new features contribute to over 50% of the accuracy improvement with a significant increase of the accuracy for vehicles and show a fair performance for all the categories.
Wenjin Wu, Xinwu Li, Huadong Guo, Lei Liang 0007
IEEE Geosci. Remote. Sens. Lett.1
2018 High-Resolution PolSAR Scene Classification With Pretrained Deep Convnets and Manifold Polarimetric Parameters
abstract
How to jointly use spatial and polarimetric information in PolSAR analysis has long been an open question. Benefiting from advanced architectures and large visual databases, deep convolutional neural networks or deep convnets (DCNNs) can generate high-level spatial features and achieve state-of-the-art performance in image analyses. However, because PolSAR data are not only multiband but also complex valued, these models cannot be easily borrowed to process them. In light of this problem, we develop a new data set to explore the abilities and potentials of DCNN on PolSAR scene classification. We observe that these models learn fixed semantic information in each layer and adapt to a different data type via changing middle-level filters. Instead of detecting colorful patterns, filters for PolSAR data tend to generate features in separate colors, which may naturally enable the network to differentiate polarimetric mechanisms. Therefore, an ensemble transfer learning framework is proposed to incorporate manifold polarimetric decompositions into a DCNN without throwing away the prelearned spatial analytic ability. Different polarimetric parameters can reflect polarization mechanisms in diverse aspects and introduce new discriminative features to enhance the object recognition. The framework achieves 99.5% validation accuracy and may benefit PolSAR applications in a wide spectrum of fields.
Wenjin Wu, Hailei Li, Lu Zhang 0017, Xinwu Li, Huadong Guo
IEEE Trans. Geosci. Remote. Sens.1
2017 A new texture feature set for ultra-high resolution SAR images
abstract
More information can be obtained with the improved spatial resolution of ultra-high resolution (UHR) SAR, whereas the increasing complexity and rich details lead to extreme difficulties in the automatic interpretation. In this paper, we propose a new texture feature set which involves four types of characteristic signatures and nine features to benefit applications of UHR SAR. Experiment based on real data shows that the new feature set can well describe the spatial patterns of UHR SAR in different aspects and has much better performance than the Gray Level Co-occurrence Matrices (GLCM) feature set.
Wenjin Wu, Xinwu Li, Huadong Guo
IGARSS1
2016 Speech enhancement using magnitude and phase spectrum compensation
abstract
Background noise is a severe problem in speech related systems. In order to solve this problem, it is important to eliminate the noise from the noisy speech, which is called speech enhancement. Typical speech enhancement algorithms only operate on the short-time magnitude spectrum, while keeping the short-time phase spectrum unchanged for synthesis. Or only compensate the phase spectrum while keeping the magnitude spectrum unchanged. In this paper, we present a novel method by changing both magnitude and phase spectra to produce a modified complex spectrum. The test of an objective speech quality measure PESQ, and spectrogram analysis had showed that the proposed method can obtain better enhancement performance.
Wenjin Wu, Qin Zhang 0009, Shilei Bai
ICIS2
2016 Unification of SAR image formation and post-processing for environmental remote sensing application
abstract
Aimed at the problems that SAR (Synthetic Aperture Radar) environment parameter inversion and data processing are lack of both overall design and synergy, this paper proposed a novel unification scheme for environmental remote sensing application. Compared to the reality that application follows data processing in the frontend, the performance of application is highlighted as the ultimate objective in the proposed scheme, where the frontend procedures should serve the backend application to get a better result. Frontend processing is divided into three parts: system design, imaging processing and post-processing procedure, while backend processing is categorized as image processing and environment parameter inversion. The concept of unification scheme focuses on the feedbacks between frontend and backend, and some demonstrations are also given in the following sections.
Jie Chen 0009, Huadong Guo, Wei Yang 0004, Xinwu Li, Lu Zhang 0017, Wenjin Wu
IGARSS8
2016 SAR information integrated processing and its application method study
abstract
Although Synthetic Aperture Radar (SAR) can capture rich land cover information as a most important advanced technique in the field of international earth observation, the application effects still limited significantly. One of the reasons is that the study of SAR imaging processing, SAR image processing and SAR applications are usually conducted respectively, and the study on integrating the three processes is lacked. Focusing on the science problem, taking the typical natural distribution targets (surface deformation, sea ice classification) and man-made targets (building complex and collapsed buildings) as examples, some application studies oriented to SAR environmental parameters inversion are conducted, and the information integrated frames and methods are proposed.
Huadong Guo, Jie Chen 0009, Xinwu Li, Chunming Han, Lu Zhang 0017, Guozhuang Shen, Guang Liu 0001, Zhuo Li 0005, Wenjin Wu
IGARSS10
2016 Noncircularity Parameters and Their Potential Applications in UHR MMW SAR Data Sets
abstract
Information containing in the complex data is seldom considered by researchers when dealing with single synthetic aperture radar (SAR) image processing. In 2015, the statistical noncircularity, which indicates the distribution consistency between the real and imaginary parts, has been found to be surprisingly effective when analyzing the ultrahigh-resolution (UHR) millimeter-wave (MMW) SAR data set, particularly for man-made structures. However, the proposed parameter can only measure the overall noncircular level, which is inadequate to separate different noncircular behaviors. Moreover, its extraction method based on the complex generalized Gaussian distribution is very time consuming and thus makes it hard to be applied. Therefore, in this letter, we present a much simpler and more universal way to compute the noncircularity level and propose three specific parameters to specifically denote different noncircular behaviors. We also give two examples based on the real Chinese UHR MMW SAR (CUM-SAR) data set to show the abilities of the noncircularity level parameter and one example to show the preliminary ability of the specific noncircular parameters. Finally, the potentials and future application directions of these parameters are discussed. We believe that noncircularity parameters will benefit various research areas and promote the applications of UHR MMW SAR systems.
Wenjin Wu, Xinwu Li, Huadong Guo, Laurent Ferro-Famil, Lu Zhang 0017
IEEE Geosci. Remote. Sens. Lett.1
2015 Urban Land Use Information Extraction Using the Ultrahigh-Resolution Chinese Airborne SAR Imagery
abstract
The rapid development of synthetic aperture radar (SAR) sensors results in the acquisition of substantial ultrahigh-resolution SAR images. In this paper, we, for the first time, present three scenes of single-polarization SAR images with decimeter resolution obtained by a millimeter-wave (MMW) Chinese airborne SAR system. An innovative framework based on the complex generalized Gaussian distribution (CGGD) model is proposed to extract land use information from them, and three CGGD parameters, including the shape parameter, the non-Gaussianity parameter, and the noncircularity parameter, are selected to identify different kinds of ground objects. It is shown that these parameters can reveal plentiful land surface information and will be extremely helpful for single-polarization SAR image interpretations. Moreover, a decision tree classifier is built to categorize these images into homogenous natural surfaces, vegetation textures, circular man-made targets, and noncircular man-made targets. Several interesting experiments are implemented, and their results well demonstrate the effectiveness of the new framework.
Wenjin Wu, Huadong Guo, Xinwu Li, Laurent Ferro-Famil, Lu Zhang 0017
IEEE Trans. Geosci. Remote. Sens.1
2013 Non-zero mean statistical models for urban area polarization SAR images
abstract
Urban area man-made target detection based on SAR images has been a challenging field for years due to the complicated scattering mechanisms of dense buildings and the poor visual quality of SAR images caused by speckle noises. To overcome the effect of speckle noise, a substantial portion of SAR image processing methods are based on statistical characteristics. In this paper, the importance of using non-zero mean models is discussed in the experiments, and two statistical models are proposed for polarization SAR image processing. Moreover, a non-zero mean index is invented to test the scattering determinacy level. This work will be very helpful for the improvement of urban area information extraction based on fully polarized SAR images.
Wenjin Wu, Huadong Guo, Xinwu Li, Jie Chen 0009, Yixing Ding
IGARSS1
2012 Nonstationary target detection method based on Rician distribution
abstract
A method based on Rician distribution is proposed in this article to detect nonstationary targets in urban area from Synthetic aperture radar (SAR) image. Rician distribution is an adaptive model which could better fit the statistical characters of urban area SAR image than Wishart distribution, and then improve the detection results. The method has been proved to be effective by the experimental results based on E-SAR data. Particularly, the result in azimuth direction is highly appropriate to discriminate man-made targets and natural targets, which will be very meaningful in urban area studies.
Wenjin Wu, Huadong Guo, Xinwu Li
IGARSS1
2011 DREX: Developer Recommendation with K-Nearest-Neighbor Search and Expertise Ranking
abstract
This paper proposes a new approach called DREX (Developer Recommendation with k-nearest-neighbor search and Expertise ranking) to developer recommendation for bug resolution based on K-Nearest-Neighbor search with bug similarity and expertise ranking with various metrics, including simple frequency and social network metrics. We collect Mozilla Fire fox open bug repository as the experimental data set and compare different ranking metrics on the performance of recommending capable developers for bugs. Our experimental results demonstrate that, when recommending 10 developers for each one of the 250 testing bugs, DREX has produced better performance than traditional methods with multi-labeled text categorization. The best performance obtained by two metrics as Out-Degree and Frequency, is with recall as 0.6 on average. Moreover, other social network metrics such as Degree and Page Rank have produced comparable performance on developer recommendation as Frequency when used for developer expertise ranking.
Wenjin Wu, Wen Zhang 0001, Qing Wang 0001
APSEC1