EDBT 2026 Demo / reviewers in the wild / expert
Wenting Xu
dblp:136/2908
· DBLP profile ↗
22ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HashRuler: Lightweight Detection of Anomalous Hash Codes for Backdoor Defense
Zhenzhu Chen, Wenting Xu, Lei Zhou 0026, Anmin Fu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID InjectionabstractRecent advances in text-to-image diffusion models have spurred significant interest in continuous story image generation. In this paper, we introduce Storynizor, a model capable of generating coherent stories with strong inter-frame character consistency, effective foreground-background separation, and diverse pose variation. The core innovation of Storynizor lies in its key modules: ID-Synchronizer and ID-Injector. The ID-Synchronizer employs an auto-mask self-attention module and a mask perceptual loss across inter-frame images to improve the consistency of character generation, vividly representing their postures and backgrounds. The ID-Injector utilize a Shuffling Reference Strategy (SRS) to integrate ID features into specific locations, enhancing ID-based consistent character generation. Additionally, to facilitate the training of Storynizor, we have curated a novel dataset called StoryDB comprising 100, 000 images. This dataset contains single and multiple-character sets in diverse environments, layouts, and gestures with detailed descriptions. Experimental results indicate that Storynizor demonstrates superior coherent story generation with high-fidelity character consistency, flexible postures, and vivid backgrounds compared to other character-specific methods. Wenting Xu, Chaoyi Zhao, Keqiang Sun, Qinfeng Jin, Xiaoda Yang, Zeng Zhao, Changjie Fan, Zhipeng Hu |
AAAI | 2 |
| 2025 | TB-HSU: Hierarchical 3D Scene Understanding with Contextual AffordancesabstractThe concept of function and affordance is a critical aspect of 3D scene understanding and supports task-oriented objectives. In this work, we develop a model that learns to structure and vary functional affordance across a 3D hierarchical scene graph representing the spatial organization of a scene. The varying functional affordance is designed to integrate with the varying spatial context of the graph. More specifically, we develop an algorithm that learns to construct a 3D hierarchical scene graph (3DHSG) that captures the spatial organization of the scene. Starting from segmented object point clouds and object semantic labels, we develop a 3DHSG with a top node that identifies the room label, child nodes that define local spatial regions inside the room with region-specific affordances, and grand-child nodes indicating object locations and object-specific affordances. To support this work, we create a custom 3DHSG dataset that provides ground truth data for local spatial regions with region-specific affordances and also object-specific affordances for each object. We employ a Transformer Based Hierarchical Scene Understanding (TB-HSU) model to learn the 3DHSG. We use a multi-task learning framework that learns both room classification and learns to define spatial regions within the room with region-specific affordances. Our work improves on the performance of state-of-the-art baseline models and shows one approach for applying transformer models to 3D scene understanding and the generation of 3DHSGs that capture the spatial organization of a room. The code and dataset are publicly available. Wenting Xu, Viorela Ila, Luping Zhou, Craig T. Jin |
AAAI | 1 |
| 2025 | ASANet: Scene Text Recognition With Alternate Self-AttentionabstractText recognition in complex scenes is a challenging task in Computer Vision. In this paper, we propose an innovative framework for scene text recognition, ASANet, which features an Alternate Attention Enhancement Encoder and a Masked Dual-modal Decoder. The encoder incorporates a 12-layer Alternating Self-Attention Module (ASAM), consisting of both Channel and Spatial Blocks, which significantly enhance the depth and breadth of feature extraction. The decoder employs a strategy that combines masking and sequence alignment modeling, effectively improving character context relevance and prediction accuracy. Extensive experimental results demonstrate that ASANet achieves state-of-the-art performance across several benchmark datasets, with a notable accuracy of 91.2% on our self-constructed Uyghur text dataset, highlighting its superior performance. Wenting Xu, Elham Eli, Alimjan Aysa, Xuebin Xu, Kurban Ubul |
ICASSP | 1 |
| 2025 | A comprehensive review of non-Latin natural scene text detection and recognition techniques
Elham Eli, Wenting Xu, Hornisa Mamat, Alimjan Aysa, Kurban Ubul |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Hybrid Reinforced Medical Report Generation With M-Linear Attention and Repetition PenaltyabstractTo reduce doctors' workload, deep-learning-based automatic medical report generation has recently attracted more and more research efforts, where deep convolutional neural networks (CNNs) are employed to encode the input images, and recurrent neural networks (RNNs) are used to decode the visual features into medical reports automatically. However, these state-of-the-art methods mainly suffer from three shortcomings: 1) incomprehensive optimization; 2) low-order and unidimensional attention; and 3) repeated generation. In this article, we propose a hybrid reinforced medical report generation method with m-linear attention and repetition penalty mechanism (HReMRG-MR) to overcome these problems. Specifically, a hybrid reward with different weights is employed to remedy the limitations of single-metric-based rewards, and a local optimal weight search algorithm is proposed to significantly reduce the complexity of searching the weights of the rewards from exponential to linear. Furthermore, we use m-linear attention modules to learn multidimensional high-order feature interactions and to achieve multimodal reasoning, while a new repetition penalty is proposed to apply penalties to repeated terms adaptively during the model's training process. Extensive experimental studies on two public benchmark datasets show that HReMRG-MR greatly outperforms the state-of-the-art baselines in terms of all metrics. The effectiveness and necessity of all components in HReMRG-MR are also proved by ablation studies. Additional experiments are further conducted and the results demonstrate that our proposed local optimal weight search algorithm can significantly reduce the search time while maintaining superior medical report generation performances. Zhenghua Xu 0001, Wenting Xu, Junyang Chen 0001, Chang Qi, Thomas Lukasiewicz |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | TextViTCNN: Enhancing Natural Scene Text Recognition with Hybrid Transformer and Convolutional Networks
Elham Eli, Wenting Xu, Alimjan Aysa, Hornisa Mamat, Kurban Ubul |
PRCV (7) | 2 |
| 2024 | Double branch synergies with modal reinforcement for weakly supervised temporal action detection
Wenting Xu |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Intra-Inter Region Adaptive Graph Convolutional Networks for skeleton-based action recognition
Wenting Xu, Guocheng Lin |
J. Vis. Commun. Image Represent. | 1 |
| 2024 | Research on visual simulation for complex weapon equipment interoperability based on MBSEabstractAbstract As military reforms continue to develop, the battlefield environment is becoming increasingly complex, and traditional single-service combat methods have evolved into integrated joint and collaborative information operations that break down service boundaries on land, sea, and air. The level of weapon system confrontation has also evolved into a system-to-system confrontation. Traditional document-based system architecture design methods can no longer address the complexity and emergent challenges of weapon system construction. In this paper, based on model-driven system engineering, an open, integrated, model-driven weapon equipment interaction system that supports human interaction was constructed using the SysML modeling language and Magicdraw modeling tool. The Unreal Engine 4 landscape building function was used to construct a virtual battlefield environment, and a communication server was developed using C# language to perform visual simulation of interoperability between weapon systems. Based on model-driven weapon equipment interoperability, visual simulation is used to ensure that the function of the weapon equipment system meets the requirements of combat and the combat effectiveness of the system is maximized. Haigen Yang, Zhun Xia, Linqun Zhu, Luohao Dai, Ruotian Xu, GuiYing Sun, HongYan Yu, Wenting Xu |
Multim. Tools Appl. | 9 |
| 2024 | Spatiotemporal information complementary modeling and group relationship reasoning for group activity recognition
Haigang Deng, Chengwei Li, Wenting Xu |
J. Supercomput. | 4 |
| 2023 | MvCo-DoT: Multi-View Contrastive Domain Transfer Network for Medical Report GenerationabstractIn clinical scenarios, multiple medical images with different views are usually generated at the same time, and they have high semantic consistency. However, the existing medical report generation methods cannot exploit the rich multi-view mutual information of medical images. Therefore, in this work, we propose the first multi-view medical report generation model, called MvCo-DoT. Specifically, MvCo-DoT first propose a multi-view contrastive learning (MvCo) strategy to help the deep reinforcement learning based model utilize the consistency of multi-view inputs for better model learning. Then, to close the performance gaps of using multi-view and single-view inputs, a domain transfer network is further proposed to ensure MvCo-DoT achieve almost the same performance as multi-view inputs using only single-view inputs. Extensive experiments on the IU X-Ray public dataset show that MvCo-DoT outperforms the SOTA medical report generation baselines in all metrics. Xiangtao Wang, Zhenghua Xu 0001, Wenting Xu, Junyang Chen 0001, Thomas Lukasiewicz |
ICASSP | 4 |
| 2023 | PSDP: Pseudo-supervised dual-processing for low-dose cone-beam computed tomography reconstruction
Lianying Chao, Wenqi Shan, Wenting Xu, Haobo Zhang 0003, Zhiwei Wang 0002, Qiang Li 0018 |
Expert Syst. Appl. | 4 |
| 2023 | Surgical action detection based on path aggregation adaptive spatial network
Zhen Chao, Wenting Xu, Ruiguo Liu, Hyosung Cho, Fucang Jia |
Multim. Tools Appl. | 2 |
| 2022 | Sparse-view cone beam CT reconstruction using dual CNNs in projection domain and image domain
Lianying Chao, Zhiwei Wang 0002, Haobo Zhang 0003, Wenting Xu, Peng Zhang 0106, Qiang Li 0018 |
Neurocomputing | 4 |
| 2022 | Dual-domain attention-guided convolutional neural network for low-dose cone-beam computed tomography reconstruction
Lianying Chao, Peng Zhang 0106, Zhiwei Wang 0002, Wenting Xu, Qiang Li 0018 |
Knowl. Based Syst. | 5 |
| 2021 | Dispersed Computing for Tactical Edge in Future Wars: Vision, Architecture, and ChallengesabstractIn the future, the tactical edge is far away from the command center, the resources of communication and computing are limited, and the battlefield situation is changing rapidly, which leads to the weak connection and fast changes of network topology in a harsh and complex battlefield environment. Thus, to meet the needs of communication and computing to build a new generation of computing architecture for real‐time sharing and service collaboration of tactical edge resources to win the future war, the dispersed computing (DCOMP) seeks a new solution to satisfy the requirements of fast and efficient sensing, transmission, integrating, scheduling, and processing of various information in the tactical edge. Through the research of a traditional computing paradigm of mobile cloud computing (MCC), fog computing (FC), mobile edge computing (MEC), mobile ad hoc network (MANET), etc., it can be found that these computations have difficulty in meeting the high changing and complex battlefield environment and we propose a novel architecture of DCOMP to build a scalable, extensible, and robust decision‐making system, to realize powerful and secure communication, computing, storage, and information processing capabilities for the tactical edge. We illustrate the fundamental principles of building a network model, channel allocation, and forwarding control mechanism of the network architecture for DCOMP called DANET and then design a new architecture, programming model, task awareness, and computing scheduling for DCOMP. Finally, we discuss the main requirements and challenges of DCOMP in future wars. Haigen Yang, GuiYing Sun, Xiangxin Meng, HongYan Yu, Wenting Xu, Xiaokun Ying |
Wirel. Commun. Mob. Comput. | 7 |
| 2004 | An effective field method of crop proportion survey in China based on GVG integrated systemabstractWith the great agriculture population and limited cropland, it is very important to estimate the output of grain produce in China by remote sensing technology. However, the smallholders of cropland can plant what they like, thus it is difficult to monitor the crop planting proportion with only RS images, even with IKONOS/QUICKBIRD data. In GVG agro-status sampling system, the video camera connected with a notebook by a video capture card and GPS receive card are integrated into the GIS environment. GVG is fixed on a motor and restore the crop pictures and their GPS data when the car is moving along the country road derived from the linear sampling frame in a plantation division unit. A great many of crop pictures along the sample lines are obtained on the field in a limited time, and then all pictures with geographical data are interpreted to calculate the ratio of each type of crop plantation. The crop proportion of plantation division unit is estimated by all picture's plantation ratio of sample line due to this unit. This literature review has demonstrated the GVG systems hardware's constitute, working principle and the case studies. GVG agro-status sampling system not only can be acquire every crop's planting proportion of large areas in short time, but also can check up the results, which get from remote sensing crop classification. Yichen Tian, Bingfang Wu, Wenbo Xu 0004, Jianxi Huang, Wenting Xu |
IGARSS | 5 |
| 2004 | A ruled-based approach to evaluate soil loss at catchments level in Miyun hilly regionabstractMiyun Reservoir is located in the northeast of Beijing and it is the most important drinking water resource of the city. Soil and water loss in this area directly affects local eco-environment and people's life. The soil and water conservation project has been launched out to combat the degrading environment in the upper reach of Miyun Reservoir basin. A ruled-based approach based on the objective of catchments, which is the unit of most soil conservation project, was applied to evaluate soil loss. For each heterogeneous hilly valley in the study area, a set of knowledge-based rules was formulated with remotely sensed images, land use map, DEM and ground investigated data. The relevant parameters, such as slope, vegetation fraction, ravine density, and rainfall distribution, which are also the input parameters of the widely applied universal soil loss equation (USLE), were scaling to the object properties to count this ruled-based model. Finally, all catchments were grouped into four grades according to the soil loss intensity, namely very severe, severe, moderate and slight, the result of the study was practicable to support to make the soil conservation planning. Bingfang Wu, Yichen Tian, Wenting Xu, Jianxi Huang, Wenbo Xu 0004 |
IGARSS | 3 |
| 2004 | Combining Spot4-vegetation and meteorological data derived land cover map in ChinaabstractThe global version of the 1km spatial resolution land cover map have been finished at the end of 2003, which is initiated by the European Commission's Joint Research Center, named Global Land Cover 2000 Project (GLC-2000). As a part of GLC2000, the China window has been developed with the 10-day composite SPOT VGT NDVI data over a period of 01 January 2000 to 31 December 2000, DEM and the Meteorological data (Multi-annual average temperature, multi-annual average precipitation data) collected from 313 weather stations distributed over the China from 1971 to 2000. In order to remove cloud contamination and interpolate the missing data masked by cloud, the Harmonic Analysis of Time Series (HANTS) was applied to NDVI data. With the assistance of Erdas ISODATA algorithm, the classification has been carried out, and 22 types of land cover has labeled in the whole China by interpreting according to the Land Cover Classification System (LCCS) developed by the UN Food and Agriculture Organization (FAO) in the framework of the AFRICOVER project. Preliminary comparisons with the statistic data from Chinese Statistics Bureau and TM data show very promising results, and the accuracv assessment of the GLC-2000 is underway. Bingfang Wu, Wenting Xu, Changzhen Yan, Wenbo Xu 0004 |
IGARSS | 2 |
| 2004 | Mapping plant diversity of broad-leaved forest ecosystem using Landsat TM dataabstractThere is an urgent need to monitor and evaluate the forest ecosystem that supports richer assemblages of biological species in order to preserve the habitats and protect the greatest number of species. Remotely sensed data hold tremendous potential for mapping species habitats and indicators of biological diversity, such as species richness. The objective of this research was to develop the empirical relationships between Landsat-5 TM spectral response and broad-leaved forest inventory parameter of species in the Longmenhe preserve of China. The regression analysis had been used to link the field-measured species richness index and remote sensing data. The accuracy of the resulting mapping was assessed by the field inventory. The results demonstrated that plant diversity can be predicted from satellite remotely sensed data, and the species richness index was highly correlated to TM Tasseled Cap wetness Wenting Xu, Bingfang Wu, Yichen Tian |
IGARSS | 1 |
| 2004 | Detecting vegetation change during the period 1998-2002 in NW China using SPOT-VGT NDVI time series dataabstractThe primary objective of this study was to assess the trends of vegetation changes in west of China since 1998 to 2002 with the 1 km/sup 2/-resolution SPOT-Vegetation 10-day maximum-value synthesized normalized difference vegetation index (NDVI) data. The framework for the analysis is the use of the coefficient of variation (COV) of the decadal NDVI as a measure of vegetative biomass change. A higher NDVI COV for a given pixel represents a greater change in vegetation biomass in the ground area represented by that pixel. And then a linear regression was used to determine the trend of COV value for each pixel over the 5-year period. The slope of the linear regression can be gained and acted as the criterion for the change direction. Pixels with a negative slope are considered to represent ground areas with decreasing amounts of vegetation, vice versa. Results showed vegetation ecosystems in west of China are undergoing accelerated change due to natural and anthropogenic disturbances since 1998 had the increase trends in the five years. Wenting Xu, Bingfang Wu, Changzhen Yan |
IGARSS | 1 |