EDBT 2026 Demo / reviewers in the wild / expert
Hongzhe Liu 0001
dblp:02/3832-1
· DBLP profile ↗
32ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0003-2314-5272ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 14 since 2021Artificial intelligence and machine learning · 10 · 6 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sgd-det:structure-guided oracle character detection in degraded rubbing images
Xuxing Qi, Cheng Xu 0005, Jiazheng Yuan, Bofeng Mo, Hongzhe Liu 0001 |
Multim. Syst. | 8 |
| 2026 | TSPPO: transformer-based sequential proximal policy optimization for multi-agent systems
Yuxiao Gao, Cheng Xu 0005, Hongzhe Liu 0001 |
Multim. Syst. | 4 |
| 2025 | YOLODF: YOLO-Based Spatial-Frequency Interaction Mining for General Deepfake Detection
Xin Li 0184, Bingxin Xu, Hongzhe Liu 0001, Weiguo Pan, Cheng Xu 0005 |
PRCV (7) | 3 |
| 2025 | Self-attention enhanced dynamic semantic multi-scale graph convolutional network for skeleton-based action recognition
Cheng Xu 0005, Songyin Dai, Nuoya Li, Weiguo Pan, Bingxin Xu, Hongzhe Liu 0001 |
Image Vis. Comput. | 7 |
| 2025 | DAN: Distortion-aware Network for fisheye image rectification using graph reasoning
Yongjia Yan, Hongzhe Liu 0001, Cheng Xu 0005, Bingxin Xu, Weiguo Pan, Songyin Dai, Yiqing Song |
Image Vis. Comput. | 2 |
| 2025 | Optimization methods in fully cooperative scenarios: a review of multiagent reinforcement learningabstractMultiagent reinforcement learning (MARL) has become a dazzling new star in the field of reinforcement learning in recent years, demonstrating its immense potential across many application scenarios. The reward function directs agents to explore their environments and make optimal decisions within them by establishing evaluation criteria and feedback mechanisms. Concurrently, cooperative objectives at the macro level provide a trajectory for agents’ learning, ensuring alignment between individual behavioral strategies and the overarching system goals. The interplay between reward structures and cooperative objectives not only bolsters the effectiveness of individual agents but also fosters interagent collaboration, offering both momentum and direction for the development of swarm intelligence and the harmonious operation of multiagent systems. This review delves deeply into the methods for designing reward structures and optimizing cooperative objectives in MARL, along with the most recent scientific advancements in this field. The article meticulously reviews the application of simulation environments in cooperative scenarios and discusses future trends and potential research directions in the field, providing a forward-looking perspective and inspiration for subsequent research efforts. Xinhao Shi, Qinghan Zeng, Cheng Xu 0005, Hongzhe Liu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2025 | Ihenet: an illumination invariant hierarchical feature enhancement network for low-light object detection
Nuoya Li, Weiguo Pan, Bingxin Xu, Hongzhe Liu 0001, Songyin Dai, Cheng Xu 0005 |
Multim. Syst. | 4 |
| 2025 | Generalization-oriented face forgery detection via discriminative feature analysis and normalization
Xin Li 0184, Bingxin Xu, Hongzhe Liu 0001, Weiguo Pan, Cheng Xu 0005 |
Multim. Syst. | 3 |
| 2025 | Flood scenarios vehicle detection algorithm based on improved YOLOv9
Jiwu Sun, Cheng Xu 0005, Hongzhe Liu 0001 |
Multim. Syst. | 6 |
| 2024 | ISE-UFDS: A Dataset for Detecting the Degree of Danger to Vehicles in Urban Flooding and Performance Assessment
Jiwu Sun, Cheng Xu 0005, Hongzhe Liu 0001 |
ICIC (7) | 5 |
| 2024 | TSD-YOLO: Small traffic sign detection based on improved YOLO v8abstractAbstract Traffic sign detection is critical for autonomous driving technology. However, accurately detecting traffic signs in complex traffic environments remains challenge despite the widespread use of one‐stage detection algorithms known for their real‐time processing capabilities. In this paper, the authors propose a traffic sign detection method based on YOLO v8. Specifically, this study introduces the Space‐to‐Depth (SPD) module to address missed detections caused by multi‐scale variations of traffic signs in traffic scenes. The SPD module compresses spatial information into depth channels, expanding the receptive field and enhancing the detection capabilities for objects of varying sizes. Furthermore, to address missed detections caused by complex backgrounds such as trees, this paper employs the Select Kernel attention mechanism. This mechanism enables the model to dynamically adjust its focus and more effectively concentrate on key features. Additionally, considering the uneven distribution of training data, the authors adopted the WIoUv3 loss function, which optimizes loss calculation through a weighted approach, thereby improving the model's detection performance across various sizes and frequencies of instances. The proposed methods were validated on the CCTSDB and TT100K datasets. Experimental results demonstrate that the authors’ method achieves substantial improvements of 3.2% and 5.1% on the mAP50 metric compared to YOLOv8s, while maintaining high detection speed, significantly enhancing the overall performance of the detection system. The code for this paper is located at https://github.com/dusongjie/TSD‐YOLO‐Small‐Traffic‐Sign‐Detection‐Based‐on‐Improved‐YOLO‐v8 Songjie Du, Weiguo Pan, Nuoya Li, Songyin Dai, Bingxin Xu, Hongzhe Liu 0001, Cheng Xu 0005, Xuewei Li 0006 |
IET Image Process. | 6 |
| 2024 | SFDiff: Diffusion model with sufficient spatial-Fourier frequency information interaction for low-light image enhancementabstractAbstract Diffusion models are increasingly applied in low‐light image enhancement tasks due to their exceptional capability to model data distributions, but most current methods focus only on the original pixel space and neglect the potential of Fourier frequency information. In this article, SFDiff is proposed, a novel low‐light image enhancement method that integrates Fourier frequency information into the diffusion process. Specifically, Fourier transforms are applied at both the image and feature levels to separately enhance the amplitude and phase components, which restores global illumination degradation and positional information. Then a Spatial‐Frequency Fusion (SFF) block is used to fully integrate and interact with the information across spatial and frequency domains. Since illumination degradation is primarily manifested in the amplitude component, a loss function based on maximum likelihood learning is employed to constrain the amplitude component at each step of the sampling process, ensuring that the reverse process maintains an optimal trajectory. Owing to the streamlined network design and the fact that the Fourier transform requires no extra parameters, SFDiff achieves a reduction in parameters of over compared to several state‐of‐the‐art (SOTA) diffusion models and delivers high‐quality enhancement results on multiple real‐world datasets. The code is available at https://github.com/MrWan001/SFDiff . Bingxin Xu, Jingli Yao, Weiguo Pan, Hongzhe Liu 0001 |
IET Image Process. | 6 |
| 2024 | FSKT-GE: Feature maps similarity knowledge transfer for low-resolution gaze estimationabstractAbstract The limited of texture details information in low‐resolution facial or eye images presents a challenge for gaze estimation. To address this, FSKT‐GE (feature maps similarity knowledge transfer for low‐resolution gaze estimation) is proposed, a gaze estimation framework consisting of both a high resolution (HR) network and low resolution (LR) network with the identical structure. Rather than mere feature imitation, this issue is addressed by assessing the cosine similarity of feature layers, emphasizing the distribution similarity between the HR and LR networks. This enables the LR network to acquire richer knowledge. This framework utilizes a combination loss function, incorporating cosine similarity measurement, soft loss based on probability distribution difference and gaze direction output, along with a hard loss from the LR network output layer. This approach on low‐resolution datasets derived from Gaze360 and RT‐Gene datasets is validated, demonstrating excellent performance in low‐resolution gaze estimation. Evaluations on low‐resolution images obtained through 2×, 4×, and 8× down‐sampling are conducted on two datasets. On the Gaze360 dataset, the lowest mean angular errors of 10.97°, 11.22°, and 13.61° were achieved, while on the RT‐Gene dataset, the lowest mean angular errors of 6.73°, 6.83°, and 7.75° were obtained. Weiguo Pan, Songyin Dai, Bingxin Xu, Cheng Xu 0005, Hongzhe Liu 0001, Xuewei Li 0006 |
IET Image Process. | 6 |
| 2024 | PSC diffusion: patch-based simplified conditional diffusion model for low-light image enhancement
Bingxin Xu, Weiguo Pan, Hongzhe Liu 0001 |
Multim. Syst. | 4 |
| 2023 | Diversity-induced consensus and structured graph learning for multi-view clustering
Zhibin Gu, Hongzhe Liu 0001, Songhe Feng |
Appl. Intell. | 2 |
| 2023 | Integrating topology beyond descriptions for zero-shot learning
Yutong Gao 0001, Congyan Lang, Yidong Li, Hongzhe Liu 0001, Fayao Liu |
Pattern Recognit. | 6 |
| 2023 | SSR-Net: A Spatial Structural Relation Network for Vehicle Re-identificationabstractVehicle re-identification (Re-ID) represents the task aiming to identify the same vehicle from images captured by different cameras. Recent years have seen various feature learning-based approaches merely focusing on feature representations including global features or local features to obtain more subtle details to identify highly similar vehicles. However, few such methods consider the spatial geometrical structure relationship among local regions or between the global and local regions. By contrast, in this study, we propose a Spatial Structural Relation Network (SSR-Net) that explores the above-mentioned two kinds of relations simultaneously to learn more discriminative features by modeling the spatial structure information and global context information. In this article, we propose to adopt a Graph Convolution Network (GCN), for modeling spatial structural relationships among characteristic features. The GCN model aggregating the local and global features is shown to be more discriminative and robust to several car image transformations. To improve the performance of our proposed network, we jointly combine the classification loss with metric learning loss. Extensive experiments conducted on the public VehicleID and VeRi-776 datasets validate the effectiveness of our approach in comparison with recent works. Zheming Xu, Congyan Lang, Songhe Feng, Tao Wang 0011, Adrian G. Bors, Hongzhe Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | W-ART: Action Relation Transformer for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) is a long-standing and challenging research problem in video signal analysis. It is to localize the action segments in the video given only video-level labels. The key to this task is understanding how the diverse actions interact. In this paper, we propose W-ART, a relation Transformer to explicitly capture the relationships between action segments. We devise a new effective Transformer architecture and construct new training loss functions for WTAL. Further, we propose a dedicated query mechanism to satisfy the different feature preferences between classification and localization. Thanks to these designs, our W-ART can accurately localize the diverse actions even in weakly-supervised setting. Extensive evaluation and empirical analysis show that our method outperforms the state of the arts on two challenging benchmarks, Charades and THUMOS14. Mengzhu Li, Hongjun Wu 0002, Yongcheng Liu, Hongzhe Liu 0001, Cheng Xu 0005, Xuewei Li 0006 |
ICASSP | 4 |
| 2022 | Transtl: Spatial-Temporal Localization Transformer for Multi-Label Video ClassificationabstractMulti-label video classification (MLVC) is a long-standing and challenging research problem in video signal analysis. Generally, there exist many complex action labels in real-world videos and these actions are with inherent dependencies at both spatial and temporal domains. Motivated by this observation, we propose TranSTL, a spatial-temporal localization Transformer framework for MLVC task. In addition to leverage global action label co-occurrence, we also propose a novel plug-and-play Spatial Temporal Label Dependency (STLD) layer in TranSTL. STLD not only dynamically models the label co-occurrence in a video by self-attention mechanism, but also fully captures spatial-temporal label dependencies using cross-attention strategy. As a result, our TranSTL is able to explicitly and accurately grasp the diverse action labels at both spatial and temporal domains. Extensive evaluation and empirical analysis show that TranSTL achieves superior performance over the state of the arts on two challenging benchmarks, Charades and Multi-Thumos. Hongjun Wu 0002, Mengzhu Li, Yongcheng Liu, Hongzhe Liu 0001, Cheng Xu 0005, Xuewei Li 0006 |
ICASSP | 4 |
| 2022 | Hybrid supervised instance segmentation by learning label noise suppression
Ying Fu 0001, Shaodi You, Hongzhe Liu 0001 |
Neurocomputing | 4 |
| 2022 | Multi-label video classification via coupling attentional multiple instance learning with label relation graph
Xuewei Li 0006, Hongjun Wu 0002, Mengzhu Li, Hongzhe Liu 0001 |
Pattern Recognit. Lett. | 4 |
| 2022 | SMART: Semantic-Aware Masked Attention Relational Transformer for Multi-label Image RecognitionabstractAs objects usually co-exist in an image, learning the label co-occurrence is a compelling approach to improving the performance of multi-label image recognition. However, the dependencies among the non-exist categories in an image cannot be effectively evaluated. These redundant label dependencies may bring noise and further decrease the performance of classification. Therefore, we proposed SMART, a Semantic-aware Masked Attention Relational Transformer for Multi-label Image Recognition for multi-label image recognition tasks. In addition to leveraging Transformer to model the inter-class dependencies, the proposed masked attention filters out the redundant dependencies among the non-exist categories. SMART is able to explicitly and accurately capture the label dependencies without extra word embeddings. Moreover, our method achieves new state-of-the-art results on two benchmarks for multi-label image recognition: MS-COCO 2014 and NUS-WIDE. In addition, extensive ablation studies and empirical analysis are provided to demonstrate the effectiveness of the essential components of our method under different factors. Hongjun Wu 0002, Cheng Xu 0005, Hongzhe Liu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | A Real-Time Complex Road AI Perception Based on 5G-V2X for Smart City SecurityabstractThe Internet of Vehicles and information security are key components of a smart city. Real‐time road perception is one of the most difficult tasks. Traditional detection methods require manual adjustment of parameters, which is difficult, and is susceptible to interference from object occlusion, light changes, and road wear. Designing a robust road perception algorithm is still challenging. On this basis, we combine artificial intelligence algorithms and the 5G‐V2X framework to propose a real‐time road perception method. First, an improved model based on Mask R‐CNN is implemented to improve the accuracy of detecting lane line features. Then, the linear and polynomial fitting methods of feature points in different fields of view are combined. Finally, the optimal parameter equation of the lane line can be obtained. We tested our method in complex road scenes. Experimental results show that, combined with 5G‐V2X, this method ultimately has a faster processing speed and can sense road conditions robustly under various complex actual conditions. Cheng Xu 0005, Hongjun Wu 0002, Yinong Zhang, Songyin Dai, Hongzhe Liu 0001 |
Wirel. Commun. Mob. Comput. | 5 |
| 2020 | Mutual authentication for vehicular network in complex and uncertain driving
Cheng Xu 0005, Hongzhe Liu 0001 |
Neural Comput. Appl. | 2 |
| 2020 | A novel deep network and aggregation model for saliency detection
Ye Liang 0003, Hongzhe Liu 0001 |
Vis. Comput. | 2 |
| 2018 | Is visual saliency useful for content-based image retrieval?
Yanzhang Wu, Hongzhe Liu 0001, Jiazheng Yuan, Qikun Zhang |
Multim. Tools Appl. | 2 |
| 2016 | Image representation via sub-dictionary based sparse codingabstractIn this paper, a sub-dictionary based sparse coding method is proposed for image representation. The novel sparse coding method substitutes a new regularization item for L1-norm in the sparse representation model. The proposed sparse coding method involves a series of sub-dictionaries. Each sub-dictionary contains all the training samples except for those from one particular category. For the test sample to be represented, all the sub-dictionaries should linearly represent it apart from the one that does not contain samples from that label, and this sub-dictionary is called irrelevant sub-dictionary. This new regularization item restricts the sparsity of each sub-dictionary's residual, and this restriction is helpful for classification. The experimental results demonstrate that the proposed method is superior to the previous related sparse representation based classification. Bingxin Xu, Qian Yin 0001, Ping Guo 0002, Hongzhe Liu 0001 |
IJCNN | 4 |
| 2016 | SubMIL: Discriminative subspaces for multi-instance learning
Jiazheng Yuan, Xiankai Huang, Hongzhe Liu 0001, Bing Li 0001, Weihua Xiong |
Neurocomputing | 3 |
| 2014 | Exposure fusion based on sparse representation using approximate K-SVD
Jinhua Wang 0004, Hongzhe Liu 0001 |
Neurocomputing | 2 |
| 2013 | Adaptive all-season image tag ranking by saliency-driven image pre-classification
Songhe Feng, Congyan Lang, Hongzhe Liu 0001, Xiankai Huang |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Concept vector for semantic similarity and relatedness based on WordNet structure
Hongzhe Liu 0001, Hong Bao, De Xu |
J. Syst. Softw. | 1 |
| 2005 | An Ontology-Based Semantic Integration for Digital Museums
Hong Bao, Hongzhe Liu 0001, Jiehua Yu |
WAIM | 2 |