EDBT 2026 Demo / reviewers in the wild / expert
Xuesong Gao
dblp:202/2854
· DBLP profile ↗
21ranked-venue papers
2as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Copy-move detection method based on Decoupled Edge Supervision and multi-domain cross correlation modeling
Niantai Jing, Jie Nie, Jingyu Wang 0005, Xiaodong Wang 0006, Xuesong Gao |
Multim. Tools Appl. | 6 |
| 2025 | High Efficient 3D Convolution Feature CompressionabstractIn this paper, a high efficient 3D convolution feature compression method is proposed. This method is mainly used to compress the three-dimensional convolution deep features extracted by video analysis. Gather the compressed features to the cloud server instead of all video features. This method can solve the problem that the data aggregation of video big data analysis requires a large amount of network bandwidth. Quantization is a common method for feature compression, but the existing quantization-based methods often carry out model training and quantization in stages, which makes the robustness of quantization results poor. To solve this problem, the method proposed in this paper is to apply the feature quantization operation directly to the network and train it with the analysis task, and use the parameter iterative optimization method to solve the non-differentiable problem of quantization operation. Different from the deep features extracted from images or single objects, the 3D convolution features extracted from video clips have high time-domain redundancy. In this paper, by serializing the three-dimensional convolution features, the time-domain prediction coding method is used to remove the time-domain redundancy of the three-dimensional convolution features, to improve the feature compression ratio. The experimental results show that this method can only use 1 bit to represent the elements in the three-dimensional convolution deep feature. When the analysis accuracy loss is no more than 1%, the feature compression ratio can reach 4500 times compared with the original feature data, and the data transmission can be reduced by 96%. Yangang Cai, Peiyin Xing, Xuesong Gao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Multi-Layer Transfer Learning for Cross-Domain Recommendation Based on Graph Node Representation Enhancement
Jie Nie, Niantai Jing, Jianliang Xu, Xiaodong Wang 0006, Xuesong Gao, Mingxing Jiang, Chihung Chi, Zhiqiang Wei 0002 |
IEEE Trans. Multim. | 6 |
| 2024 | One Step Learning, One Step ReviewabstractVisual fine-tuning has garnered significant attention with the rise of pre-trained vision models. The current prevailing method, full fine-tuning, suffers from the issue of knowledge forgetting as it focuses solely on fitting the downstream training set. In this paper, we propose a novel weight rollback-based fine-tuning method called OLOR (One step Learning, One step Review). OLOR combines fine-tuning with optimizers, incorporating a weight rollback term into the weight update term at each step. This ensures consistency in the weight range of upstream and downstream models, effectively mitigating knowledge forgetting and enhancing fine-tuning performance. In addition, a layer-wise penalty is presented to employ penalty decay and the diversified decay rate to adjust the weight rollback levels of layers for adapting varying downstream tasks. Through extensive experiments on various tasks such as image classification, object detection, semantic segmentation, and instance segmentation, we demonstrate the general applicability and state-of-the-art performance of our proposed OLOR. Code is available at https://github.com/rainbow-xiao/OLOR-AAAI-2024. Xiaolong Huang 0001, Qiankun Li 0004, Xueran Li, Xuesong Gao |
AAAI | 4 |
| 2024 | Generalized Sampling of Non-Local Textural Clues Multi-View Stereo Framework
Jingyuan Tang, Yangang Cai, Xuesong Gao, Songlin Sun |
ACM Multimedia | 3 |
| 2024 | Strong robust copy-move forgery detection network based on layer-by-layer decoupling refinement
Jingyu Wang 0005, Xuesong Gao, Jie Nie, Xiaodong Wang 0006, Lei Huang 0010, Weizhi Nie, Mingxing Jiang, Zhiqiang Wei 0002 |
Inf. Process. Manag. | 2 |
| 2023 | Outage Performance Analysis of STAR-RIS Assisted CR-NOMA NetworksabstractTo achieve low-cost, low energy consumption green Internet of Things (IoT) communication and meet 360oarea full-coverage, we propose a novel simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) assisted cognitive radio (CR)-non-orthogonal multiple access (NOMA) network. Specifically, the secondary transmitter serves as a relay of the primary network to forward the messages, and the secondary transmitter communicates with the user with the assistance of the STAR-RIS. To evaluate the performance of the considered network, we derive the outage probability (OP) for the users under the Nakagami-m fading channels. In addition, the asymptotic behavior at high signal-to-noise ratio (SNR) regions is analyzed. The following meaningful insights are obtained from the simulation experiments: 1) The OPs of users decrease continuously with the transmit power Ps, and increasing Psat high SNR is no longer effective for system reliability; 2) The increase in the components number of STAR-RIS has a positive impact on the reliability for the STAR-RIS assisted overlay CR-NOMA network and saturates after a certain value; 3) The scheme we considered has superior reliable performance by comparing with orthogonal multiple access. Baowang Lian, Xuesong Gao, Xingwang Li 0001, Ming Zeng 0002 |
GLOBECOM | 3 |
| 2023 | K Asynchronous Federated Learning with Cosine Similarity Based Aggregation on Non-IID Data
Yizhi Zhou, Xuesong Gao, Heng Qi |
ICA3PP (6) | 3 |
| 2023 | Population-Based Evolutionary Gaming for Unsupervised Person Re-identification
Yunpeng Zhai, Peixi Peng, Mengxi Jia, Xuesong Gao, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 6 |
| 2022 | High-Resolution Image Harmonization via Collaborative Dual TransformationsabstractGiven a composite image, image harmonization aims to adjust the foreground to make it compatible with the background. High-resolution image harmonization is in high demand, but still remains unexplored. Conventional image harmonization methods learn global RGB-to-RGB transformation which could effortlessly scale to high resolution, but ignore diverse local context. Recent deep learning methods learn the dense pixel-to-pixel transformation which could generate harmonious outputs, but are highly constrained in low resolution. In this work, we propose a high-resolution image harmonization network with Collaborative Dual Transformation (CDTNet) to combine pixel-to-pixel transformation and RGB-to-RGB transformation coherently in an end-to-end network. Our CDTNet consists of a low-resolution generator for pixel-to-pixel transformation, a color mapping module for RGB-to-RGB transformation, and a refinement module to take advantage of both. Extensive experiments on high-resolution bench-mark dataset and our created high-resolution real composite images demonstrate that our CDTNet strikes a good balance between efficiency and effectiveness. Our used datasets can be found in https://github.com/bcmi/CDTNet-High-Resolution-Image-Harmonization. Wenyan Cong, Xinhao Tao, Li Niu 0002, Jing Liang 0007, Xuesong Gao, Qihao Sun, Liqing Zhang 0001 |
CVPR | 5 |
| 2022 | Trace-Level Invisible Enhanced Network for 6D Pose EstimationabstractEstimating 6D pose of the object from a single image is es-sential for robotic manipulation. Many recent learning-based methods directly regress the pose from 2D-3D points corre-spondence. The problem is that, these methods only make use of visible information from the single-view image, resulting ambiguity for the network to solve pose from the limited cor-responding pairs. To overcome this problem, this paper intro-duces INVNet, integrating invisible information into the visi-ble 2D-3D correspondence to model geometry features of the 3D object. Instead of directly reconstruct the coordinate of in-visible points, we propose Trace-level Geometry Path, which estimates the trace-level depth of the object model for each image pixel. Specifically, our INVNet generates dense visible correspondence as well as Trace-level Geometry Path map, then learn to solve 6D pose from them. Meanwhile, each cam-era ray along with Trace-level Geometry Path is transformed to the object space by the predicted pose to compute invisi-ble correspondence loss from visible one, back to enhance its learning. Extensive experiments show that our approach out-performs state-of-the-art methods on the benchmark LM and LM-O datasets. Hanbo Sang, Zelin Ni, Huanyu He, Xuesong Gao, Qihao Sun, Sihai Zhang, Supavadee Aramvith, Weiyao Lin |
ICME | 4 |
| 2022 | Context-Dependent Text-to-SQL Generation with Intermediate Representation
Xuesong Gao, Junfeng Zhao 0005 |
PRICAI (1) | 1 |
| 2022 | Re-EnGAN: Unsupervised image-to-image translation based on reused feature encoder in CycleGANabstractAbstract As an advanced task for image synthesis without labelled data, unsupervised image‐to‐image translation refers to the overall conversion of a certain characteristic image domain X to another domain Y. The key point is to learn a mapping relationship that can be transformed between different image domains. Existing methods mainly adopt GANs to generate authentic images. While the discriminators will be abandoned after the training process is completed. In order to avoid this waste of training resources, a feature encoder reusing method is proposed, which could reduce the number of parameters and accelerate the training speed. In addition, we add the adaptive perceptual loss for the purpose of paying more attention to the quality of generated images. This loss directly uses the encoder during training to perform feature‐level constraints, and applies the L1‐norm on the intermediate feature layer of the generated samples. The experiments illustrate that our framework can generate more natural images and provide an effective solution for unsupervised translation. Lin Lv, Xuesong Gao |
IET Image Process. | 4 |
| 2022 | Robust object tracking via deformation samples generator
Xuesong Gao, Yuan Zhou 0006, Shuwei Huo, Zizi Li, Keqiu Li |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Towards 6DoF live video streaming system for immersive media
Yangang Cai, Xuesong Gao, Ronggang Wang |
Multim. Tools Appl. | 2 |
| 2022 | Enhancing adversarial transferability with partial blocks on vision transformer
Yanyang Han, Lingchen Gu, Xuesong Gao |
Neural Comput. Appl. | 6 |
| 2022 | Adaptive Progressive Continual LearningabstractContinual learning paradigm learns from a continuous stream of tasks in an incremental manner and aims to overcome the notorious issue: the catastrophic forgetting. In this work, we propose a new adaptive progressive network framework including two models for continual learning: Reinforced Continual Learning (RCL) and Bayesian Optimized Continual Learning with Attention mechanism (BOCL) to solve this fundamental issue. The core idea of this framework is to dynamically and adaptively expand the neural network structure upon the arrival of new tasks. RCL and BOCL employ reinforcement learning and Bayesian optimization to achieve it, respectively. An outstanding advantage of our proposed framework is that it will not forget the knowledge that has been learned through adaptively controlling the architecture. We propose effective ways of employing the learned knowledge in the two methods to control the size of the network. RCL employs previous knowledge directly while BOCL selectively utilizes previous knowledge (e.g., feature maps of previous tasks) via attention mechanism. The experiments on variants of MNIST, CIFAR-100 and Sequence of 5-Datasets demonstrate that our methods outperform the state-of-the-art in preventing catastrophic forgetting and fitting new tasks better under the same or less computing resource. Ju Xu, Xuesong Gao, Zhanxing Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | HCISNet: Higher-capacity invisible image steganographic networkabstractAbstract Information hiding as a crucial method for multimedia security and privacy protection, has received a substantial amount of attention. However, the most existing methods focus on the resistance of steganalysis and the robustness of information, resulting in low capacity. In this paper, to further increase the capacity of information hiding, a high‐capacity adversarial image steganography model with end‐to‐end manner, termed as HCISNet is proposed. First, an enhanced Dense Atrous Spatial Pyramid Pooling module is presented to learn more semantic information in cover image, which can adaptively embed more message in redundant regions with rich textures. Then, a modified discriminator network with lightweight residual block is employed to assist the encoder network with optimal high‐capacity solution. Meanwhile, the multiple objective function with the perceptual loss and several training tricks are developed, which can improve the visual and perceptual consistency of the original and the steganographic image, further enhancing the capacity. Finally, experiments and analysis are conducted on three public datasets, where this model has better imperceptibility, higher security and greater capacity. Under same experimental conditions, the cover image can embed 5.68 bits per pixel (BPP), far exceeding the previous highest value of 4.4 BPP in state‐of‐the‐art methods in the authors' knowledge. Xuejing Wang, Xuesong Gao |
IET Image Process. | 5 |
| 2021 | Attention adjacency matrix based graph convolutional networks for skeleton-based action recognition
Qiguang Miao, Ruyi Liu 0001, Wentian Xin, Sheng Zhong 0006, Xuesong Gao |
Neurocomputing | 7 |
| 2021 | Clone Detection Based on BPNN and Physical Layer Reputation for Industrial Wireless CPSabstractIndustrial wireless cyber-physical systems are vulnerable to malicious node attacks, for example, clone node attack. The existing clone detection schemes are either based on upper layer observations or physical layer channel state information. The schemes based on upper layer observations are vulnerable to defamation while the schemes based on channel state information perform better against defamation, but are badly affected by channel conditions. This article applies physical layer reputation and back propagation neural network to clone detection, aiming at improving the detection accuracy. The proposed scheme accumulates the physical layer reputations by channel state information and input them to the neural network. The cloud server performs attack detection by group detection first. If a certain group is classified as attacked, the corresponding edge processor will perform attack tracing to identify the specific clone nodes. During the attack tracing stage, multiple reputations of each node is adopted for a comprehensive inspection. Extensive experiments are conducted on the Universal Software Radio Peripheral platform. The numerical results show that the proposed scheme significantly improves the detection accuracy. Hong Wen 0001, Xuesong Gao, Haibo Pu, Zhibo Pang |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Simplifying Graph Attention Networks with Source-Target SeparationabstractWe present a novel Graph Neural Networks (GNN) architecture as an simplification of Graph Attentional Network (GAT) model with implicit computation of edge attention coefficients and shared sparse-dense matrix multiplication between heads. These improvements reduce training time and memory consumption while keeping the model capacity of GAT. On several established benchmarks, our model has a performance on par with state-of-the-art, yet with improved efficiency and scalability similar to simpler models including Graph Convolutional Network (GCN). Notably, we are able to apply the model to the large-scale Reddit social network dataset within a reasonable training time and memory constraint, which is previously infeasible for models with similar complexity including GAT. Hantao Guo, Rui Yan 0001, Yansong Feng 0002, Xuesong Gao, Zhanxing Zhu |
ECAI | 4 |