Rubo Zhang

dblp:65/3509 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-3211-6273ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Edge-aware Affinity Enhancement for Image Manipulation Localization
abstract
Image manipulation localization (IML) refers to the task of identifying regions in images that have been altered by specific tampering techniques, such as copy-move, splicing, or inpainting. Transformers have been applied to IML tasks due to their ability to model long-range correlations between pixels through the self-attention mechanism. However, the inherent self-attention mechanism may not accurately model or sufficiently enhance these correlations, particularly in detailed edge traces, due to its limitations. In this paper, we propose an Edge-Aware Affinity Enhancement approach for the IML task. Specifically, we introduce an Affinity Regularization Module to establish inter-patch correlations for feature regularization via random walk propagation. Based on the extracted correlation representation, we propose an Edge-Affinity Guidance strategy to further refine the correlation accuracy, particularly in ambiguous edge regions. Extensive experimental results demonstrate that our method outperforms state-of-the-art image manipulation localization techniques in terms of localization accuracy.
Tianyi Zhang 0004, Qinglong Lin, Pengming Feng, Rubo Zhang
ACM Multimedia5
2024 FPN with GMM Based Feature Enhancement Strategy for Object Detection in Remote Sensing Images
abstract
In the realm of object detection, the age-old challenge of accommodating large variations in target scales, particularly in the intricate domain of remote sensing imagery, has long perplexed computer vision aficionados. Feature Pyramid Network (FPN) family, a widely-used stalwart, strives to tame this scale variation challenge by harmoniously fusing features across different levels. However, this typical feature fusion strategy often leads us astray. Noise introduction and feature smoothing problems due to different semantic information from high/low resolution feature maps, which results in semantic misalignment and inconspicuous gradient discrepancy between targets and background. This, in turn, leads to the difficulty in locating and distinguishing target from complex background in remote sensing images. In this paper, a GMM Feature Enhancement Module (GFEM) is proposed to address the problem by generating and enhancing feature of target with Gaussian Mixture Model (GMM), hence avoiding the gradient smoothing problem. Moreover, we introduce a generic feature fusion network named GFEM-FPN, elevating our approach to the next level. GFEM-FPN extracts multi-scale target enhancement features to enhance the ability of discriminating targets and background. The proposed methods are evaluated on NWPU VHR-10 and DIOR-R datasets, and the outperformance in results verify the effectiveness of the proposed method.
Hongning Liu, Pengming Feng, Mingjie Xie, Dongli Xu, Jian Guan 0001, Guangjun He, Rubo Zhang
ICASSP7
2024 Weakly supervised temporal action localization: a survey
Ronglu Li, Tianyi Zhang 0004, Rubo Zhang
Multim. Tools Appl.3
2024 Mutual learning generative adversarial network
Lin Mao, Rubo Zhang
Multim. Tools Appl.4
2024 Integration of Global and Local Knowledge for Foreground Enhancing in Weakly Supervised Temporal Action Localization
abstract
Weakly Supervised Temporal Action Localization (WTAL) aims to identify the temporal duration of actions and classify the action categories with only video-level labels in the training stage. Motivated by the intuition that the attention maps generated from various views will assist in enhancing the foreground action temporal segments, in this paper we propose a WTAL pipeline based on a novel attention mechanism that effectively integrates global and local knowledge. Our attention mechanism is mainly composed of a global attention branch and a local attention branch. Specifically, the global attention branch is built on the inter-segment similarity to sparsely mine out the correlation knowledge within the entire video, while the local attention branch is built on the convolutional structure to densely aggregate the information within the fixed local respective field. Experiments on THUMOS14 and ActivityNet v1.3 datasets demonstrate the effectiveness of our proposed WTAL pipeline compared to state-of-the-art methods.
Tianyi Zhang 0004, Ronglu Li, Pengming Feng, Rubo Zhang
IEEE Trans. Multim.4
2023 A temporal and channel-combined attention block for action segmentation
Lin Mao, Rubo Zhang
Appl. Intell.4
2023 ChaInNet: Deep Chain Instance Segmentation Network for Panoptic Segmentation
Lin Mao, Fengzhi Ren, Rubo Zhang
Neural Process. Lett.4
2021 Convolutional Feature Frequency Adaptive Fusion Object Detection Network
Lin Mao, Xuemeng Li, Rubo Zhang
Neural Process. Lett.4
2018 Inverse Reinforcement Learning via Neural Network in Driver Behavior Modeling
abstract
Inverse Reinforcement Learning(IRL) is formulated within the framework of Markov decision process(MDP) where we are not explicitly given a reward function, but where instead we can observe an expert demonstrating the task that we want to learn to perform. Then the expert as trying to maximize a reward function that is expressible as a linear combination of known features specifying the reward function. However, in autonomous driving tasks, due to the difference of scene factor, such as obstacle and weather, the state spaces are frequently large and demonstrations can hardly visit all the states. it's hard to get an optimal policy with RL method to express driver behavior model based on the reward which recovered with IRL method in this large-scale state space. In this paper, we focus on driving behavior modeling with IRL method which introduces the convolutional neural network to extract the associated state feature automatically, and express the policy by neural network to generalize the expert's behaviors. Experimental results compared with the traditional end-to-end method on simulated vehicle show that the accuracy of decision-making greatly improved in the train curve, and in the new curve scene with a large number of unvisited state, this method shows a perfect generalization efficiency.
QiJie Zou, Rubo Zhang
Intelligent Vehicles Symposium3
2017 Swarm-based intelligent optimization approach for layout problem
Fengqiang Zhao, Guangqiang Li, Rubo Zhang, Jialu Du, Chen Guo 0001, Yiran Zhou, Zhihan Lyu
Multim. Tools Appl.3
2008 Response threshold model of aggregation in a swarm: A theoretical and simulative comparison
abstract
Swarm Intelligence(SI) which is inspired by social animals has been paid more and more attention. It always appeals to the collective behaviors observed in social animals. Aiming at the feature and factors in self-organization of SI system, the aggregation behavior is studied. Firstly the response threshold model of the system is built according to the rules in aggregation. Then the stability of the steady-state solutions of the model is analyzed and the bifurcation of the steady-state solution is obtained. Finally, the effects of the parameter are analyzed based on the theory model. And the Monte Carlo simulations which give certain differences against theory results are also analyzed. All of the theoretical and simulative results show that the aggregation behavior is impacted by the relationship between the swarm size and the response threshold and sensitivity significantly. It is also proved that complex behavior emerges from local interaction of individuals. The work of this paper gives the mechanism in the emergent complex pattern of self-organized aggregation and the factors which affect the system evolution.
Bailong Liu, Rubo Zhang, Changting Shi
IEEE Congress on Evolutionary Computation2
2008 A New Decision Rule for Statistical Word Sense Disambiguation
Dongmei Fan, Zhimao Lu, Rubo Zhang
ICIC (1)3
2008 A Vicarious Words Method for Word Sense Discrimination
Zhimao Lu, Dongmei Fan, Rubo Zhang
ICIC (1)3
2005 A Novel Anomaly Detection Using Small Training Sets
Qingbo Yin, Liran Shen, Rubo Zhang, Xueyao Li
IDEAL3
2004 Speech Hiding Based on Auditory Wavelet
Liran Shen, Xueyao Li, Rubo Zhang
ICCSA (4)4
2003 A new approach for structural credit assignment in distributed reinforcement learning systems
abstract
Most existing algorithm for structural credit assignment are developed for competitive reinforcement learning systems. In competitive reinforcement learning system, agents are activated one by one, so there is only one active agent at a time and structural credit assignment could be implemented by some temporal credit assignment algorithms. In collaborated reinforcement learning systems, agents are activated simultaneously, so how to transform the global reinforcement signal fed back from the environment to a reinforcement vector is a crucial difficulty that could not be slide over. In this article, the first really feasible and efficient structural credit assignment difficulty in collaborated reinforcement learning systems is primarily solved. The experiments show that the algorithm converges very rapidly and the assignment result is quite satisfying.
Guochang Gu, Rubo Zhang
ICRA3