Hong Zhang 0018

dblp:24/6914-18 · DBLP profile ↗
← Back
51ranked-venue papers
21as first author
31since 2021 · last 2026
0000-0002-1282-3755ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 12 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Towards Explainable Video Camouflaged Object Detection: SAM2 with Eventstream-Inspired Data
abstract
Video Camouflaged Object Detection (VCOD) poses significant challenges due to the subtle appearance of camouflaged objects, especially under dynamic motion and occlusion. Existing methods predominantly rely on optical flow or black-box features for motion modeling, which often entail substantial computational costs and suffer from limited interpretability. Inspired by the human strategy of identifying abnormal movements between frames and the principle of event camera image formation, we propose an eventstream-inspired dual-branch framework for VCOD. Specifically, we design an eventstream-like data extraction module to capture pixel-level motion variations, effectively distinguishing object motion from background dynamics. This event-based representation is integrated into SAM2 through a dual-branch memory-augmented framework, consisting of Time Bridge Attention and Visual Bridge Attention, enabling joint modeling of motion and appearance cues. In addition, we introduce a Prompt Embedding Generator to eliminate the need for human-provided interactive prompts, facilitating fully automatic VCOD. Extensive experiments on MoCA-Mask and CAD2016 demonstrate that our approach significantly outperforms state-of-the-art methods, achieving both superior segmentation accuracy and interpretable motion modeling. To the best of our knowledge, this is the first work to incorporate eventstream-inspired representations into the VCOD task.
Hong Zhang 0018, Yixuan Lyu, Jianbo Song, Ding Yuan 0001, Yifan Yang 0003
AAAI1
2026 MCTrack: Multi-cue spatio-temporal object tracking
Jianbo Song, Hong Zhang 0018, Yachun Feng, Yifan Yang 0003
Expert Syst. Appl.2
2026 MGFNet: Meta Global Filter Network for multi-size image feature extraction
Hong Zhang 0018, Jiaxu Wan, Jianbo Song, Ding Yuan 0001, Yifan Yang 0003
Int. J. Comput. Vis.1
2026 SDP-GS: Sparse-view Gaussian splatting via segmentation-aware depth priors
Qi Zhao 0037, Yangyan Deng, Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001
Neurocomputing4
2026 STCTracker: Enhancing sequential temporal consistency in referring multi-object tracking
Hong Zhang 0018, Jiabi Zhao, Ding Yuan 0001, Yifan Yang 0003
Image Vis. Comput.1
2026 LKTrack: a novel tracking framework with large kernel network
Hong Zhang 0018, Huakao Lin, Ding Yuan 0001, Jianbo Song, Yifan Yang 0003
Multim. Syst.1
2026 An end-to-end shadow removal framework with an intuitive interaction scheme
Ding Yuan 0001, Yuqian Meng, Yachun Feng, Hong Zhang 0018, Yifan Yang 0003
Pattern Recognit.5
2026 E2IGB: Enhanced effective-information-guided class-balanced loss for long-tailed object recognition
Hong Zhang 0018, Zhigang Li 0005, Yangyan Deng, Yachun Feng, Ding Yuan 0001, Yifan Yang 0003
Pattern Recognit.1
2026 TransSTC: transformer tracker meets efficient spatial-temporal cues
Hong Zhang 0018, Wanli Xing 0004, Yifan Yang 0003, Ding Yuan 0001
Pattern Recognit.1
2026 CCMamba: A Multi-Head Criss-Cross Mamba for Pavement Crack Segmentation
abstract
Automatic and accurate detection of pavement cracks on highways and main roads is very important for intelligent pavement maintenance and traffic safety. However, due to the irregular shapes and uncertain characteristics of cracks, current methods face serious challenges on crack segmentation task. To solve the above issues, we propose a Multi-Head Criss-Cross Mamba (CCMamba) to efficiently enhance crack segmentation performance. Based on the sufficient details retained in low-level features and semantic information with strong discriminative ability in high-level features, CCMamba utilizes high-level features to extract crack information from low-level features. First, because of the feature representation capacity of latent state in State Space Model (SSM), we introduce a Multi-Head Latent State Module (Multi-Head LSM) to criss-cross study high-level local features and generate multiple Dynamic Convolution kernels. Second, these Dynamic Convolution kernels are applied to low-level features in a convolutional way, filtering crack information from massive background interferences. Third, the horizontal and vertical features output from Dynamic Convolution layers are fused with head attention mechanism, producing crack sensitive features. Finally, we use the obtained multi-scale features to predict segmentation masks. Comprehensive experiments on five public datasets, Crack500, GAPs384, CFD, CrackVision12K and CPRID, are conducted and our CCMamba achieves state-of-the-art (SOTA) performances compared to current crack segmentation methods. Meanwhile, ablation studies, visualization analysis and real scene testing also validate the effectiveness of CCMamba. Codes of this paper are public available athttps://github.com/cv516Buaa/BinghaoLiu/tree/main/CCMamba
Binghao Liu, Qi Zhao 0037, Hongbo Xie, Hong Zhang 0018
IEEE Trans. Intell. Transp. Syst.5
2025 SP2T: Sparse Proxy Attention for Dual-Stream Point Transformer
Jiaxu Wan, Hong Zhang 0018, Ziqi He, Yangyan Deng, Qishu Wang, Ding Yuan 0001, Yifan Yang 0003
ICCV2
2025 EMA-GS: Improving sparse point cloud rendering with EMA gradient and anchor upsampling
Ding Yuan 0001, Sizhe Zhang, Hong Zhang 0018, Yangyan Deng, Yifan Yang 0003
Image Vis. Comput.3
2025 AwareTrack: Object awareness for visual tracking via templates interaction
Hong Zhang 0018, Jianbo Song, Yifan Yang 0003, Huimin Ma 0001
Image Vis. Comput.1
2025 CODdiff: Prior leading diffusion model for Camouflage Object Detection
Hong Zhang 0018, Yixuan Lyu, Xuliang Li 0005, Yawei Li 0003, Ding Yuan 0001, Yifan Yang 0003
Knowl. Based Syst.1
2025 Online Self-Training Driven Attention-Guided Self-Mimicking Network for Semantic Segmentation
abstract
In the realm of semantic segmentation tasks, knowledge distillation (KD) has emerged as a prominent strategy, leveraging the transfer of mature knowledge from large teacher networks to enhance the performance of smaller student networks. However, existing methods often rely heavily on high-quality yet cumbersome teacher networks, leading to a complex training process. To address this challenge, we introduce a novel approach termed self-training driven attention-guided self-mimicking online ensemble network. Our proposed method begins by employing intermediate channel-joint attention maps to guide image augmentation. Both the original and augmented images are then input into the networks. Leveraging intermediate feature maps and predictive predictions generated from the two images, we employ KD to uncover invariant features. To further harness representation potential through learning from credible predictions, we introduce a self-training mechanism. This mechanism utilizes an exponential moving average (EMA)-teacher network constructed using the exponential moving average technique to generate feature maps and predicted posterior probabilities. The knowledge of the EMA-teacher is subsequently transferred to the student network through distillation. Extensive experiments and visualization analyses conducted on multiple benchmark datasets, including Cityscapes, Pascal VOC, CamVid, and ADE20k, validate the effectiveness of self-training driven attention-guided self-mimicking network (ST-ASMNet). The interpretability of our method is further validated through visualization and analysis. Our code will be publicly available.
Shuchang Lyu, Qi Zhao 0037, Hong Zhang 0018, Chenguang Yang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Language-guided Visual Tracking: Comprehensive and Effective Multimodal Information Fusion
abstract
Current vision-language trackers often struggle to fuse multimodal information comprehensively and effectively, leading to suboptimal performance in multimodal tasks. This study introduces LGTrack, a novel language-guided visual tracking framework designed to achieve a more comprehensive and efficient fusion of vision and language information. In the encoding stage, an Enhanced Multimodal Interaction Module is proposed to achieve full multimodal fusion, and it is used to construct Early Language Multilevel-guided Multimodal Encoding, which leverages deep semantic information for early and multilevel guidance of vision encoding. In the decoding stage, a multimodal decoding based on Joint Query is proposed, utilizing global features from both vision and language modalities, guiding the efficient operation of the decoding layers. These innovations achieve a more comprehensive fusion of multimodal information. Additionally, a contrastive learning strategy is introduced to align vision-language features in the semantic space, further enhancing the fusion effectiveness. Extensive experiments on multiple benchmarks such as LaSOT, \(\rm{LaSOT_{ext}}\) , TNL2K, and OTB99-Lang demonstrate that our approach outperforms existing state-of-the-art trackers.
Jianbo Song, Hong Zhang 0018, Yachun Feng, Yifan Yang 0003
ACM Trans. Multim. Comput. Commun. Appl.2
2025 P2FTrack: Multi-Object Tracking with Motion Prior and Feature Posterior
abstract
Multiple object tracking (MOT) has emerged as a crucial component of the rapidly developing computer vision. However, existing multi-object tracking methods often overlook the relationship between features and motion, hindering the ability to strike a performance balance between coupled motion and complex scenes. In this work, we propose a novel end-to-end multi-object tracking method that integrates motion and feature information. To achieve this, we introduce a motion prior generator that transforms motion information into attention masks. Additionally, we leverage prior-posterior fusion multi-head attention to combine the motion-derived priors and attention-based posteriors. Our proposed method is extensively evaluated on MOT17 and DanceTrack datasets through comprehensive experiments and ablation studies, demonstrating state-of-the-art performance in the feature-based method with reasonable speed.
Hong Zhang 0018, Jiaxu Wan, Jing Zhang 0017, Ding Yuan 0001, Xuliang Li 0005, Yifan Yang 0003
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and Visual Analysis Strategy
Hong Zhang 0018, Yixuan Lyu, Qian Yu 0008, Huimin Ma 0001, Ding Yuan 0001, Yifan Yang 0003
ECCV (55)1
2024 Feature Block-Aware Correlation Filters for Real-Time UAV Tracking
abstract
Recently, by virtue of the high computational efficiency and accuracy, discriminative correlation filter (DCF)- based tracking methods have gained attraction in the field of unmanned aerial vehicle (UAV). However, conventional DCF-based methods merely rely on cyclic shift to produce training samples. As a result, the filter trained by these samples owns limited discriminative ability, ineffectively addressing various challenges in the tracking stage. Here, to promote the filter's discriminative ability, we develop a feature block-aware correlation filter (CF) method. Specifically, the extracted feature is divided into two blocks, i.e., target and background feature blocks. These blocks only contain target and background features, respectively, by using different mask matrixes. Then, two regularization terms are proposed to combine both feature blocks into the DCF framework. In addition, we employ effective channel reliability weights to generate target response for precise positioning. Furthermore, substantial experiments have been accomplished on multiple public UAV benchmarks, proving that our tracker possesses superior tracking capabilities and operates at ∼40 frames per second (FPS) on the CPU platform.
Hong Zhang 0018, Yan Li 0094, Ding Yuan 0001, Yifan Yang 0003
IEEE Signal Process. Lett.1
2024 Efficient Template Distinction Modeling Tracker With Temporal Contexts for Aerial Tracking
abstract
In recent years, trackers based on neural networks have demonstrated excellent tracking performance. Compared to general tracking tasks, aerial tracking tasks require more stringent operational efficiency of the tracker. Furthermore, since aerial devices such as unmanned aerial vehicles usually shoot from a high altitude, the tracked target occupies fewer pixels, resulting in scarce target discrimination information and more susceptibility to interference from cluttered backgrounds. Unfortunately, existing trackers usually model the entire template region nondifferently, which tends to confuse target and background information in the template and reduce the robustness of the tracker in complex scenes. Additionally, most tracking networks refuse or only refer to a single feature layer to learn the risky temporal context information, which makes it difficult to accurately supplement the scarce target discrimination information. To this end, we propose an efficient template distinction modeling tracker with temporal contexts, called ETDMT, designed to improve the complex scene robustness of aerial tracking through a combination of template distinction modeling and temporal context analysis. The tracker employs a template distinction modeling transformer network, which distinguishes between target and background elements and adopts different modeling approaches for different elements to alleviate the background interference problem prevalent in aerial tracking. Then, the temporal contexts of the tracker are complemented by a global-local spatial awareness update module, which enhances the tracker’s understanding of the latest target state by comprehensively evaluating and adjusting dynamic templates. Extensive experiments demonstrate that the proposed ETDMT achieves advanced aerial tracking performance and efficient running speed.
Hong Zhang 0018, Wanli Xing 0004, Huakao Lin, Yifan Yang 0003
IEEE Trans. Geosci. Remote. Sens.1
2024 UEDG:Uncertainty-Edge Dual Guided Camouflage Object Detection
abstract
According to Darwinian evolutionary theory, numerous species in the wild have developed remarkable adaptive mechanisms, involving pattern rearrangement and environmental assimilation, to evade predators. These obfuscation strategies pose significant challenges for both individuals and algorithms when performing the Camouflage Object Detection (COD) task in complex and intricate scenarios. Inspired by human strategies in the COD task, which involve assigning uncertainties to the entire input and then focusing on highly uncertain areas with the aid of prior knowledge such as boundary information, we propose the Uncertainty-Edge Dual Guide (UEDG) architecture. UEDG effectively combines probabilistic-derived uncertainty and deterministic-derived edge information to accurately detect concealed objects. The architecture consists of two independent branches dedicated to uncertainty reasoning and edge inference, which are subsequently integrated into a feature fusion module utilizing recursion feedback and feature-reuse techniques. This novel COD framework leverages the benefits of Bayesian learning and convolution-based learning, resulting in a powerful multi-task guided approach. Extensive experiments conducted on four widely employed datasets demonstrate the superior performance of UEDG compared to 12 state-of-the-art approaches, while maintaining an acceptable level of computational complexity. Overall, UEDG presents a promising solution for addressing the challenges of COD in complex environments by combining evolutionary-inspired strategies with advanced computer vision techniques.
Yixuan Lyu, Hong Zhang 0018, Yan Li 0094, Yifan Yang 0003, Ding Yuan 0001
IEEE Trans. Multim.2
2023 Redefined target sample-based background-aware correlation filters for object tracking
Wanli Xing 0004, Hong Zhang 0018, Yujie Wu 0001, Yawei Li 0003, Ding Yuan 0001
Appl. Intell.2
2023 SiamST: Siamese network with spatio-temporal awareness for object tracking
Hong Zhang 0018, Wanli Xing 0004, Yifan Yang 0003, Yan Li 0094, Ding Yuan 0001
Inf. Sci.1
2023 GradQuant: Low-Loss Quantization for Remote-Sensing Object Detection
abstract
Convolutional neural network based methods have shown remarkable performance in remote sensing object detection. However, their deployment on resource-limited embedded devices is hindered by their high computational complexity. Neural network quantization methods have been proven effective in compressing and accelerating CNN models by clipping outlier activations and utilizing low-precision values to represent weights and clipped activations. Nonetheless, the clipping of outlier activations leads to distortion of object local features. Furthermore, the lack of enhanced overall feature mining exacerbates the degradation of detection accuracy. To address the limitations above, we propose an innovative clipping-free quantization method called GradQuant, which mitigates model’s quantization accuracy loss caused by clipping outlier activations and the lack of overall feature mining. Specifically, a bounded activation function (sigmoid-weighted tanh, SiTanh) is carefully designed to ensure that object features are represented within a limited range without clipping. On the basis of this, an activation substitute training (AST) method is co-designed to prompt models to focus more on non-outlier object features instead of outlier-like local ones. Extensive experiments on public remote-sensing datasets demonstrate the effectiveness of GradQuant method compared with other state-of-the-art quantization methods.
Chenwei Deng, Yuqi Han, Donglin Jing, Hong Zhang 0018
IEEE Geosci. Remote. Sens. Lett.5
2023 RISTrack: Learning Response Interference Suppression Correlation Filters for UAV Tracking
abstract
With the high computation efficiency and tracking accuracy, discriminative correlation filters (DCF) have been applied to UAV tracking. However, in the scenarios (i.e., complex background and temporary occlusion), DCF-based trackers usually generate low credibility response under the influence of background distractors, which contains multiple side peaks and declines the tracking performance. Motivated by the response consistency in adjacent frames and background information penalization, we propose learning a response interference suppression (RIS) correlation filter to tackle this problem. Specifically, we introduce a RIS regularization into the DCF-based framework, which aims to keep the target area response consistent in adjacent frames and repress distractors’ response in the background. Besides, we adopt a response auxiliary strategy (RAS) to smooth the target response, which intends to obtain the precise location and avoid target drift. Furthermore, extensive experiments on three UAV benchmarks demonstrate the excellent performance of the proposed method against other 19 state-of-the-art trackers. Moreover, the tracking speed of the proposed method can reach 42 FPS on a single CPU.
Yan Li 0094, Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001
IEEE Geosci. Remote. Sens. Lett.2
2023 Cross-modality complementary information fusion for multispectral pedestrian detection
Chaoqi Yan, Hong Zhang 0018, Xuliang Li 0005, Yifan Yang 0003, Ding Yuan 0001
Neural Comput. Appl.2
2022 R-SSD: refined single shot multibox detector for pedestrian detection
Chaoqi Yan, Hong Zhang 0018, Xuliang Li 0005, Ding Yuan 0001
Appl. Intell.2
2022 Feature adaptation-based multipeak-redetection spatial-aware correlation filter for object tracking
Wanli Xing 0004, Hong Zhang 0018, Hao Chen 0052, Yifan Yang 0003, Ding Yuan 0001
Neurocomputing2
2022 MSAGNet: Multi-Stream Attribute-Guided Network for Occluded Pedestrian Detection
abstract
Pedestrian detection plays an indispensable role in human-centric applications. Although having enjoyed the merits of generic object detectors based on deep learning frameworks, pedestrian detection is still a persistent crucial task since the pedestrians often gather together and occlude each other. In this study, we propose a simple yet effective Multi-Stream Attribute-Guided Network (MSAGNet) to regard occluded pedestrian detection as a standard central point and height estimation problem. Specifically, we focus on searching for the central points of the pedestrians and predicting the scales and offsets of the corresponding pedestrians. Meanwhile, an adaptive weighting parameter, i.e., Intersection over the Visible part region of ground truth (IoV), is utilized to conduct accurate bounding box regression. Furthermore, a novel nonlinear Non-Maximum Suppression (NMS) is proposed to flexibly prune false positives and decrease the miss rate of adjacent overlapping pedestrians. Experimental results on Caltech-USA, CityPersons, CrowdHuman and WiderPerson pedestrian datasets show that the proposed MSAGNet can obtain significant performance boosts, while maintaining a reasonable run-time speed.
Hong Zhang 0018, Chaoqi Yan, Xuliang Li 0005, Yifan Yang 0003, Ding Yuan 0001
IEEE Signal Process. Lett.1
2022 Semantic Segmentation With Attention Mechanism for Remote Sensing Images
abstract
Semantic segmentation for high-resolution remote sensing images is one of the most significant tasks in the field of remote sensing applications. Remote sensing images contain substantial detailed information of ground objects, such as shape, location, and texture. Therefore, these objects make the images exhibit large intraclass variance and small interclass variance, which makes it very difficult to be recognized. In this study, an end-to-end attention-based semantic segmentation network (SSAtNet) is proposed. A pyramid attention pooling module is proposed to introduce the attention mechanism into the multiscale module for adaptive features refinement. To correct the detailed information, the pooling index correction module integrates pooling index maps from the encoder with high-level feature maps, which can help recover the fine-grained features. In the encoder phase, a more effective ResNet-101 backbone is designed to capture detailed features. What is more, a series of data augmentation methods are proposed to enhance the model’s robustness. The proposed model is compared with several previous advanced networks and achieves the state of the art on the ISPRS Vaihingen dataset. The experiment results prove the effectiveness of the SSAtNet.
Qi Zhao 0037, Hong Zhang 0018
IEEE Trans. Geosci. Remote. Sens.4
2021 Dynamic Scene Video Deblurring Using Robust Incremental Weighted Fourier Aggregation
abstract
Motion blur is an inevitable problem when shooting a video in motion, and it will cause serious degradation of images. In this letter, we propose a new robust incremental weighted Fourier aggregation algorithm for dynamic scene video deblurring. This is motivated by the fact that existing multiframe aggregation methods usually require a tedious iterative process to obtain sharp information in distant frames. We propose to perform alignment and aggregation once for each frame using the current frame and the previous frame result, which includes the sharp information of all the previous frames. Moreover, a criterion for evaluating the alignment of blurred images is proposed. Based on this, a robust aggregation approach is also proposed to help reduce artifacts caused by image patch alignment failure. Specifically, we decompose the current frame into a series of image patches, find consistent image patches in the deblurring results of the previous frame by a coarse-to-fine alignment algorithm, and aggregate them in the Fourier domain. Experiments show that our method can obtain improved deblurring results and reduce the computational complexity compared to the traditional aggregation method.
Yawei Li 0003, Hong Zhang 0018, Yujie Wu 0001, Ding Yuan 0001
IEEE Signal Process. Lett.2
2020 Image deblurring using tri-segment intensity prior
Hong Zhang 0018, Yujie Wu 0001, Lei Zhang 0098, Yawei Li 0003
Neurocomputing1
2020 Effective Data-Driven Technology for Efficient Vision-Based Outdoor Industrial Systems
abstract
Vision systems are the core information collection module in outdoor industrial systems such as factory inspection robots. However, haze greatly reduces working efficiency. Existing dehazing methods have two problems-first, they are not specifically designed for the industrial systems; second, these methods include several assumptions in their design processes and imaging models, leading to unsatisfactory results. In this article, an approach for single image dehazing is proposed to improve the efficiency of outdoor vision-based systems. First, a novel haze imaging model is proposed based on the dichromatic atmospheric scattering model. It considers the effects of multiple scattering and involves fewer assumptions. Then a data-driven technique called sparse representation is used to solve this model. Considering a haze image, a distorted and blurred version of a fine image, every patch is presented using dedicatedly prepared over-complete dictionaries and is traced back to a haze-free image. Quantitative and qualitative comparisons on a number of real-world haze images demonstrate that the proposed approach not only is more stable but also leads to better dehazing results.
Jiafeng Li 0001, Li Zhuo 0001, Hong Zhang 0018, Guoqiang Li 0001, Naixue Xiong
IEEE Trans. Ind. Informatics3
2019 Video denoising for security and privacy in fog computing
abstract
Summary To reduce heavy noise from degraded video in low or predictable latency and preserve privacy, a powerful and efficient video denoising algorithm is proposed based on fog computing for Visual Internet of Things. The conventional method is to remove noise in the cloud; however, this may overload computation and communication and raise security and privacy issues. The proposed denoising algorithm is distributed to heterogeneous devices at network edges to preserve privacy and avoid security risks as noise can be reduced in the fog rather than the cloud. To address the problems of latency, communication rate, and extremely heavy noise, structure registration, inter‐frame and inner‐frame filters, and distribution compensation are applied in the proposed algorithm. A scheme for encrypting the denoised data at network edges is provided so that security and privacy issues may be avoided during transmission and storage. Compared with other denoising approaches under extremely heavy noise conditions, the experimental results demonstrate that the proposed approach achieves superior denoising performance in terms of peak signal‐noise ratio and visual quality at low computational cost, high bandwidth efficiency, and low‐latency response in a fog computing manner.
Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001, Daniel Sun 0004, Jun Zhang 0010, Guoqiang Li 0001, Mingui Sun
Concurr. Comput. Pract. Exp.1
2019 Leveraging semantic segmentation with learning-based confidence measure
Feiyang Cheng, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun
Neurocomputing2
2018 End-to-end temporal attention extraction and human action recognition
Hong Zhang 0018, Miao Xin, Shuhang Wang, Yifan Yang 0003, Lei Zhang 0098, Helong Wang
Mach. Vis. Appl.1
2018 Learning to refine depth for robust stereo estimation
Feiyang Cheng, Xuming He 0001, Hong Zhang 0018
Pattern Recognit.3
2017 Learning discriminative action and context representations for action recognition in still images
abstract
Action recognition in still images is a challenging task in computer vision. Recent successes in deep feature-learning advance this research, employing robust and rich-semantic feature representation. However, the issue that recognition fails when two action images share similar contexts is long-standing. In this paper, we employ metric learning method to address within-class and between-class confusions in action recognition. We propose a novel loss function, named composite-triplet loss. Supervised by this loss function, our method directly learns a similarity function from data. Employing a customized human action and context detection network, we obtain highly discriminative action image embeddings, which can be used in action image recognition and other tasks. Our approach is evaluated on the still-image action recognition task and the image caption generation task. On the PASCAL VOC dataset, our approach outperforms state-of-the-art methods, achieving 90.6% mean AP.
Miao Xin, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun
ICME2
2017 Stacked Learning to Search for Scene Labeling
abstract
Search-based structured prediction methods have shown promising successes in both computer vision and natural language processing recently. However, most existing search-based approaches lead to a complex multi-stage learning process, which is ill-suited for scene labeling problems with a high-dimensional output space. In this paper, a stacked learning to search method is proposed to address scene labeling tasks. We design a simplified search process consisting of a sequence of ranking functions, which are learned based on a stacked learning strategy to prevent over-fitting. Our method is able to encode rich prior knowledge by incorporating a variety of local and global scene features. In addition, we estimate a labeling confidence map to further improve the search efficiency from two aspects: first, it constrains the search space more effectively by pruning out low-quality solutions based on confidence scores; second, we employ the confidence map as an additional ranking feature to improve its prediction performance and thus reduce the search steps. Our approach is evaluated on both semantic segmentation and geometric labeling tasks, including the Stanford Background, Sift Flow, Geometric Context and NYUv2 RGB-D dataset. The competitive results demonstrate that our stacked learning to search method provides an effective alternative paradigm for scene labeling.
Feiyang Cheng, Xuming He 0001, Hong Zhang 0018
IEEE Trans. Image Process.3
2016 Recurrent Temporal Sparse Autoencoder for attention-based action recognition
abstract
Visual context is fundamental to understand human actions in videos. However, to efficiently employ temporal context information presents an enormous challenge to this area. Two main problems are long-standing: (1) video frames are redundant while discriminative information is sparse; (2) large amount of interference information is mixed in frame sequences. These factors results in redundant computation and recognition failures. In this paper, we propose a learnable temporal attention mechanism to automatically select important time points from action sequences. We design an unsupervised Recurrent Temporal Sparse Autoencoder (RTSAE) network, which learns to extract sparse key-frames to sharpen discriminative yet to retain descriptive capability, as well to shield interfere information. By applying this technique to a recent proposed action recognition model Adaptive Recurrent-convolutional Hybrid network (ARCH), we significantly improve its performance in both speed and accuracy. Experiments demonstrate that, with the help of the RTSAE, ARCH outperforms most state-of-the-art methods on UCF101 and HMDB51 datasets.
Miao Xin, Hong Zhang 0018, Mingui Sun, Ding Yuan 0001
IJCNN2
2016 ARCH: Adaptive recurrent-convolutional hybrid networks for long-term action recognition
Miao Xin, Hong Zhang 0018, Helong Wang, Mingui Sun, Ding Yuan 0001
Neurocomputing2
2015 Single image dehazing using the change of detail prior
Jiafeng Li 0001, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun
Neurocomputing2
2015 Cross-trees, edge and superpixel priors-based cost aggregation for stereo matching
Feiyang Cheng, Hong Zhang 0018, Mingui Sun, Ding Yuan 0001
Pattern Recognit.2
2014 Cross-Trees for Stereo Matching with Priors
abstract
We propose a cross-trees structure to perform the non-local cost aggregation for dense stereo matching. The cross-trees structure consists of a horizontal-tree and a vertical-tree. Compared to other spanning trees, the significant superiority of the cross-trees is that the trees' constructions are efficient and independent on any local or global property. Moreover, the trees are exactly unique. By traversing the two crossed trees successively, a fast non-local cost aggregation algorithm is performed to filter the matching cost volume and then the disparity maps are established with the Winner-Take-All (WTA) strategy. Additionally, two different priors: edge prior and super pixel prior, are proposed to tackle the false smoothing at the depth boundaries. Hence, our method contains two different algorithms in terms of the cross-trees prior in this paper. Performance evaluation on the 27 Middlebury data sets shows that both our algorithms outperform the other two tree-based methods, namely minimum spanning tree (MST) and segment-tree (ST). By performing the non-local cost aggregation on different trees, MST, ST and our method all have competitive rankings on the Middlebury website compared to the local cost aggregation methods.
Feiyang Cheng, Hong Zhang 0018, Mingui Sun, Helong Wang, Ding Yuan 0001
ICPR2
2014 Stereo matching by using the global edge constraint
Feiyang Cheng, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun
Neurocomputing2
2012 Stereo matching with Global Edge Constraint and Graph Cuts
Hong Zhang 0018, Feiyang Cheng, Ding Yuan 0001, Yuecheng Li, Mingui Sun
ICPR1
2012 Modified Clipped Histogram Equalization for Contrast Enhancement
abstract
Histogram equalization (HE) based methodologies are popular and effective ways to improve image contrast and visual quality, but standard HE is not directly applied on consumer electronics. This paper proposes a modified Clipped Histogram Equalization (CHE) for contrast enhancement based on the fundamental of histogram modification. It improves the visual quality by enhancing details both in dark and bright regions through automatically detecting the dynamic range of the image. Experimental results show that the proposed algorithm outperforms state-of-the-art HE based algorithms in improving contrast and giving a visually-pleasing result while avoiding over-enhancement and noise amplification. Its high efficiency makes the proposed method practical for applications on consumer electronic.
Yuecheng Li, Hong Zhang 0018
PDCAT2
2011 Physical activity recognition based on motion in images acquired by a wearable camera
Hong Zhang 0018, Wenyan Jia, John D. Fernstrom, Robert J. Sclabassi, Zhi-Hong Mao, Mingui Sun
Neurocomputing1
2010 Multi-scale sparse feature point correspondence by graph cuts
Hong Zhang 0018, Yuhu You
Sci. China Inf. Sci.1
2007 Automatic Video Object Segmentation using Graph Cut
abstract
This paper presents an algorithm for automatic video object planes extraction from coarse to fine. For the case of single moving object in a scene, block-based segmentation defines regions of foreground, background and boundary blocks. Then, the segmentation problem is formulated as an energy minimization problem which is settled by using graph cut algorithm. Automatic segmentation can be realized by obtaining prior knowledge from foreground and background blocks and computation complex is reduced by restricting the refined segmentation region to boundary blocks. Experimental results show the effectiveness of proposed algorithm. It is can be implemented in the head-and-shoulder video sequence segmentation applications.
Hong Zhang 0018, Helong Wang, Wei Zuo
ICIP (3)2
2007 The Study of Detecting for IR Weak and Small Targets Based on Fractal Features
Hong Zhang 0018, Zhu Zhenfu
MMM (2)1