VLDB 2026 Research / reviewers in the wild / expert
Haosheng Chen 0001
dblp:251/3427-1
· DBLP profile ↗
15ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-6834-2136ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
3D vision · 40% Video understanding and tracking · 29% Face, body and person analysis · 16% | |
| Computer graphics and multimedia
1 paper |
Computational photography and imaging · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
camera pose estimation |
0.9 | 1 | 2025 | MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning · IJCAI 2025 |
Computer vision › 3D vision › pose estimation
correspondence-based pose estimation |
0.9 | 1 | 2025 | MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning · IJCAI 2025 |
Computer vision › 3D vision
outlier rejection |
0.9 | 1 | 2025 | MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning · IJCAI 2025 |
Computer vision › Face, body and person analysis
person re-identification |
0.9 | 1 | 2025 | Dual-Space Video Person Re-identification · Int. J. Comput. Vis. 2025 |
Computer vision › 3D vision › feature matching › two-view correspondence
two-view correspondence learning |
0.9 | 1 | 2025 | MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning · IJCAI 2025 |
Computer vision › Face, body and person analysis › person re-identification
video-based person re-identification |
0.9 | 1 | 2025 | Dual-Space Video Person Re-identification · Int. J. Comput. Vis. 2025 |
Computer vision › Video understanding and tracking › object tracking
event-based tracking |
0.8 | 2 | 2020 | End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object Tracking · AAAI 2020 Asynchronous Tracking-by-Detection on Adaptive Time Surfaces for Event-based Object Tracking · ACM Multimedia 2019 |
Computer vision › Video understanding and tracking
object tracking |
0.8 | 2 | 2020 | End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object Tracking · AAAI 2020 Asynchronous Tracking-by-Detection on Adaptive Time Surfaces for Event-based Object Tracking · ACM Multimedia 2019 |
Computer vision › Video understanding and tracking › video anomaly detection
violence detection |
0.8 | 1 | 2024 | Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence Detection · NeurIPS 2024 |
Computer vision › 3D vision
motion estimation |
0.4 | 1 | 2020 | End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object Tracking · AAAI 2020 |
Computer vision › 3D vision › motion estimation
object motion estimation |
0.4 | 1 | 2020 | End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object Tracking · AAAI 2020 |
Computer vision › Video understanding and tracking
video object detection |
0.4 | 1 | 2020 | Dual Semantic Fusion Network for Video Object Detection · ACM Multimedia 2020 |
Computer vision › Video understanding and tracking › multi-object tracking
tracking-by-detection |
0.4 | 1 | 2019 | Asynchronous Tracking-by-Detection on Adaptive Time Surfaces for Event-based Object Tracking · ACM Multimedia 2019 |
Machine learning › Graph learning › graph neural network › attention-based graph neural network
graph attention network |
0.3 | 1 | 2025 | MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning · IJCAI 2025 |
Machine learning › Graph learning › graph neural network › graph attention
multi-graph attention |
0.3 | 1 | 2025 | MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning · IJCAI 2025 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning |
0.2 | 1 | 2024 | Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence Detection · NeurIPS 2024 |
Computer vision › 3D vision › event-based vision
event camera |
0.1 | 1 | 2020 | End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object Tracking · AAAI 2020 |
Computer vision › Image recognition and object detection
object detection |
0.1 | 1 | 2020 | Dual Semantic Fusion Network for Video Object Detection · ACM Multimedia 2020 |
Computational photography and imaging
event camera |
0.1 | 1 | 2019 | Asynchronous Tracking-by-Detection on Adaptive Time Surfaces for Event-based Object Tracking · ACM Multimedia 2019 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 0.9sparse graph network · 0.9group relative policy optimization · 0.9graph of thoughts · 0.9graph attention network · 0.9dual-space learning · 0.9deep neural network · 0.9contextual attention · 0.9dual-space representation learning · 0.8cross-space attention · 0.8linear time decay · 0.4adaptive time surface · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transformer-Based 3-D Hand Pose Estimation via Bidirectional Multiscale Fusion and Learnable Anchor GuidanceabstractAccurate 3D hand pose estimation faces inherent challenges, including self-occlusions, joint similarities, and high degrees of freedom. Although most existing CNN-based or Transformer-based methods leverage global contexts, they often fail to capture fine-grained local details and robustly handle occluded joints. To address these limitations, we propose a novel Transformer framework for depth-based 3D hand pose estimation, which incorporates two key designs: the Bidirectional Multiscale Fusion and Learnable Anchor Guidance. Firstly, we propose bidirectional multiscale fusion that sequentially propa-gates features from different encoder levels in both top-down and bottom-up directions, followed by a final aggregation of all scale features. Such an intricate feature interaction eventually enables effective joint modeling of fine-grained local details (e.g., fingertip positions) and high-level semantic context information (e.g., palm orientation). Secondly, we introduce the learnable anchor query as prior guidance to dynamically guide the decoder to localize ambiguous joints better. The learnable anchors are derived from joint-specific attention maps under 3D ground-truth supervision and then are concatenated with static grid anchors to form hybrid anchors, which effectively enable more precise 3D hand pose estimation, especially for occlusions. To further demonstrate the practicality of our framework for IoT-oriented deployment, we conduct edge-device experiments to validate its deployment feasibility. Experiments on benchmark datasets (including NYU, ICVL, MSRA, and DexYCB) demonstrate the superiority of our method over previous state-of-the-art approaches. Ji Gan, Weiqiang Wang 0001, Feng Gao 0005, Jiaxu Leng, Haosheng Chen 0001, Xinbo Gao 0001 |
IEEE Internet Things J. | 6 |
| 2026 | Spatially Aware Adaptive Diffusion: Unifying Low-Resolution Image Fusion and Super-ResolutionabstractLow-resolution visible-infrared image fusion and super-resolution (LRVIF) are critical for enhancing image quality in low-resolution scenarios, yet limited information in the input images often constrains performance. To address these challenges, we propose SaDiff, a spatially-aware adaptive diffusion model that introduces diffusion processes into LRVIF for the first time, representing a major breakthrough in the field. Leveraging the generative capabilities of diffusion models, our approach unifies and enhances image fusion and super-resolution within a cohesive framework. A key component of SaDiff is the Spatial Residual Adaptation Block, which extends the diffusion process by dynamically adapting feature representations to spatial variations in the local regions of the input images. This module maximally preserves crucial information from the input images, such as texture details and contrast, while effectively suppressing noise, ensuring robust and context-aware feature refinement. Then we further propose Direct Diffusion Synthesis, a novel mechanism that utilizes noise predictions during diffusion to generate fused images, enabling joint training of the fusion and super-resolution networks. Additionally, a Cross-Feature Fusion Module integrates texture and contrast details, producing super-resolution fused images with improved clarity and structural integrity. Extensive experiments show that SaDiff achieves state-of-the-art performance, offering a robust and unified solution to infrared-visible image fusion and super-resolution. The code for the proposed method will be made available at https://github.com/guobaoxiao/SaDiff. Jiajia Fu, Zhenni Yu, Haosheng Chen 0001, Songlin Du, Changcai Yang, Lianghua He, Guobao Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence LearningabstractTwo-view correspondence learning is a key task in computer vision, which aims to establish reliable matching relationships for applications such as camera pose estimation and 3D reconstruction. However, existing methods have limitations in local geometric modeling and cross-stage information optimization, which make it difficult to accurately capture the geometric constraints of matched pairs and thus reduce the robustness of the model. To address these challenges, we propose a Multi-Graph Contextual Attention Network (MGCA-Net), which consists of a Contextual Geometric Attention (CGA) module and a Cross-Stage Multi-Graph Consensus (CSMGC) module. Specifically, CGA dynamically integrates spatial position and feature information via an adaptive attention mechanism and enhances the capability to capture both local and global geometric relationships. Meanwhile, CSMGC establishes geometric consensus via a cross-stage sparse graph network, ensuring the consistency of geometric information across different stages. Experimental results on two representative YFCC100M and SUN3D datasets show that MGCA-Net significantly outperforms existing SOTA methods in the outlier rejection and camera pose estimation tasks. Source code is available at http://www.linshuyuan.com. Shuyuan Lin, Mengtin Lo, Haosheng Chen 0001, Qiangqiang Wu |
IJCAI | 3 |
| 2025 | A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly UnderstandingabstractWhile unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditions, leading to significant performance drops in drone-view scenarios.To bridge this gap, we introduce A2Seek (Aerial Anomaly Seek), a large-scale, reasoning-centric benchmark dataset for aerial anomaly understanding. This dataset covers various scenarios and environmental conditions, providing high-resolution real-world aerial videos with detailed annotations, including anomaly categories, frame-level timestamps, region-level bounding boxes, and natural language explanations for causal reasoning. Building on this dataset, we propose A2Seek-R1, a novel reasoning framework that generalizes R1-style strategies to aerial anomaly understanding, enabling a deeper understanding of “Where” anomalies occur and “Why” they happen in aerial frames.To this end, A2Seek-R1 first employs a graph-of-thought (GoT)-guided supervised fine-tuning approach to activate the model's latent reasoning capabilities on A2Seek. Then, we introduce Aerial Group Relative Policy Optimization (A-GRPO) to design rule-based reward functions tailored to aerial scenarios. Furthermore, we propose a novel “seeking” mechanism that simulates UAV flight behavior by directing the model's attention to informative regions.Extensive experiments demonstrate that A2Seek-R1 achieves up to a 22.04\% improvement in AP for prediction accuracy and a 13.9\% gain in mIoU for anomaly localization, exhibiting strong generalization across complex environments and out-of-distribution scenarios. Our dataset and code are released at https://2-mo.github.io/A2Seek/. Mengjingcheng Mo, Xinyang Tong, Mingpi Tan, Jiaxu Leng, Jiankang Zheng, Haosheng Chen 0001, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001 |
NeurIPS | 7 |
| 2025 | Dual-Space Video Person Re-identification
Jiaxu Leng, Changjiang Kuang, Ji Gan, Haosheng Chen 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 5 |
| 2024 | Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence DetectionabstractWhile numerous Video Violence Detection (VVD) methods have focused on representation learning in Euclidean space, they struggle to learn sufficiently discriminative features, leading to weaknesses in recognizing normal events that are visually similar to violent events (i.e., ambiguous violence). In contrast, hyperbolic representation learning, renowned for its ability to model hierarchical and complex relationships between events, has the potential to amplify the discrimination between visually similar events. Inspired by these, we develop a novel Dual-Space Representation Learning (DSRL) method for weakly supervised VVD to utilize the strength of both Euclidean and hyperbolic geometries, capturing the visual features of events while also exploring the intrinsic relations between events, thereby enhancing the discriminative capacity of the features. DSRL employs a novel information aggregation strategy to progressively learn event context in hyperbolic spaces, which selects aggregation nodes through layer-sensitive hyperbolic association degrees constrained by hyperbolic Dirichlet energy. Furthermore, DSRL attempts to break the cyber-balkanization of different spaces, utilizing cross-space attention to facilitate information interactions between Euclidean and hyperbolic space to capture better discriminative features for final violence detection. Comprehensive experiments demonstrate the effectiveness of our proposed DSRL. Jiaxu Leng, Zhanjie Wu, Mingpi Tan, Ji Gan, Haosheng Chen 0001, Xinbo Gao 0001 |
NeurIPS | 6 |
| 2024 | Joint Spatio-Temporal Similarity and Discrimination Learning for Visual TrackingabstractVisual tracking is a task of localizing a target unceasingly in a video with an initial target state at the first frame. The limited target information makes this problem an extremely challenging task. Existing tracking methods either perform matching based similarity learning or optimization based discrimination reasoning. However, these two types of tracking methods suffer from the problem of ineffectiveness for distinguishing target objects from background distractors and the problem of insufficiency in maintaining spatio-temporal consistency among successive frames, respectively. In this paper, we design a joint spatio-temporal similarity and discrimination learning (STSDL) framework for accurate and robust tracking. The designed framework is composed of two complementary branches: a similarity learning branch and a discrimination learning branch. The similarity learning branch uses an effective transformer encoder-decoder to gather rich spatio-temporal context information to generate a similarity map. In parallel, the discrimination learning branch exploits an efficient model predictor to train a target model to produce a discriminative map. Finally, the similarity map and the discriminative map are adaptively fused for accurate and robust target localization. Experimental results on six prevalent datasets demonstrate that the proposed STSDL can obtain satisfactory results, while it retains a real-time tracking speed of 50 FPS on a single GPU. Haosheng Chen 0001, Qiangqiang Wu, Changqun Xia, Jia Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Guided Sampling Based Feature Aggregation for Video Object DetectionabstractVideo object detection is a challenging task due to the presence of appearance deterioration in video frames. Recently, feature aggregation based methods which aggregate context information from object proposals in different frames to improve the performance, have dominated the task. However, much invalid information may be introduced during feature aggregation since frames and proposals are usually selected at random. In this paper, we propose a guided sampling based feature aggregation network (GSFA) to perform more effective feature aggregation. Specifically, we introduce a frame-level sampling module and a proposal-level sampling module to sample informative frames and proposals from a video sequence adaptively. As a result, the proposed GSFA can effectively aggregate context information from the semantically rich frames and proposals to boost the performance. Experimental results on the ImageNet VID dataset show the proposed GSFA achieves the state-of-the-art performance of 84.8% mAP with ResNet-101 and 85.8% mAP with ResNeXt-101. Haosheng Chen 0001, Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
ICIP | 2 |
| 2021 | Semantic Loop Closure Detection With Instance-Level Inconsistency Removal in Dynamic Industrial ScenesabstractA novel semantic loop closure detection (SLCD) method is proposed in this article for visual simultaneous localization and mapping systems. SLCD aims to relieve the instance-level semantic inconsistency issue that arose from dynamic industrial scenes (e.g., autonomous driving in big cities). As the first step in this direction, SLCD fully exploits both low- and high-level video frame information, in a coarse-to-fine way. In SLCD, we adopt a convolutional neural network based object detection to acquire object information from the consecutive frames. Meanwhile, we perform a bag of visual words based similarity calculation to narrow the frames to coarse loop closure candidates. For these candidates, we perform an object matching on them to find their semantic inconsistency cases and remove involved semantic inconsistencies according to their cases. Then, we recalculate the similarity scores for these candidates. Finally, loop closures are determined by the similarity scores and a geometrical verification. Favorable performance of the proposed method is demonstrated by comparing it to other state-of-the-art methods using data from several public datasets and our new Dynamic Scenes dataset. Haosheng Chen 0001, Ge Zhang 0008, Yangdong Ye |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object TrackingabstractEvent cameras, which are asynchronous bio-inspired vision sensors, have shown great potential in computer vision and artificial intelligence. However, the application of event cameras to object-level motion estimation or tracking is still in its infancy. The main idea behind this work is to propose a novel deep neural network to learn and regress a parametric object-level motion/transform model for event-based object tracking. To achieve this goal, we propose a synchronous Time-Surface with Linear Time Decay (TSLTD) representation, which effectively encodes the spatio-temporal information of asynchronous retinal events into TSLTD frames with clear motion patterns. We feed the sequence of TSLTD frames to a novel Retinal Motion Regression Network (RMRNet) to perform an end-to-end 5-DoF object motion regression. Our method is compared with state-of-the-art object tracking methods, that are based on conventional cameras or event cameras. The experimental results show the superiority of our method in handling various challenging environments such as fast motion and low illumination conditions. Haosheng Chen 0001, David Suter, Qiangqiang Wu, Hanzi Wang |
AAAI | 1 |
| 2020 | Learning Target-Specific Response Attention for Siamese Network Based Visual Tracking
Penghui Zhao, Haosheng Chen 0001, Yan Yan 0001, Hanzi Wang |
ACIVS | 2 |
| 2020 | Dual Semantic Fusion Network for Video Object DetectionabstractVideo object detection is a tough task due to the deteriorated quality of video sequences captured under complex environments. Currently, this area is dominated by a series of feature enhancement based methods, which distill beneficial semantic information from multiple frames and generate enhanced features through fusing the distilled information. However, the distillation and fusion operations are usually performed at either frame level or instance level with external guidance using additional information, such as optical flow and feature memory. In this work, we propose a dual semantic fusion network (abbreviated as DSFNet) to fully exploit both frame-level and instance-level semantics in a unified fusion framework without external guidance. Moreover, we introduce a geometric similarity measure into the fusion process to alleviate the influence of information distortion caused by noise. As a result, the proposed DSFNet can generate more robust features through the multi-granularity fusion and avoid being affected by the instability of external guidance. To evaluate the proposed DSFNet, we conduct extensive experiments on the ImageNet VID dataset. Notably, the proposed dual semantic fusion network achieves, to the best of our knowledge, the best performance of 84.1% mAP among the current state-of-the-art video object detectors with ResNet-101 and 85.4% mAP with ResNeXt-101 without using any post-processing steps. Lijian Lin, Haosheng Chen 0001, Honglun Zhang, Yu Li 0003, Ying Shan, Hanzi Wang |
ACM Multimedia | 2 |
| 2020 | Learning intra-inter semantic aggregation for video object detectionabstractVideo object detection is a challenging task due to the appearance deterioration problems in video frames. Thus, object features extracted from different frames of a video are usually deteriorated in varying degrees. Currently, some state-of-the-art methods enhance the deteriorated object features in a reference frame by aggregating the undeteriorated object features extracted from other frames, simply based on their learned appearance relation among object features. In this paper, we propose a novel intra-inter semantic aggregation method (ISA) to learn more effective intra and inter relations for semantically aggregating object features. Specifically, in the proposed ISA, we first introduce an intra semantic aggregation module (Intra-SAM) to enhance the deteriorated spatial features based on the learned intra relation among the features at different positions of an individual object. Then, we present an inter semantic aggregation module (Inter-SAM) to enhance the deteriorated object features in the temporal domain based on the learned inter relation among object features. As a result, by leveraging Intra-SAM and Inter-SAM, the proposed ISA can generate discriminative features from the novel perspective of intra-inter semantic aggregation for robust video object detection. We conduct extensive experiments on the ImageNet VID dataset to evaluate ISA. The proposed ISA obtains 84.5% mAP and 85.2% mAP with ResNet-101 and ResNeXt-101, and it achieves superior performance compared with several state-of-the-art video object detectors. Haosheng Chen 0001, Kaiwen Du, Yan Yan 0001, Hanzi Wang |
MMAsia | 2 |
| 2019 | Asynchronous Tracking-by-Detection on Adaptive Time Surfaces for Event-based Object TrackingabstractEvent cameras, which are asynchronous bio-inspired vision sensors, have shown great potential in a variety of situations, such as fast motion and low illumination scenes. However, most of the event-based object tracking methods are designed for scenarios with untextured objects and uncluttered backgrounds. There are few event-based object tracking methods that support bounding box-based object tracking. The main idea behind this work is to propose an asynchronous Event-based Tracking-by-Detection (ETD) method for generic bounding box-based object tracking. To achieve this goal, we present an Adaptive Time-Surface with Linear Time Decay (ATSLTD) event-to-frame conversion algorithm, which asynchronously and effectively warps the spatio-temporal information of asynchronous retinal events to a sequence of ATSLTD frames with clear object contours. We feed the sequence of ATSLTD frames to the proposed ETD method to perform accurate and efficient object tracking, which leverages the high temporal resolution property of event cameras. We compare the proposed ETD method with seven popular object tracking methods, that are based on conventional cameras or event cameras, and two variants of ETD. The experimental results show the superiority of the proposed ETD method in handling various challenging environments. Haosheng Chen 0001, Qiangqiang Wu, Xinbo Gao 0001, Hanzi Wang |
ACM Multimedia | 1 |
| 2019 | Robust Visual Tracking via Statistical Positive Sample Generation and Gradient Aware LearningabstractIn recent years, Convolutional Neural Network (CNN) based trackers have achieved state-of-the-art performance on multiple benchmark datasets. Most of these trackers train a binary classifier to distinguish the target from its background. However, they suffer from two limitations. Firstly, these trackers cannot effectively handle significant appearance variations due to the limited number of positive samples. Secondly, there exists a significant imbalance of gradient contributions between easy and hard samples, where the easy samples usually dominate the computation of gradient. In this paper, we propose a robust tracking method via Statistical Positive sample generation and Gradient Aware learning (SPGA) to address the above two limitations. To enrich the diversity of positive samples, we present an effective and efficient statistical positive sample generation algorithm to generate positive samples in the feature space. Furthermore, to handle the issue of imbalance between easy and hard samples, we propose a gradient sensitive loss to harmonize the gradient contributions between easy and hard samples. Extensive experiments on three challenging benchmark datasets including OTB50, OTB100 and VOT2016 demonstrate that the proposed SPGA performs favorably against several state-of-the-art trackers. Lijian Lin, Haosheng Chen 0001, Yan Yan 0001, Hanzi Wang |
MMAsia | 2 |