VLDB 2026 Research / reviewers in the wild / expert
Beihang Song
dblp:313/9633
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-2179-4949ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Image recognition and object detection · 50% Deep learning architectures and training · 33% Video understanding and tracking · 17% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
object detection |
1.0 | 1 | 2026 | Efficient Oriented Object Detection via Wavelet-Based Energy Label Reassignment and Dual Prediction Strategy · IEEE Trans. Multim. 2026 |
Computer vision › Image recognition and object detection › object detection
oriented object detection |
1.0 | 1 | 2026 | Efficient Oriented Object Detection via Wavelet-Based Energy Label Reassignment and Dual Prediction Strategy · IEEE Trans. Multim. 2026 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.7 | 1 | 2023 | Tracking With Mutual Attention Network · IEEE Trans. Multim. 2023 |
Machine learning › Deep learning architectures and training › attention mechanism
mutual attention |
0.7 | 1 | 2023 | Tracking With Mutual Attention Network · IEEE Trans. Multim. 2023 |
Computer vision › Video understanding and tracking
object tracking |
0.7 | 1 | 2023 | Tracking With Mutual Attention Network · IEEE Trans. Multim. 2023 |
Methods — techniques the papers use, named apart from their topics
wavelet transform · 1.0heatmap keypoint prediction · 1.0feature fusion · 1.0energy weighting · 1.0foreground reinforcement learning · 0.7background training enhancement · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-modal category-aware gating for oriented object detection with single-category experts
Beihang Song, Hongquan Sun, Tong Liu 0039, Kai Zhu 0009, Jun Wan 0005 |
Neurocomputing | 1 |
| 2026 | Efficient Oriented Object Detection via Wavelet-Based Energy Label Reassignment and Dual Prediction StrategyabstractArbitrary-oriented object detection remains a pivotal research focus due to its practical significance and inherent challenges. Existing methods often extend frameworks and sampling strategies designed for horizontal object detectors, which struggle to handle the arbitrary orientations, high aspect ratios, and diverse scales of oriented objects. To overcome these limitations, we propose a novel and efficient method for arbitrary-oriented object detection. This approach dynamically assigns prediction layers by object pixel area, then leverages wavelet transform-based energy weighting for bottom-up sample reassignment, optimizing feature representation for oriented targets. In addition, a robust framework integrates heatmap keypoint prediction on feature maps of a quarter-sized image, along with sparse predictions on other scales. By querying small-object regions within deep feature maps, a progressive top-down feature fusion strategy further enhances the perception of fine-grained details. Extensive evaluations on four benchmark datasets demonstrate the method's substantial improvements in detection performance, establishing its potential for broader applications in oriented object detection. Beihang Song, Jing Li 0055, Jia Wu 0001, Xuefei Li 0001, Jun Wan 0005 |
IEEE Trans. Multim. | 1 |
| 2025 | HEART: Historically Information Embedding and Subspace Re-Weighting Transformer-Based TrackingabstractTransformers-based trackers offer significant potential for integrating semantic interdependence between template and search features in tracking tasks. Transformers possess inherent capabilities for processing long sequences and extracting correlations within them. Several researchers have explored the feasibility of incorporating Transformers to model continuously changing search areas in tracking tasks. However, their approach has substantially increased the computational cost of an already resource-intensive Transformer. Additionally, existing Transformers-based trackers rely solely on mechanically employing multi-head attention to obtain representations in different subspaces, without any inherent bias. To address these challenges, we propose HEART (Historical Information Embedding And Subspace Re-weighting Tracker). Our method embeds historical information into the queries in a lightweight and Markovian manner to extract discriminative attention maps for robust tracking. Furthermore, we develop a multi-head attention distribution mechanism to retrieve the most promising subspace weights for tracking tasks. HEART has demonstrated its effectiveness on five datasets, including OTB-100, LaSOT, UAV123, TrackingNet, and GOT-10k. Tianpeng Liu, Jing Li 0055, Amin Beheshti, Jia Wu 0001, Beihang Song, Lezhi Lian |
IEEE Trans. Big Data | 6 |
| 2024 | Single-stage oriented object detection via Corona Heatmap and Multi-stage Angle Prediction
Beihang Song, Jing Li 0055, Jia Wu 0001, Shan Xue 0001, Jun Wan 0005 |
Knowl. Based Syst. | 1 |
| 2024 | Direction Prediction Redefinition: Transfer Angle to Scale in Oriented Object DetectionabstractOriented object detection has garnered significant attention. However, rotational symmetry and discontinuity at boundaries can confuse networks, leading to discontinuous loss and regression inconsistency. In this paper, we propose an efficient multi-directional object detection framework named Direction Prediction Redefinition (DPR). We describe the angle variation of rotated bounding boxes ($B_{r}$) as changes in the dimensions of horizontal bounding boxes ($B_{h}$). Specifically, we generate two sets of horizontal bounding boxes by predicting the center points of the corresponding boundaries within the rotated bounding box, thereby avoiding boundary issues caused by angle prediction. To further achieve robust rotated boundary representation, we propose the Joint Scale Representation method and the State Feature Encoding module, which are used to eliminate outliers in rotated boundaries and guide the correct selection of horizontal bounding box vertices, respectively. Moreover, we further abstract DPR as Multiple Trigonometric functions based DPR (DPR-MT). This method maps a single angle into four sets of trigonometric functions and considers them as the four sides of the horizontal bounding box. This approach predicts angles in the form of horizontal bounding boxes without complex operations, making it plug-and-play. Experimental results and visual analysis on challenging datasets further verify the effectiveness and competitiveness of our proposed method. Beihang Song, Jing Li 0055, Jia Wu 0001, Jun Wan 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | SRDF: Single-Stage Rotate Object Detector via Dense Prediction and False Positive SuppressionabstractOriented object detection has made astonishing progress. However, existing methods neglect to address the issue of false positives caused by the background or nearby clutter objects. Meanwhile, class imbalance and boundary overflow issues caused by the predicting rotation angles may affect the accuracy of rotated bounding box predictions. To address the above issues, we propose a Single-stage Rotate object detector via Dense prediction and False positive suppression (SRDF). Specifically, we design an Instance-level False Positive Suppression Module (IFPSM), IFPSM acquires the weight information of target and non-target regions by supervised learning of spatial feature encoding, and applies these weight values to the deep feature map, thereby attenuating the response signals of non-target regions within the deep feature map. Compared to commonly used attention mechanisms, this approach more accurately suppresses false positive regions. Then, we introduce a hybrid classification and regression method to represent the object orientation, the proposed mothed divide the angle into two segments for prediction, reducing the number of categories and narrowing the range of regression. This alleviates the issue of class imbalance caused by treating one degree as a single category in classification prediction, as well as the problem of boundary overflow caused by directly regressing the angle. In addition, we transform the traditional post-processing steps based on matching and searching to a two-dimensional probability distribution mathematical model, which accurately and quickly extracts the bounding boxes from dense prediction results. Extensive experiments on Remote Sensing, Synthetic Aperture Radar, and Scene Text benchmarks demonstrate the superiority of the proposed SRDF method over state-of-the-art rotated object detection methods. Our codes are available at https://github.com/TomZandJerryZ/SRDF. Beihang Song, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Jun Wan 0005, Tianpeng Liu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Tracking With Mutual Attention NetworkabstractVisual tracking is a visual task that tracks a specific target by only giving its first frame location and size. To punish the low-quality but high-scoring tracking results, researchers resorted to foreground reinforcement learning to suppress the scores of positive samples near edges. However, for training with negative samples, all backgrounds are equally labeled as false. In this way, the interdependence and difference between the foreground and the background are not considered. We interpret the underlying reason for drifts as the imbalance between the embedding of background and foreground information. Specifically, some catastrophic tracking results and common tracking errors should not be treated equally but should strengthen the implicit connection between the foreground and background. In this paper, we propose a Mutual Attention (MA) module to strengthen the interdependence between positive and negative samples. It can aggregate the rich contextual interdependence between the target template and the search area, thereby providing an implicit way to update the target template accordingly. As for the difference, we design a background training enhancement (BTE) mechanism to distinguish negative samples with varying degrees of error, that is, to down-weight outrageous and absurd tracking results to improve the robustness of the tracker. The results on a large number of benchmarks indicate the validity of our results, such as OTB-100, VOT-2018, VOT-2019, and LaSOT. Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Beihang Song |
IEEE Trans. Multim. | 5 |