VLDB 2026 Research / reviewers in the wild / expert
Runbo Hu
dblp:245/9084
· DBLP profile ↗
15ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0002-3105-9772ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ARPDR++: Exploiting local-global temporal modeling for smartphone-based indoor pedestrian localization
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Yulan Guo, Pengfei Xu 0013, Runbo Hu |
Comput. Networks | 6 |
| 2025 | Universal Federated Domain Adaptation Through One-vs-All Self-Supervision for Internet of ThingsabstractIn practical Internet of Things (IoT) applications, deep neural networks (DNNs) often encounter challenges arising from covariate shifts (differences in feature distributions) and category shifts (discrepancies in label spaces), which significantly degrade their generalization performance. To mitigate these issues, universal federated domain adaptation (UFDA) techniques have been proposed to train a global model that can classify known and unknown categories while keeping data private. Nevertheless, most existing methods still struggle to precisely identify samples belonging to unknown classes in the target domain due to the unavailability of data from the source domain clients. To address these challenges, we propose a novel method, termed one-vs-all self-supervision (OSS) for IoT scenario. Specifically, OSS mainly consists of following three components. First, one-vs-all pseudo-label generation is proposed to generate high-quality pseudo-labels by leveraging source client models. Subsequently, we design a category-diverse strategy to aggregate the source models by assigning appropriate weights to each source domain client. Finally, we implement a target self-supervised learning strategy to refine feature alignment with respect to cluster centers. Comprehensive experiments are performed on four benchmark datasets: Office-31, Office-Home, VisDA-2017+ImageCLEF-DA, and Digits. The results show that our proposed OSS method achieves state-of-the-art performance in UFDA, significantly enhancing the recognition accuracy. Haojin Liao, Qiang Wang 0051, Sicheng Zhao, Tengfei Xing, Runbo Hu |
IEEE Internet Things J. | 5 |
| 2024 | LDTR: Transformer-based lane detection with anchor-chain representationabstractDespite recent advances in lane detection methods, scenarios with limited- or no-visual-clue of lanes due to factors such as lighting conditions and occlusion remain challenging and crucial for automated driving. Moreover, current lane representations require complex post-processing and struggle with specific instances. Inspired by the DETR architecture, we propose LDTR, a transformer-based model to address these issues. Lanes are modeled with a novel anchor-chain, regarding a lane as a whole from the beginning, which enables LDTR to handle special lanes inherently. To enhance lane instance perception, LDTR incorporates a novel multi-referenced deformable attention module to distribute attention around the object. Additionally, LDTR incorporates two line IoU algorithms to improve convergence efficiency and employs a Gaussian heatmap auxiliary branch to enhance model representation capability during training. To evaluate lane detection models, we rely on Fréchet distance, parameterized Fl-score, and additional synthetic metrics. Experimental results demonstrate that LDTR achieves state-of-the-art performance on well-known datasets. Zhongyu Yang, Tengfei Xing, Runbo Hu, Pengfei Xu 0013, Ruini Xue |
Comput. Vis. Media | 5 |
| 2023 | CANet: Curved Guide Line Network with Adaptive Decoder for Lane DetectionabstractLane detection is challenging due to the complicated onroad scenarios and line deformation from different camera perspectives. Lots of solutions were proposed, but can not deal with "corner lanes" well. To address this problem, this paper proposes a new top-down deep learning lane detection approach, CANet. A lane instance is first responded by the heatmap on the U-shaped "curved guide line" at global semantic level, thus the corresponding features of each lane are aggregated at the response point. Then CANet obtains the heatmap response of the entire lane through conditional convolution, and finally decodes the point set to describe lanes via adaptive decoder. The prototype is implemented with Pytorch, and evaluated against 3 well-known datasets extensively. The experimental results show that CANet reaches SOTA in different metrics. Zhongyu Yang, Tengfei Xing, Runbo Hu, Pengfei Xu 0013, Ruini Xue |
ICASSP | 5 |
| 2023 | Decoupling with Entropy-based Equalization for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation methods are the main solution to alleviate the problem of high annotation consumption in semantic segmentation. However, the class imbalance problem makes the model favor the head classes with sufficient training samples, resulting in poor performance of the tail classes. To address this issue, we propose a Decoupled Semi-Supervise Semantic Segmentation (DeS4) framework based on the teacher-student model. Specifically, we first propose a decoupling training strategy to split the training of the encoder and segmentation decoder, aiming at a balanced decoder. Then, a non-learnable prototype-based segmentation head is proposed to regularize the category representation distribution consistency and perform a better connection between the teacher model and the student model. Furthermore, a Multi-Entropy Sampling (MES) strategy is proposed to collect pixel representation for updating the shared prototype to get a class-unbiased head. We conduct extensive experiments of the proposed DeS4 on two challenging benchmarks (PASCAL VOC 2012 and Cityscapes) and achieve remarkable improvements over the previous state-of-the-art methods. Chuanghao Ding, Jianrong Zhang, Henghui Ding, Tengfei Xing, Runbo Hu |
IJCAI | 7 |
| 2023 | Understanding the Semantics of GPS-based Trajectories for Road Closure DetectionabstractThe accurate detection of road closures is of great value for real-time updating of digital maps. The existing methods mainly follow the paradigm of detecting the drastic changes in traffic statistical values (e.g., traffic flow), but they may lead to misidentifying since 1) drastic changes of traffic statistical values are hard to be observed in low-heat roads where the passing vehicles are sparse; 2) statistical values are sensitive to noise (e.g., traffic flow for tiny roads and tunnels is prone to miscounting); and 3) statistical values are naturally delayed, and misidentifying may occur when they have not yet shown significant changes. Surprisingly, since GPS-based trajectories can also exhibit significant abnormal patterns for road closures and have the superiority in fine granularity and timeliness, they can naturally tackle the above challenges. In this paper, we present a novel road closure detection framework based on mining the semantics of trajectories, called T-Closure. We first construct a heterogeneous graph based on the trajectory and the planned route to extract the spatial-topological property of each trajectory, where a node-level auxiliary task is proposed to guide the learning of feature encoders. A multi-view heterogeneous graph neural network (MVH-GNN) with a graph-level auxiliary task is then introduced to capture the semantics of trajectories, where intra-category relevance and inter-category interaction are both considered. Finally, a sequence-level auxiliary task refines the ability of LSTM in modeling the semantic relevance among trajectories while enhancing the robustness of our framework. Experiments on four real-world road closure datasets demonstrate the superiority of T-Closure. Online performance shows that T-Closure can detect 7000+ closure events monthly, with a delay within 1.5 hours. Kaiqiang An, Runbo Hu, Jie Shao 0001 |
KDD | 5 |
| 2023 | Domain consensual contrastive learning for few-shot universal domain adaptation
Haojin Liao, Qiang Wang 0051, Sicheng Zhao, Tengfei Xing, Runbo Hu |
Appl. Intell. | 5 |
| 2021 | Spatio-temporal Contrastive Domain Adaptation for Action RecognitionabstractCompared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial representation and temporal dynamics. Most previous works focus on short-term modeling and alignment with frame-level or clip-level features, which is not discriminative sufficiently for video-based UDA tasks. To address these problems, in this paper we propose to establish the cross-modal domain alignment via self-supervised contrastive framework, i.e., spatio-temporal contrastive domain adaptation (STCDA), to learn the joint clip-level and video-level representation alignment. Since the effective representation is modeled from unlabeled data by self-supervised learning (SSL), spatio-temporal contrastive learning (STCL) is proposed to explore the useful long-term feature representation for classification, using self-supervision setting trained from the contrastive clip/video pairs with positive or negative properties. Besides, we involve a novel domain metric scheme, i.e., video-based contrastive alignment (VCA), to optimize the category-aware video-level alignment and generalization between source and target. The proposed STCDA achieves stat-of-the-art results on several UDA benchmarks for action recognition. Sicheng Zhao, Jing-Yu Yang 0002, Huanjing Yue, Pengfei Xu 0013, Runbo Hu |
CVPR | 6 |
| 2021 | Dual Metric Discriminator for Open Set Video Domain AdaptationabstractExisting video domain adaptation methods focus on addressing closed set problems. However, it is nearly impossible to guarantee different domains share exactly the same set of categories in realistic scenarios. Hence, open set video domain adaptation (OSVDA) problem, which involves unknown categories, has achieved increasingly close attention. In this paper, we propose a seminal framework, which involves spatial and temporal information to address OSVDA problem. Besides, we design a novel discrimination module, i.e., Dual Metric Discriminator (DMD), to separate known and unknown categories based on implicit and explicit similarity metrics. We conduct comprehensive experiments on several benchmarks and achieve state-of-the-art performance with 40.4%, 33.7%, and 79.2% accuracy on UCF to HMDB, HMDB to UCF, and Kinetics to UCF scenarios respectively. Yatian Wang, Yezhen Wang, Pengfei Xu 0013, Runbo Hu |
ICASSP | 5 |
| 2020 | An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated VideosabstractEmotion recognition in user-generated videos plays an important role in human-centered computing. Existing methods mainly employ traditional two-stage shallow pipeline, i.e. extracting visual and/or audio features and training classifiers. In this paper, we propose to recognize video emotions in an end-to-end manner based on convolutional neural networks (CNNs). Specifically, we develop a deep Visual-Audio Attention Network (VAANet), a novel architecture that integrates spatial, channel-wise, and temporal attentions into a visual 3D CNN and temporal attentions into an audio 2D CNN. Further, we design a special classification loss, i.e. polarity-consistent cross-entropy loss, based on the polarity-emotion hierarchy constraint to guide the attention generation. Extensive experiments conducted on the challenging VideoEmotion-8 and Ekman-6 datasets demonstrate that the proposed VAANet outperforms the state-of-the-art approaches for video emotion recognition. Our source code is released at: https://github.com/maysonma/VAANet. Sicheng Zhao, Yunsheng Ma, Jufeng Yang, Tengfei Xing, Pengfei Xu 0013, Runbo Hu, Kurt Keutzer |
AAAI | 7 |
| 2020 | Multi-Source Distilling Domain AdaptationabstractDeep neural networks suffer from performance decay when there is domain shift between the labeled source domain and unlabeled target domain, which motivates the research on domain adaptation (DA). Conventional DA methods usually assume that the labeled data is sampled from a single source distribution. However, in practice, labeled data may be collected from multiple sources, while naive application of the single-source DA algorithms may lead to suboptimal solutions. In this paper, we propose a novel multi-source distilling domain adaptation (MDDA) network, which not only considers the different distances among multiple sources and the target, but also investigates the different similarities of the source samples to the target ones. Specifically, the proposed MDDA includes four stages: (1) pre-train the source classifiers separately using the training data from each source; (2) adversarially map the target into the feature space of each source respectively by minimizing the empirical Wasserstein distance between source and target; (3) select the source training samples that are closer to the target to fine-tune the source classifiers; and (4) classify each encoded target feature by corresponding source classifier, and aggregate different predictions using respective domain weight, which corresponds to the discrepancy between each source and target. Extensive experiments are conducted on public DA benchmarks, and the results demonstrate that the proposed MDDA significantly outperforms the state-of-the-art approaches. Our source code is released at: https://github.com/daoyuan98/MDDA. Sicheng Zhao, Guangzhi Wang, Shanghang Zhang, Yaxian Li, Zhichao Song, Pengfei Xu 0013, Runbo Hu, Kurt Keutzer |
AAAI | 8 |
| 2020 | ROAM: Recurrently Optimizing Tracking ModelabstractIn this paper, we design a tracking model consisting of response generation and bounding box regression, where the first component produces a heat map to indicate the presence of the object at different positions and the second part regresses the relative bounding box shifts to anchors mounted on sliding-window locations. Thanks to the resizable convolutional filters used in both components to adapt to the shape changes of objects, our tracking model does not need to enumerate different sized anchors, thus saving model parameters. To effectively adapt the model to appearance variations, we propose to offline train a recurrent neural optimizer to update tracking model in a meta-learning setting, which can converge the model in a few gradient steps. This improves the convergence speed of updating the tracking model while achieving better performance. We extensively evaluate our trackers, ROAM and ROAM++, on the OTB, VOT, LaSOT, GOT-10K and TrackingNet benchmark and our methods perform favorably against state-of-the-art algorithms. Tianyu Yang 0003, Pengfei Xu 0013, Runbo Hu, Antoni B. Chan |
CVPR | 3 |
| 2020 | Automatic Calibration of Road Intersection Topology using TrajectoriesabstractThe inaccuracy of road intersection in digital road map easily brings serious effects on the mobile navigation and other applications. Massive traveling trajectories of thousands of vehicles enable frequent updating of road intersection topology. In this paper, we first expand the road intersection detection issue into a topology calibration problem for road intersection influence zone. Distinct from the existing road intersection update methods, we not only determine the location and coverage of road intersection, but figure out incorrect or missing turning paths within whole influence zone based on unmatched trajectories as compared to the existing map. The important challenges of calibration issue include that trajectories are mixing with exceptional data, and road intersections are of different sizes and shapes, etc. To address above challenges, we propose a three-phase calibration framework, called CITT. It is composed of trajectory quality improving, core zone detection, and topology calibration within road intersection influence zone. From such components it can automatically obtain high quality topology of road intersection influence zone. Extensive experiments compared with the state-of-the-art methods using trajectory data obtained from Didi Chuxing and Chicago campus shuttles demonstrate that CITT method has strong stability and robustness and significantly outperforms the existing methods. Lisheng Zhao, Jiali Mao, Min Pu, Cheqing Jin, Weining Qian, Aoying Zhou, Runbo Hu |
ICDE | 9 |
| 2020 | ARPDR: An Accurate and Robust Pedestrian Dead Reckoning System for Indoor Localization on Handheld SmartphonesabstractThe proliferation of mobile computing has prompted Pedestrian Dead Reckoning (PDR) to be one of the most attractive and promising indoor localization techniques for ubiquitous applications. The existing PDR approaches either suffer position drifts caused by accumulative errors or are sensitive to various users. This paper presents ARPDR, an accurate and robust PDR approach to improve the accuracy and robustness of indoor localization methods. Particularly, we propose a novel step counting algorithm based on motion models by deeply exploiting inertial sensor data. We then combine step counting with adaptive thresholding to personalize the PDR system for different users. Furthermore, we propose a novel stride-heading model with a deep neural network to predict stride lengths and walking orientations, thus the displacement errors are significantly reduced. Extensive experiments on public datasets demonstrate that ARPDR outperforms the state-of-the-art PDR methods. Xiaoqiang Teng, Pengfei Xu 0013, Deke Guo, Yulan Guo, Runbo Hu, Didi Chuxing |
IROS | 5 |
| 2019 | Multi-source Domain Adaptation for Semantic SegmentationabstractSimulation-to-real domain adaptation for semantic segmentation has been actively studied for various applications such as autonomous driving. Existing methods mainly focus on a single-source setting, which cannot easily handle a more practical scenario of multiple sources with different distributions. In this paper, we propose to investigate multi-source domain adaptation for semantic segmentation. Specifically, we design a novel framework, termed Multi-source Adversarial Domain Aggregation Network (MADAN), which can be trained in an end-to-end manner. First, we generate an adapted domain for each source with dynamic semantic consistency while aligning at the pixel-level cycle-consistently towards the target. Second, we propose sub-domain aggregation discriminator and cross-domain cycle discriminator to make different adapted domains more closely aggregated. Finally, feature-level alignment is performed between the aggregated domain and target domain while training the segmentation network. Extensive experiments from synthetic GTA and SYNTHIA to real Cityscapes and BDDS datasets demonstrate that the proposed MADAN model outperforms state-of-the-art approaches. Our source code is released at: https://github.com/Luodian/MADAN. Sicheng Zhao, Bo Li 0080, Xiangyu Yue 0001, Pengfei Xu 0013, Runbo Hu, Kurt Keutzer |
NeurIPS | 6 |