VLDB 2026 Research / reviewers in the wild / expert
Hu Cao
dblp:64/1651
· DBLP profile ↗
32ranked-venue papers
12as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeuralSEIR: Modeling uncertainty in non-pharmaceutical interventions with neural epidemic dynamicsabstractUnderstanding how non-pharmaceutical interventions (NPIs) influence epidemic trajectories is critical for evidence-based public health planning. Most COVID-19 models either treat NPIs as fixed effects, ignoring behavioral fatigue, or use black-box learning approaches that lack epidemiological transparency. We propose NeuralSEIR , a hybrid modeling framework that links daily human mobility to time-varying disease transmission in a mechanistically interpretable way. The model integrates a vaccination-aware compartmental core (SVEIC), which estimates a baseline transmission rate, with a shallow multilayer perceptron (MLP) that adjusts this baseline using Google mobility data to produce a behavior-aware rate. These components are coupled via ordinary differential equations, ensuring epidemiological consistency while allowing adaptive learning. Applied to nine COVID-19 waves across Germany, Japan, and the Philippines, NeuralSEIR reduces 14-day mean absolute percentage error by up to 60 % compared to a mobility-free baseline. It also reveals country-specific correlations between mobility patterns and transmission, highlighting the role of retail, transit, and residential activity in shaping NPI effectiveness. Benchmarks on U.S. state-level data against six CDC-tracked models further demonstrate its accuracy. By capturing mobility-driven changes in transmission, NeuralSEIR offers a transparent, data-informed tool for tailoring NPIs to local behavioral dynamics-bridging mechanistic epidemiology and explainable AI. Hu Cao, Longbing Cao |
Pattern Recognit. | 1 |
| 2026 | I2EKD: Efficient and Versatile Image-to-Event Knowledge DistillationabstractRecently, general-purpose features for event camera data have become increasingly important in advancing event-based vision applications. Current methods typically adopt pre-training paradigms, yielding promising performance. However, the limited data and sparse spatial information of events hinder effective use of pretraining for rich semantic learning. In this paper, we tackle semantic scarcity by transferring knowledge from large pre-trained image models, without increasing event training data. Concretely, we propose a novel image-to-event knowledge distillation method named I2EKD. Acknowledging that different backbones suit different applications, we fix the teacher and keep the student architecture flexible. To improve versatility, we equip I2EKD with two model-agnostic objectives at the logit and feature levels. Additionally, without task-specific objectives or labels, I2EKD avoids re-distillation and transfers well to downstream applications. Furthermore, leveraging DINOv2 as the teacher, whose feature distribution is built from billions of data, the student can swiftly mimic the superior distribution in a data-efficient manner. Compared with the SOTA pre-training method, I2EKD generates outperforming or comparable features with 1/15 training cost (1/10 data × 2/3 epochs). Extensive experiments on different vision tasks (object recognition, semantic segmentation, and monocular depth) verify the effectiveness of our method. Notably, I2EKD achieves top-1 object recognition accuracy of 70.72%, leading the pre-training SOTA by 5.89%. Hu Cao, Sanqing Qu, Fan Lu 0001, Yan Zhong 0001, Zhichao Lu, Luziwei Leng, Guang Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | An Online-Training-Free Adaptor for Open Heterogeneous Collaborative Perception via Diffusion ModelabstractCollaborative perception seeks to mitigate the limitations of single-vehicle perception, such as occlusions, by facilitating communication and information sharing among connected vehicles. However, most existing works assume a homogeneous scenario where all vehicles share identity sensor types and perception model architectures. In contrast, real-world systems often involve heterogeneous agents with diverse sensor configurations and independently developed models. In such settings, directly exchanging features without proper alignment can significantly degrade performance and hinder effective collaboration. While some methods have been proposed to address heterogeneity, they typically require retraining or access to internal model parameters, making them impractical for scalable deployment. To address these challenges, we propose DiffAlign, a plug-and-play adapter that enables feature alignment across heterogeneous agents in a training-free and model-agnostic manner. DiffAlign treats received BEV features as noisy latent representations and progressively refines them through a pretrained diffusion process. This alignment strategy does not require access to model internals or any retraining, which makes it both scalable and privacy-preserving while supporting diverse sensor modalities and perception backbones. Extensive experiments on simulated OPV2V and real-world V2V4Real datasets demonstrate that DiffAlign consistently improves detection performance in heterogeneous settings, improving CoBEVT by 132.01% and 91.95%, respectively. Our method provides a practical path toward scalable, generalizable, and deployment-ready collaborative perception. Tianhang Wang, Fan Lu 0001, Sanqing Qu, Bin Li 0087, Hu Cao, Alois C. Knoll, Guang Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous Drivingabstract28031 Rui Song 0007, Chenwei Liang, Yan Xia 0003, Walter Zimmer, Hu Cao, Holger Caesar, Andreas Festag, Alois C. Knoll |
ICCV | 5 |
| 2025 | TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesabstractWe present TUMTraf VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatio-temporal object expressions, TUMTraf VideoQA unifies three essential tasks—multiple-choice video question answering, referred object captioning, and spatio-temporal object grounding—within a cohesive evaluation framework. We further introduce the TraffiX-Qwen baseline model, enhanced with visual token sampling strategies, providing valuable insights into the challenges of fine-grained spatio-temporal reasoning. Extensive experiments demonstrate the dataset’s complexity, highlight the limitations of existing models, and position TUMTraf VideoQA as a robust foundation for advancing research in intelligent transportation systems. The dataset and benchmark are publicly available to facilitate further exploration. Xingcheng Zhou, Konstantinos Larintzakis, Walter Zimmer, Hu Cao, Venkatnarayanan Lakshminarasimhan, Leah Strand, Alois C. Knoll |
ICML | 6 |
| 2025 | Feature-aligned Fisheye Object Detection Network for Autonomous DrivingabstractFisheye cameras, renowned for their panoramic field of view (FOV) of 360°, are crucial for surround-view perception in autonomous driving. However, research on object perception in fisheye images lags behind that of standard images. To address this gap, we propose a feature-aligned fisheye object detection network specifically tailored for autonomous driving. Current fisheye perception algorithms often overlook the misalignment issues that typically arise in object detectors. To tackle these challenges in the feature pyramid network (FPN), we introduce a feature-aligned pyramid module (FaPM), which learns pixel transformation offsets to contextually align feature maps. Additionally, we present a location-aligned detection head (LaDH) to align the spatial distribution of classification and regression localization. Integrating these modules into a detection framework results in a novel feature-aligned fisheye object detector. Our method undergoes extensive evaluation on the WoodScape dataset, achieving a mean average precision (mAP) of 32.2%, surpassing the performance of existing methods. Hu Cao, Dongyi Sun, Rui Song 0007, Yan Xia 0003, Alois C. Knoll |
IROS | 1 |
| 2024 | Strong but Simple: A Baseline for Domain Generalized Dense Perception by CLIP-Based Transfer Learning
Christoph Hümmer, Manuel Schwonberg, Liangwei Zhou, Hu Cao, Alois C. Knoll, Hanno Gottschalk |
ACCV (10) | 4 |
| 2024 | BiSeg-SAM: Weakly-Supervised Post-Processing Framework for Boosting Binary Segmentation in Segment Anything ModelsabstractAccurate segmentation of polyps and skin lesions is essential for diagnosing colorectal and skin cancers. While various segmentation methods for polyps and skin lesions using fully supervised deep learning techniques have been developed, the pixel-level annotation of medical images by doctors is both time-consuming and costly. Foundational vision models like the Segment Anything Model (SAM) have demonstrated superior performance; however, directly applying SAM to medical segmentation may not yield satisfactory results due to the lack of domain-specific medical knowledge. In this paper, we propose BiSeg-SAM, a SAM-guided weakly supervised prompting and boundary refinement network for the segmentation of polyps and skin lesions. Specifically, we fine-tune SAM combined with a CNN module to learn local features. We introduce a WeakBox with two functions: automatically generating box prompts for the SAM model and using our proposed Multi-choice Mask-to-Box (MM2B) transformation for rough mask-to-box conversion, addressing the mismatch between coarse labels and precise predictions. Additionally, we apply scale consistency (SC) loss for prediction scale alignment. Our DetailRefine module enhances boundary precision and segmentation accuracy by refining coarse predictions using a limited amount of ground truth labels. This comprehensive approach enables BiSeg-SAM to achieve excellent multi-task segmentation performance. Our method demonstrates significant superiority over state-of-the-art (SOTA) methods when tested on five polyp datasets and one skin cancer dataset. The code for this work is open-sourced and available at https://github.com/suencgo/BiSeg-SAM. Encheng Su, Hu Cao, Alois C. Knoll |
BIBM | 2 |
| 2024 | Collaborative Semantic Occupancy Prediction with Hybrid Feature Fusion in Connected Automated VehiclesabstractCollaborative perception in automated vehicles lever-ages the exchange of information between agents, aiming to elevate perception results. Previous camera-based collabo-rative 3D perception methods typically employ 3D bounding boxes or bird's eye views as representations of the en-vironment. However, these approaches fall short in offering a comprehensive 3D environmental prediction. To bridge this gap, we introduce the first method for collaborative 3D semantic occupancy prediction. Particularly, it improves local 3D semantic occupancy predictions by hybrid fusion of (i) semantic and occupancy task features, and (ii) Compressed orthogonal attention features shared between vehi-cles. Additionally, due to the lack of a collaborative perception dataset designed for semantic occupancy prediction, we augment a current collaborative perception dataset to include 3D collaborative semantic occupancy labels for a more robust evaluation. The experimental findings highlight that: (i) our collaborative semantic occupancy predictions excel above the results from single vehicles by over 30%, and (ii) models anchored on semantic occupancy outpace state-of-the-art collaborative 3D detection techniques in subsequent perception applications, showcasing enhanced accuracy and enriched semantic-awareness in road environments. Rui Song 0007, Chenwei Liang, Hu Cao, Zhiran Yan, Walter Zimmer, Markus Gross 0003, Andreas Festag, Alois C. Knoll |
CVPR | 3 |
| 2024 | Embracing Events and Frames with Hierarchical Feature Refinement Network for Object Detection
Hu Cao, Zehua Zhang 0009, Yan Xia 0003, Jiahao Xia 0001, Guang Chen 0001, Alois C. Knoll |
ECCV (84) | 1 |
| 2024 | Dataset Distillation by Automatic Training Trajectories
Dai Liu, Jindong Gu, Hu Cao, Carsten Trinitis, Martin Schulz 0001 |
ECCV (87) | 3 |
| 2024 | Lightweight Fisheye Object Detection Network with Transformer-based Feature Enhancement for Autonomous DrivingabstractFisheye cameras, offering a wide field of view (FOV) of 360◦, are extensively employed for surround-view perception in autonomous driving. Compared with the object detection on the standard images, it lacks studies for fisheye images. Moreover, efficient perception is crucial for autonomous vehicles with limited computational capability. In this work, we introduce a lightweight fisheye object detection network with transformer-based feature enhancement for autonomous driving. Specifically, we leverage ShuffleNet V2 as a feature extraction network to reduce computation complexity and develop a transformer-based feature enhancement module (TFEM) to integrate multi-level features. Notably, we observe that data augmentation methods like mix-up and mosaic, effective on standard images, do not yield positive results on fisheye images. The results on the WoodScape dataset demonstrate that our method can achieve better performance with fewer parameters and floating-point operations per second (FLOPs). Extending our evaluation to the Microsoft Common Objects in Context (MS COCO) dataset shows that the proposed method has excellent generalization capability. Hu Cao, Yinlong Liu, Guang Chen 0001, Alois C. Knoll |
IROS | 1 |
| 2024 | Cyclostationary harmonic product spectrum with its application for rolling bearing fault resonance frequency band adaptive location
Cai Yi, Hu Cao, Lei Yan 0004, Qiuyang Zhou, Guiting Tang, Le Ran, Jianhui Lin |
Expert Syst. Appl. | 3 |
| 2024 | Transformation Decoupling Strategy Based on Screw Theory for Deterministic Point Cloud Registration With Gravity PriorabstractPoint cloud registration is challenging in the presence of heavy outlier correspondences. This paper focuses on addressing the robust correspondence-based registration problem with gravity prior that often arises in practice. The gravity directions are typically obtained by inertial measurement units (IMUs) and can reduce the degree of freedom (DOF) of rotation from 3 to 1. We propose a novel transformation decoupling strategy by leveraging the screw theory. This strategy decomposes the original 4-DOF problem into three sub-problems with 1-DOF, 2-DOF, and 1-DOF, respectively, enhancing computation efficiency. Specifically, the first 1-DOF represents the translation along the rotation axis, and we propose an interval stabbing-based method to solve it. The second 2-DOF represents the pole which is an auxiliary variable in screw theory, and we utilize a branch-and-bound method to solve it. The last 1-DOF represents the rotation angle, and we propose a global voting method for its estimation. The proposed method solves three consensus maximization sub-problems sequentially, leading to efficient and deterministic registration. In particular, it can even handle the correspondence-free registration problem due to its significant robustness. Extensive experiments on both synthetic and real-world datasets demonstrate that our method is more efficient and robust than state-of-the-art methods, even when dealing with outlier rates exceeding 99%. Zijian Ma, Yinlong Liu, Walter Zimmer, Hu Cao, Feihu Zhang, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Unsupervised Deep Hashing With Fine-Grained Similarity-Preserving Contrastive Learning for Image RetrievalabstractUnsupervised deep hashing has demonstrated significant advancements with the development of contrastive learning. However, most of previous methods have been hindered by insufficient similarity mining using global-only image representations. This has led to interference from background or non-interest objects during similarity reconstruction and contrastive learning. To address this limitation, we propose a novel unsupervised deep hashing framework named Fine-grained Similarity-preserving Contrastive learning Hashing (FSCH), which explores fine-grained semantic similarity among different images and their augmented views more comprehensively. It mainly comprises two modules: the global-local fine-grained similarity consistency preservation module and the local fine-grained similarity contrast preservation module. Specifically, we reconstruct local pairwise similarity structures by matching fine-grained patches, in conjunction with global similarity structures based on global hash codes cosine similarity, to generate hash codes with the ability to preserve global-local similarity consistency. Moreover, the preservation of local fine-grained similarity among augmented views is accomplished through the common regional features mutual representation between patches, then we enhance the discriminability of hash codes by mitigating the potential features difference during contrastive learning. Experimental results on four benchmark datasets demonstrate that our FSCH achieves an excellent retrieval performance compared to state-of-the-art unsupervised hashing methods. Hu Cao, Lei Huang 0010, Jie Nie, Zhiqiang Wei 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | SDPT: Semantic-Aware Dimension-Pooling Transformer for Image SegmentationabstractImage segmentation plays a critical role in autonomous driving by providing vehicles with a detailed and accurate understanding of their surroundings. Transformers have recently shown encouraging results in image segmentation. However, transformer-based models are challenging to strike a better balance between performance and efficiency. The computational complexity of the transformer-based models is quadratic with the number of inputs, which severely hinders their application in dense prediction tasks. In this paper, we present the semantic-aware dimension-pooling transformer (SDPT) to mitigate the conflict between accuracy and efficiency. The proposed model comprises an efficient transformer encoder for generating hierarchical features and a semantic-balanced decoder for predicting semantic masks. In the encoder, a dimension-pooling mechanism is used in the multi-head self-attention (MHSA) to reduce the computational cost, and a parallel depth-wise convolution is used to capture local semantics. Simultaneously, we further apply this dimension-pooling attention (DPA) to the decoder as a refinement module to integrate multi-level features. With such a simple yet powerful encoder-decoder framework, we empirically demonstrate that the proposed SDPT achieves excellent performance and efficiency on various popular benchmarks, including ADE20K, Cityscapes, and COCO-Stuff. For example, our SDPT achieves 48.6$\%$mIOU on the ADE20K dataset, which outperforms the current methods with fewer computational costs. The codes can be found at https://github.com/HuCaoFighting/SDPT. Hu Cao, Guang Chen 0001, Hengshuang Zhao, Dongsheng Jiang, Xiaopeng Zhang 0008, Qi Tian 0001, Alois C. Knoll |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Robust Face Alignment via Inherent Relation Learning and Uncertainty EstimationabstractHuman tends to locate the facial landmarks with heavy occlusion by their relative position to the easily identified landmarks. The clue is defined as the landmark inherent relation while it is ignored by most existing methods. In this paper, we present Dynamic Sparse Local Patch Transformer (DSLPT), a novel face alignment framework for the inherent relation learning and uncertainty estimation. Unlike most existing methods that regress facial landmarks directly from global features, the DSLPT first generates a rough representation of each landmark from a local patch cropped from the feature map and then adaptively aggregates them by a case dependent inherent relation. Finally, the DSLPT predicts the coordinate and uncertainty of each landmark by regressing their probability distribution from the output features. Moreover, we introduce a coarse-to-fine framework to incorporate with DSLPT for an improved result. In the framework, the position and size of each patch are determined by the probability distribution of the corresponding landmark predicted in the previous stage. The dynamic patches will ensure a fine-grained landmark representation for inherent relation learning so that a rough prediction result can gradually converge to the target facial landmarks. We integrate the coarse-to-fine model into an end-to-end training pipeline and carry out experiments on the mainstream benchmarks. The results demonstrate that the DSLPT achieves state-of-the-art performance with much less computational complexity. The codes and models are available at https://github.com/Jiahao-UTS/DSLPT. Jiahao Xia 0001, Min Xu 0001, Haimin Zhang 0001, Jianguo Zhang 0001, Wenjian Huang 0001, Hu Cao, Shiping Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Globally Optimal Robust Radar Calibration in Intelligent Transportation SystemsabstractRadar is among the most popular sensors in modern Intelligent Transportation Systems (ITSs), enabling weather-robust perception. The orientation and position of the traffic radar relative to the ITS coordinate system are necessary for the perception fusion in ITSs. However, due to the unknown target association, sparseness and noisiness of traffic radar measurements, the robust and accurate extrinsic calibration of traffic radar is challenging. In this paper, we propose a targetless traffic radar calibration method based on GPS to overcome the inconvenience during ITS operation, because the installation of a dedicated calibration target on the highway is impractical and dangerous. On the other hand, the high-precision GPS device installed on the moving vehicle can provide traffic radar with accurate positioning information of the detection target. Furthermore, during the optimization process of extrinsic calibration, we propose a globally optimal registration method, which is robust to noise and outliers in radar measurements, and is called Gaussian Mixture Robust Branch and Bound (GMRBnB). Specifically, we first construct the robust objective function by utilizing the Gaussian Mixture Model (GMM). Then, we derive novel relaxation bounds and present the GMRBnB algorithm that overcomes the susceptibility to local minima and the dependence on initialization of traditional optimization methods. Compared with existing methods, extensive experiments in synthetic and real-world data demonstrate that our method is not only globally optimal, but also more accurate and robust. Yinlong Liu, Venkatnarayanan Lakshminarasimhan, Hu Cao, Feihu Zhang, Alois C. Knoll |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Spatio-temporal tendency reasoning for human body pose and shape estimation from videos
Suping Wu, Hu Cao, Kehua Ma |
BMVC | 3 |
| 2022 | OneEE: A One-Stage Framework for Fast Overlapping and Nested Event ExtractionabstractEvent extraction (EE) is an essential task of information extraction, which aims to extract structured event information from unstructured text. Most prior work focuses on extracting flat events while neglecting overlapped or nested ones. A few models for overlapped and nested EE includes several successive stages to extract event triggers and arguments,which suffer from error propagation. Therefore, we design a simple yet effective tagging scheme and model to formulate EE as word-word relation recognition, called OneEE. The relations between trigger or argument words are simultaneously recognized in one stage with parallel grid tagging, thus yielding a very fast event extraction speed. The model is equipped with an adaptive event fusion module to generate event-aware representations and a distance-aware predictor to integrate relative distance information for word-word relation recognition, which are empirically demonstrated to be effective mechanisms. Experiments on 3 overlapped and nested EE benchmarks, namely FewFC, Genia11, and Genia13, show that OneEE achieves the state-of-the-art (SOTA) results. Moreover, the inference speed of OneEE is faster than those of baselines in the same condition, and can be further substantially improved since it supports parallel inference. Hu Cao, Fangfang Su, Fei Li 0021, Hao Fei 0001, Shengqiong Wu, Bobo Li 0001, Liang Zhao 0001, Donghong Ji |
COLING | 1 |
| 2022 | A Biologically-Inspired Global Localization System for Mobile Robots Using LiDAR SensorabstractLocalization in the environment is an essential navigational capability for animals and indoor robotic vehicles. In indoor environments, it is still challenging to perfectly solve the global localization problem using probabilistic methods. However, animals are able to instinctively localize themselves with much less effort. Therefore, an intriguing and promising approach is to seek biological inspiration from animals. In this paper, we present a biologically-inspired global localization system using a LiDAR sensor that utilizes a hippocampal model and a landmark-based relocalization approach. The experiment results show that the proposed method is competitive with Monte Carlo Localization, and the results demonstrate the high accuracy, applicability, and reliability of the proposed biologically-inspired localization system in various localization scenarios. Genghang Zhuang, Carlo Cagnetta, Zhenshan Bing, Hu Cao, Kai Huang 0001, Alois C. Knoll |
IV | 4 |
| 2021 | Frame-level Feature Tokenization Learning for Human Body Pose and Shape EstimationabstractIn this paper, we propose a frame-level feature tokenization method for human body pose and shape es-timation(FTHE). Despite conventional 3D human pose and shape estimation methods have achieved success based on a single image, recovering accurate and smooth 3D human motion from a video is still challenging. Different from existing methods, our FTHE aims to pay attention to the meaningful detailed temporal feature between different granular tokens of video objects, and reduce the dominance of the current static frame. To this end, we carefully design an accurate and interpretable temporal encoding module for feature extraction and motion reconstruction. More specifically, our model captures temporal features and static features of different granular tokens, and simultaneously enhances their correlation and multi-granularity consistency. Extensive experimental results on large-scale publicly available datasets demonstrate that our FTHE achieves compelling performance compared to the state-of-the-art. Code has been made available at: https://githuh.com/chriful/FG_2020_FTHE. Hu Cao, Meining Jia, Suping Wu |
FG | 1 |
| 2021 | Towards Rich-Detail 3D Face Reconstruction and Dense Alignment via Multi-Scale Detail Augmentationabstract3D face reconstruction based on a single image is a longstanding challenging problem in computer vision. Existing end-to-end methods are difficult to reconstruct rich 3D face details. To solve this problem, we propose a two-stream convolutional neural network combined with a face super-resolution method, which can effectively restore the image’s 3D position information. Our method combines an attention fusion mechanism, which can learn the individual attention mapping of each feature subspace, and effectively learn cross-channel information while learning multi-scale and multi-frequency features. Meanwhile, our module obtains the most discriminative features in different local areas, and enhances the consistency and correlation between the attention areas. Experimental results show that our SRCNet has made significant improvements in the 3D face reconstruction and face alignment of the AFLW2000-3D and AFLW datasets. Suping Wu, Lei Li 0044, Kui Lin, Xing Zheng, Hu Cao |
ICME | 6 |
| 2021 | Residual Squeeze-and-Excitation Network with Multi-scale Spatial Pyramid Module for Fast Robotic Grasping DetectionabstractThis paper proposes an efficient, fully convolutional neural network to generate robotic grasps by using 300×300 depth images as input. Specifically, a residual squeeze-and-excitation network (RSEN) is introduced for deep feature extraction. Following the RSEN block, a multi-scale spatial pyramid module (MSSPM) is developed to obtain multi-scale contextual information. The outputs of each RSEN block and MSSPM are combined as inputs for hierarchical feature fusion. Then, the fused global features are upsampled to perform pixel-wise learning for grasping pose estimation. The experimental results on Cornell and Jacquard grasping datasets indicate that the proposed method has a fast inference speed of 5ms while achieving high grasp detection accuracy of 96.4% and 94.8% on Cornell and Jacquard, respectively, which strikes a balance between accuracy and running speed. Our method also gets a 90% physical grasp success rate with a UR5 robot arm. Hu Cao, Guang Chen 0001, Zhijun Li 0001, Jianjie Lin, Alois C. Knoll |
ICRA | 1 |
| 2016 | Non-sequential protein structure alignment based on variable length AFPs using the maximal cliqueabstractProtein structure alignment plays an important role in the study of bioinformatics. Many protein structure alignment methods have been proposed. However, most of them are sequential alignment methods which are not able to detect non-sequential similarities between proteins. Although some non-sequential protein structure alignment methods have been proposed recently, their results are still not satisfactory. In this work, a new non-sequential protein structure alignment method based on Aligned Fragment Pairs (AFPs) and the maximal clique is proposed. Different from other methods, our method is based on variable length AFPs which can better represent the local structure similarities and can also greatly speed up the computations. Moreover, the spatial information of the AFPs is used to select “good” AFPs. Then, a graph is built to represent the relationship between all the “good” AFPs. If two AFPs can be aligned at the same time, an edge is added to the graph. A high quality maximal clique of the graph is found to produce the initial alignment. Finally, a greedy-like refinement algorithm is executed to get the final alignment. The experiments show that, compared with seven non-sequential methods, except DEDAL which is a non-rigid body method, and MICAN which uses secondary structure information, the proposed method usually produces more correctly aligned residue pairs than the other non-sequential alignment methods. Xingmei Liu, Yonggang Lu, Hu Cao |
BIBM | 3 |
| 2015 | Flexible protein structure alignment by variable-length Aligned Fragment PairsabstractWith the growing number of known protein 3D structures, how to efficiently compare protein structures becomes an important and challenging problem in computational structural biology. So, many protein structure alignment methods have been developed in recent years. Flexible structure alignment methods are shown to be superior to rigid structure alignment methods in identifying the structure similarities between proteins which have gone through conformational changes. It is also found that the methods based on Aligned Fragment Pairs (AFP) own special advantages in balancing global structure similarities and local structure similarities. In this work, we propose a new flexible protein 3D structure alignment method based on variable-length AFPs. Different from other methods, our method owns three special features: firstly, it uses a new AFP identification algorithm based on a local coordinate system; secondly, it allows different AFPs separated by other AFPs to share the same transformation during the concatenation; thirdly, it allows different AFPs to have different lengths, which can not only reduce the total number of AFPs, and also can improve the representation of the local structure similarities. The experiments show that, compared to three other flexible structure alignment methods, FlexProt, FATCAT and FlexSnap, the proposed method can achieve similar or better results by introducing fewer twists and in much less running time due to the reduced number of AFPs. Hu Cao, Yonggang Lu |
BIBM | 1 |
| 2012 | Tripzoom: an app to improve your mobility behaviorabstractMobile devices can help to solve urban traffic problems by improving personal mobility and making transport, traveling, and commuting for individual users more flexible, sustainable, and rewarding. For that purpose, the tripzoom application combines mobility data and patterns from mobile sensing, a dynamic incentive system, and community feedback from social networks. This paper gives an overview of tripzoom, its features, and technical realization, and explains how users can take advantage of it to monitor, manage, and improve their mobility behavior. Gregor Broll, Hu Cao, Peter Ebben, Paul Holleis, Koen Jacobs, Johan Koolwaaij, Marko Luther, Bertrand Souville |
MUM | 2 |
| 2006 | Search-and-Discover in Mobile P2P Network DatabasesabstractIn this paper we propose a novel algorithm called Rank-Based Broadcast (RBB) for discovery of local resources in mobile P2P networks. With RBB, each moving object periodically broadcasts the most relevant resource reports and queries it knows to its neighbors, and the contribution is in determining how to rank the reports and queries in terms of their relevance, when to broadcast them, and how many to broadcast. A major difference between RBB and many existing algorithms in the resource discovery and publish/subscribe literature is that RBB does not rely on any pre-established routing structure, and therefore is able to adapt to both high mobility environments. In the paper we experimentally compare RBB with flooding and PSTree, a publish/subscribe algorithm for wireless ad-hoc networks. The results show that RBB by far outperforms the other two algorithms. Ouri Wolfson, Bo Xu 0001, Huabei Yin, Hu Cao |
ICDCS | 4 |
| 2006 | Searching Local Information in Mobile DatabasesabstractA mobile ad-hoc network (MANET) is a set of moving objects that communicate with each other via unregulated, short-range wireless technologies such as IEEE 802.11, Bluetooth, or Ultra Wide Band (UWB). No fixed infrastructure is assumed or relied upon. An important application domain of MANET’s is local resource discovery. In a local resource discovery application, a user finds local resources that satisfy specified criteria. For example, a driver finds an available parking slot in a region by receiving information generated by the parking meter, or gets the traffic conditions on a highway segment a mile ahead; a cab driver finds a near-by customer, or a participant at a convention finds another participant with a matching profile. Ouri Wolfson, Bo Xu 0001, Huabei Yin, Hu Cao |
ICDE | 4 |
| 2006 | Spatio-temporal data reduction with deterministic error bounds
Hu Cao, Ouri Wolfson, Goce Trajcevski |
VLDB J. | 1 |
| 2005 | Nonmaterialized Motion Information in Transport Networks
Hu Cao, Ouri Wolfson |
ICDT | 1 |
| 2002 | Management of Dynamic Location Information in DOMINO
Ouri Wolfson, Hu Cao, Goce Trajcevski, Fengli Zhang, Naphtali Rishe |
EDBT | 2 |