VLDB 2026 Research / reviewers in the wild / expert
Jianming Hu
dblp:69/4427
· DBLP profile ↗
59ranked-venue papers
13as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 8 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Style-Based Profiling Framework for Quantifying the Synthetic-to-Real Gap in Autonomous Driving DatasetsabstractEnsuring the reliability of autonomous driving perception systems requires extensive environment-based testing, yet real-world execution is often impractical. Synthetic datasets have therefore emerged as a promising alternative, offering advantages such as cost-effectiveness, bias free labeling, and controllable scenarios. However, the domain gap between synthetic and real-world datasets remains a major obstacle to model generalization. To address this challenge from a data-centric perspective, this paper introduces a profile extraction and discovery framework for characterizing the style profiles underlying both synthetic and real image datasets. We propose Style Embedding Distribution Discrepancy (SEDD) as a novel evaluation metric. Our framework combines Gram matrix-based style extraction with metric learning optimized for intra-class compactness and inter-class separation to extract style embeddings. Furthermore, we establish a benchmark using publicly available datasets. Experiments are conducted on a variety of datasets and sim-to-real methods, and the results show that our method is capable of quantifying the synthetic-to-real gap. This work provides a standardized profiling-based quality control paradigm that enables systematic diagnosis and targeted enhancement of synthetic datasets, advancing future development of data-driven autonomous driving systems. Dingyi Yao, Xinyao Han, Ruibo Ming, Zhihang Song, Lihui Peng, Jianming Hu, Danya Yao, Yi Zhang 0029 |
IV | 6 |
| 2026 | A novel graph learning framework for interpretable and imbalance financial fraud detection
Junhao Lu, Qiupeng Xu, Jianming Hu |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | PFL-TEnet: Personalized federated learning for wind power forecasting via time-series embedding and hypernetwork modeling
Yi Li 0072, Jianming Hu |
Expert Syst. Appl. | 4 |
| 2026 | Few-shot medical image segmentation via dual-stream feature extractor and detail-enhanced prototype transformer
Wenfeng Zhang, Jianming Hu, Qibing Qin |
Knowl. Based Syst. | 4 |
| 2026 | Enhanced radiology report generation via comprehensive sequence rearrangement and multi-scale cross-region attention
Qibing Qin, Jianming Hu, Dengwei Yan, Wenfeng Zhang, Jing Qiao |
Vis. Comput. | 4 |
| 2025 | Are Expressive Models Truly Necessary for Offline RL?abstractAmong various branches of offline reinforcement learning (RL) methods, goal-conditioned supervised learning (GCSL) has gained increasing popularity as it formulates the offline RL problem as a sequential modeling task, therefore bypassing the notoriously difficult credit assignment challenge of value learning in conventional RL paradigm. Sequential modeling, however, requires capturing accurate dynamics across long horizons in trajectory data to ensure reasonable policy performance. To meet this requirement, leveraging large, expressive models has become a popular choice in recent literature, which, however, comes at the cost of significantly increased computation and inference latency. Contradictory yet promising, we reveal that lightweight models as simple as shallow 2-layer MLPs, can also enjoy accurate dynamics consistency and significantly reduced sequential modeling errors against large expressive models by adopting a simple recursive planning scheme: recursively planning coarse-grained future sub-goals based on current and target information, and then executes the action with a goal-conditioned policy learned from data relabeled with these sub-goal ground truths. We term our method as Recursive Skip-Step Planning (RSP). Simple yet effective, RSP enjoys great efficiency improvements thanks to its lightweight structure, and substantially outperforms existing methods, reaching new SOTA performances on the D4RL benchmark, especially in multi-stage long-horizon tasks. Li Jiang 0008, Jianming Hu, Xianyuan Zhan |
AAAI | 5 |
| 2025 | Efficient Robotic Policy Learning via Latent Space Backward PlanningabstractCurrent robotic planning methods often rely on predicting multi-frame images with full pixel details. While this fine-grained approach can serve as a generic world model, it introduces two significant challenges for downstream policy learning: substantial computational costs that hinder real-time deployment, and accumulated inaccuracies that can mislead action extraction. Planning with coarse-grained subgoals partially alleviates efficiency issues. However, their forward planning schemes can still result in off-task predictions due to accumulation errors, leading to misalignment with long-term goals. This raises a critical question: Can robotic planning be both efficient and accurate enough for real-time control in long-horizon, multi-stage tasks?
To address this, we propose a **B**ackward **P**lanning scheme in **L**atent space (**LBP**), which begins by grounding the task into final latent goals, followed by recursively predicting intermediate subgoals closer to the current state. The grounded final goal enables backward subgoal planning to always remain aware of task completion, facilitating on-task prediction along the entire planning horizon. The subgoal-conditioned policy incorporates a learnable token to summarize the subgoal sequences and determines how each subgoal guides action extraction.
Through extensive simulation and real-robot long-horizon experiments, we show that LBP outperforms existing fine-grained and forward planning methods, achieving SOTA performance. Project Page: [https://lbp-authors.github.io](https://lbp-authors.github.io). Dongxiu Liu, Jinliang Zheng, Yinan Zheng, Zhonghong Ou, Jianming Hu, Xianyuan Zhan |
ICML | 7 |
| 2025 | H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics GapsabstractSolving real-world complex tasks using reinforcement learning (RL) without high-fidelity simulation environments or large amounts of offline data can be quite challenging. Online RL agents trained in imperfect simulation environments can suffer from severe sim-to-real issues. Offline RL approaches although bypass the need for simulators, often pose demanding requirements on the size and quality of the offline datasets. The recently emerged hybrid offline-and-online RL provides an attractive framework that enables joint use of limited offline data and imperfect simulator for transferable policy learning. In this paper, we develop a new algorithm, called$\mathrm{H} 2 \mathrm{O}+$, which offers great flexibility to bridge various choices of offline and online learning methods, while also accounting for dynamics gaps between the real and simulation environments. Through extensive simulation and real-world robotics experiments, we demonstrate superior performance and flexibility of$\mathbf{H 2 O}+$over advanced cross-domain online and offline RL algorithms. Tianying Ji, Bingqi Liu, Haocheng Zhao, Jianying Zheng, Guyue Zhou, Jianming Hu, Xianyuan Zhan |
ICRA | 9 |
| 2025 | Progressive class-aware instance enhancement for aircraft detection in remote sensing imagery
Tianjun Shi, Jinnan Gong, Jianming Hu, Yu Sun 0028, Guangzhen Bao, Pengfei Zhang 0011, Xiyang Zhi, Wei Zhang 0220 |
Pattern Recognit. | 3 |
| 2025 | StyleFormer: Spatial-Temporal Style Projecting Bidirectional Interactive Transformer for Change DetectionabstractRemote sensing image change detection is an important means for Earth monitoring task, which has a wide application prospect. In multitemporal optical remote sensing, there are inherent differences in factors such as lighting and sensors. This leads to the coupling of content change and image style change, making it difficult to distinguish. Therefore, a meaningful thinking for change detection is to decouple and capture the real changes of ground objects from multitemporal images. Based on this motivation, a novel general change detection architecture is explored, StyleFormer. It first proposes the concept of spatial–temporal style base and no longer constrains to semantic representation in a single image style. Instead, it introduces a spatial–temporal interactive style projection layer between bitemporal images, which projects the unseen diverse styles into the consistent expression space for change detection. Furthermore, an iterative interaction strategy of Transformer and CNN features is proposed to mine spatial–temporal context information more finely. It solves the lack of local perception and nonhierarchical features in ViT, and improves the model expression ability. After that, a change prior-guided cross-attention is introduced to fuse bitemporal features. It can adaptively enhance the change feature and improve the perception ability for small changes in remote sensing scenes. Sufficient experiments on four typical change detection datasets show that the proposed method is superior to the state-of-the-art methods. Especially on the datasets CDD-CD and SYSU-CD, the F1 score improved to 96.08% and 83.29%. The code of this work will be available athttps://github.com/Tom-Dongfang/change-detection-StyleFormer. Qichao Han, Xiyang Zhi, Jianming Hu, Shuqing Zhang, Wenbin Chen 0007, Yuanxin Huang, Shikai Jiang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Complementarity-Aware Feature Fusion for Aircraft Detection via Unpaired Opt2SAR Image TranslationabstractDetecting aircraft in complex remote sensing scene has significant value for military and civilian applications. To overcome the interference of complex environmental factors and achieve high-accuracy detection performance, the comprehensive utilization of optical and SAR images for object detection has become a promising research direction. However, currently there are problems with optical and SAR fusion detection, such as difficulty in obtaining paired registration training data and incomplete consideration of feature elements in fusion model. To tackle these challenges, we present an aircraft detection method based on optical-SAR complementarity-aware feature fusion. Firstly, an unpaired image translation model based on scattering feature enhancement GAN (SFEG) is designed to generate SAR images that are pixel-level registered with the input optical image. On this basis, a complementarity-aware feature fusion detection network (CFFDNet) combining differential feature spatial-aware complementary (DFSC) units and gate-generated weighted fusion (GWF) units is proposed to enhance the effective features of single source image while improving the complementary fusion effect of multimodal features. Experiments on CORS-ADD and MAR20 datasets demonstrate that our method outperforms the compared classical single-modal and multimodal detection models. The latest code is available soon at: https://github.com/JimmyRSlab/Complementarity-aware-Feature-Fusion-for-Aircraft-Detection-via-Unpaired-Opt2SAR-Image-Translation. Jianming Hu, Xiyang Zhi, Tianjun Shi, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Dataset and Benchmark for Fine-Grained Ship Recognition in Complex Optical Remote Sensing ScenariosabstractShip recognition in remote sensing imagery is crucial for numerous applications such as monitoring maritime security, preventing illegal activities and implementing environmental protection. However, most existing research rarely focuses on both the complexity of the scene and the fine-grained recognition of targets, which limits the accuracy and applicability of ship recognition technology. To propel advancements in ship recognition methodology, we propose a new dataset named Fine-Grained Ship Recognition in Complex Scenarios (FGSRCS), which contains 17 typical categories of large and medium-sized ships from 280 widely distributed ports. For the richness of the image characteristics, the dataset images are collected from multiple satellite platforms, such as Ikonos, Jilin-1, OrbView, Pleiades and WorldView, and the time span of the data covers the past 20 years. To ensure the diversity of scenarios, the dataset mainly considers five complex scenarios, including thick clouds, mists, shadows, sea clutter and ground facilities, which helps to train and improve the algorithm applicability to practical application scenarios. Furthermore, we conduct experiments with ten state-of-the-art recognition algorithms on FGSRCS dataset, providing a benchmark for algorithm application. The research result can furnish both theoretical insights and practical guidance for the development of future ship recognition models. The dataset is available on https://github.com/dwddw/FGSRCS. Jianming Hu, Xiyang Zhi, Tianjun Shi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Few-Shot Testing of Autonomous Vehicles With Scenario Similarity LearningabstractTesting and evaluation are critical to the development and deployment of autonomous vehicles (AVs). Given the rarity of safety-critical events such as crashes, millions of tests are typically needed to accurately assess AV safety performance. Although techniques like importance sampling can accelerate this process, it usually still requires too many tests for field testing. This severely hinders the testing and evaluation process, especially for third-party testers and governmental bodies with very limited testing budgets. The rapid development cycles of AV technology further exacerbate this challenge. To fill this research gap, this paper introduces the few-shot testing (FST) problem and proposes a methodological framework to tackle it. As the testing budget is very limited, usually smaller than 100, the FST method transforms the testing scenario generation problem from probabilistic sampling to deterministic optimization, reducing the uncertainty of testing results. To optimize the selection of testing scenarios, a cross-attention similarity mechanism is proposed to extract the information of AV’s testing scenario space. This allows iterative searches for scenarios with the smallest evaluation error, ensuring precise testing within budget constraints. Experimental results in cut-in scenarios demonstrate the effectiveness of the FST method, significantly enhancing accuracy and enabling efficient, precise AV testing. Honglin He, Jianming Hu, Yi Zhang 0029, Shuo Feng 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Adaptive Testing Environment Generation for Connected and Automated Vehicles With Dense Reinforcement LearningabstractThe assessment of safety performance plays a pivotal role in the development and deployment of connected and automated vehicles (CAVs). A common approach involves designing testing scenarios based on prior knowledge of CAVs (e.g., surrogate models), conducting tests in these scenarios, and subsequently evaluating CAVs’ safety performances. However, substantial differences between CAVs and the prior knowledge can significantly diminish the evaluation efficiency. In response to this issue, existing studies predominantly concentrate on the adaptive design of testing scenarios during the CAV testing process. Yet, these methods have limitations in their applicability to high-dimensional scenarios. To overcome this challenge, we develop an adaptive testing environment that bolsters evaluation robustness by incorporating multiple surrogate models and optimizing the combination coefficients of these surrogate models to enhance evaluation efficiency. We formulate the optimization problem as a regression task utilizing quadratic programming. To efficiently obtain the regression target via reinforcement learning, we propose the dense reinforcement learning method and devise a new adaptive policy with high sample efficiency. Essentially, our approach centers on learning the values of critical scenes displaying substantial surrogate-to-real gaps. The effectiveness of our method is validated in high-dimensional overtaking scenarios, demonstrating that our approach achieves notable evaluation efficiency. Ruoxuan Bai, Haoyuan Ji, Yi Zhang 0029, Jianming Hu, Shuo Feng 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Distributed Input Mapping Control for Multiple Mixed Platoons With a Flexible Structure ModelabstractResearch on the longitudinal control of mixed platoons has garnered significant attention, particularly addressing the uncertainties of human-driven vehicles (HDVs) and the computational inefficiency of centralized control for large platoons. This paper presents a novel distributed control strategy for multiple mixed platoons, utilizing a flexible model structure to account for dynamic uncertainties and lateral disturbances. A distributed control framework is introduced, where each platoon’s states and constraints are coupled with neighboring platoons, improving scalability and resilience in large-scale mixed traffic environments. Each platoon’s dynamic uncertainties are characterized by convex hulls in the subsystem model. A distributed data-driven input mapping method is proposed, where subsystems dynamically map the process data linearly from both local and neighboring platoons to local control law, eliminating the need for pre-collected datasets or offline training. To reduce iterations and communication requirements, a distributed implementation algorithm is proposed by designing subsystem basic laws with robust model predict control, then the control problem is solved online using convex quadratic programming, enabling fast computation of control actions and enhancing adaptability to real-time traffic variations. Additionally, a theoretical analysis guarantees the stability and optimality of the system, providing conditions that assist in subsystem parameter design for practical implementation. Simulation results validate the effectiveness and efficiency of the proposed approach, particularly in comparison to traditional methods that rely on pre-collected data. Shaoying He, Jianming Hu, Yunwen Xu, Dewei Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | An Online Calibration Method for Robust Multi-Modality 3D Object DetectionabstractMulti-modality sensor fusion for 3D perception is a significant part for autonomous driving perception, which enables a comprehensive integration of different sensors and obtains a holistic understanding of the surrounding environment to improve accuracy, stability and reliability. Nevertheless, vibrations, collisions, and acceleration/deceleration in motion may result in a minor disturbance to the position of sensors, leading to offsets in sensor calibration. To address the limitations of the original method demanding considerable manual effort and time for hand-annotating checkerboards, we introduce an online Camera-LiDAR calibration method, which can perform calculations during the operation of autonomous vehicles to eliminate sensor biases. Different from the previous CNN-based online calibration algorithm, our approach differs in utilizing multi-scale features to accomplish alignment with a foundational ResNet and FPN Backbone alongside an attention-based multi-scale feature fusion module. Moreover, in the design of the loss function, we incorporate depth loss and point cloud loss in addition to the original smooth L1 norm loss to facilitate network backpropagation. Ultimately, our proposed method achieves an average translation error of 0.88 cm and a rotation error of 0.073° on the KITTI odometry dataset, which can significantly limit the discrepancies between sensors and thus enhance detection accuracy. Yige Yao, Jianming Hu, Zhidong Deng |
DSAA | 3 |
| 2024 | Continual Driving Policy Optimization with Closed-Loop Individualized CurriculaabstractThe safety of autonomous vehicles (AV) has been a long-standing top concern, stemming from the absence of rare and safety-critical scenarios in the long-tail naturalistic driving distribution. To tackle this challenge, a surge of research in scenario-based autonomous driving has emerged, with a focus on generating high-risk driving scenarios and applying them to conduct safety-critical testing of AV models. However, limited work has been explored on the reuse of these extensive scenarios to iteratively improve AV models. Moreover, it remains intractable and challenging to filter through gigantic scenario libraries collected from other AV models with distinct behaviors, attempting to extract transferable information for current AV improvement. Therefore, we develop a continual driving policy optimization framework featuring Closed-Loop Individualized Curricula (CLIC), which we factorize into a set of standardized sub-modules for flexible implementation choices: AV Evaluation, Scenario Selection, and AV Training. CLIC frames AV Evaluation as a collision prediction task, where it estimates the chance of AV failures in these scenarios at each iteration. Subsequently, by re-sampling from historical scenarios based on these failure probabilities, CLIC tailors individualized curricula for downstream training, aligning them with the evaluated capability of AV. Accordingly, CLIC not only maximizes the utilization of the vast pre-collected scenario library for closed-loop driving policy optimization but also facilitates AV improvement by individualizing its training with more challenging cases out of those poorly organized scenarios. Experimental results clearly indicate that CLIC surpasses other curriculum-based training strategies, showing substantial improvement in managing risky scenarios, while still maintaining proficiency in handling simpler cases. Xingjian Jiang, Jianming Hu |
ICRA | 4 |
| 2024 | A Comprehensive Survey of Cross-Domain Policy Transfer for Embodied Agents
Jianming Hu, Guyue Zhou, Xianyuan Zhan |
IJCAI | 2 |
| 2024 | Few-Shot Scenario Testing for Autonomous Vehicles Based on Neighborhood Coverage and SimilarityabstractTesting and evaluating the safety performance of autonomous vehicles (AVs) is essential before the large-scale deployment. Practically, the number of testing scenarios permissible for a specific AV is severely limited by tight constraints on testing budgets and time. With the restrictions imposed by strictly restricted numbers of tests, existing testing methods often lead to significant uncertainty or difficulty to quantifying evaluation results. In this paper, we formulate this problem for the first time the "few-shot testing" (FST) problem and propose a systematic framework to address this challenge. To alleviate the considerable uncertainty inherent in a small testing scenario set, we frame the FST problem as an optimization problem and search for the testing scenario set based on neighborhood coverage and similarity. Specifically, under the guidance of better generalization ability of the testing scenario set on AVs, we dynamically adjust this set and the contribution of each testing scenario to the evaluation result based on coverage, leveraging the prior information of surrogate models (SMs). With certain hypotheses on SMs, a theoretical upper bound of evaluation error is established to verify the sufficiency of evaluation accuracy within the given limited number of tests. The experiment results on cut-in scenarios demonstrate a notable reduction in evaluation error and variance of our method compared to conventional testing methods, especially for situations with a strict limit on the number of scenarios. Honglin He, Yi Zhang 0029, Jianming Hu, Shuo Feng 0002 |
IV | 5 |
| 2024 | A novel time series probabilistic prediction approach based on the monotone quantile regression neural network
Jianming Hu, Jingwei Tang, Zhi Liu 0005 |
Inf. Sci. | 1 |
| 2024 | A Method for Detecting Aircraft Small Targets in Remote Sensing Images by Using CNNs Fused With Handcrafted FeaturesabstractAircraft target detection is a challenging task in remote sensing images, especially for aircraft small target detection. The most advanced object detection framework currently processes all information in the image uniformly through a deep neural network. In the past, in the process of detecting aircraft small targets, the feature extraction process was carefully designed, and hand-crafted features were derived from expert knowledge or historical data, which included prior knowledge that was conducive to object detection. Embedding prior features into deep neural networks can enhance the saliency of target information, improve the detection performance of the model. Accordingly, this paper proposes a Hand-crafted Feature Fusion Stream (HFFS) for embedding prior knowledge. We obtain hand-crafted features based on the grayscale co-occurrence matrix and edge extraction operator, and generate an attention map in deep convolutional neural networks (CNNs) to achieve the fusion of hand-crafted feature maps and high-level feature maps in deep convolutional networks. The experimental results show that using HFFS on the baseline model improves the detection performance of the model for aircraft small targets. Compared with the baseline model, our detection model achieves improvements of 1.1% AR, 1.6% [email protected], and 1.6% [email protected]:0.95 in the proposed dataset. Lijian Yu, Xiyang Zhi, Shuqing Zhang, Shikai Jiang, Jianming Hu, Wei Zhang 0220, Yuanxin Huang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Visual-Textual Cross-Modal Interaction Network for Radiology Report GenerationabstractThe radiology report generation task generates diagnostic descriptions from radiology images, aiming to alleviate the onerous task for radiologists and alerting them to abnormalities. However, the data bias problem poses a persistent challenge, since the abnormal regions usually occupy a small portion of radiology image, while the report generation process should pay greater attention to the abnormal regions. Moreover, the data volume is relatively small compared to large language models, posing challenges during training. To address these issues effectively, we propose a Visual-textual Cross-model Interaction Network (VCIN) to enhance the quality of generated reports. VCIN comprises two key modules: Abundant Clinical Information Embedding (ACIE), which gathers rich cross-modal interaction information to promote the report generation of abnormal regions; and a Bert-based Decoder-only Generator (BDG), built on Bert architecture to mitigate training difficulties. The superior performance of our proposed model is demonstrated through experimental results obtained from two public benchmark datasets. The code is available athttps://github.com/QinLab-WFU/VCIN. Wenfeng Zhang, Baoning Cai, Jianming Hu, Qibing Qin, Kezhen Xie |
IEEE Signal Process. Lett. | 3 |
| 2024 | Self-Supervised Denoising via Blind Feature Extraction and Diffusion-Based Texture GenerationabstractIn the field of remote sensing, detection in dimly lit or shadowed areas has traditionally been difficult because of detector noise. Given that noise in real-world images of remote sensing exhibits spatial correlation, existing self-supervised methods encounter difficulties in reconciling the suppression of spatially correlated noise with the preservation of local texture details. To address this challenge, we propose a self-supervised model that combines blind-spot feature extraction with diffusion-based texture generation to fine denoising of real-world images under adverse conditions. We first introduce a blind-spot feature extraction structure based on the fusion of U-Net with blind-spot net (UBSN) and blind transformer (BTF). In UBSN, we integrate multistride blind-spot convolution (BSC) + dilated convolution (DC) feature extraction nodes and employ a Reshuffle strategy in skip-layer connections to maintain large-scale blind-spot characteristics. Additionally, we design a transformer structure for blind spot between patches to remove the noise with spatial correlations while ensuring global feature acquisition. Subsequently, to restore texture details blurred by the blind-spot structures, we introduce a texture generation diffusion structure during model training, achieving a balance between large-scale blind-spot characteristics and local rich texture details. Experimental results demonstrate that our approach outperforms other self-supervised denoising methods, even some methods leveraging unpaired images, without the need for parameters related to on-orbit satellite detectors. Guangzhen Bao, Xiyang Zhi, Pengfei Zhang 0011, Jianming Hu, Tianjun Shi, Shikai Jiang, Yayun Wu, Jinnan Gong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Dataset and Benchmark for Ship Detection in Complex Optical Remote Sensing ImageabstractShip detection plays a pivotal role in numerous military and civil applications, yet detecting ships in complex maritime and aerial environments remains a challenging task. While several publicly available datasets for ship detection have been introduced by researchers, most of them do not adequately address the impacts of diverse and intricate environmental factors, which makes the trained algorithms difficult to apply for practical application scenes involving clouds, sea clutter, complex lighting, and facility interferences, limiting the effectiveness and robustness of the detection models. To advance the field of ship detection method research, we propose a high-quality dataset named ship collection in complex optical scene (SCCOS), which is obtained from multiple platform sources including Google Earth, Microsoft map, Worldview-3, Pleiades, Orbview-3, Jilin-1, and Ikonos satellites. The dataset comprehensively considers complex scenes such as thin clouds, mist, thick clouds, light shadows, sea clutter, and port facilities. Additionally, we conduct experiments on this dataset with 11 representative detection algorithms and establish a performance benchmark, which can provide the theoretical basis and practical reference for the design and optimization of subsequent ship detection models. The latest dataset is available at:https://github.com/JimmyRSlab/Dataset-and-Benchmark-for-Ship-Detection-in-Complex-Optical-Remote-Sensing-Image. Jianming Hu, Xiyang Zhi, Tianjun Shi, Xiaogang Sun |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | FDDBA-NET: Frequency Domain Decoupling Bidirectional Interactive Attention Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) involves determining the coordinate position of the target in complex infrared images. However, challenges arise due to the absence of internal texture structure, edge dispersion, weak energy characteristics of the target, and a significant amount of background clutter resembling the target’s morphology, impeding precise target location. To address these challenges, we propose an IRSTD network, named frequency domain decoupling bidirectional interactive attention network (FDDBA-NET), designed from the perspective of frequency domain decoupling (FDD). To suppress backgrounds that are similar in shape and structure to the target, exploiting the spectral differences between the target and background in the frequency domain, we adopt two learnable masks to extract the target-specific spectrum and the target-background-consistent spectrum detrimental to detection. The specific spectrum aids in target detection. A target-level contrast loss is designed to maximize the disparity between these two spectra, ensuring optimal detection results. In addition, to preserve target details in high-level semantic information, we introduce a bidirectional interactive attention module that leverages mutual modulation of deep global and shallow local features, facilitating deep and shallow feature fusion. To validate our approach, we conduct experiments comparing our proposed network with state-of-the-art conventional methods and deep learning methods on public datasets. The results demonstrate the superior performance of our method. Yuanxin Huang, Xiyang Zhi, Jianming Hu, Lijian Yu, Qichao Han, Wenbin Chen 0007, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | LMAFormer: Local Motion Aware Transformer for Small Moving Infrared Target DetectionabstractIn temporal infrared small target detection, it is crucial to leverage the disparities in spatiotemporal characteristics between the target and the background to distinguish the former. However, remote imaging and the relative motion between the detection platform and the background cause significant coupling of spatiotemporal characteristics, making target detection highly challenging. To address these challenges, we propose a network named LMAFormer. First, we introduce a local motion-aware spatiotemporal attention mechanism that aligns and enhances multiframe features to extract local spatiotemporal salient features of targets while avoiding interference from moving backgrounds. Second, we employ a multiscale fusion transformer encoder that computes self-attention weights across and within scales during encoding, to establish multiscale correlations among different regions of temporal images, enabling motion background modeling. Last, we propose a multiframe joint query decoder. The shallowest feature map after multiscale feature propagation is mapped to initial query weights, which are refined through grouped convolutions to generate grouped query vectors. These are jointly optimized to encapsulate rich multiframe details, strengthening motion background modeling and target feature representation, improving prediction accuracy. Experimental results on the NUDT-MIRSDT, IRDST, and the established TSIRMT datasets demonstrate that our network outperforms state-of-the-art (SOTA) methods. Our code and dataset will be available athttps://github.com/lifier/LMAFormer. Yuanxin Huang, Xiyang Zhi, Jianming Hu, Lijian Yu, Qichao Han, Wenbin Chen 0007, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multiscale Progressive Fusion Filter Network for Infrared Small Target DetectionabstractInfrared small target detection is widely used in remote sensing fields. However, the application scenes of space-based remote sensing imaging often lead to problems such as small target scale, weak energy, and serious influence by strong clutters. At present, traditional methods are often difficult to adapt to the change of target scale. And deep learning methods are often difficult to extract small target features, and the change in imaging characteristics also brings challenges to the generalization ability. To complement each other’s advantages, we propose an infrared small target detection method that combines the traditional methods with the deep learning methods. First, we construct a multi-stage feature extraction network for guiding the typical multi-scale traditional filtering results to progressively fuse. Secondly, we propose a multi-scale attention supervision module to adjust the semantic consistency of different stages, improving the network generalization ability. Next, a dynamic weight convolution module is utilized to obtain the optimal distribution of grayscale in the neighborhood. Finally, we use the background modelling results to suppress the background, effectively weakening the influence of background clutters and enhancing the target contrast. Experimental results show that the proposed method has good detection results for targets with different scales and signal-to-clutter ratios in a variety of complex scenes. Compared with the typical methods, our method has better detection performance and generalization ability. Pengfei Zhang 0011, Zhile Wang, Guangzhen Bao, Jianming Hu, Tianjun Shi, Guanjie Sun, Jinnan Gong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | (Re)2H2O: Autonomous Driving Scenario Generation via Reversely Regularized Hybrid Offline-and-Online Reinforcement LearningabstractAutonomous driving and its widespread adoption have long held tremendous promise. Nevertheless, without a trustworthy and thorough testing procedure, not only does the industry struggle to mass-produce autonomous vehicles (AV), but neither the general public nor policymakers are convinced to accept the innovations. Generating safety-critical scenarios that present significant challenges to AV is an essential first step in testing. Real-world datasets include naturalistic but overly safe driving behaviors, whereas simulation would allow for unrestricted exploration of diverse and aggressive traffic scenarios. Conversely, higher-dimensional searching space in simulation disables efficient scenario generation without real-world data distribution as implicit constraints. In order to marry the benefits of both, it seems appealing to learn to generate scenarios from both offline real-world and online simulation data simultaneously. Therefore, we tailor a Reversely Regularized Hybrid Offline-and-Online ((Re)2H2O) Reinforcement Learning recipe to additionally penalize Q-values on real-world data and reward Q-values on simulated data, which ensures the generated scenarios are both varied and adversarial. Through extensive experiments, our solution proves to produce more risky scenarios than competitive baselines and it can generalize to work with various autonomous driving models. In addition, these generated scenarios are also corroborated to be capable of fine-tuning AV performance. Ziyuan Yang 0004, Yichen Lin, Yi Zhang 0029, Jianming Hu |
IV | 7 |
| 2023 | Local Adaptive Prior-Based Image Restoration Method for Space Diffraction Imaging SystemsabstractThin-film diffractive optical elements (DOEs) have considerable potential to be used in the field of high-resolution remote sensing imaging satellites because of advantages such as a large aperture, small volume, lightness, wide tolerance range of surface shape, and easy replication. However, there are problems associated with thin-film diffraction imaging, including space variation, serious blur, and low contrast, which result in insufficient imaging quality with regard to traditional optical system requirements. To address this, a local adaptive prior-based image restoration method is proposed for thin-film diffraction imaging systems. An entire degraded image was divided into several isohalo regions based on imaging characteristics. Then, the regularization constraints were adaptively selected and updated according to the local scene prior characteristics. Additionally, the system parameters in the corresponding field of view were used as input to restore each subregion. In particular, the diffraction efficiency (DIE) was introduced into the model to remove the nondesign level background radiation. The experimental results show that the proposed algorithm can effectively improve the image quality of a thin-film diffraction imaging system, including space variation correction, clarity enhancement, and background radiation suppression. Furthermore, a DIE of less than 60% was found to significantly impact the final image products. Shikai Jiang, Jianming Hu, Xiyang Zhi, Wei Zhang 0220, Dawei Wang 0007, Xiaogang Sun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Adaptive Feature Fusion With Attention-Guided Small Target Detection in Remote Sensing ImagesabstractSmall target detection in remote sensing images has considerable significance in practical applications such as military dynamic discrimination and traffic monitoring. However, the limited appearance features of small-scale targets and the widespread false alarm sources make small target detection in remote sensing images a tough challenge. To address these problems, we propose a novel small detection method by employing an adaptive multi-level feature fusion module (AMFFM) and an attention-augmented high-resolution head (AAHRH). Specifically, AMFFM is designed to suppress the interference of false alarm sources in complicated scenes. We upsample the high-level features by the context modeling of semantic information and refine the low-level features for noise removal. Then the enhanced multi-level features are fused based on the spatial and channel significance. After that, AAHRH is put forward to enhance the perception of small targets by embedding cross-dimension interaction with the attention mechanism. The prediction heads are reconstructed with high-resolution layers to improve the detection performance in densely distributed scenes. We conduct dilated and comparison experiments on a constructed small car dataset, a public small ship dataset, and the VEDAI dataset. The experimental results on two datasets verify the effectiveness and robustness of the proposed method with the state-of-the-art performance. Tianjun Shi, Jinnan Gong, Jianming Hu, Xiyang Zhi, Guiyi Zhu, Binhuan Yuan, Yu Sun 0028, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningabstractLearning effective reinforcement learning (RL) policies to solve real-world complex tasks can be quite challenging without a high-fidelity simulation environment. In most cases, we are only given imperfect simulators with simplified dynamics, which inevitably lead to severe sim-to-real gaps in RL policy learning. The recently emerged field of offline RL provides another possibility to learn policies directly from pre-collected historical data. However, to achieve reasonable performance, existing offline RL algorithms need impractically large offline data with sufficient state-action space coverage for training. This brings up a new question: is it possible to combine learning from limited real data in offline RL and unrestricted exploration through imperfect simulators in online RL to address the drawbacks of both approaches? In this study, we propose the Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning (H2O) framework to provide an affirmative answer to this question. H2O introduces a dynamics-aware policy evaluation scheme, which adaptively penalizes the Q function learning on simulated state-action pairs with large dynamics gaps, while also simultaneously allowing learning from a fixed real-world dataset. Through extensive simulation and real-world tasks, as well as theoretical analysis, we demonstrate the superior performance of H2O against other cross-domain online and offline RL algorithms. H2O provides a brand new hybrid offline-and-online RL paradigm, which can potentially shed light on future RL algorithm design for solving practical real-world tasks. Yiwen Qiu, Guyue Zhou, Jianming Hu, Xianyuan Zhan |
NeurIPS | 6 |
| 2022 | Influence of Space Variability on Remote Sensing Image Restoration PerformancesabstractWith the continuous increase in the resolution of optical remote sensing satellites, the influence of space variations on the image quality cannot be ignored, especially in new imaging systems such as thin-film diffraction and rectangular rotating pupils. This paper was conducted to analyze the influence of space variability on restoration performances of different methods, then a new processing strategy of space-variant images is proposed. According to the analytical experiment results, we suggest using the block method when the PSV < 0.20% and otherwise selecting the global method. In order to ensure the final image quality, we also suggest controlling the PSV within 0.28% when designing optical systems. This study can provide a foundation for optimizing the design of front-end optical systems and selecting back-end processing methods in engineering applications. Shikai Jiang, Xiyang Zhi, Tianjun Shi, Jianming Hu, Wei Zhang 0220, Jinnan Gong |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Supervised Multi-Scale Attention-Guided Ship Detection in Optical Remote Sensing ImagesabstractShip detection in optical remote sensing images plays a significant role in a wide range of civilian and military tasks. However, it is still a challenging issue owing to complex environmental interferences and a large variety of target scales and positions. To overcome these limitations, we propose a supervised multi-scale attention-guided detection framework, which can effectively detect ships of different scales both in complex pure ocean and port scenes. Specifically, a multi-scale supervision module is first proposed to adjust the semantic consistency of different feature levels, obtaining extracted features with small semantic gaps. Next, an attention-guided module is utilized to aggregate context information from both spatial and channel dimensions by calculating map correlations, adaptively enhancing the feature representation. Moreover, to preserve the attribute and spatial relationship of the optimized features, we adopt a capsule-based module as the classifier and obtain satisfactory classification performance. Experimental results conducted on two public high-quality datasets demonstrate that the proposed method obtains state-of-the-art performance in comparison with several advanced methods. Jianming Hu, Xiyang Zhi, Shikai Jiang, Hao Tang 0005, Wei Zhang 0220, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Global Information Transmission Model-Based Multiobjective Image Inversion Restoration Method for Space Diffractive Membrane Imaging SystemsabstractDiffractive membrane imaging systems have been an important development trend for high-orbit satellite cameras owing to their advantages of large aperture, light weight, rapid manufacture, and low cost. However, caused by the cross-coupling effects of diffraction imaging, membrane properties, subaperture stitching, on-orbit disturbances, and other physical factors, lager-aperture space diffractive membrane imaging systems have specific and complex degradation characteristics: the modulation transfer function (MTF) and signal-to-noise ratio (SNR) have more prominent degradation and serious space-variant characteristics over fields of view, with obvious background radiation properties that seriously affect the application of imaging products. To address this problem, this study established a global information transmission model by characterizing the PSF and background radiation in a full field of view to represent the imaging law of an on-orbit system. Aiming at the inverse problem of the information transmission model, we also propose a novel image inversion restoration method for the special degradation characteristics. In particular, the effect of diffraction efficiency is introduced into the inversion restoration method to solve the background radiation problem. Moreover, we innovatively designed matrix regularization parameters to further improve the correction ability of spatial variation. When the diffraction efficiency was experimentally higher than 60% and the mean measured spatial variability was less than 0.2, the proposed method exhibited a satisfactory processing performance, and could improve multiobjective comprehensive processing, such as transfer function compensation, spatial variation correction, and background radiation removal. Shikai Jiang, Xiyang Zhi, Wei Zhang 0220, Dawei Wang 0007, Jianming Hu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Harmonious Lane Changing via Deep Reinforcement LearningabstractIn this paper, we study how to learn a harmonious deep reinforcement learning (DRL) based lane-changing strategy for autonomous vehicles without Vehicle-to-Everything (V2X) communication support. The basic framework of this paper can be viewed as a multi-agent reinforcement learning in which different agents will exchange their strategies after each round of learning to reach a zero-sum game state. Unlike cooperation driving, harmonious driving only relies on individual vehicles’ limited sensing results to balance overall and individual efficiency. Specifically, we propose a well-designed reward that combines individual efficiency with overall efficiency for harmony, instead of only emphasizing individual interests like competitive strategy. Testing results show that competitive strategy often leads to selfish lane change behaviors, anarchy of crowd, and thus the degeneration of traffic efficiency. In contrast, the proposed harmonious strategy can promote traffic efficiency in both free flow and traffic jam than the competitive strategy. This interesting finding indicates that we should take care of the reward setting for reinforcement learning-based AI robots (e.g., automated vehicles) design, when the utilities of these robots are not strictly in alignment. Jianming Hu, Zhiheng Li 0001, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Bi-directional Attention Feature Enhancement for Video Instance SegmentationabstractAs a recently proposed task, video instance segmentation (VIS) can classify, detect, segment and track each instance in a given video, which is very useful for driving environment perception of autonomous vehicles. In this paper, we propose a novel method called bi-directional attention feature enhancement (BAFE) to make up for the lack of feature processing in existing works of VIS. BAFE contains a top-down attention branch and a bottom-up attention branch. It uses attention to perform two-way information transmission between high-level, semantic features and low-level, local features to combine them efficiently and can be utilized in any network that uses both high-level and low-level features. We add BAFE before the mask-specialized regression branch of the latest work, spatial information preservation for VIS (SipMask- VIS), to enhance features for mask generating. Besides, we introduce modified path aggregation network (Mod-PAN) to further enhance features. Different from existing works which first train their models on image datasets and then use the pre-trained models to finish the VIS task, we obtain VIS results in an end-to-end way without any pre-training. Our method outperforms SipMask-VIS by an absolute gain of 2.5%, which strongly proves its effectiveness. Tianyun Fu, Jianming Hu |
IV | 2 |
| 2018 | On-Road i-Vics Management for Blockage of Abreast Low Speed Vehicles Near Signalized IntersectionsabstractDriving at velocity much lower than the speed limit greatly restraints the movements of the followers, resulting in considerable waste of road resource every day. When two slow vehicles are moving abreast on the city road, the result is devastating for their followers. To solve such problem by the idea the i-Vics (intelligent vehicular infrastructure cooperative systems), this paper proposes a dedicated on-road traffic management plan for abreast low speed connected vehicles, especially for scenarios near signalized intersections. The management plan takes three steps to erase the traffic blockage caused by those abreast slow drivers, including detection of abreast slow drivers, appraisal of environmental factors, and traffic guidance notification for all relevant drivers. In the proposed plan, the detection model initially evaluates the whole street and picks out all the abreast low speed vehicles based on the i-Vics. For every blockage target, a appraisal model of the surroundings determines the strategy how to dissolve the traffic blockage, including the feasibility of overtaking, and possibility of passing through the intersection and etc. Finaly, once the dissolving strategy is approved by the appraisal program, the management system sends notification of speed guidance to all relevant drivers to guide them to pass through the signalized intersection within the remaining green time. Kaizhe Hou, Jianming Hu, Yi Zhang 0029 |
Intelligent Vehicles Symposium | 2 |
| 2018 | Adaptive Traffic Signal Control with Deep Recurrent Q-learningabstractThe application of modern technologies makes it possible for a transportation system to collect real-time data of some specific traffic scenes, helping traffic control center to improve the traffic efficiency. Based on such consideration, we introduce a variant deep reinforcement learning agent that might take advantage of the real-time GPS data and learn how to control the traffic lights in an isolated intersection. We combine the recurrent neural network (RNN) with Deep Q-Network, namely DRQN and compare its performance with standard Deep Q-Network (DQN) in partially observed traffic situations. The agent is trained by using Q-learning with experience replay in traffic simulator SUMO, so as to generate traffic signal control policy. Based on the experiments, both DQN and DRQN method are able to adjust its traffic signal timing policy to specific traffic environment and achieve lower average vehicle delay than fixed time control. In addition, the recurrent Q-learning method gets better simulation result than standard Q-learning method in the environment of different probe vehicle proportion. Jinghong Zeng, Jianming Hu, Yi Zhang 0029 |
Intelligent Vehicles Symposium | 2 |
| 2017 | Centralized cooperative intersection control under automated vehicle environmentabstractWith the rapid development in vehicular communication technologies, cooperative driving of intelligent vehicles can provide promising efficiency, safety and sustainability to the intelligent transportation systems. In this paper, a centralized cooperative intersection control (CCIC) approach is proposed for the non-signalized intersections under automated vehicle environment. The cooperative intersection control problem is converted to a nonlinear constrained programming problem considering vehicle delay, fuel consumption, emission and driver comfort level. Furthermore, a simulation-based case study is carried out on a four-legged, two-lane non-signalized intersection under different traffic volume scenarios to compare CCIC with the actuated intersection control (AIC) system. The results indicate that the CCIC approach shows significant potential improvements on the traffic efficiency (i.e., nearly 14% of traffic flow increase, nearly 90% of travelling time saving), emission (nearly 60% of CO2reduction) and driver comfort level (nearly 2% of comfort level increase). Jishiyu Ding, Huile Xu, Jianming Hu, Yi Zhang 0029 |
Intelligent Vehicles Symposium | 3 |
| 2017 | Queue length estimation at isolated intersections based on intelligent vehicle infrastructure cooperation systemsabstractWith the advance of intelligent vehicle infrastructure cooperation systems (i-VICS), many traffic parameters can be inferred given data from the OBU (On Board Unit of i-VICS)-equipped vehicles to calculate some information used in the modern urban traffic management. As a typical application, the estimation of investigated queue length at isolated intersections is proposed in the paper. Firstly, the queue length estimation at isolated intersections is explored to be transformed into the problem of deriving the number of queued vehicles. Two models, the improved and advanced interpolation methods, are then introduced to the condition of single cycle and multiple cycles, respectively. The microscopic simulation software VISSIM is adopted to evaluate the effect of the derived models. Its analysis shows that, under the condition of single cycle, the mean absolute error (MAE) of the improved interpolation method is less than 2 vehicles when the equipped rate of i-VICS OBU is higher than 50% and the maximum MAE with 30% equipped rate will be no more than 4 vehicles. On the other hand, under the condition of multiple cycles, the MAE of the advanced interpolation method with very low OBU-equipped rates will be less than 6 vehicles. Finally, numerical results of proposed approach have shown the relations of MAE with the volume-to-capacity ratio and the OBU-equipped rate. Huile Xu, Jishiyu Ding, Yi Zhang 0029, Jianming Hu |
Intelligent Vehicles Symposium | 4 |
| 2017 | Person re-identification by unsupervised video matching
Xiatian Zhu, Shaogang Gong, Xudong Xie, Jianming Hu, Kin-Man Lam 0001, Yisheng Zhong |
Pattern Recognit. | 5 |
| 2015 | Saliency detection based on singular value decomposition
Xudong Xie, Kin-Man Lam 0001, Jianming Hu, Yisheng Zhong |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | A SIFT-based mean shift algorithm for moving vehicle trackingabstractThe classical mean shift algorithm is easy to pass into local maxima, which is caused by the lack of appropriate target model updating mechanism. In this paper, a SIFT-based mean shift algorithm is proposed, which can be used for continuous vehicle tracking in complex situations, such as the shape and the illumination of the vehicle object change. In our algorithm, the mean shift algorithm is utilized to determine the candidate target region, and then a judgment on the tracking effect is made according to the Bhattacharyya coefficient. If tracking fails, the candidate area is matched with the target model by SIFT feature, and a new track position is determined. Otherwise, the target model is periodically updated by SIFT feature matching, and the target model can be constantly updated according to the state change of the moving vehicle. In the scenes of moving vehicle target deformations, such as the variation of scale and illumination, the algorithm is tested and compared with other algorithms. The experimental results show that the proposed method can effectively track an object under the condition of varying illumination and shape deformation. Xudong Xie, Yi Zhang 0029, Jianming Hu |
Intelligent Vehicles Symposium | 5 |
| 2011 | Short-time traffic flow prediction with ARIMA-GARCH modelabstractShort-time traffic flow prediction is a significant interest in transportation study, and it is essential in congestion control and traffic network management. In this paper, we propose an Autoregressive Integrated Moving Average with Generalized Autoregressive Conditional Heteroscedasticity (ARIMA-GARCH) model for traffic flow prediction. The model combines linear ARIMA model with nonlinear GARCH model, so it can capture both the conditional mean and conditional heteroscedasticity of traffic flow series. The model is calibrated, validated and used for prediction based on PeMS single loop detector data. The performance of the hybrid model is compared with that of standard ARIMA model. The results show that the introduction of conditional heteroscedasticity cannot bring satisfactory improvement to prediction accuracy, in some cases the general GARCH(1,1) model may even deteriorate the performance. Thus for ordinary traffic flow prediction, the standard ARIMA model is sufficient. Chenyi Chen, Jianming Hu, Yi Zhang 0029 |
Intelligent Vehicles Symposium | 2 |
| 2009 | Urban Road Network Modeling and Real-Time Prediction Based on Householder Transformation and Adjacent Vector
Shuo Deng, Jianming Hu, Yi Zhang 0029 |
ISNN (3) | 2 |
| 2009 | PPCA-Based Missing Data Imputation for Traffic Flow Volume: A Systematical ApproachabstractThe missing data problem greatly affects traffic analysis. In this paper, we put forward a new reliable method called probabilistic principal component analysis (PPCA) to impute the missing flow volume data based on historical data mining. First, we review the current missing data-imputation method and why it may fail to yield acceptable results in many traffic flow applications. Second, we examine the statistical properties of traffic flow volume time series. We show that the fluctuations of traffic flow are Gaussian type and that principal component analysis (PCA) can be used to retrieve the features of traffic flow. Third, we discuss how to use a robust PCA to filter out the abnormal traffic flow data that disturb the imputation process. Finally, we recall the theories of PPCA/Bayesian PCA-based imputation algorithms and compare their performance with some conventional methods, including the nearest/mean historical imputation methods and the local interpolation/regression methods. The experiments prove that the PPCA method provides significantly better performance than the conventional methods, reducing the root-mean-square imputation error by at least 25%. Li Qu, Jianming Hu, Li Li 0013, Yi Zhang 0029 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2007 | A Reliable Logo and Replay Detector for Sports VideoabstractReplay is one of the key cues indicating highlights in sports videos. A replay is usually sandwiched by two identical logos which prompt the start and end of a replay. A logo transition usually contains 10-30 frames, describes a flying or varying object(s). In this paper, a reliable logo and replay detecting approach is proposed. It contains two main stages: first, a logo transition template is unsupervised learned, a key frame (K-frame) and a set of pixels that describes logo object (logo pixels, L-pixels) accurately are also extracted; second, the learned information are used jointly to detect logos and replays in the video. In addition to traditional color analysis, optical flow feature is employed to depict the movement of the logo object(s). Extensive experiments show that the proposed approach can reliably detect logos and replays regardless of the types of sports videos. Jianming Hu, Wei Hu 0002, Tao Wang 0003, Hongliang Bai, Yimin Zhang 0002 |
ICME | 2 |
| 2006 | A Novel Calibration System for a Space ManipulatorabstractIt is important to calibrate the position and attitude precision for space manipulator in laboratory environment. The construction goal of the calibration system for a space manipulator is to simulate the microgravity environment in space, to counterbalance the impact of gravity on manipulator and to realize the test of 3D position precision, attitude and trajectory precision in the terrestrial environment. This paper introduces a novel calibration system. Based on the definitions of five kinds of reference frames, this paper illustrates the relationships among the reference frames and the coordinate transformation method between the relevant reference frames. Furthermore, the paper defines four calibration items. Finally, a calibration case study is presented in detail, which verifies the availability of mechanical interface, air-bearing test-bed and air-bearing feet Jingyan Song, Danya Yao, Jianming Hu, Jingran Ma, Jingchun Wang |
IROS | 3 |
| 2006 | Radial Basis Function Network for Traffic Scene Classification in Single Image Mode
Jianming Hu, Jingyan Song, Tianliang Gao |
ISNN (2) | 2 |
| 2003 | A novel networked traffic parameter forecasting method based on Markov chain modelabstractThis paper introduces a novel networked traffic parameter forecasting method. Based on the detailed analysis of the literature, the paper describes the fundamental ideas. Then we select a typical traffic network in Beijing City. In order to simplify the problem, we classify the links using clustering analysis and find the representative links in each group. Furthermore, we introduce the Markov chain model to predict the traffic parameter. EM algorithm is applied to estimate the parameters of mixed Gaussian distributions, i.e., means, covariances and mixing coefficients. According to the regression equations between the representative links and the other links in the same group, we can obtain all the predicted traffic parameters of all the link in the road network. The case studies using real data from UTC-SCOOT system in Beijing have proved the effectiveness and applicability of the proposed method. Jianming Hu, Jingyan Song, Guoqiang Yu, Yi Zhang 0029 |
SMC | 1 |
| 2002 | Automatic Detection and Verification of Text Regions in News Video FramesabstractTextual information in a video is very useful for video indexing and retrieving. Detecting text blocks in video frames is the first important procedure for extracting the textual information. Automatic text location is a very difficult problem due to the large variety of character styles and the complex backgrounds. In this paper, we describe the various steps of the proposed text detection algorithm. First, the gray scale edges are detected and smoothed horizontally. Second, the edge image is binarized, and run length analysis is applied to find candidate text blocks. Finally, each detected block is verified by an improved logical level technique (ILLT). Experiments show this method is not sensitive to color/texture changes of the characters, and can be used to detect text lines in news videos effectively. Jianming Hu, Jie Xi, Lide Wu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2002 | Page segmentation of Chinese newspapers
Jie Xi, Jianming Hu, Lide Wu |
Pattern Recognit. | 2 |
| 1999 | Locating head and face boundaries for head-shoulder images
Jianming Hu, Hong Yan 0001, Mustafa Sakalli |
Pattern Recognit. | 1 |
| 1999 | Construction of partitioning paths for touching handwritten characters
Jianming Hu, Donggang Yu, Hong Yan 0001 |
Pattern Recognit. Lett. | 1 |
| 1998 | Algorithms for partitioning path construction of handwritten numeral stringsabstractFor connected handwritten characters, strokes are merged together in the touching area. Artifacts are often present in the separated characters and therefore the recognition performance will deteriorate. The paper describes an approach to find a partitioning path which can reduce the artifacts significantly. Jianming Hu, Donggang Yu, Hong Yan 0001 |
ICPR | 1 |
| 1998 | A Model-Based Segmentation Method for Handwritten Numeral Strings
Jianming Hu, Hong Yan 0001 |
Comput. Vis. Image Underst. | 1 |
| 1998 | Structural primitive extraction and coding for handwritten numeral recognition
Jianming Hu, Hong Yan 0001 |
Pattern Recognit. | 1 |
| 1998 | A multiple point boundary smoothing algorithm,
Jianming Hu, Donggang Yu, Hong Yan 0001 |
Pattern Recognit. Lett. | 1 |
| 1996 | Structural decomposition and description of printed and handwritten charactersabstractA structural method for describing both printed and handwritten characters is presented in this paper. In this approach, each character is decomposed into primitives based on feature points detection, where a new directional point is introduced and an improved scheme for the bend point detection is proposed. Each primitive is then characterized by a so-called primitive code. A global code is derived from the primitive codes, which is used to describe the topological structure of the character. Each character is thus completely described by the primitive codes and the global code. This method has been successfully applied in handwritten numeral and printed character recognition. Jianming Hu, Hong Yan 0001 |
ICPR | 1 |