EDBT 2026 Demo / reviewers in the wild / expert
Jian Pu
dblp:43/6295
· DBLP profile ↗
83ranked-venue papers
10as first author
55since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 4 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 3 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 7 since 2021Systems, architecture and hardware · 10 · 10 since 2021Computer networks · 6 · 3 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoupling Scene Perception and Ego Status: A Multi-Context Fusion Approach for Enhanced Generalization in End-to-End Autonomous DrivingabstractModular design of planning-oriented autonomous driving has markedly advanced end-to-end systems. However, existing architectures remain constrained by an over-reliance on ego status, hindering generalization and robust scene understanding. We identify the root cause as an inherent design within these architectures that allows ego status to be easily leveraged as a shortcut. Specifically, the premature fusion of ego status in the upstream BEV encoder allows an information flow from this strong prior to dominate the downstream planning module. To address this challenge, we propose AdaptiveAD, an architectural-level solution based on a multi-context fusion strategy. Its core is a dual-branch structure that explicitly decouples scene perception and ego status. One branch performs scene-driven reasoning based on multi-task learning, but with ego status deliberately omitted from the BEV encoder, while the other conducts ego-driven reasoning based solely on the planning task. A scene-aware fusion module then adaptively integrates the complementary decisions from the two branches to form the final planning trajectory. To ensure this decoupling does not compromise multi-task learning, we introduce a path attention mechanism for ego-BEV interaction and add two targeted auxiliary tasks: BEV unidirectional distillation and autoregressive online mapping. Extensive evaluations on the nuScenes dataset demonstrate that AdaptiveAD achieves state-of-the-art open-loop planning performance. Crucially, it significantly mitigates the over-reliance on ego status and exhibits impressive generalization capabilities across diverse scenarios. Jiacheng Tang, Mingyue Feng, Jiachao Liu, Yaonong Wang, Jian Pu |
AAAI | 5 |
| 2026 | GS-Net: Point cloud sampling with graph neural networks
Xiaolei Chen 0002, Shoumeng Qiu, Xiangyang Xue 0001, Jian Pu |
Pattern Recognit. | 5 |
| 2026 | EC-SLAM: Effectively constrained neural RGB-D SLAM with TSDF hash encoding and joint optimization
Guanghao Li 0001, Qi Chen 0025, Yuxiang Yan 0002, Jian Pu |
Pattern Recognit. | 4 |
| 2026 | Natural Language to Code for Automated Annotation in Autonomous DrivingabstractThe fast expansion of deep learning models has led to an increasing need for well-annotated datasets, while traditional manual annotation cannot meet this requirement. Current research on annotation mainly focuses on automating the annotation process. These studies typically rely on a set of predefined functionalities. However, in complex scenarios, for example, autonomous driving, annotation workflow, and postprocessing functions must be tailored to specific tasks. The challenge here lies in ensuring that newly generated functions integrate with the existing function set of the system, which requires the pipeline to understand the user requirement and real-time system context to generate appropriate input-output data structures. Previous annotation methods relied on human programming to meet this requirement. This dependency on professional assistance restricts the generalizability of annotation methods. Drawing on modern software engineering principles, we introduce an interactive code generation pipeline based on natural language input to address this challenge. Our approach supports real-time code generation by natural language input. To the best of our knowledge, this is one of the latest applications that apply customizable functional extensions in the annotation pipeline. Our evaluations on public datasets and a self-built real-world dataset demonstrate that our method significantly enhances the range of application scenarios for annotation tools while reducing manual intervention. Upon acceptance, the code will be open source. Aoxiang Qin, Xiangyang Xue 0001, Jian Pu |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2025 | PC-BEV: An Efficient Polar-Cartesian BEV Fusion Framework for LiDAR Semantic SegmentationabstractAlthough multiview fusion has demonstrated potential in LiDAR segmentation, its dependence on computationally intensive point-based interactions, arising from the lack of fixed correspondences between views such as range view and Bird's-Eye View (BEV), hinders its practical deployment. This paper challenges the prevailing notion that multiview fusion is essential for achieving high performance. We demonstrate that significant gains can be realized by directly fusing Polar and Cartesian partitioning strategies within the BEV space. Our proposed BEV-only segmentation model leverages the inherent fixed grid correspondences between these partitioning schemes, enabling a fusion process that is orders of magnitude faster (170x speedup) than conventional point-based methods. Furthermore, our approach facilitates dense feature fusion, preserving richer contextual information compared to sparse point-based alternatives. To enhance scene understanding while maintaining inference efficiency, we also introduce a hybrid Transformer-CNN architecture. Extensive evaluation on the SemanticKITTI and nuScenes datasets provides compelling evidence that our method outperforms previous multiview fusion approaches in terms of both performance and inference speed, highlighting the potential of BEV-based fusion for LiDAR segmentation. Shoumeng Qiu, Xinrun Li, Xiangyang Xue 0001, Jian Pu |
AAAI | 4 |
| 2025 | VSFusion: Video-Sensor Multimodal Fusion with Gated Geometry Network for Freezing of Gait Detection in Parkinson's DiseaseabstractFreezing of Gait (FoG) is a transient and debilitating motor symptom of Parkinson's Disease (PD), marked by sudden interruptions in walking despite continued motor intention. Timely and accurate detection of FoG is critical for early intervention and clinical decision-making. However, existing approaches are often limited by single-modality inputs or lack temporal precision. In this study, we propose Video-Sensor Multimodal Fusion (VSFusion), a novel framework for frame-level FoG detection by leveraging both visual and inertial modalities. The core of VSFusion is our specially designed Gated-Delta Geometry Multimodal Fusion Network, which integrates skeletal sequences extracted from videos with wearable inertial sensor signals. Our network employs Gated DeltaNet which captures long-range temporal dependencies through outer-product memory and dynamic gating, while a geometry fusion mechanism modulates cross-modal interactions based on spatial distances across time. Experimental evaluations on a clinical FoG dataset demonstrate the superiority of our method, achieving an AUC of 0.961, surpassing several state-of-the-art baselines. These results highlight the potential of our model as a robust tool for reliable and fine-grained FoG assessment in real-world clinical settings. Qizheng Hao, Fangjin Liu, Fumin Jia, Jian Pu |
BIBM | 4 |
| 2025 | Detecting OOD Samples via Optimal Transport Scoring FunctionabstractTo deploy machine learning models in the real world, researchers have proposed many OOD detection algorithms to help models identify unknown samples during the inference phase and prevent them from making untrustworthy predictions. Unlike methods that rely on extra data for outlier exposure training, post hoc methods detect Out-of-Distribution (OOD) samples by developing scoring functions, which are model agnostic and do not require additional training. However, previous post hoc methods may fail to capture the geometric cues embedded in network representations. Thus, in this study, we propose a novel score function based on the optimal transport theory, named OTOD, for OOD detection. We utilize information from features, logits, and the softmax probability space to calculate the OOD score for each test sample. Our experiments show that combining this information can boost the performance of OTOD with a certain margin. Experiments on the CIFAR-10 and CIFAR-100 benchmarks demonstrate the superior performance of our method. Notably, OTOD outperforms the state-of-the-art method GEN by 7.19% in the mean FPR@95 on the CIFAR-10 benchmark using ResNet-18 as the backbone, and by 12.51% in the mean FPR@95 using WideResNet-28 as the backbone. In addition, we provide theoretical guarantees for OTOD. The code is available in https://github.com/HengGao12/OTOD. Zhuolin He, Jian Pu |
ICASSP | 3 |
| 2025 | Dark-ISP: Enhancing RAW Image Processing for Low-Light Object Detection
Jiasheng Guo, Yuxiang Yan 0002, Guanghao Li 0001, Jian Pu |
ICCV | 5 |
| 2025 | Deep Incomplete Multi-view Learning via Cyclic Permutation of VAEsabstractMulti-View Representation Learning (MVRL) aims to derive a unified representation from multi-view data by leveraging shared and complementary information across views. However, when views are irregularly missing, the incomplete data can lead to representations that lack sufficiency and consistency. To address this, we propose Multi-View Permutation of Variational Auto-Encoders (MVP), which excavates invariant relationships between views in incomplete data. MVP establishes inter-view correspondences in the latent space of Variational Auto-Encoders, enabling the inference of missing views and the aggregation of more sufficient information. To derive a valid Evidence Lower Bound (ELBO) for learning, we apply permutations to randomly reorder variables for cross-view generation and then partition them by views to maintain invariant meanings under permutations. Additionally, we enhance consistency by introducing an informational prior with cyclic permutations of posteriors, which turns the regularization term into a similarity measure across distributions. We demonstrate the effectiveness of our approach on seven diverse datasets with varying missing ratios, achieving superior performance in multi-view clustering and generation tasks. Jian Pu |
ICLR | 2 |
| 2025 | JointDeblur-Gs: Joint Blur-Aware Gaussian SplattingabstractMotion blur poses a critical challenge for 3D scene reconstruction, leading to degraded geometric consistency and visual clarity, severely impacting reconstruction quality. Although 3D Gaussian-Based approaches excel in static scenes, they suffer significant quality degradation under motion blur. To address this, we propose JointDeblur-Gs, a joint optimization framework that integrates a blur-aware network to jointly optimize the image enhancement module and 3D Gaussian parameters for end-to-end motion blur removal and multiview consistency. The framework leverages photometric consistency and structural similarity losses, achieving substantial reconstruction quality improvements while maintaining real-time performance. Our experiments demonstrate that JointDeblur-Gs outperforms state-of-the-art deblurring approaches on both synthetic and real datasets. Sijia Hu, Xinxiao Wang, Luyue Sun, Jian Pu |
ICME | 7 |
| 2025 | A2DO: Adaptive Anti-Degradation Odometry with Deep Multi-Sensor Fusion for Autonomous NavigationabstractAccurate localization is essential for the safe and effective navigation of autonomous vehicles, and Simultaneous Localization and Mapping (SLAM) is a cornerstone technology in this context. However, The performance of the SLAM system can deteriorate under challenging conditions such as low light, adverse weather, or obstructions due to sensor degradation. We present A2DO, a novel end-to-end multi-sensor fusion odometry system that enhances robustness in these scenarios through deep neural networks. A2DO integrates LiDAR and visual data, employing a multilayer, multi-scale feature encoding module augmented by an attention mechanism to mitigate sensor degradation dynamically. The system is pretrained extensively on simulated datasets covering a broad range of degradation scenarios and fine-tuned on a curated set of real-world data, ensuring robust adaptation to complex scenarios. Our experiments demonstrate that A2DO maintains superior localization accuracy and robustness across various degradation conditions, showcasing its potential for practical implementation in autonomous vehicle systems. Hui Lai, Junping Zhang, Jian Pu |
ICRA | 4 |
| 2025 | CasPoinTr: Point Cloud Completion with Cascaded Networks and Knowledge DistillationabstractPoint clouds collected from real-world environments are often incomplete due to factors such as limited sensor resolution, single viewpoints, occlusions, and noise. These challenges make point cloud completion essential for various applications. A key difficulty in this task is predicting the overall shape and reconstructing missing regions from highly incomplete point clouds. To address this, we introduce CasPoinTr, a novel point cloud completion framework using cascaded networks and knowledge distillation. CasPoinTr decomposes the completion task into two synergistic stages: Shape Reconstruction, which generates auxiliary information, and Fused Completion, which leverages this information alongside knowledge distillation to generate the final output. Through knowledge distillation, a teacher model trained on denser point clouds transfers incomplete-complete associative knowledge to the student model, enhancing its ability to estimate the overall shape and predict missing regions. Together, the cascaded networks and knowledge distillation enhance the model’s ability to capture global shape context while refining local details, effectively bridging the gap between incomplete inputs and complete targets. Experiments on ShapeNet-55 under different difficulty settings demonstrate that CasPoinTr outperforms existing methods in shape recovery and detail preservation, highlighting the effectiveness of our cascaded structure and distillation strategy. Yuxiang Yan 0002, Boda Liu 0002, Jian Pu |
IROS | 4 |
| 2025 | Learning Spatial-Aware Manipulation OrderingabstractManipulation in cluttered environments is challenging due to spatial dependencies among objects, where an improper manipulation order can cause collisions or blocked access. Existing approaches often overlook these spatial relationships, limiting their flexibility and scalability. To address these limitations, we propose OrderMind, a unified spatial-aware manipulation ordering framework that directly learns object manipulation priorities based on spatial context. Our architecture integrates a spatial context encoder with a temporal priority structuring module. We construct a spatial graph using k-Nearest Neighbors to aggregate geometric information from the local layout and encode both object-object and object-manipulator interactions to support accurate manipulation ordering in real-time. To generate physically and semantically plausible supervision signals, we introduce a spatial prior labeling method that guides a vision-language model to produce reasonable manipulation orders for distillation. We evaluate OrderMind on our Manipulation Ordering Benchmark, comprising 163,222 samples of varying difficulty. Extensive experiments in both simulation and real-world environments demonstrate that our method significantly outperforms prior approaches in effectiveness and efficiency, enabling robust manipulation in cluttered scenes. Yuxiang Yan 0002, Guanghao Li 0001, Qunyan Pu, Jian Pu |
NeurIPS | 8 |
| 2025 | GMV: A Unified and Efficient Graph Multi-View Learning FrameworkabstractGraph Neural Networks (GNNs) are pivotal in graph classification but often struggle with generalization and overfitting. We introduce a unified and efficient Graph Multi-View (GMV) learning framework that integrates multi-view learning into GNNs to enhance robustness and efficiency. Leveraging the lottery ticket hypothesis, GMV activates diverse sub-networks within a single GNN through a novel training pipeline, which includes mixed-view generation, and multi-view decomposition and learning. This approach simultaneously broadens "views" from the data, model, and optimization perspectives during training to enhance the generalization capabilities of GNNs. During inference, GMV only incorporates additional prediction heads into standard GNNs, thereby achieving multi-view learning at minimal cost. Our experiments demonstrate that GMV surpasses other augmentation and ensemble techniques for GNNs and Graph Transformers across various graph classification scenarios. Qipeng Zhu, Jian Pu, Junping Zhang |
NeurIPS | 3 |
| 2025 | Efficient Spiking PointNet via Progressive Attention Decay
Yikai Pan, Shuai Zuo, Jian Pu |
PRCV (10) | 4 |
| 2025 | Advancing 3D Object Detection With Depth-Aware Spatial Knowledge DistillationabstractAccurate 3D object detection from images can be hindered by inherent depth ambiguity. While knowledge distillation (KD) from privileged sensors such as LiDAR offers a promising direction, it often suffers from a critical cross-sensor domain gap. To address this, we introduce DK3D, a novel depth-aware knowledge distillation framework for 3D detection. Our core strategy involves providing the teacher with privileged ground-truth depth during training. This directly avoids the feature representation mismatch and subsequent inefficient knowledge transfer required when distilling from a LiDAR teacher (sparse, geometric) to a camera-based student (dense, semantic). DK3D introduces specialized modules tailored for two primary student paradigms. For depth-assisted models, we employ a channel-wise projection layer (CPL) and an adversarial scoring block (ASB) to align intermediate features at both the pixel and distribution levels. For depth-independent models, a novel vision-depth association module allows the student to implicitly reason about geometry by fusing depth cues with visual features. Both approaches are further enhanced by target-aware spatial response distillation, which captures complex inter-object spatial relationships. Extensive experiments on the KITTI and nuScenes benchmarks demonstrate that DK3D significantly improves performance for both monocular and multi-view 3D detection, outperforming state-of-the-art methods. As a versatile, plug-and-play framework, DK3D boosts existing models without requiring additional training data or increasing the computational cost at inference. Zizhang Wu, Yuanzhu Gan, Yunzhe Wu, Tianhao Xu, Xiaoquan Wang, Jian Pu |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | UL-SLAM: A Universal Monocular Line-Based SLAM Via Unifying Structural and Non-Structural ConstraintsabstractLeveraging structural line features to complement sparse point features has been studied in recent years. However, this approach relies on a Manhattan world assumption and does not incorporate non-structural lines due to the triangulation degeneracy problem and tracking process instability. To address these problems, we propose a general line-based SLAM system that combines points, structural and non-structural lines. First, an efficient line matching algorithm for multi-scale is designed to obtain more accurate matching pairs through a divide-and-conquer approach. In addition, a novel line triangulation strategy utilizing spatial-temporal consistency and degeneracy identification is proposed to improve the quality of line generation in a sliding window. Finally, universal structural constraints based on measurement of the vanishing directions are implemented to complement the information missing from the Plücker line projection in local mapping optimization. Extensive experiments are conducted on the public EuRoC and TUM datasets as well as a self-collected dataset, and the results show that UL-SLAM achieves cutting-edge performance among recent state-of-the-art methods in both accuracy and speed. Ablation experiments also demonstrate that the integration of different line features can improve the robustness and accuracy of a visual SLAM system in challenging scenarios with low texture and weak illumination. Our implementation of the UL-SLAM will be open-sourced to benefit the community (https://github.com/jhch1995/UL-SLAM).Note to Practitioners—This article was motivated by the challenges of visual localization problems in human-made indoor scenes. Visual localization has been widely used in various robotic fields such as self-driving vehicles, augmented reality (AR), and virtual reality (VR). In real-world scenes, the localization accuracy will be significantly decreased because of the sparse and uncertain visual features in low-texture or weak illumination environments, which reduces the robustness of robot tracking. To address this problem, this article proposes a novel universal line-based SLAM system (UL-SLAM) that unifies structural and non-structural constraints within a general framework unrestricted by the strong global Manhattan world assumption. UL-SLAM can not only improve the accuracy of pose estimation due to the proposed methods for the line features but also achieve real-time performance. In addition, for 3D mapping construction, UL-SLAM can also enrich the geometric structure information of the indoor scenes. Extensive experiments are conducted on various indoor datasets for autonomous robots, and the results demonstrate the efficiency, accuracy, and robustness of the proposed system in different complex scenarios. Haochen Jiang, Rui Qian 0004, Liang Du 0004, Jian Pu, Jianfeng Feng |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | TripleNet: Exploiting Complementary Features and Pseudo-Labels for Semi-Supervised Salient Object DetectionabstractDue to the limited output categories, semi-supervised salient object detection faces challenges in adapting conventional semi-supervised strategies. To address this limitation, we propose a multi-branch architecture that extracts complementary features from labeled data. Specifically, we introduce TripleNet, a three-branch network architecture designed for contour, content, and holistic saliency prediction. The supervision signals for the contour and content branches are derived by decomposing the limited ground truths. After training on the labeled data, the model produces pseudo-labels for unlabeled images, including contour, content, and salient objects. By leveraging the complementarity between the contour and content branches, we construct coupled pseudo-saliency labels by integrating the pseudo-contour and pseudo-content labels, which differ from the model-inferred pseudo-saliency labels. We further develop an enhanced pseudo-labeling mechanism that generates enhanced pseudo-saliency labels by combining reliable regions from both pseudo-saliency labels. Moreover, we incorporate a partial binary cross-entropy loss function to guide the learning of the saliency branch to focus on effective regions within the enhanced pseudo-saliency labels, which are identified through our adaptive thresholding approach. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance using only 329 labeled training images. Ming-Hsuan Yang 0001, Jian Pu, Zhonglong Zheng |
IEEE Trans. Image Process. | 3 |
| 2025 | A Survey on Autonomous and Intelligent Swarms of Uncrewed Aerial Vehicles (UAVs)abstractUAV swarms have attracted much attention due to their high potential to execute complex missions more robustly and effectively. Essential technologies for swarms are the family of algorithms that allow the individual agents to undertake tasks intelligently, localize their relative positions, perceive surroundings, and plan and track collision-free and low-cost trajectories cooperatively so that the swarm’s overall objectives are efficiently achieved. There is still a lack of corresponding surveys that provide a systematic summary covering the control layer to task allocation and guide application-driven researchers in leveraging these capabilities for diverse UAV swarm applications. This survey debates the essential technologies of UAV swarms, including swarm trajectory planning, task assignment, control approaches, localization, perception, and communications. State-of-the-art algorithms and recent technical advancements have been investigated to expose the potential for developing highly autonomous and intelligent swarm systems. It further explores the use cases of UAV swarms in civil applications and critically analyzes existing technologies. The paper concludes by emphasizing the challenges for autonomous and intelligent UAV swarms and outlining potential future research directions. Overall, this paper provides a contemporary and comprehensive review of UAV swarm technologies and investigates their potential to transform civil application fields and support future technology advancement. Zhenpeng Du, Chunbo Luo, Geyong Min, Cai Luo, Jian Pu, Shuai Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Toward Camera Open-Set 3D Object Detection for Autonomous Driving ScenariosabstractConventional camera-based 3D object detectors in autonomous driving are limited to recognizing a predefined set of objects, which poses a safety risk when encountering novel or unseen objects in real-world scenarios. To address this limitation, we present OS-Det3D, a two-stage training framework designed for camera-based open-set 3D object detection. In the first stage, our proposed 3D object discovery network (ODN3D) uses geometric cues from LiDAR point clouds to generate class-agnostic 3D object proposals, each of which are assigned a 3D objectness score. This approach allows the network to discover objects beyond known categories, allowing for the detection of unfamiliar objects. However, due to the absence of class constraints, ODN3D-generated proposals may include noisy data, particularly in cluttered or dynamic scenes. To mitigate this issue, we introduce a joint selection (JS) module in the second stage. The JS module uses both camera bird’s eye view (BEV) feature responses and 3D objectness scores to filter out low-quality proposals, yielding high-quality pseudo ground truth for unknown objects. OS-Det3D significantly enhances the ability of camera 3D detectors to discover and identify unknown objects while also improving the performance on known objects, as demonstrated through extensive experiments on the nuScenes and KITTI datasets. Zhuolin He, Xinrun Li, Jiacheng Tang, Shoumeng Qiu, Wenfu Wang, Xiangyang Xue 0001, Jian Pu |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Unleashing the Potential of Hierarchical Region Clues for Open-Vocabulary Multi-Label ClassificationabstractOpen-vocabulary multi-label classification (OV-MLC) aims to leverage the rich multi-modal knowledge from Vision-language pre-training (VLP) models to further improve the recognition ability for unseen (novel) classes beyond the training set in multi-label scenarios. Existing OV-MLC methods only perform predictions on single hierarchical regions, and aggregate the prediction scores of these regions through simpletop-kmean pooling. This fails to unleash the potential of rich hierarchical region clues in multi-label images and does not fully exploit the discriminative information from all regions in the image, resulting in sub-optimal performance. In this work, we propose a novel OV-MLC framework to fully harness the power of multiple hierarchical region clues. Specifically, we first design a hierarchical clue gathering (HCG) module to gather different hierarchical clues, enabling more precise recognition of multiple object categories with different sizes in a multi-label image. Then, by viewing multi-label classification as single-label classification of each region within the image, we present a novel hierarchical score aggregation (HSA) approach, thereby better utilizing the predictions of each image region for each class. We also utilize a well-designed region selection strategy (RSS) to eliminate noise or background regions in an image that are irrelevant to classification, achieving higher multi-label classification accuracy. In addition, we propose a hybrid prompt learning (HPL) strategy to enhance visual-semantic consistency while preserving the generalization capability of label embeddings for unseen classes. Extensive experiments on public benchmark datasets demonstrate that our method significantly outperforms the current state-of-the-art. Peirong Ma, Wu Ran, Zhiquan He, Jian Pu, Hong Lu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | FD3D: Exploiting Foreground Depth Map for Feature-Supervised Monocular 3D Object DetectionabstractMonocular 3D object detection usually adopts direct or hierarchical label supervision. Recently, the distillation supervision transfers the spatial knowledge from LiDAR- or stereo-based teacher networks to monocular detectors, but remaining the domain gap. To mitigate this issue and pursue adequate label manipulation, we exploit Foreground Depth map for feature-supervised monocular 3D object detection named FD3D, which develops the high-quality instructive intermediate features to conduct desirable auxiliary feature supervision with only the original image and annotation foreground object-wise depth map (AFOD) as input. Furthermore, we build up our instructive feature generation network to create instructive spatial features based on the sufficient correlation between image features and pre-processed AFOD, where AFOD provides the attention focus only on foreground objects to achieve clearer guidance in the detection task. Moreover, we apply the auxiliary feature supervision from the pixel and distribution level to achieve comprehensive spatial knowledge guidance. Extensive experiments demonstrate that our method achieves state-of-the-art performance on both the KITTI and nuScenes datasets, with no external data and no extra inference computational cost. We also conduct quantitative and qualitative studies to reveal the effectiveness of our designs. Zizhang Wu, Yuanzhu Gan, Yunzhe Wu, Xiaoquan Wang, Jian Pu |
AAAI | 6 |
| 2024 | Make a Strong Teacher with Label Assistance: A Novel Knowledge Distillation Approach for Semantic Segmentation
Shoumeng Qiu, Xinrun Li, Ru Wan, Xiangyang Xue 0001, Jian Pu |
ECCV (84) | 6 |
| 2024 | G-MIMO: Empowering GNNs with Diverse Sub-Networks for Graph ClassificationabstractGraph neural networks (GNNs) demonstrate impressive performance in graph classification, albeit exhibiting challenges in terms of generalizability and robustness. Despite its proven efficacy in enhancing the robustness and generalizability of GNNs, ensemble learning encounters practical limitations due to its extensive computational and memory requirements. This paper presents a novel low-cost ensemble learning method for graph classification, which utilizes a Graph Multi-Input and Multi-Output framework (G-MIMO) to allow multiple sub-networks within a single GNN simultaneously treated. Extensive experiments demonstrate that G-MIMO effectively enhances GNN performance across multiple graph classification tasks without introducing much computational overhead. G-MIMO outperforms existing graph augmentations and ensemble approaches, delivering a 3.33% and 4.25% increase in average accuracy over standard GCN and GIN models, respectively. The source code has been released at https://github.com/smurf-1119/GMIMO. Qipeng Zhu, Junping Zhang, Jian Pu |
ICME | 4 |
| 2024 | PointSSC: A Cooperative Vehicle-Infrastructure Point Cloud Benchmark for Semantic Scene CompletionabstractSemantic Scene Completion (SSC) aims to jointly generate space occupancies and semantic labels for complex 3D scenes. Most existing SSC models focus on volumetric representations, which are memory-inefficient for large outdoor spaces. Point clouds provide a lightweight alternative but existing benchmarks lack outdoor point cloud scenes with semantic labels. To address this, we introduce PointSSC, the first cooperative vehicle-infrastructure point cloud benchmark for semantic scene completion. These scenes exhibit long-range perception and minimal occlusion. We develop an automated annotation pipeline leveraging Semantic Segment Anything to efficiently assign semantics. To benchmark progress, we propose a LiDAR-based model with a Spatial-Aware Transformer for global and local feature extraction and a Completion and Segmentation Cooperative Module for joint completion and segmentation. PointSSC provides a challenging testbed to drive advances in semantic point cloud completion for real-world navigation. The code and datasets are available at https://github.com/yyxssm/PointSSC. Yuxiang Yan 0002, Boda Liu 0002, Jianfei Ai, Qinbu Li, Ru Wan, Jian Pu |
ICRA | 6 |
| 2024 | Multi-LIO: A Lightweight Multiple LiDAR-Inertial Odometry SystemabstractThe integration of multiple LiDAR sensors has the potential to significantly enhance odometry systems by providing comprehensive environmental measurements. However, current multiple LiDAR-inertial odometry frameworks face challenges in real-time processing due to the voluminous data generated. This paper introduces a real-time, computationally efficient multiple LiDAR-inertial odometry system (Multi-LIO) that outperforms existing state-of-the-art solutions in accuracy and scalability. Utilizing a novel parallel strategy for state updates and a voxelized map format, Multi-LIO optimizes computational efficiency. Furthermore, we introduce a point-wise uncertainty estimation method to augment the accuracy of scan-to-map registration, particularly in large-scale and complex scenarios. We validate our system’s performance through extensive experiments on various challenging sequences. Multi-LIO emerges as a robust, scalable, and extensible solution, adaptable to various LiDAR configurations. Qi Chen 0025, Guanghao Li 0001, Xiangyang Xue 0001, Jian Pu |
ICRA | 4 |
| 2024 | Mitigating Causal Confusion in Vector-Based Behavior Cloning for Safer Autonomous PlanningabstractThe utilization of vector-based deep learning techniques has great prospects in the realm of autonomous driving, particularly in the domains of prediction and planning tasks. However, the application of vector-based backbones for prediction and planning tasks may lead to the occurrence of causal confusion. Previous studies have explored the phenomenon of causal confusion, with a specific emphasis on the context of visual imitation learning. As for the vector-based model, we observe that the states of surrounding vehicles can be a nuisance shortcut. In our work, an off-policy approach is proposed to alleviate the issue by incorporating de-confounding supervision. Additionally, to better capture the environmental cues, such as route and traffic lights, in vectorized representation, a decoder utilizing iterative route fusion is devised. By incorporating auxiliary supervision and employing a dedicated decoder, we demonstrate the effectiveness of our methods in reducing causal confusion and improving performance in planning tasks through reactive and nonreactive closed-loop simulations on the nuPlan dataset. Jiayu Guo 0001, Mingyue Feng, Jinsheng Dou, Di Feng, Chengjun Li, Ru Wan, Jian Pu |
ICRA | 8 |
| 2024 | FastOcc: Accelerating 3D Occupancy Prediction by Fusing the 2D Bird's-Eye View and Perspective ViewabstractIn autonomous driving, 3D occupancy prediction outputs voxel-wise status and semantic labels for more comprehensive understandings of 3D scenes compared with traditional perception tasks, such as 3D object detection and bird’s-eye view (BEV) semantic segmentation. Recent researchers have extensively explored various aspects of this task, including view transformation techniques, ground-truth label generation, and elaborate network design, aiming to achieve superior performance. However, the inference speed, crucial for running on an autonomous vehicle, is neglected. To this end, a new method, dubbed FastOcc, is proposed. By carefully analyzing the network effect and latency from four parts, including the input image resolution, image backbone, view transformation, and occupancy prediction head, it is found that the occupancy prediction head holds considerable potential for accelerating the model while keeping its accuracy. Targeted at improving this component, the time-consuming 3D convolution network is replaced with a novel residual-like architecture, where features are mainly digested by a lightweight 2D BEV convolution network and compensated by integrating the 3D voxel features interpolated from the original image features. Experiments on the Occ3D-nuScenes benchmark demonstrate that our FastOcc achieves state-of-the-art results with a fast inference speed. Wenhao Guan, Di Feng, Yuheng Du, Xiangyang Xue 0001, Jian Pu |
ICRA | 8 |
| 2024 | HP3: Hierarchical Prediction-Pretrained Planning for Unprotected Left Turn
Zhihao Ou, Yue Hua, Jinsheng Dou, Di Feng, Jian Pu |
IROS | 6 |
| 2024 | Automated Label Unification for Multi-Dataset Semantic Segmentation with GNNsabstractDeep supervised models possess significant capability to assimilate extensive training data, thereby presenting an opportunity to enhance model performance through training on multiple datasets. However, conflicts arising from different label spaces among datasets may adversely affect model performance. In this paper, we propose a novel approach to automatically construct a unified label space across multiple datasets using graph neural networks. This enables semantic segmentation models to be trained simultaneously on multiple datasets, resulting in performance improvements. Unlike existing methods, our approach facilitates seamless training without the need for additional manual reannotation or taxonomy reconciliation. This significantly enhances the efficiency and effectiveness of multi-dataset segmentation model training. The results demonstrate that our method significantly outperforms other multi-dataset training methods when trained on seven datasets simultaneously, and achieves state-of-the-art performance on the WildDash 2 benchmark. Our code can be found in https://github.com/Mrhonor/AutoUniSeg. Xiangyang Xue 0001, Jian Pu |
NeurIPS | 4 |
| 2024 | A Robust Diffusion Modeling Framework for Radar Camera 3D Object DetectionabstractRadar-camera 3D object detection aims at interacting radar signals with camera images for identifying objects of interest and localizing their corresponding 3D bounding boxes. To overcome the severe sparsity and ambiguity of radar signals, we propose a robust framework based on probabilistic denoising diffusion modeling. We design our framework to be easily implementable on different multi-view 3D detectors without the requirement of using LiDAR point clouds during either the training or inference. In specific, we first design our framework with a denoised radar-camera encoder via developing a lightweight denoising diffusion model with semantic embedding. Secondly, we develop the query denoising training into 3D space via introducing the reconstruction training at depth measurement for the transformer detection decoder. Our framework achieves new state-of-the-art performance on the nuScenes 3D detection benchmark but with few computational cost increases compared to the baseline detectors. Zizhang Wu, Yunzhe Wu, Xiaoquan Wang, Yuanzhu Gan, Jian Pu |
WACV | 5 |
| 2024 | Guided contrastive boundary learning for semantic segmentation
Shoumeng Qiu, Haiqiang Zhang, Ru Wan, Xiangyang Xue 0001, Jian Pu |
Pattern Recognit. | 6 |
| 2024 | Embrace sustainable AI: Dynamic data subset selection for image classification
Zimo Yin, Jian Pu, Ru Wan, Xiangyang Xue 0001 |
Pattern Recognit. | 2 |
| 2024 | VPL-SLAM: A Vertical Line Supported Point Line Monocular SLAM SystemabstractTraditional monocular visual simultaneous localization and mapping (SLAM) systems rely on point features or line features to estimate and optimize the camera trajectory and build a map of the surrounding environment. However, in complex scenarios such as underground parking, the performance of traditional point-line SLAM systems tends to degrade due to mirror reflection, illumination change, poor texture, and other interference. This paper proposes VPL-SLAM, a structural vertical line supported point-line monocular SLAM system that works well in complex environments such as underground parking or campus. The proposed system leverages structural vertical lines at all instances of the process. With the assistance of the structural vertical lines and global vertical direction, our system can output a more accurate visual odometry result. Furthermore, the resulting map of our system is a more reasonable structural line feature map than the previous point-line-based monocular SLAM systems. Our system has been tested with the popular autonomous driving dataset Kitti Odometry. In addition, to fully test the proposed SLAM system, we also test our system using a self-collected dataset, including underground parking and campus scenarios. As a result, our proposal reveals a more accurate navigation result and a more reasonable structural resulting map compared to state-of-the-art point-line SLAM systems such as Structure PLP-SLAM. Qi Chen 0025, Yu Cao 0024, Guanghao Li 0001, Shoumeng Qiu, Xiangyang Xue 0001, Hong Lu 0001, Jian Pu |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2024 | Exploiting Neighbor Effect: Conv-Agnostic GNN Framework for Graphs With HeterophilyabstractDue to the homophily assumption in graph convolution networks (GCNs), a common consensus in the graph node classification task is that graph neural networks (GNNs) perform well on homophilic graphs but may fail on heterophilic graphs with many interclass edges. However, the previous interclass edges' perspective and related homo-ratio metrics cannot well explain the GNNs' performance under some heterophilic datasets, which implies that not all the interclass edges are harmful to GNNs. In this work, we propose a new metric based on the von Neumann entropy to reexamine the heterophily problem of GNNs and investigate the feature aggregation of interclass edges from an entire neighbor identifiable perspective. Moreover, we propose a simple yet effective Conv-Agnostic GNN framework (CAGNNs) to enhance the performance of most GNNs on the heterophily datasets by learning the neighbor effect for each node. Specifically, we first decouple the feature of each node into the discriminative feature for downstream tasks and the aggregation feature for graph convolution (GC). Then, we propose a shared mixer module to adaptively evaluate the neighbor effect of each node to incorporate the neighbor information. The proposed framework can be regarded as a plug-in component and is compatible with most GNNs. The experimental results over nine well-known benchmark datasets indicate that our framework can significantly improve performance, especially for the heterophily graphs. The average performance gain is 9.81%, 25.81%, and 20.61% compared with graph isomorphism network (GIN), graph attention network (GAT), and GCN, respectively. Extensive ablation studies and robustness analysis further verify the effectiveness, robustness, and interpretability of our framework. Code is available at https://github.com/JC-202/CAGNN. Shouzhen Chen, Junbin Gao, Zengfeng Huang, Junping Zhang, Jian Pu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object DetectionabstractMonocular 3D object detection is a low-cost but challenging task, as it requires generating accurate 3D localization solely from a single image input. Recent developed depth-assisted methods show promising results by using explicit depth maps as intermediate features, which are either precomputed by monocular depth estimation networks or jointly evaluated with 3D object detection. However, inevitable errors from estimated depth priors may lead to misaligned semantic information and 3D localization, hence resulting in feature smearing and suboptimal predictions. To mitigate this issue, we propose ADD, an Attention-based Depth knowledge Distillation framework with 3D-aware positional encoding. Unlike previous knowledge distillation frameworks that adopt stereo- or LiDAR-based teachers, we build up our teacher with identical architecture as the student but with extra ground-truth depth as input. Credit to our teacher design, our framework is seamless, domain-gap free, easily implementable, and is compatible with object-wise ground-truth depth. Specifically, we leverage intermediate features and responses for knowledge distillation. Considering long-range 3D dependencies, we propose 3D-aware self-attention and target-aware cross-attention modules for student adaptation. Extensive experiments are performed to verify the effectiveness of our framework on the challenging KITTI 3D object detection benchmark. We implement our framework on three representative monocular detectors, and we achieve state-of-the-art performance with no additional inference computational cost relative to baseline models. Our code is available at https://github.com/rockywind/ADD. Zizhang Wu, Yunzhe Wu, Jian Pu, Xiaoquan Wang |
AAAI | 3 |
| 2023 | From Node Interaction to Hop Interaction: New Effective and Scalable Graph Learning ParadigmabstractExisting Graph Neural Networks (GNNs) follow the message-passing mechanism that conducts information interaction among nodes iteratively. While considerable progress has been made, such node interaction paradigms still have the following limitation. First, the scalability limitation precludes the broad application of GNNs in large-scale industrial settings since the node interaction among rapidly expanding neighbors incurs high computation and memory costs. Second, the over-smoothing problem restricts the discrimination ability of nodes, i.e., node representations of different classes will converge to indistinguishable after repeated node interactions. In this work, we propose a novel hop interaction paradigm to address these limitations simultaneously. The core idea is to convert the interaction target among nodes to pre-processed multi-hop features inside each node. We design a simple yet effective HopGNN framework that can easily utilize existing GNNs to achieve hop interaction. Furthermore, we propose a multi-task learning strategy with a self-supervised learning objective to enhance HopGNN. We conduct extensive experiments on 12 benchmark datasets in a wide range of domains, scales, and smoothness of graphs. Experimental results show that our methods achieve superior performance while maintaining high scalability and efficiency. The code is at https://github.com/JC-202/HopGNN. Jie Chen 0001, Zilong Li 0001, Junping Zhang, Jian Pu |
CVPR | 5 |
| 2023 | Instance-Aware Diffusion Implicit Process for Box-Based Instance SegmentationabstractThe diffusion model has demonstrated impressive performance in image generation, but its potential for discriminative tasks such as instance segmentation remains unexplored. In this paper, we propose an Instance-aware Diffusion Implicit Process (IDIP) framework for instance segmentation based on boxes. During training, IDIP diffuses ground-truth boxes across various time steps, extracting corresponding Region of Interest (RoI) features. Dynamic convolution is then used to predict boxes and categories for each RoI, and the mask head generates masks from these predictions. During inference, IDIP iteratively refines randomly generated boxes with the denoising diffusion implicit model, while the mask head derives final masks from RoIs based on the refined boxes. Our method surpasses existing approaches on the COCO benchmark, requiring fewer training steps and less memory resources due to its dynamic design and instance-aware characteristic. Hao Ren 0002, Xingsong Liu, Junjian Huang, Ru Wan, Jian Pu, Hong Lu 0001 |
ECAI | 5 |
| 2023 | Knowledge Distillation from 3D to Bird's-Eye-View for LiDAR Semantic SegmentationabstractLiDAR point cloud segmentation is one of the most fundamental tasks for autonomous driving scene understanding. However, it is difficult for existing models to achieve both high inference speed and accuracy simultaneously. For example, voxel-based methods perform well in accuracy, while Bird’s-Eye-View (BEV)-based methods can achieve real-time inference. To overcome this issue, we develop an effective 3D-to-BEV knowledge distillation method that transfers rich knowledge from 3D voxel-based models to BEV-based models. Our framework mainly consists of two modules: the voxel-to-pillar distillation module and the label-weight distillation module. Voxel-to-pillar distillation distills sparse 3D features to BEV features for middle layers to make the BEV-based model aware of more structural and geometric information. Label-weight distillation helps the model pay more attention to regions with more height information. Finally, we conduct experiments on the SemanticKITTI dataset and Paris-Lille-3D. The results on SemanticKITTI show more than 5% improvement on the test set, especially for classes such as motorcycle and person, with more than 15% improvement. The code can be accessed at https://github.com/fengjiang5/Knowledge-Distillation-from-Cylinder3D-to-PolarNet. Shoumeng Qiu, Haiqiang Zhang, Ru Wan, Jian Pu |
ICME | 6 |
| 2023 | Multi-to-Single Knowledge Distillation for Point Cloud Semantic Segmentationabstract3D point cloud semantic segmentation is one of the fundamental tasks for environmental understanding. Although significant progress has been made in recent years, the performance of classes with few examples or few points is still far from satisfactory. In this paper, we propose a novel multi-to-single knowledge distillation framework for the 3D point cloud semantic segmentation task to boost the performance of those hard classes. Instead of fusing all the points of multi-scans directly, only the instances that belong to the previously defined hard classes are fused. To effectively and sufficiently distill valuable knowledge from multi-scans, we leverage a multilevel distillation framework, i.e., feature representation distillation, logit distillation, and affinity distillation. We further develop a novel instance-aware affinity distillation algorithm for capturing high-level structural knowledge to enhance the distillation efficacy for hard classes. Finally, we conduct experiments on the SemanticKITTI dataset, and the results on both the validation and test sets demonstrate that our method yields substantial improvements compared with the baseline method. The code is available at https://github.com/skyshoumeng/M2SKD. Shoumeng Qiu, Haiqiang Zhang, Xiangyang Xue 0001, Jian Pu |
ICRA | 5 |
| 2023 | MVFusion: Multi-View 3D Object Detection with Semantic-aligned Radar and Camera FusionabstractMulti-view radar-camera fused 3D object detection provides a farther detection range and more helpful features for autonomous driving, especially under adverse weather. The current radar-camera fusion methods deliver kinds of designs to fuse radar information with camera data. However, these fusion approaches usually adopt the straightforward concatenation operation between multi-modal features, which ignores the semantic alignment with radar features and sufficient correlations across modals. In this paper, we present MVFusion, a novel Multi-View radar-camera Fusion method to achieve semantic-aligned radar features and enhance the cross-modal information interaction. To achieve so, we inject the semantic alignment into the radar features via the semantic-aligned radar encoder (SARE) to produce image-guided radar features. Then, we propose the radar-guided fusion transformer (RGFT) to fuse our radar and image features to strengthen the two modals' correlation from the global scope via the cross-attention mechanism. Extensive experiments show that MVFusion achieves state-of-the-art performance (51.7% NDS and 45.3% mAP) on the nuScenes dataset. We shall release our code and trained networks upon publication. Zizhang Wu, Guilian Chen, Yuanzhu Gan, Jian Pu |
ICRA | 5 |
| 2023 | MonoPGC: Monocular 3D Object Detection with Pixel Geometry ContextsabstractMonocular 3D object detection reveals an economical but challenging task in autonomous driving. Recently center-based monocular methods have developed rapidly with a great trade-off between speed and accuracy, where they usually depend on the object center's depth estimation via 2D features. However, the visual semantic features without sufficient pixel geometry information, may affect the performance of clues for spatial 3D detection tasks. To alleviate this, we propose MonoPGC, a novel end-to-end Monocular 3D object detection framework with rich Pixel Geometry Contexts. We introduce the pixel depth estimation as our auxiliary task and design depth cross-attention pyramid module (DCPM) to inject local and global depth geometry knowledge into visual features. In addition, we present the depth-space-aware transformer (DSAT) to integrate 3D space position and depth-aware features efficiently. Besides, we design a novel depth-gradient positional encoding (DGPE) to bring more distinct pixel geometry contexts into the transformer for better object detection. Extensive experiments demonstrate that our method achieves the state-of-the-art performance on the KITTI dataset. Zizhang Wu, Yuanzhu Gan, Guilian Chen, Jian Pu |
ICRA | 5 |
| 2023 | Learning Monocular Depth in Dynamic Environment via Context-aware Temporal AttentionabstractThe monocular depth estimation task has recently revealed encouraging prospects, especially for the autonomous driving task. To tackle the ill-posed problem of 3D geometric reasoning from 2D monocular images, multi-frame monocular methods are developed to leverage the perspective correlation information from sequential temporal frames. However, moving objects such as cars and trains usually violate the static scene assumption, leading to feature inconsistency deviation and misaligned cost values, which would mislead the optimization algorithm. In this work, we present CTA-Depth, a Context-aware Temporal Attention guided network for multi-frame monocular Depth estimation. Specifically, we first apply a multi-level attention enhancement module to integrate multi-level image features to obtain an initial depth and pose estimation. Then the proposed CTA-Refiner is adopted to alternatively optimize the depth and pose. During the CTA-Refiner process, context-aware temporal attention (CTA) is developed to capture the global temporal-context correlations to maintain the feature consistency and estimation integrity of moving objects. In particular, we propose a long-range geometry embedding (LGE) module to produce a long-range temporal geometry prior. Our approach achieves significant improvements (e.g., 13.5% for the Abs Rel metric on the KITTI dataset) over state-of-the-art approaches on three benchmark datasets. Zizhang Wu, Zhuozheng Li, Zhi-Gang Fan, Yunzhe Wu, Yuanzhu Gan, Jian Pu |
IJCAI | 6 |
| 2023 | Understanding Depth Map Progressively: Adaptive Distance Interval Separation for Monocular 3d Object DetectionabstractMonocular 3D object detection aims to locate objects in different scenes with just a single image. Due to the absence of depth information, several monocular 3D detection techniques have emerged that rely on auxiliary depth maps from the depth estimation task. There are multiple approaches to understanding the representation of depth maps, including treating them as pseudo-LiDAR point clouds, leveraging implicit end-to-end learning of depth information, or considering them as an image input. However, these methods have certain drawbacks, such as their reliance on the accuracy of estimated depth maps and suboptimal utilization of depth maps due to their image-based nature. While LiDAR-based methods and convolutional neural networks (CNNs) can be utilized for pseudo point clouds and depth maps, respectively, it is always an alternative. In this paper, we propose a framework named the Adaptive Distance Interval Separation Network (ADISN) that adopts a novel perspective on understanding depth maps, as a form that lies between LiDAR and images. We utilize an adaptive separation approach that partitions the depth map into various subgraphs based on distance and treats each of these subgraphs as an individual image for feature extraction. After adaptive separations, each subgraph solely contains pixels within a learned interval range. If there is a truncated object within this range, an evident curved edge will appear, which we can leverage for texture extraction using CNNs to obtain rich depth information in pixels. Meanwhile, to mitigate the inaccuracy of depth estimation, we designed an uncertainty module. To take advantage of both images and depth maps, we use different branches to learn localization detection tasks and appearance tasks separately. Our approach significantly enhances the baseline and outperforms depth-assisted techniques, as shown by our extensive experiments on the KITTI monocular 3D object detection benchmark. Xianhui Cheng, Shoumeng Qiu, Zhikang Zou, Jian Pu, Xiangyang Xue 0001 |
IJCNN | 4 |
| 2023 | V2Depth: Monocular Depth Estimation via Feature-Level Virtual-View Simulation and RefinementabstractDue to the lack of spatial cues giving merely a single image, many monocular depth estimation methods have been developed to leverage stereo or multi-view images to learn the spatial information of a scene in a self-supervised manner. However, these methods have limited performance gain since they are not able to exploit sufficient 3D geometry cues during inference, where only monocular images are available. In this work, we present V2Depth, a novel coarse-to-fine framework with Virtual View feature simulation for supervised monocular Depth estimation. Specifically, we first design a virtual-view feature simulator by leveraging the technique of novel view synthesis and contrastive learning to generate virtual view feature maps. In this way, we explicitly provide representative spatial geometry for subsequent depth estimation in both the training and inference stages. Then we introduce a 3DVA-Refiner to iteratively optimize the predicted depth map. During the optimization process, 3D-aware virtual attention is developed to capture the global spatial-context correlations to maintain the feature consistency of different views and estimation integrity of the 3D scene such as objects with occlusion relationships. Decisive improvements over state-of-the-art approaches on three benchmark datasets across all metrics demonstrate the superiority of our method. Zizhang Wu, Zhuozheng Li, Zhi-Gang Fan, Yunzhe Wu, Jian Pu, Xianzhi Li 0001 |
ACM Multimedia | 5 |
| 2023 | Embedding expert demonstrations into clustering buffer for effective deep reinforcement learningabstractAs one of the most fundamental topics in reinforcement learning (RL), sample efficiency is essential to the deployment of deep RL algorithms. Unlike most existing exploration methods that sample an action from different types of posterior distributions, we focus on the policy sampling process and propose an efficient selective sampling approach to improve sample efficiency by modeling the internal hierarchy of the environment. Specifically, we first employ clustering methods in the policy sampling process to generate an action candidate set. Then we introduce a clustering buffer for modeling the internal hierarchy, which consists of on-policy data, off-policy data, and expert data to evaluate actions from the clusters in the action candidate set in the exploration stage. In this way, our approach is able to take advantage of the supervision information in the expert demonstration data. Experiments on six different continuous locomotion environments demonstrate superior reinforcement learning performance and faster convergence of selective sampling. In particular, on the LGSVL task, our method can reduce the number of convergence steps by 46.7% and the convergence time by 28.5%. Furthermore, our code is open-source for reproducibility. The code is available at https://github.com/Shihwin/SelectiveSampling . Shihmin Wang, Binqi Zhao, Zhengfeng Zhang, Junping Zhang, Jian Pu |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2023 | Graph Decoupling Attention Markov Networks for Semisupervised Graph Node ClassificationabstractGraph neural networks (GNNs) have been ubiquitous in graph node classification tasks. Most GNN methods update the node embedding iteratively by aggregating its neighbors' information. However, they often suffer from negative disturbances, due to edges connecting nodes with different labels. One approach to alleviate this negative disturbance is to use attention to learn the weights of aggregation, but current attention-based GNNs only consider feature similarity and suffer from the lack of supervision. In this article, we consider label dependency of graph nodes and propose a decoupling attention mechanism to learn both hard and soft attention. The hard attention is learned on labels for a refined graph structure with fewer interclass edges so that the aggregation's negative disturbance can be reduced. The soft attention aims to learn the aggregation weights based on features over the refined graph structure to enhance information gains during message passing. Particularly, we formulate our model under the expectation-maximization (EM) framework, and the learned attention is used to guide label propagation in the M-step and feature propagation in the E-step, respectively. Extensive experiments are performed on six well-known benchmark graph datasets to verify the effectiveness of the proposed method. Shouzhen Chen, Mingyuan Bai, Jian Pu, Junping Zhang, Junbin Gao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Memory-Based Message Passing: Decoupling the Message for Propagation from DiscriminationabstractMessage passing is a fundamental procedure for graph neural networks in the field of graph representation learning. Based on the homophily assumption, the current message passing always aggregates features of connected nodes, such as the graph Laplacian smoothing process. However, real-world graphs tend to be noisy and/or non-smooth. The homophily assumption does not always hold, leading to sub-optimal results. A revised message passing method needs to maintain each node’s discriminative ability when aggregating the message from neighbors. To this end, we propose a Memory-based Message Passing (MMP) method to decouple the message of each node into a self-embedding part for discrimination and a memory part for propagation. Furthermore, we develop a control mechanism and a decoupling regularization to control the ratio of absorbing and excluding the message in the memory for each node. More importantly, our MMP is a general skill that can work as an additional layer to help improve traditional GNNs performance. Extensive experiments on various datasets with different homophily ratios demonstrate the effectiveness and robustness of the proposed method. Jian Pu |
ICASSP | 3 |
| 2022 | Compressed Video Quality Enhancement with Motion Approximation and Blended AttentionabstractIn recent years, various methods have been proposed to tackle the compressed video quality enhancement problem. It aims at restoring the distorted information in low-quality target frames from high-quality reference frames in the compressed video. Most methods for video quality enhancement contain two key stages, i.e., the synchronization and the fusion stages. The synchronization stage synchronizes the input frames by compensating the estimated motion vector to reference frames. The fusion stage reconstructs each frame with the compensated frames. However, the synchronization stage in previous works merely estimates the motion vector between the reference frame and the target frame. Due to the quality fluctuation of frames and region occlusion of objects, the missing detail information cannot be adequately replenished. To make full use of the temporal motion between input frames, we propose a motion approximation scheme to utilize the motion vector between the reference frames. It is able to generate additional compensated frames to further refine the missing details in the target frame. In the fusion stage, we propose a deep neural network to extract frame features with blended attention to the texture details and the quality discrepancy at different times. The experimental results show the effectiveness and robustness of our method. Xiaohao Han, Jian Pu |
ICPR | 3 |
| 2022 | Exploring Spatiotemporal Relationships for Improving Compressed Video QualityabstractThe compressed video quality enhancement problem aims at restoring the distorted information of the compressed video. Details in high-quality reference frames are usually used to improve the detailed information and reduce the artifacts of the low-quality target frames. In this paper, we propose a novel network to extract helpful information from peak quality frames and then incorporate it into low-quality frames. The proposed network first integrates the information from both the target and reference frames, and then fully spatiotemporal deformable convolutions are used to better adapt the sampling strategy of the input frames. In addition, the temporal offset calibration procedure is also applied to reduce the loss of high-frequency components in the interpolation process. The results on benchmark datasets show that our method achieves competitive performance compared with other existing methods. Xiaohao Han, Jian Pu |
ICPR | 3 |
| 2021 | Selfgait: A Spatiotemporal Representation Learning Method for Self-Supervised Gait RecognitionabstractGait recognition plays a vital role in human identification since gait is a unique biometric feature that can be perceived at a distance. Although existing gait recognition methods can learn gait features from gait sequences in different ways, the performance of gait recognition suffers from insufficient labeled data, especially in some practical scenarios associated with short gait sequences or various clothing styles. It is unpractical to label the numerous gait data. In this work, we propose a self-supervised gait recognition method, termed SelfGait, which takes advantage of the massive, diverse, unlabeled gait data as a pre-training process to improve the representation abilities of spatiotemporal backbones. Specifically, we employ the horizontal pyramid mapping (HPM) and micro-motion template builder (MTB) as our spatiotemporal backbones to capture the multi-scale spatiotemporal representations. Experiments on CASIA-B and OU-MVLP benchmark gait datasets demonstrate the effectiveness of the proposed SelfGait compared with four state-of-the-art gait recognition methods. The source code has been released at https://github.com/EchoItLiu/SelfGait. Yiqun Liu 0009, Jian Pu, Hongming Shan, Peiyang He, Junping Zhang |
ICASSP | 3 |
| 2021 | Asymmetric Loss for Positive-Unlabeled LearningabstractPositive-unlabeled (PU) learning is a learning paradigm when only positive and unlabeled data are available in the training stage. This paradigm is particularly useful for the applications that the negative samples are hard to define or expensive to obtain. We propose a simple yet effective and scalable method to address PU learning with a novel asymmetric loss. The proposed asymmetric loss behaviors differently for the prediction errors of the labeled and unlabeled samples, and thus encourages the identification of the negative examples from the unlabeled set. For the PU learning with SCAR assumption, neither hyper-parameter nor class prior is required to be tuned or known. For the situation with selection bias on the labeled samples, we propose a heuristic method to automatically choose the hyper-parameter according to the class prior on the training data. Compared with previous approaches, our method only requires a slight modification of the conventional cross-entropy loss and is compatible with various deep neural networks in an end-to-end way. Extensive experiments on synthetic and real-world datasets with and without SCAR assumption verify the superior performance of the proposed method. Jian Pu, Junping Zhang |
ICME | 2 |
| 2021 | Defense against Adversarial Attacks with an Induced ClassabstractThough deep neural networks have succeeded in various real applications, the prediction performance is significantly degraded when facing adversarial attacks. In this work, we investigate the alternation of the prediction distribution pattern under adversarial attacks and argue that such alternation is the primary reason for performance drop. To this end, we propose a simple yet effective method by introducing an induced class to attract the adversarial attack and thus protect the original classes' prediction order. Experiments on two real-world datasets demonstrate that the proposed method can maintain the prediction performance for both natural and adversarial examples. Jun Wang 0006, Jian Pu |
IJCNN | 3 |
| 2021 | Parallel pathway dense neural network with weighted fusion structure for brain tumor segmentation
Fangyan Ye, Yingbin Zheng, Hao Ye 0005, Xiaohao Han, Jun Wang 0006, Jian Pu |
Neurocomputing | 7 |
| 2021 | A self-supervised method for treatment recommendation in sepsisabstractSepsis treatment is a highly challenging effort to reduce mortality in hospital intensive care units since the treatment response may vary for each patient. Tailored treatment recommendations are desired to assist doctors in making decisions efficiently and accurately. In this work, we apply a self-supervised method based on reinforcement learning (RL) for treatment recommendation on individuals. An uncertainty evaluation method is proposed to separate patient samples into two domains according to their responses to treatments and the state value of the chosen policy. Examples of two domains are then reconstructed with an auxiliary transfer learning task. A distillation method of privilege learning is tied to a variational auto-encoder framework for the transfer learning task between the low- and high-quality domains. Combined with the self-supervised way for better state and action representations, we propose a deep RL method called high-risk uncertainty (HRU) control to provide flexibility on the trade-off between the effectiveness and accuracy of ambiguous samples and to reduce the expected mortality. Experiments on the large-scale publicly available real-world dataset MIMIC-III demonstrate that our model reduces the estimated mortality rate by up to 2.3% in total, and that the estimated mortality rate in the majority of cases is reduced to 9.5%. Sihan Zhu, Jian Pu |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2020 | Knowledge Distillation with a Precise Teacher and Prediction with Abstention
Jian Pu |
ICPR | 2 |
| 2020 | Aggregated context network for crowd countingabstractCrowd counting has been applied to a variety of applications such as video surveillance, traffic monitoring, assembly control, and other public safety applications. Context information, such as perspective distortion and background interference, is a crucial factor in achieving high performance for crowd counting. While traditional methods focus merely on solving one specific factor, we aggregate sufficient context information into the crowd counting network to tackle these problems simultaneously in this study. We build a fully convolutional network with two tasks, i.e., main density map estimation and auxiliary semantic segmentation. The main task is to extract the multi-scale and spatial context information to learn the density map. The auxiliary semantic segmentation task gives a comprehensive view of the background and foreground information, and the extracted information is finally incorporated into the main task by late fusion. We demonstrate that our network has better accuracy of estimation and higher robustness on three challenging datasets compared with state-of-the-art methods. Si-yue Yu, Jian Pu |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2019 | The Kelly Growth Optimal Portfolio with Ensemble LearningabstractAs a competitive alternative to the Markowitz mean-variance portfolio, the Kelly growth optimal portfolio has drawn sufficient attention in investment science. While the growth optimal portfolio is theoretically guaranteed to dominate any other portfolio with probability 1 in the long run, it practically tends to be highly risky in the short term. Moreover, empirical analysis and performance enhancement studies under practical settings are surprisingly short. In particular, how to handle the challenging but realistic condition with insufficient training data has barely been investigated. In order to fill voids, especially grappling with the difficulty from small samples, in this paper, we propose a growth optimal portfolio strategy equipped with ensemble learning. We synergically leverage the bootstrap aggregating algorithm and the random subspace method into portfolio construction to mitigate estimation error. We analyze the behavior and hyperparameter selection of the proposed strategy by simulation, and then corroborate its effectiveness by comparing its out-of-sample performance with those of 10 competing strategies on four datasets. Experimental results lucidly confirm that the new strategy has superiority in extensive evaluation criteria. Weiwei Shen, Jian Pu, Jun Wang 0006 |
AAAI | 3 |
| 2019 | False positive rate control for positive unlabeled learning
Shuchen Kong, Weiwei Shen, Yingbin Zheng, Jian Pu, Jun Wang 0006 |
Neurocomputing | 5 |
| 2018 | Tau-FPL: Tolerance-Constrained Learning in Linear Time
Nan Li 0019, Jian Pu, Jun Wang 0006, Junchi Yan, Hongyuan Zha |
AAAI | 3 |
| 2018 | Learning Multiviewpoint Context-Aware Representation for RGB-D Scene ClassificationabstractEffective visual representation plays an important role in the scene classification systems. While many existing methods are focused on the generic descriptors extracted from the RGB color channels, we argue the importance of depth context, since scenes are composed with spatial variability and depth is an essential component in understanding the geometry. In this letter, we present a novel depth representation for RGB-D scene classification based on a specific designed convolutional neural network (CNN). Contrast to previous deep models that transfer from pretrained RGB CNN models, we harness model by using the multiviewpoint depth image augmentation to overcome the data scarcity problem. The proposed CNN framework contains the dilated convolutions to expand the receptive field and a subsequent spatial pooling to aggregate multiscale contextual information. The combination of contextual design and multiviewpoint depth images are important toward a more compact representation, compared to directly using original depth images or off-the-shelf networks. Through extensive experiments on SUN RGB-D dataset, we demonstrate that the representation outperforms recent state of the arts, and combining it with standard CNN-based RGB features can lead to further improvements. Yingbin Zheng, Hao Ye 0005, Li Wang 0033, Jian Pu |
IEEE Signal Process. Lett. | 4 |
| 2017 | Boosting Alzheimer diagnosis accuracy with the help of incomplete privileged informationabstractEarly and accurate diagnosis of Alzheimer's disease is beneficial to both preserve daily functioning and test possible new treatments. However, current diagnosis depends on dozens of factors, including the family member, past medical problems, tests of memory, blood and urine tests, brain scans and even cerebrospinal fluid specimens. Among them, regular features (e.g., blood and urine tests, brain scans) are simple and accurate. Privileged features (e.g., tests of memory, cerebrospinal fluid specimens) are inconvenient and uncomfortable to acquire, and thus incomplete in most cases. In this work, we propose a two-stage learning framework to predict merely using the regular feature with the aid of incomplete privileged data in the training time. In particular, we first complement the missing data of privileged features by exploring the relationship with regular features and labels. The recovered privileged features, regular features, as well as labels, are combined as a new set of privileged features. Then privileged learning via matching logits is applied to boost the diagnosis accuracy. Our experiments and comparison studies with competing techniques on both synthetic data and real benchmarks have corroborated the effectiveness and superiority of the proposed framework for biomedical applications. Jian Pu, Jun Wang 0006, Yingbin Zheng, Hao Ye 0005, Weiwei Shen, Hongyuan Zha |
BIBM | 1 |
| 2017 | Glioma grading based on 3D multimodal convolutional neural network and privileged learningabstractBrain tumors, especially high-grade gliomas, are one of the most lethal cancers for humankind today. Early and accurate diagnosis of tumor grading is the key for subsequent therapy and treatment. In the past, conventional computer-aided diagnosis relies on handcrafted features from magnetic resonance images (MRI), which are usually inaccurate and laborious. Recently, deep neural networks have been developed and applied for tumor segmentation and classification. However, most existing methods consider 3D MRI as a series of 2D images and use a simple modality fusion method via feature concatenation. In this paper, we propose an end-to-end 3-dimensional convolutional neural network (3D CNN) with gated multimodal unit (GMU) fusion to integrate the information both in three dimensions and in multiple modalities. Specifically, 3D convolutional kernels are directly applied to the whole MRI images, gathering the abnormalities in sagittal, axial and coronal directions. GMU with hidden states is proposed to fuse the information of multiple MRI modalities in both feature and decision level. Based on these, privilege information extracted by GMU fusion model is utilized to train a novel network called distilled-CNN, which significantly improves the performance of classification using single modality. Empirical studies on BRATS datasets corroborate the effectiveness of the proposed 3D CNN with GMU fusion and distilled-CNN to distinguish benign gliomas and malignant gliomas. Fangyan Ye, Jian Pu, Jun Wang 0006, Hongyuan Zha |
BIBM | 2 |
| 2016 | Multiple task learning with flexible structure regularization
Jian Pu, Jun Wang 0006, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
Neurocomputing | 1 |
| 2016 | Flexible multi-task learning with latent task grouping
Jian Pu, Yu-Gang Jiang 0001, Rui Feng 0001, Xiangyang Xue 0001 |
Neurocomputing | 2 |
| 2015 | Discriminative Structured Feature Engineering for Macroscale Brain ConnectomesabstractNeuroimaging techniques can measure structural and functional brain connectivity with unprecedented detail in vivo. This so-called brain connectome can be represented as high dimensional matrices corresponding to edge weights in graphs. After measuring the matrices of two cohorts (i.e., patients and healthy controls), one is often required to formulate computational network models for effective feature engineering to draw discriminative distinctions between the cohorts, as well as estimate the associated statistical significance. We designed a novel method to reveal the intrinsic features of functional matrices of discriminative power for group comparison. More specifically, by encouraging co-selection of edges connected to the same node, we preserved the discriminative edges to maximum extent. To reduce the false positive rate of the extracted discriminative edges, an optimization procedure was developed to evaluate the significance of these edges and remove trivial ones. We validated the proposed method using both synthetic data and real benchmarks, and compared it to ℓ1 regularized logistic regression, univariate t-test and stability selection. The experimental results clearly showed that the proposed approach outperformed the three competing methods under various settings. In addition to increasing the F-measure of feature selection, our approach captured the endogenous, discriminative connectivity patterns consistent with recent findings in biomedical literature. This data-driven method paves a new avenue of enquiry into the inherent nature of network models for functional brain connectomes. Jian Pu, Jun Wang 0006, Wenwen Yu, Zhuangming Shen, Kristina Zeljic, Bomin Sun, Zheng Wang 0035 |
IEEE Trans. Medical Imaging | 1 |
| 2014 | Which Looks Like Which: Exploring Inter-class Relationships in Fine-Grained Visual Categorization
Jian Pu, Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001 |
ECCV (3) | 1 |
| 2014 | Exploring Inter-feature and Inter-class Relationships with Deep Neural Networks for Video ClassificationabstractVideos contain very rich semantics and are intrinsically multimodal. In this paper, we study the challenging task of classifying videos according to their high-level semantics such as human actions or complex events. Although extensive efforts have been paid to study this problem, most existing works combined multiple features using simple fusion strategies and neglected the exploration of inter-class semantic relationships. In this paper, we propose a novel unified framework that jointly learns feature relationships and exploits the class relationships for improved video classification performance. Specifically, these two types of relationships are learned and utilized by rigorously imposing regularizations in a deep neural network (DNN). Such a regularized DNN can be efficiently launched using a GPU implementation with an affordable training cost. Through arming the DNN with better capability of exploring both the inter-feature and the inter-class relationships, the proposed regularized DNN is more suitable for identifying video semantics. With extensive experimental evaluations, we demonstrate that the proposed framework exhibits superior performance over several state-of-the-art approaches. On the well-known Hollywood2 and Columbia Consumer Video benchmarks, we obtain to-date the best reported results: 65.7% and 70.6% respectively in terms of mean average precision. Zuxuan Wu, Yu-Gang Jiang 0001, Jun Wang 0006, Jian Pu, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2013 | Multiple Task Learning Using Iteratively Reweighted Least Square
Jian Pu, Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001 |
IJCAI | 1 |
| 2012 | Human Identification Using Temporal Information Preserving Gait TemplateabstractGait Energy Image (GEI) is an efficient template for human identification by gait. However, such a template loses temporal information in a gait sequence, which is critical to the performance of gait recognition. To address this issue, we develop a novel temporal template, named Chrono-Gait Image (CGI), in this paper. The proposed CGI template first extracts the contour in each gait frame, followed by encoding each of the gait contour images in the same gait sequence with a multichannel mapping function and compositing them to a single CGI. To make the templates robust to a complex surrounding environment, we also propose CGI-based real and synthetic temporal information preserving templates by using different gait periods and contour distortion techniques. Extensive experiments on three benchmark gait databases indicate that, compared with the recently published gait recognition approaches, our CGI-based temporal information preserving approach achieves competitive performance in gait recognition with robustness and efficiency. Junping Zhang, Liang Wang 0001, Jian Pu, Xiaoru Yuan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | Chrono-Gait Image: A Novel Temporal Template for Gait Recognition
Junping Zhang, Jian Pu, Xiaoru Yuan, Liang Wang 0001 |
ECCV (1) | 3 |
| 2010 | Utility-based fair bandwidth sharing in vehicular networksabstractThe time-constraint data traffic is very common in vehicle-to-infrastructure (V2I) and vehicle-to-vehicle (V2V) communications. Time-constraint flows are those who have fixed start and stop times and try to maximize the transferred data volume during the limited connection time. From the view of users, the total data transmitted by each flow should be proportional to each one's holding time if the network resources are allocated fairly. However, we find that the traditional fairness concept solely based on flow rates is not suitable for this scenario. Therefore, we redefine the fairness concept regarding the application utility for time-constraint flows. According to this utility-based fairness definition, we propose a new practical bandwidth sharing scheme using TCP parameter tuning for transferring data with fast-moving wireless nodes such as vehicles. Simulations through ns-2 show that the proposed algorithm can achieve better utility fairness than the standard TCP and it is also friendly to TCP in the long term. Jian Pu, Mounir Hamdi |
IWCMC | 1 |
| 2010 | Low-Resolution Gait RecognitionabstractUnlike other biometric authentication methods, gait recognition is noninvasive and effective from a distance. However, the performance of gait recognition will suffer in the low-resolution (LR) case. Furthermore, when gait sequences are projected onto a nonoptimal low-dimensional subspace to reduce the data complexity, the performance of gait recognition will also decline. To deal with these two issues, we propose a new algorithm called superresolution with manifold sampling and backprojection (SRMS), which learns the high-resolution (HR) counterparts of LR test images from a collection of HR/LR training gait image patch pairs. Then, we incorporate SRMS into a new algorithm called multilinear tensor-based learning without tuning parameters (MTP) for LR gait recognition. Our contributions include the following: 1) With manifold sampling, the redundancy of gait image patches is remarkably decreased; thus, the superresolution procedure is more efficient and reasonable. 2) Backprojection guarantees that the learned HR gait images and the corresponding LR gait images can be more consistent. 3) The optimal subspace dimension for dimension reduction is automatically determined without introducing extra parameters. 4) Theoretical analysis of the algorithm shows that MTP converges. Experiments on the USF human gait database and the CASIA gait database show the increased efficiency of the proposed algorithm, compared with previous algorithms. Junping Zhang, Jian Pu, Changyou Chen, Rudolf Fleischer |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2010 | Utility-based fair bandwidth sharing in vehicular networks (extended)abstractAbstract The time‐constraint data traffic is very common in vehicle‐to‐infrastructure (V2I) and vehicle‐to‐vehicle (V2V) communications. Time‐constraint flows are those who have fixed start and stop times and try to maximize the transferred data volume during the limited connection time. From the view of users, the total data transmitted by each flow should be proportional to each one's holding time if the network resources are allocated fairly. However, we find that the traditional fairness concept solely based on flow rates is not suitable for this scenario. Therefore, we redefine the fairness concept regarding the application utility for time‐constraint flows. According to this new utility‐based fairness definition, we propose two practical bandwidth sharing schemes for transferring data by fast‐moving wireless nodes such as vehicles. Simulations through ns‐2 show that the proposed algorithms can achieve better utility fairness than the standard TCP and are also friendly to TCP in the long term. Copyright © 2010 John Wiley & Sons, Ltd. Jian Pu, Mounir Hamdi |
Wirel. Commun. Mob. Comput. | 1 |
| 2009 | Interactive Super-Resolution through Neighbor Embedding
Jian Pu, Junping Zhang, Peihong Guo, Xiaoru Yuan |
ACCV (3) | 1 |
| 2009 | Model-Tree-Based Rate Adaptation Scheme for Vehicular NetworksabstractRate adaptation techniques have been extensively studied for traditional wireless LANs as way to adjust the data transmission bandwidth as a function of the channel quality. However, existing solutions which are mostly based on statistics collection inherently experience a large delay in response to the wireless channel fluctuations, which is unsuitable for the rapid changing of the channel conditions in vehicular networks. In this paper, we propose an efficient self-adaptive model-tree-based rate adaptation scheme termed MTRA that can predict the packet error rate (PER) and adapt the data rate in real time. We also present a detailed methodology for the PER model tree building and update. We have performed comprehensive experimentations using ns-2 simulations which demonstrate that MTRA can achieve much better performance than the traditional rate adaptation approaches under various scenarios in vehicular networks. Qiuyan Xia, Jian Pu, Mounir Hamdi |
ICC | 2 |
| 2009 | Practical and efficient open-loop rate/link adaptation algorithm for high-speed IEEE 802.11n WLANsabstractIn this paper, we propose a new open-loop rate/link adaptation algorithm (ARFHT) for the emerging high-speed IEEE 802.11n WLANs. ARFHT extends the legacy rate adaptation algorithms for SISO WLANs to make it applicable in the context of MIMO-based 802.11n WLANs. It adapts the MIMO mode in terms of spatial multiplexing and spatial diversity, the two fundamental characteristics of the 802.11n MIMO PHY. It also modifies the link estimation and probing behavior of legacy SISO algorithms. The combined adaptation to the appropriate MIMO mode as well as the appropriate modulation coding scheme selection achieves high channel utilization. In this paper, we provide the intuition and the design details of the ARFHT algorithm. A comprehensive simulation study using ns-2 will demonstrate that ARFHT achieves excellent throughput performance in most scenarios and is highly responsive to varying link conditions, with minimum overhead. Qiuyan Xia, Jian Pu, Mounir Hamdi, Khaled Ben Letaief |
ISCC | 2 |
| 2009 | QoS based scheduling in the downlink of multi-user wireless systems (extended)
Jian Pu, Mounir Hamdi |
Comput. Commun. | 2 |
| 2009 | Neighbor embedding based super-resolution algorithm through edge detection and feature selection
Tak-Ming Chan, Junping Zhang, Jian Pu, Hua Huang 0001 |
Pattern Recognit. Lett. | 3 |
| 2008 | Enhancements on Router-Assisted Congestion Control for Wireless NetworksabstractThis paper addresses two challenges that are encountered when applying router-assisted explicit-feedback congestion control to wireless networks. The first challenge is how to distinguish between the two kinds of packet loss (non- congestion loss and congestion loss) in wireless networks and how to react to them properly. The second challenge is how to probe the unknown bandwidth capacity of wireless links which is required in calculating the router feedback. Some practical and novel enhancements on router-assisted congestion control in such environments are introduced in this paper. We have implemented these enhancements in a router-assisted congestion control protocol termed QFCP. Extensive simulation experiments using ns-2 will demonstrate that QFCP can fairly and efficiently allocate wireless bandwidth resources among competing flows in heterogeneous networks. Jian Pu, Mounir Hamdi |
IEEE Trans. Wirel. Commun. | 1 |
| 2007 | Improving Quality of Service for Congestion Control in High-Speed Wired-cum-Wireless NetworksabstractTCP is currently the dominate congestion control protocol for the Internet. However, as the Internet evolves into a high-speed wired-cum-wireless hybrid network, performance degradation problems of TCP have appeared, such as underutilizing high-speed links, regarding wireless loss as congestion signal, and unfairness among flows with different RTTs. In order to improve the quality of service for such high-speed hybrid networks, we propose a router-assisted congestion control protocol called quick flow control protocol (QFCP). Performance evaluation using network simulator NS-2 shows that QFCP can significantly shorten flow completion time, fairly allocate bandwidth resource, and be robust to non-congestion- related loss. Jian Pu, Mounir Hamdi |
GLOBECOM | 1 |
| 2006 | Roundup: a multi-genome repository of orthologs and evolutionary distancesabstractSUMMARY: We have created a tool for ortholog and phylogenetic profile retrieval called Roundup. Roundup is backed by a massive repository of orthologs and associated evolutionary distances that was built using the reciprocal smallest distance algorithm, an approach that has been shown to improve upon alternative approaches of ortholog detection, such as reciprocal blast. Presently, the Roundup repository contains all possible pair-wise comparisons for over 250 genomes, including 32 Eukaryotes, more than doubling the coverage of any similar resource. The orthologs are accessible through an intuitive web interface that allows searches by genome or gene identifier, presenting results as phylogenetic profiles together with gene and molecular function annotations. Results may be downloaded as phylogenetic matrices for subsequent analysis, including the construction of whole-genome phylogenies based on gene-content data. AVAILABILITY: http://rodeo.med.harvard.edu/tools/roundup. Todd F. DeLuca, I-Hsien Wu, Jian Pu, Thomas Monaghan, Leonid Peshkin, Saurav Singh, Dennis P. Wall |
Bioinform. | 3 |
| 2002 | A general model for machinable features and its application to machinability evaluation of mechanical parts
Jian Pu, Xiankui Wang |
Comput. Aided Des. | 2 |