EDBT 2026 Demo / reviewers in the wild / expert
Yanyong Zhang
dblp:44/2799
· DBLP profile ↗
178ranked-venue papers
13as first author
84since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 67 · 13 since 2021Artificial intelligence and machine learning · 46 · 1 first-author · 43 since 2021Systems, architecture and hardware · 41 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 26 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Security and privacy · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Versatile Vision-Language Model for 3D Computed TomographyabstractRepresentation learning serves as a foundational component of medical vision-language models (MVLMs), enabling cross-modal alignment, semantic consistency, and enhanced generalization capabilities for downstream tasks. As generalist models rapidly evolve, there is a pressing need to unify diverse downstream tasks, such as diagnosis, segmentation, report generation, and multiple choice within a cohesive framework, demanding more efficient and versatile visual representation learning. However, current MVLMs predominately follow CLIP-style vision pretraining, failing to leverage heterogeneous data resources with multi-dimensional imaging and diverse annotation forms. And there lacks systematic analysis of efficient vision encoder design across varied downstream applications, including diagnosis, segmentation, and text generation tasks, particularly for volumetric imaging like Computed Tomography (CT). Besides, current MVLMs exhibit constrained voxel-level capabilities, lacking effective multi-task instruction tuning framework capable of achieving robust performance across various downstream tasks. To address these challenges, we propose CTInstruct, a novel MVLM employing a hybrid ResNet-ViT encoder with multi-granular vision-language pretraining for efficient heterogeneous data modeling, and unified instruction tuning that jointly optimizes discriminative, generative, and voxel-level reasoning for volumetric medical imaging. CTInstruct achieves SOTA performance across 8 CT benchmarks, setting a new standard for data-efficient multimodal learning in medical imaging. Jiayu Lei, Ziqing Fan, Yanyong Zhang, Weidi Xie, Ya Zhang 0002, Yanfeng Wang 0001 |
AAAI | 3 |
| 2026 | Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent SpaceabstractThe rapid proliferation of Large Language Models (LLMs) has led to a fragmented and inefficient ecosystem, a state of ``model lock-in'' where seamlessly integrating novel models remains a significant bottleneck. Current routing frameworks require exhaustive, costly retraining, hindering scalability and adaptability. We introduce ZeroRouter, a new paradigm for LLM routing that breaks this lock-in. Our approach is founded on a universal latent space, a model-agnostic representation of query difficulty that fundamentally decouples the characterization of a query from the profiling of a model. This allows for zero-shot onboarding of new models without full-scale retraining. ZeroRouter features a context-aware predictor that maps queries to this universal space and a dual-mode optimizer that balances accuracy, cost, and latency. Our framework consistently outperforms all baselines, delivering higher accuracy at lower cost and latency. Wuyang Zhang, Ziyang Tao, Yanyong Zhang |
AAAI | 8 |
| 2026 | TransLLM: A Unified Multi-Task Large Language Model for Urban Transportation via Learnable PromptingabstractJiaming Leng, Yunying Bi, Chuan Qin, Zhenya Huang, Bing Yin, Haojie Ren, Yanyong Zhang, Chao Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiaming Leng, Yunying Bi, Chuan Qin 0002, Zhenya Huang, Haojie Ren, Yanyong Zhang, Chao Wang 0086 |
ACL (1) | 7 |
| 2026 | FlowGait: Enabling Robust Long-Term Gait Recognition Across Real-World Covariates with mmWave RadarabstractGait recognition enables proactive and personalized smart home interactions, but its long-term reliability is challenged by the non-static nature of gait. Covariates like carrying items and clothing induce a persistent domain shift that degrades traditional, static models. To solve this, we introduce FlowGait, a mmWave-based framework designed for robust, long-term adaptation. It combines self-training with continual learning, allowing the model to daily align with a user’s evolving gait by learning from readily available unlabeled data. It features a specialized transformer network for radar spectrogram analysis and a novel two-stage labeling algorithm that leverages the gait’s hierarchical nature to assign pseudo-labels to the unlabeled data accurately. Evaluated on three challenging datasets from 47 volunteers (covering 12 gait-covariates, 11 routes, and two weeks), FlowGait achieves high accuracies of 94.8 (cross-covariate), 98.6% (cross-route), and 95.5% (cross-day). Notably, for the long-term dataset, it reduced performance decay from 13.6% to just 1.4%, demonstrating its real-world robustness. Dequan Wang, Chenming He, Chengzhen Meng, Xiaoran Fan, Yanyong Zhang |
CHI | 6 |
| 2026 | FeelWave: Enabling Emotion-Aware Voice Interaction through Noise-Robust mmWave Emotion SensingabstractVoice has been a primary interaction mode with LLM-powered assistants. Beyond semantics, voice carries emotional cues with potential to guide empathetic system responses. Yet, robust vocal emotion sensing in noise and its use in optimizing interactions remain underexplored. In response, we present FeelWave, which achieves empathetic voice interaction through noise-robust mmWave emotion sensing and structured LLM prompts. It extracts robust vocal information from mmWave signals, applies audio-to-mmWave transfer learning for efficient emotion recognition, and employs chain-of-thought-based query optimization to enable emotion-adaptive responses. Evaluations show that FeelWave achieves 92.3% emotion recognition accuracy and remains robust in noisy environments, yielding a 62.9 percentage-point gain over audio-based models. In voice interaction studies, 74.3% of users prefer FeelWave, reporting significantly higher satisfaction than a baseline without emotion sensing (4.37 vs. 3.22). A SUS score of 88.3 confirms FeelWave’s high usability in real-world deployment. We hope this work will inspire more empathetic, user-centered AI-driven assistants. You Zuo, Dequan Wang, Chenming He, Chengzhen Meng, Xiaoran Fan, Yanyong Zhang |
CHI | 8 |
| 2026 | Rethinking Popularity Bias in Collaborative Filtering via Analytical Vector DecompositionabstractPopularity bias fundamentally undermines the personalization capabilities of collaborative filtering (CF) models, causing them to disproportionately recommend popular items while neglecting users' genuine preferences for niche content. While existing approaches treat this as an external confounding factor, we reveal that popularity bias is an intrinsic geometric artifact of Bayesian Pairwise Ranking (BPR) optimization in CF models. Through rigorous mathematical analysis, we prove that BPR systematically organizes item embeddings along a dominant "popularity direction" where embedding magnitudes directly correlate with interaction frequency. This geometric distortion forces user embeddings to simultaneously handle two conflicting tasks-expressing genuine preference and calibrating against global popularity-trapping them in suboptimal configurations that favor popular items regardless of individual tastes. We propose Directional Decomposition and Correction (DDC), a universally applicable framework that surgically corrects this embedding geometry through asymmetric directional updates. DDC guides positive interactions along personalized preference directions while steering negative interactions away from the global popularity direction, disentangling preference from popularity at the geometric source. Extensive experiments across multiple BPR-based architectures demonstrate that DDC significantly outperforms state-of-the-art debiasing methods, reducing training loss to less than 5% of heavily-tuned baselines while achieving superior recommendation quality and fairness. Code is available in https://github.com/LingFeng-Liu-AI/DDC. Yixin Song 0004, Dazhong Shen, Yanyong Zhang, Chao Wang 0086 |
KDD (1) | 6 |
| 2026 | RAG-Based Enterprise Knowledge Platform: Design and Evaluation
Zhijiang Zhang, Yanyong Zhang, Hao Zhou 0001 |
KSEM (2) | 2 |
| 2026 | Needle in a Haystack: Tracking UAVs from Massive Noise in Real-World 5G-A Base Station Data
Chengzhen Meng, Chenming He, Yidong Jiang, Xiaoran Fan, Dequan Wang, Jianmin Ji, Yanyong Zhang |
MobiSys | 8 |
| 2026 | MCLMR: A Model-Agnostic Causal Learning Framework for Multi-Behavior Recommendation
Ranxu Zhang, Junjie Meng, Ying Sun 0006, Ziqi Xu 0001, Yanyong Zhang, Chao Wang 0086 |
WWW | 7 |
| 2026 | Interpretable Brain MRI Report Generation Anchored by Lesion TopographyabstractRadiologists face increasing workloads that make accurate and timely report generation both critical and challenging. This paper presents a novel system for grounded automatic brain MRI report generation, with contributions in three key areas: First, we release RadGenome-Brain MRI, a benchmark dataset featuring multi-modal scans, expert-annotated abnormality masks, and radiology reports with region-level grounding to support fine-grained, explainable report generation. Second, we propose AutoRG-Brain, the first brain MRI report generation framework that combines automatic anomaly segmentation with a visual prompting-based language model to produce structured, anatomically grounded findings. Third, we conduct extensive quantitative and expert evaluations across segmentation and reporting tasks, and demonstrate in real clinical settings that our system significantly enhances junior radiologists' ability to detect subtle abnormalities and compose high-quality reports, narrowing the gap with senior doctors. All code, models, and datasets will be publicly released to facilitate future research and development. Jiayu Lei, Xiaoman Zhang, Chaoyi Wu, Lisong Dai, Ya Zhang 0002, Yanyong Zhang, Yanfeng Wang 0001, Weidi Xie |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Hard Sample Mining: A New Paradigm of Efficient and Robust Model TrainingabstractOver the past two decades, deep learning (DL) has achieved unprecedented breakthroughs across diverse application domains spanning computer vision (CV) to natural language processing (NLP). However, despite significant advances in computational resources and algorithmic frameworks, the training of deep neural networks continues to present formidable challenges due to persistent issues of training inefficiency and inherent data distribution biases. Recent years have witnessed the emergence of hard sample mining (HSM) as a promising paradigm to mitigate training inefficiencies and enhance model robustness through representative sample selection. Although HSM is reshaping contemporary AI research, its critical role in enabling efficient and robust model training has not yet been systematically explored. This article presents a comprehensive survey of HSM methodologies by: 1) establishing unified definitions of hard samples through rigorous sample complexity quantification criteria; 2) proposing a systematic taxonomy of HSM approaches with in-depth technical analysis; and 3) identifying pivotal research frontiers in this evolving field. This survey not only consolidates the foundations of HSM but also provides a roadmap for advancing efficient, robust, and generalizable deep learning models. Lei Liu 0073, Yunji Liang, Xiaokai Yan, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001, Yanyong Zhang, Daniel Dajun Zeng |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Language Adaptation of Large Language Models: An Empirical Study on LLaMA2abstractThere has been a surge of interest regarding language adaptation of Large Language Models (LLMs) to enhance the processing of texts in low-resource languages. While traditional language models have seen extensive research on language transfer, modern LLMs still necessitate further explorations in language adaptation. In this paper, we present a systematic review of the language adaptation process for LLMs, including vocabulary expansion, continued pre-training, and instruction fine-tuning, which focuses on empirical studies conducted on LLaMA2 and discussions on various settings affecting the model’s capabilities. This study provides helpful insights covering the entire language adaptation process, and highlights the compatibility and interactions between different steps, offering researchers a practical guidebook to facilitate the effective adaptation of LLMs across different languages. Yuexiang Xie, Bolin Ding, Jinyang Gao, Yanyong Zhang |
COLING | 5 |
| 2025 | RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera FusionabstractWe propose Radar-Camera fusion transformer (RaC-Former) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation-if the depth of pixels is not accurately estimated, the naive combination of BEV features actually integrates unaligned visual content. To avoid this problem, we propose a query-based framework that enables adaptive sampling of instance-relevant features from both the bird’s-eye view (BEV) and the original image view. Furthermore, we enhance system performance by two key designs: optimizing query initialization and strengthening the representational capacity of BEV. For the former, we introduce an adaptive circular distribution in polar coordinates to refine the initialization of object queries, allowing for a distance-based adjustment of query density. For the latter, we initially incorporate a radar-guided depth head to refine the transformation from image view to BEV. Subsequently, we focus on leveraging the Doppler effect of radar and introduce an implicit dynamic catcher to capture the temporal elements within the BEV. Extensive experiments on nuScenes and View-of-Delft (VoD) datasets validate the merits of our design. Remarkably, our method achieves superior results of 64.9% mAP and 70.2% NDS on nuScenes. RaCFormer also secures the state-of-the-art performance on the VoD dataset. Code is available at https://github.com/cxmomo/RaCFormer. Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Houqiang Li, Yanyong Zhang |
CVPR | 6 |
| 2025 | OccMamba: Semantic Occupancy Prediction with State Space ModelsabstractTraining deep learning models for semantic occupancy prediction is challenging due to factors such as a large number of occupancy cells, severe occlusion, limited visual cues, complicated driving scenarios, etc. Recent methods often adopt transformer-based architectures given their strong capability in learning input-conditioned weights and long-range relationships. However, transformer-based networks are notorious for their quadratic computation complexity, seriously undermining their efficacy and deployment in semantic occupancy prediction. Inspired by the global modeling and linear computation complexity of the Mamba architecture, we present the first Mamba-based network for semantic occupancy prediction, termed OccMamba. Specifically, we first design the hierarchical Mamba module and local context processor to better aggregate global and local contextual information, respectively. Besides, to relieve the inherent domain gap between the linguistic and 3D domains, we present a simple yet effective 3D-to-1D reordering scheme, i.e., height-prioritized 2D Hilbert expansion. It can maximally retain the spatial structure of 3D voxels as well as facilitate the processing of Mamba blocks. Endowed with the aforementioned designs, our OccMamba is capable of directly and efficiently processing large volumes of dense scene grids, achieving state-of-the-art performance across three prevalent occupancy prediction benchmarks, including OpenOccupancy, SemanticKITTI, and SemanticPOSS. Notably, on OpenOc-cupancy, our OccMamba outperforms the previous state-of-the-art Co-Occ by 5.1% IoU and 4.3% mIoU, respectively. Our implementation is open-sourced and available at: https://github.com/USTCLH/OccMamba. Yuenan Hou, Xiaohan Xing, Yuexin Ma, Xiao Sun 0001, Yanyong Zhang |
CVPR | 6 |
| 2025 | Killing Two Birds with One Stone: A Spatio-temporal Prompt for the Inductive Traffic Extrapolation
Leilei Ding, Zhipeng Tang, Le Zhang 0010, Dazhong Shen, Chao Wang 0086, Ziyang Tao, Jingbo Zhou 0003, Yanyong Zhang, Hui Xiong 0001 |
DASFAA (2) | 8 |
| 2025 | T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on EdgeabstractThe deployment of Large Language Models (LLMs) on edge devices is increasingly important to enhance on-device intelligence. Weight quantization is crucial for reducing the memory footprint of LLMs on devices. However, low-bit LLMs necessitate mixed precision matrix multiplication (mpGEMM) of low precision weights and high precision activations during inference. Existing systems, lacking native support for mpGEMM, resort to dequantize weights for high precision computation. Such an indirect way can lead to a significant inference overhead. Jianyu Wei, Shijie Cao, Ting Cao 0003, Lingxiao Ma, Lei Wang 0222, Yanyong Zhang, Mao Yang 0004 |
EuroSys | 6 |
| 2025 | GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping Under Flexible Language InstructionsabstractFlexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities of the large language models (LLMs) to establish mappings between expressions and targets, allowing robots to comprehend users' intentions in the instructions. However, the LLM's knowledge about objects' physical properties remains underexplored despite its tight relevance to grasping. In this work, we propose GraspCoT, a 6-DoF grasp detection framework that integrates a Chain-of-Thought (CoT) reasoning mechanism oriented to physical properties, guided by auxiliary question-answering (QA) tasks. Particularly, we design a set of QA templates to enable hierarchical reasoning that includes three stages: target parsing, physical property analysis, and grasp action selection. Moreover, GraspCoT presents a unified multimodal LLM architecture, which encodes multi-view observations of 3D scenes into 3D-aware visual tokens, and then jointly embeds these visual tokens with CoT-derived textual tokens within LLMs to generate grasp pose predictions. Furthermore, we present IntentGrasp, a large-scale benchmark that fills the gap in public datasets for multi-object grasp detection under diverse and indirect verbal commands. Extensive experiments on IntentGrasp demonstrate the superiority of our method, with additional validation in real-world robotic applications confirming its practicality. The code is available at https://github.com/cxmomo/GraspCoT. Xiaomeng Chu, Jiajun Deng, Guoliang You, Jianmin Ji, Yanyong Zhang |
ICCV | 7 |
| 2025 | SpatialSplat: Efficient Semantic 3D from Sparse Unposed ImagesabstractA major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods extend this paradigm by associating each primitive with a compressed semantic feature vector. However, these methods have two major limitations: (a) the naively compressed feature compromises expressiveness, affecting the model's ability to capture fine-grained semantics, and (b) the pixel-wise primitive prediction introduces redundancy in overlapping areas, causing unnecessary memory overhead. To this end, we introduce \textbf{SpatialSplat}, a feedforward framework that produces redundancy-aware Gaussians and capitalizes on a dual-field semantic representation. Particularly, with the insight that primitives within the same instance exhibit high semantic consistency, we decompose the semantic representation into a coarse feature field that encodes uncompressed semantics with minimal primitives, and a fine-grained yet low-dimensional feature field that captures detailed inter-instance relationships. Moreover, we propose a selective Gaussian mechanism, which retains only essential Gaussians in the scene, effectively eliminating redundant primitives. Our proposed Spatialsplat learns accurate semantic information and detailed instances prior with more compact 3D Gaussians, making semantic 3D reconstruction more applicable. We conduct extensive experiments to evaluate our method, demonstrating a remarkable 60\% reduction in scene representation parameters while achieving superior performance over state-of-the-art methods. The code is available at https://github.com/shengyuuu/SpatialSplat.git Yu Sheng, Jiajun Deng, Yu Zhang 0086, Bei Hua, Yanyong Zhang, Jianmin Ji |
ICCV | 6 |
| 2025 | InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation
Zhuoran Yang, Chenjing Ding, Chiyu Wang, Wei Wu 0021, Yanyong Zhang |
ICCV | 6 |
| 2025 | S3R-GS: Streamlining the Pipeline for Large-Scale Street Scene ReconstructionabstractRecently, 3D Gaussian Splatting (3DGS) has reshaped the field of photorealistic 3D reconstruction, achieving impressive rendering quality and speed. However, when applied to large-scale street scenes, existing methods suffer from rapidly escalating per-viewpoint reconstruction costs as scene size increases, leading to significant computational overhead. After revisiting the conventional pipeline, we identify three key factors accounting for this issue: unnecessary local-to-global transformations, excessive 3D-to-2D projections, and inefficient rendering of distant content. To address these challenges, we propose S3R-GS, a 3DGS framework that Streamlines the pipeline for large-scale Street Scene Reconstruction, effectively mitigating these limitations. Moreover, most existing street 3DGS methods rely on ground-truth 3D bounding boxes to separate dynamic and static components, but 3D bounding boxes are difficult to obtain, limiting real-world applicability. To address this, we propose an alternative solution with 2D boxes, which are easier to annotate or can be predicted by off-the-shelf vision foundation models. Such designs together make S3R-GS readily adapt to large, in-the-wild scenarios. Extensive experiments demonstrate that S3R-GS enhances rendering quality and significantly accelerates reconstruction. Remarkably, when applied to videos from the challenging Argoverse2 dataset, it achieves state-of-the-art PSNR and SSIM, reducing reconstruction time to below 50%--and even 20%--of competing methods. Guangting Zheng, Jiajun Deng, Xiaomeng Chu, Houqiang Li, Yanyong Zhang |
ICCV | 6 |
| 2025 | Modalities Contribute Unequally: Enhancing Medical Multi-modal Learning through Adaptive Modality Token Re-balancingabstractMedical multi-modal learning requires an effective fusion capability of various heterogeneous modalities. One vital challenge is how to effectively fuse modalities when their data quality varies across different modalities and patients. For example, in the TCGA benchmark, the performance of the same modality can differ between types of cancer. Moreover, data collected at different times, locations, and with varying reagents can introduce inter-modal data quality differences ($i.e.$, $\textbf{Modality Batch Effect}$). In response, we propose ${\textbf{A}}$daptive ${\textbf{M}}$odality Token Re-Balan${\textbf{C}}$ing ($\texttt{AMC}$), a novel top-down dynamic multi-modal fusion approach. The core of $\texttt{AMC}$ is to quantify the significance of each modality (Top) and then fuse them according to the modality importance (Down). Specifically, we access the quality of each input modality and then replace uninformative tokens with inter-modal tokens, accordingly. The more important a modality is, the more informative tokens are retained from that modality. The self-attention will further integrate these mixed tokens to fuse multi-modal knowledge. Comprehensive experiments on both medical and general multi-modal datasets demonstrate the effectiveness and generalizability of $\texttt{AMC}$. Jie Peng 0002, Jenna L. Ballard, Mohan Zhang, Sukwon Yun, Jiayi Xin, Qi Long, Yanyong Zhang, Tianlong Chen 0001 |
ICML | 7 |
| 2025 | Hierarchical Masked Autoregressive Models with Low-Resolution Token PivotsabstractAutoregressive models have emerged as a powerful generative paradigm for visual generation. The current de-facto standard of next token prediction commonly operates over a single-scale sequence of dense image tokens, and is incapable of utilizing global context especially for early tokens prediction. In this paper, we introduce a new autoregressive design to model a hierarchy from a few low-resolution image tokens to the typical dense image tokens, and delve into a thorough hierarchical dependency across multi-scale image tokens. Technically, we present a Hierarchical Masked Autoregressive models (Hi-MAR) that pivot on low-resolution image tokens to trigger hierarchical autoregressive modeling in a multi-phase manner. Hi-MAR learns to predict a few image tokens in low resolution, functioning as intermediary pivots to reflect global structure, in the first phase. Such pivots act as the additional guidance to strengthen the next autoregressive modeling phase by shaping global structural awareness of typical dense image tokens. A new Diffusion Transformer head is further devised to amplify the global context among all tokens for mask token prediction. Extensive evaluations on both class-conditional and text-to-image generation tasks demonstrate that Hi-MAR outperforms typical AR baselines, while requiring fewer computational costs. Guangting Zheng, Yehao Li, Yingwei Pan, Jiajun Deng, Ting Yao 0003, Yanyong Zhang, Tao Mei 0001 |
ICML | 6 |
| 2025 | CELLmap: Enhancing LiDAR SLAM Through Elastic and Lightweight Spherical Map RepresentationabstractSLAM is a fundamental capability of unmanned systems, with LiDAR-based SLAM gaining widespread adoption due to its high precision. Current SLAM systems can achieve centimeter-level accuracy within a short period. However, there are still several challenges when dealing with largescale mapping tasks including significant storage requirements and difficulty of reusing the constructed maps. To address this, we first design an elastic and lightweight map representation called CELLmap, composed of several CELLS, each representing the local map at the corresponding location. Then, we design a general backend including CELL-based bidirectional registration module and loop closure detection module to improve global map consistency. Our experiments have demonstrated that CELLmap can represent the precise geometric structure of large-scale maps of KITTI dataset using only about 60 MB. Additionally, our general backend achieves up to a 26.88% improvement over various LiDAR odometry methods. Yifan Duan, Yao Li 0016, Guoliang You, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang |
ICRA | 7 |
| 2025 | OG-Gaussian: Occupancy Based Street Gaussians for Autonomous DrivingabstractAccurate and realistic 3D scene reconstruction enables the lifelike creation of autonomous driving simulation environments. With advancements in 3D Gaussian Splatting (3DGS), previous studies have applied it to reconstruct complex dynamic driving scenes. These methods typically require expensive LiDAR sensors and pre-annotated datasets of dynamic objects. To address these challenges, we propose OG-Gaussian, a novel approach that replaces LiDAR point clouds with Occupancy Grids (OGs) generated from surround-view camera images using Occupancy Prediction Network (ONet). Our method leverages the semantic information in OGs to separate dynamic vehicles from static street background, converting these grids into two distinct sets of initial point clouds for reconstructing both static and dynamic objects. Additionally, we estimate the trajectories and poses of dynamic objects through a learning-based approach, eliminating the need for complex manual annotations. Experiments on Waymo Open dataset demonstrate that OG-Gaussian is on par with the current state-of-the-art in terms of reconstruction quality and rendering speed, achieving an average PSNR of 35.13 and a rendering speed of 143 FPS, while significantly reducing computational costs and economic overhead. Yedong Shen, Yifan Duan, Yilong Wu, Jianmin Ji, Yanyong Zhang, Huiqing Jin |
ICRA | 8 |
| 2025 | MT-PCR: Leveraging Modality Transformation for Large-Scale Point Cloud Registration with Limited OverlapabstractLarge-scale scene point cloud registration with limited overlap is a challenging task due to computational load and constrained data acquisition. To tackle these issues, we propose a point cloud registration method, MT-PCR, based on Modality Transformation. MT-PCR leverages a Bird's Eye View (BEV) capturing the maximal overlap information to improve the accuracy and utilizes images to provide complementary spatial features. Specifically, MT-PCR converts 3D point clouds to BEV images and estimates correspondence by 2D image keypoints extraction and matching. Subsequently, the 2D correspondence estimates are then transformed back to 3D point clouds using inverse mapping. We have applied MT-PCR to Terrestrial Laser Scanning (TLS) and Aerial Laser Scanning (ALS) point cloud registration on the GrAco dataset, involving 8 low-overlap, square-kilometer scale registration scenarios. Experiments and comparisons with commonly used methods demonstrate that MT-PCR can achieve superior accuracy and robustness in large-scale scenes with limited overlap. Yilong Wu, Yifan Duan, Yedong Shen, Jianmin Ji, Yanyong Zhang |
ICRA | 7 |
| 2025 | Large-Scale Gaussian Splatting SLAMabstractThe recently developed Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have shown encour-aging and impressive results for visual SLAM. However, most representative methods require RGBD sensors and are only available for indoor environments. The robustness of reconstruction in largescale outdoor scenarios remains unexplored. This paper introduces a large-scale 3DGS-based visual SLAM with stereo cameras, termed LSG-SLAM. The proposed LSG-SLAM employs a multi-modality strategy to estimate prior poses under large view changes. In tracking, we introduce feature-alignment warping constraints to alleviate the adverse effects of appearance similarity in rendering losses. For the scalability of large-scale scenarios, we introduce continuous Gaussian Splatting submaps to tackle unbounded scenes with limited memory. Loops are detected between GS sub maps by place recognition and the relative pose between looped keyframes is optimized utilizing rendering and feature warping losses. After the global optimization of camera poses and Gaussian points, a structure refinement module enhances the reconstruction quality. With extensive evaluations on the EuRoc and KITTI datasets, LSG-SLAM achieves superior performance over existing Neural, 3DGS-based, and even traditional approaches. Project page: https://lsg-slam.github.io. Zhe Xin, Penghui Huang, Yanyong Zhang, Yinian Mao, Guoquan Huang 0003 |
ICRA | 4 |
| 2025 | CAFE-AD: Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous DrivingabstractImitation learning based planning tasks on the nuPlan dataset have gained great interest due to their potential to generate human-like driving behaviors. However, open-loop training on the nuPlan dataset tends to cause causal confusion during closed-loop testing, and the dataset also presents a longtail distribution of scenarios. These issues introduce challenges for imitation learning. To tackle these problems, we introduce CAFE-AD, a Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous Driving method, designed to enhance feature representation across various scenario types. We develop an adaptive feature pruning module that ranks feature importance to capture the most relevant information while reducing the interference of noisy information during training. Moreover, we propose a cross-scenario feature interpolation module that enhances scenario information to introduce diversity, enabling the network to alleviate overfitting in dominant scenarios. We evaluate our method CAFEAD, on the challenging public nuPlan Test14-Hard closed-loop simulation benchmark. The results demonstrate that CAFEAD outperforms state-of-the-art methods including rule-based and hybrid planners, and exhibits the potential in mitigating the impact of long-tail distribution within the dataset. Additionally, we further validate its effectiveness in real-world environments. The code and models will be made available at https://github.com/AlniyatRui/CAFE-AD. Junrui Zhang 0012, Chenjie Wang, Jie Peng 0002, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang |
ICRA | 7 |
| 2025 | Improving Efficiency of Answer Set Planning with Rough Solutions from Large Language Models for Robotic Task PlanningabstractAnswer Set Programming (ASP) planning can be used to refine the rough solutions generated by Large Language Models (LLMs) to handle specific restrictions of actions, i.e., reconstruct the rough solutions to be executable, for robotic task planning. However, it is still challenging to efficiently solve ASP programs that have multiple variables with large domains, which prevents the above application of ASP planning from real-world task planning problems. In this paper, we consider how to reduce the domains of variables without losing possible solutions for ASP planning, while given these rough solutions from LLMs. Based on the above reduction, we introduce CLMASP, an approach that couples LLMs with ASP for robotic task planning. We evaluate CLMASP on the VirtualHome platform for common indoor tasks, demonstrating a significant improvement in the executable rate from under 10% to nearly 90% and reducing average ASP planning time from over 2 hours to under 5 seconds. Code is available at https://github.com/CLMASP/CLMASP. Xinrui Lin, Yangfan Wu, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
IJCAI | 7 |
| 2025 | CalibWorkflow: A General MLLM-Guided Workflow for Centimeter-Level Cross-Sensor CalibrationabstractExtrinsic calibration is a fundamental step in sensor fusion systems. However, existing methods often lack generalization capabilities when facing diverse hardware configurations, sensor poses, and environmental conditions, hindering their large-scale deployment. To address this limitation, we propose a general extrinsic calibration method, CalibWorkflow. Our core innovation lies in positioning multimodal large language models (MLLMs) as ''visual guides'' for the calibration process, leveraging their powerful vision-language understanding capabilities to guide parameter search and refinement. This reliance on visual scene understanding, rather than specific geometric features or sensor characteristics, enables the method to generalize effectively across diverse hardware and environmental conditions. Specifically, CalibWorkflow employs a three-stage calibration pipeline: initial parameter search, coarse optimization, and fine optimization. First, it utilizes the MLLM to assess the visual consistency between the projected point cloud and the image, rapidly determining an initial range for the extrinsic parameters. Next, the MLLM serves as a differential evaluator, giving simple ''better'' or ''worse'' feedback on parameter changes to guide the search through the parameter space. Finally, the method refines the calibration by matching edge features and performing non-linear optimization. Extensive experiments are conducted across six diverse scenarios and four heterogeneous sensor combinations. CalibWorkflow achieves state-of-the-art sub-degree and centimeter-level accuracy on four datasets and demonstrates highly competitive performance on others. These results thoroughly validate the generalization and robustness when facing various scenarios. Codes will be available. Wuyang Zhang, Guoliang You, Xiaomeng Chu, Wenhao Yu 0010, Yifan Duan, Yanyong Zhang |
ACM Multimedia | 8 |
| 2025 | VLMPlanner: Integrating Visual Language Models with Motion PlanningabstractIntegrating large language models (LLMs) into autonomous driving motion planning has recently emerged as a promising direction, offering enhanced interpretability, better controllability, and improved generalization in rare and long-tail scenarios. However, existing methods often rely on abstracted perception or map-based inputs, missing crucial visual context, such as fine-grained road cues, accident aftermath, or unexpected obstacles, which are essential for robust decision-making in complex driving environments. To bridge this gap, we propose VLMPlanner, a hybrid framework that combines a learning-based real-time planner with a vision-language model (VLM) capable of reasoning over raw images. The VLM processes multi-view images to capture rich, detailed visual information and leverages its common-sense reasoning capabilities to guide the real-time planner in generating robust and safe trajectories. Furthermore, we develop the Context-Adaptive Inference Gate (CAI-Gate) mechanism that enables the VLM to mimic human driving behavior by dynamically adjusting its inference frequency based on scene complexity, thereby achieving an optimal balance between planning performance and computational efficiency. We evaluate our approach on the large-scale, challenging nuPlan benchmark, with comprehensive experimental results demonstrating superior planning performance in scenarios with intricate road conditions and dynamic elements. Zhipeng Tang, Sha Zhang 0002, Jiajun Deng, Chenjie Wang, Guoliang You, Xinrui Lin, Yanyong Zhang |
ACM Multimedia | 8 |
| 2025 | PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum
Sha Zhang 0002, Jiajun Deng, Yedong Shen, Yanyong Zhang |
ACM Multimedia | 6 |
| 2025 | Ghost Points Matter: Far-Range Vehicle Detection with a Single mmWave Radar in TunnelabstractVehicle detection in tunnels is crucial for traffic monitoring and accident response, yet remains underexplored. In this paper, we develop mmTunnel, a millimeter-wave radar system that achieves far-range vehicle detection in tunnels. The main challenge here is coping with ghost points caused by multi-path reflections, which lead to severe localization errors and false alarms. Instead of merely removing ghost points, we propose correcting them to true vehicle positions by recovering their signal reflection paths, thus reserving more data points and improving detection performance, even in occlusion scenarios. However, recovering complex 3D reflection paths from limited 2D radar points is highly challenging. To address this problem, we develop a multi-path ray tracing algorithm that leverages the ground plane constraint and identifies the most probable reflection path based on signal path loss and spatial distance. We also introduce a curve-to-plane segmentation method to simplify tunnel surface modeling such that we can significantly reduce the computational delay and achieve real-time processing. Chenming He, Chengzhen Meng, Xiaoran Fan, Dequan Wang, Haojie Ren, Jianmin Ji, Yanyong Zhang |
MobiCom | 8 |
| 2025 | UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous DrivingabstractThe rapid advancements in autonomous driving have introduced increasingly complex, real-time GPU-bound tasks critical for reliable vehicle operation. However, the proprietary nature of these autonomous systems and closed-source GPU drivers hinder fine-grained control over GPU executions, often resulting in missed deadlines that compromise vehicle performance. To address this, we present UrgenGo, a non-intrusive, urgency-aware GPU scheduling system that operates without access to application source code. UrgenGo implicitly prioritizes GPU executions through transparent kernel launch manipulation, employing task-level stream binding, delayed kernel launching, and batched kernel launch synchronization. We conducted extensive real-world evaluations in collaboration with a self-driving startup, developing 11 GPU-bound task chains for a realistic autonomous navigation application and implementing our system on a self-driving bus. Our results show a significant 61% reduction in the overall deadline miss ratio, compared to the state-of-the-art GPU scheduler that requires source code modifications. Hanqi Zhu, Wuyang Zhang, Ziyang Tao, Xinrui Lin, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
MobiCom | 8 |
| 2025 | UniSense: Spatial-Uncertainty-Aware Collaborative Sensing for Autonomous DrivingabstractVehicle-to-vehicle collaborative perception faces fundamental deployment barriers: raw LiDAR data sharing requires over 300 Mbps per vehicle - far exceeding V2X network capacities, while network delays of 80-200ms create dangerous temporal misalignments at highway speeds. We present UniSense, a distributed collaborative perception system that enables efficient and reliable multi-vehicle perception through uncertainty-driven sensor data exchange. Instead of sharing raw sensor data, vehicles exchange compact uncertainty maps that identify regions requiring additional perceptual information. Our key innovations include: (1) a lightweight uncertainty quantification pipeline that runs in real-time on automotive hardware, identifying perception-critical regions while reducing bandwidth requirements by more than 10×, (2) a bandwidth-aware protocol that dynamically adapts data sharing based on network conditions and perception uncertainty, and (3) a selective motion compensation scheme that maintains temporal consistency. We evaluate UniSense through a year-long deployment with 16 roadside LiDAR nodes and autonomous vehicles across our campus. Our experimental results show that UniSense extends reliable perception range from local 80m to 140m, improving accuracy by 1.33× on average, up to 1.73×, over the state-of-the-art baselines, under communication constraints. The code and dataset are available at https://github.com/LetStarFly/UniSense. Haojie Ren, Wuyang Zhang, Shuyao Shi, Yanyong Zhang |
MobiSys | 6 |
| 2025 | FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM TokensabstractRecently, large language models (LLMs) have been explored for integration with collaborative filtering (CF)-based recommendation systems, which are crucial for personalizing user experiences. However, a key challenge is that LLMs struggle to interpret the latent, non-semantic embeddings produced by CF approaches, limiting recommendation effectiveness and further applications. To address this, we propose FACE, a general interpretable framework that maps CF embeddings into pre-trained LLM tokens. Specifically, we introduce a disentangled projection module to decompose CF embeddings into concept-specific vectors, followed by a quantized autoencoder to convert continuous embeddings into LLM tokens (descriptors). Then, we design a contrastive alignment objective to ensure that the tokens align with corresponding textual signals. Hence, the model-agnostic FACE framework achieves semantic alignment without fine-tuning LLMs and enhances recommendation performance by leveraging their pre-trained capabilities. Empirical results on three real-world recommendation datasets demonstrate performance improvements in benchmark models, with interpretability studies confirming the interpretability of the descriptors. Code is available in \url{https://github.com/YixinRoll/FACE}. Chao Wang 0086, Yixin Song 0004, Jinhui Ye, Chuan Qin 0002, Dazhong Shen, Yanyong Zhang |
NeurIPS | 8 |
| 2025 | GS-Share: Enabling High-fidelity Map Sharing with Incremental Gaussian SplattingabstractAbstract Constructing and sharing 3D maps is essential for many applications, including autonomous driving and augmented reality. Recently, 3D Gaussian splatting has emerged as a promising approach for accurate 3D reconstruction. However, a practical map‐sharing system that features high‐fidelity, continuous updates, and network efficiency remains elusive. To address these challenges, we introduce GS‐Share, a photorealistic map‐sharing system with a compact representation. The core of GS‐Share includes anchor‐based global map construction, virtual‐image‐based map enhancement, and incremental map update. We evaluate GS‐Share against state‐of‐the‐art methods, demonstrating that our system achieves higher fidelity, particularly for extrapolated views, with improvements of 11%, 22%, and 74% in PSNR, LPIPS, and Depth L1, respectively. Furthermore, GS‐Share is significantly more compact, reducing map transmission overhead by 36%. Hanqi Zhu, Yifan Duan, Yanyong Zhang |
Comput. Graph. Forum | 4 |
| 2025 | OA-DET3D: Embedding Object Awareness As A General Plug-in for Multi-Camera 3D Object Detection
Xiaomeng Chu, Jiajun Deng, Jianmin Ji, Yu Zhang 0086, Houqiang Li, Yanyong Zhang |
Int. J. Comput. Vis. | 6 |
| 2025 | $\gamma$-Razor: Hardness-Aware Dataset Pruning for Efficient Neural Network TrainingabstractTraining deep neural networks (DNNs) on large-scale datasets is often inefficient with large computational needs and significant energy consumption. Although great efforts have been taken to optimize DNNs, few studies focused on the inefficiency caused by the data samples with less value for model training. In this article, we empirically demonstrate that sample complexity is important for model efficiency and selecting representative samples is constructive to the model efficiency. In particular, we propose hardness-aware dataset pruning method ($\gamma$-Razor) to select representative samples from large-scale datasets to remove the less valuable data samples for model training.$\gamma$-Razor is a two-stage framework that includes interclass sampling and intraclass sampling. First, we introduce the inverse self-paced learning strategy to learn hard samples and adjust their weights adaptively according to the inverse frequency of effective samples of each class. For intraclass sampling, hardness-aware cluster sampling algorithm is proposed to downsample easy samples within each class. To evaluate the performance of$\gamma$-Razor, we conducted extensive experiments on three large-scale datasets for image classification tasks. The experimental results show that models trained with the pruned datasets show competitive performances against their counterparts trained with the original large-scale datasets in terms of robustness and efficiency. Furthermore, models trained with the pruned datasets converge faster with lower energy consumption. Lei Liu 0073, Peng Zhang 0139, Yunji Liang, Lia Morra, Bin Guo 0001, Zhiwen Yu 0001, Yanyong Zhang, Daniel Dajun Zeng |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2025 | Blockchain-Enabled Secure Offloading for VEC: A Multi-Agent Reinforcement Learning ApproachabstractVehicular edge computing (VEC) helps improve the task computational performance of vehicles on roads but has difficulty in defending against eavesdropping and selfish attacks simultaneously. In this paper, we design a reputation-based smart contract with blockchain and propose a multi-agent reinforcement learning (RL) based secure offloading scheme for VEC against both eavesdropping and selfish attacks. This scheme has a three-level hierarchical structure for each vehicle and uses the reputations obtained from the blockchain as the basis to optimize the edge node selection, offloading ratio, and power allocation, which aims to reduce the task computational latency, the vehicle energy consumption and eavesdropping rate. By using a punishment function based on the constraints, this scheme avoids exploring dangerous policies that can cause task failure or severe data leakage. A multi-agent deep RL-based secure offloading scheme is proposed for vehicles with sufficient resources, which evaluates the long-term risk rather than the punishment function to further improve the secure offloading performance. The regret bound is analyzedand the cumulative reward upper bound is provided. Simulation results verify the effectiveness of our schemes as compared with the benchmark. Xiaozhen Lu, Liang Xiao 0003, Yilin Xiao 0001, Zehui Xiong, Zhe Liu 0001, Yanyong Zhang, Weihua Zhuang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Edge-Assisted Collaborative Perception Against Jamming and Interference in Vehicular NetworksabstractCollaborative perception of connected autonomous vehicles (CAVs) that offload the sensing data, such as the feature map extracted from light detection and ranging (LiDAR) point clouds, to an edge device such as the roadside unit (RSU) to detect traffic objects has severe performance degradation due to the offloading latency and packet loss rate (PLR) under jamming and interference. In this paper, we propose an edge-assisted reinforcement learning (RL)-based collaborative perception scheme for CAVs to enhance the accuracy and speed against jamming and interference in LiDAR-based object detection. Based on the spatial confidence score of the feature map, the data size, the channel gains, the received jamming power and interference level, this scheme chooses the critical regions of the feature map, radio channel and transmit power with the hierarchical structure to enhance the learning efficiency. The risk level of the selected policy evaluates the time asynchronization and information loss of the shared feature map using the multi-level risk function based on multiple thresholds of the offloading latency and PLR, with assigning different penalties to mitigate the selection of high-risk policies that degrade perception performance. The upper performance bound in terms of the perception accuracy, latency and utility is provided based on the Stackelberg equilibrium of the game between the jammer and CAVs. Experimental results based on the Robosense RS-LiDAR-16 sensors and the Raspberry Pi to detect 10 vehicles in an$8.5\times 4\times 3.5$m3area show the performance gain with 22.4% higher perception accuracy and 41.3% less latency compared with the benchmark against a smart jammer. Zhiping Lin 0002, Liang Xiao 0003, Zefang Lv, Yunjun Zhu, Yanyong Zhang, Yong-Jin Liu 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2024 | SDAC: A Multimodal Synthetic Dataset for Anomaly and Corner Case Detection in Autonomous DrivingabstractNowadays, closed-set perception methods for autonomous driving perform well on datasets containing normal scenes. However, they still struggle to handle anomalies in the real world, such as unknown objects that have never been seen while training. The lack of public datasets to evaluate the model performance on anomaly and corner cases has hindered the development of reliable autonomous driving systems. Therefore, we propose a multimodal Synthetic Dataset for Anomaly and Corner case detection, called SDAC, which encompasses anomalies captured from multi-view cameras and the LiDAR sensor, providing a rich set of annotations for multiple mainstream perception tasks. SDAC is the first public dataset for autonomous driving that categorizes anomalies into object, scene, and scenario levels, allowing the evaluation under different anomalous conditions. Experiments show that closed-set models suffer significant performance drops on anomaly subsets in SDAC. Existing anomaly detection methods fail to achieve satisfactory performance, suggesting that anomaly detection remains a challenging problem. We anticipate that our SDAC dataset could foster the development of safe and reliable systems for autonomous driving. Yu Zhang 0086, Yingqing Xia, Yanyong Zhang, Jianmin Ji |
AAAI | 4 |
| 2024 | Graph-Specific Schema-Guided Query Optimization
Chaijun Xu, Yunlong Liang, Yu Zhang 0086, Hairong Hu, Yanyong Zhang |
DASFAA (1) | 8 |
| 2024 | Agent3D-Zero: An Agent for Zero-Shot 3D Understanding
Sha Zhang 0002, Jiajun Deng, Shixiang Tang, Wanli Ouyang, Tong He 0001, Yanyong Zhang |
ECCV (19) | 7 |
| 2024 | Reinforcement Learning Based Collaborative Perception for Vehicular NetworksabstractReinforcement learning (RL)-based collaborative perception in vehicular networks chooses the sub-frame of radio channel resources for connected autonomous vehicles (CAVs) to exchange sensing data to enhance the perception performance, but leads to inaccurate detection in the light detection and ranging (LiDAR)-based object detection due to the asynchronous scan period of the LiDAR point clouds. This paper proposes a RL-based collaborative perception scheme to choose the transmit power and sub-frame to share the feature maps extracted from point clouds. Based on the estimated packet timestamp, the network topology and the channel gains among CAVs, this scheme enhances the perception accuracy and latency against path-loss and interference. The collaborative risk in the policy distribution is formulated as a weighted sum of the perception latency and packet loss rate to avoid the time asynchronization and information loss of the feature map exchange. The performance bound of the perception accuracy and latency is provided based on a Nash equilibrium of the cooperative game among CAVs. Simulation results based on five CAVs show the performance gain of the perception accuracy and latency over the benchmarks. Zhiping Lin 0002, Yunjun Zhu, Jieling Li, Liang Xiao 0003, Yuliang Tang, Yanyong Zhang |
GLOBECOM | 7 |
| 2024 | OmniCache: A Unified Cache for Efficient Query Handling in LSM-tree Based Key-Value StoresabstractKey-value (KV) stores built on Log-Sturctured Merge (LSM) trees have become fundamental in handling large-scale data in modern applications. While LSM-trees are optimized for high write throughput, they suffer from degraded read performance. Caching is one of the main techniques to improve the performance of read operations. Traditional block caches have several issues: 1. They store redundant data for point lookups by caching entire blocks even when only a single KV pair is needed. 2. Frequent compaction operations cause cached blocks to become invalid, leading to increased disk I/Os and higher costs. 3. Block cache does not store global ordering information, so for range queries, the system must repeatedly fetch, sort, and merge data from multiple blocks to generate results, leading to substantial overhead.In this paper, we introduce OmniCache, a unified cache system that stores the results of both point lookups and range queries in LSM-tree based KV stores to optimize read performance. Omni-Cache employs a hybrid data structure, combining a hash table and skip-list, to dynamically store fine-grained results with global ordering. Our system reduces redundant data, mitigates cache invalidation from compaction, and improves query performance. Experimental evaluations show that OmniCache avoids 39% disk read traffic by reducing block cache invalidation. OmniCache also improves throughput up to 2.72x in read-heavy workloads and up to 1.85x in write-heavy workloads. Yiyang Geng, Huai Xu, Yanyong Zhang, Fuxin Zhang |
HPCC | 3 |
| 2024 | BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D SegmentationabstractRecent research has demonstrated the advantages of Bird’s-eye-view (BEV) representation in the field of 3D perception. However, due to the lack of height information, BEV representation alone is insufficient to accurately reconstruct the complete surrounding 3D scene. On the other hand, voxel representation excels in describing 3D structures, but their memory and computational cost pose challenges for fast inference. To tackle these limitations, we propose an innovative method dubbed BEVoxSeg, which leverages the computational efficiency of BEV methods while incorporating essential geometric information from voxel features. By combining the advantages from both representations, our approach achieved state-of-the-art results for LiDAR semantic segmentation on nuScenes and demonstrated a superior performance in the occupancy prediction tasks on Occ3D-nuScenes dataset. Jianmin Ji, Yanyong Zhang |
ICASSP | 5 |
| 2024 | Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint OptimizationabstractIn multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adaptive ASR framework that possesses more implicit target speaker enhancement capability by efficiently joint-optimized speaker recognition (SR) and ASR models. Our framework introduces sharing self-supervised learning representation, optimization transfer and hierarchy speaker-gated attention. In this manner, it can maximize effectiveness of embedding bias and emphasize target speaker corresponding to semantic units. In the CHiME-7 DASR sub-track, the proposed method achieves a 28.19% relative reduction in word error rate (WER) on the development sets when compared to the official baseline. Notably, this framework has also been employed in the champion system for the CHiME-7 DASR. Haitao Tang 0001, Jiahuan Fan, Ruoyu Wang 0029, Hang Chen 0001, Yanyong Zhang, Jun Du 0002, Hengshun Zhou, Lei Sun 0010, Tian Gao 0005, Genshun Wan, Jianqing Gao |
ICASSP | 6 |
| 2024 | OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous DrivingabstractVisual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupancy, thereby circumventing the traditional need for concurrent estimation of ego poses and landmark locations. Within this framework, we utilize the TPV-Former to convert surround view cameras’ images into 3D semantic occupancy. Addressing the challenges presented by this transformation, we have specifically tailored a pose estimation and mapping algorithm that incorporates Semantic Label Filter, Dynamic Object Filter, and finally, utilizes Voxel PFilter for maintaining a consistent global semantic map. Evaluations on the Occ3D-nuScenes not only showcase a 20.6% improvement in Success Ratio and a 29.6% enhancement in trajectory accuracy against ORB-SLAM3, but also emphasize our ability to construct a comprehensive map. Our implementation is open-sourced and available at: https://github.com/USTCLH/OCC-VO. Yifan Duan, Jianmin Ji, Yanyong Zhang |
ICRA | 6 |
| 2024 | CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration NetworkabstractThe fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different data modalities. Previously, many calibration methods involved specific targets and/or manual intervention, which has proven to be cumbersome and costly. Learning-based online calibration methods have been proposed, but their performance is barely satisfactory in most cases. These methods usually suffer from issues such as sparse feature maps, unreliable cross-modality association, inaccurate calibration parameter regression, etc. In this paper, to address these issues, we propose CalibFormer, an end-to-end network for automatic LiDAR-camera calibration. We aggregate multiple layers of camera and LiDAR image features to achieve high-resolution representations. A multi-head correlation module is utilized to identify correlations between features more accurately. Lastly, we employ transformer architectures to estimate accurate calibration parameters from the correlation information. Our method achieved a mean translation error of 0.8751cm and a mean rotation error of 0.0562° on the KITTI dataset, surpassing existing state-of-the-art methods and demonstrating strong robustness, accuracy, and generalization capabilities. Yao Li 0016, Chengzhen Meng, Jianmin Ji, Yanyong Zhang |
ICRA | 6 |
| 2024 | DGR: A General Graph Desmoothing Framework for Recommendation via Global and Local Perspectives
Leilei Ding, Dazhong Shen, Chao Wang 0086, Tianfu Wang 0002, Le Zhang 0010, Yanyong Zhang |
IJCAI | 6 |
| 2024 | LDP: A Local Diffusion Planner for Efficient Robot Navigation and Collision AvoidanceabstractThe conditional diffusion model has been demonstrated as an efficient tool for learning robot policies, owing to its advancement to accurately model the conditional distribution of policies. The intricate nature of real-world scenarios, characterized by dynamic obstacles and maze-like structures, underscores the complexity of robot local navigation decision-making as a conditional distribution problem. Nevertheless, leveraging the diffusion model for robot local navigation is not trivial and encounters several under-explored challenges: (1) Data Urgency The complex conditional distribution in local navigation needs training data to include diverse policy in diverse real-world scenarios; (2) Myopic Observation Due to the diversity of the perception scenarios, diffusion decisions based on the local perspective of robots may prove suboptimal for completing the entire task, as they often lack foresight. In certain scenarios requiring detours, the robot may become trapped. To address these issues, our approach begins with an exploration of a diverse data generation mechanism that encompasses multiple agents exhibiting distinct preferences through target selection informed by integrated global-local insights. Then, based on this diverse training data, a diffusion agent is obtained, capable of excellent collision avoidance in diverse scenarios. Subsequently, we augment our Local Diffusion Planner, also known as LDP by incorporating global observations in a lightweight manner. This enhancement broadens the observational scope of LDP, effectively mitigating the risk of becoming ensnared in local optima and promoting more robust navigational decisions. Our experimental results demonstrated that the LDP outperforms other baseline algorithms in navigation performance, exhibiting enhanced robustness across diverse scenarios with different policy preferences and superior generalization capabilities for unseen scenarios. Moreover, we highlighted the competitive advantage of the LDP within real-world settings. Wenhao Yu 0010, Jie Peng 0002, Junrui Zhang 0012, Yifan Duan, Jianmin Ji, Yanyong Zhang |
IROS | 7 |
| 2024 | CRPlace: Camera-Radar Fusion with BEV Representation for Place RecognitionabstractThe integration of complementary characteristics from camera and radar data has emerged as an effective approach in 3D object detection. However, such fusion-based methods remain unexplored for place recognition, an equally important task for autonomous systems. Given that place recognition relies on the similarity between a query scene and the corresponding candidate scene, the stationary background of a scene is expected to play a crucial role in the task. As such, current well-designed camera-radar fusion methods for 3D object detection can hardly take effect in place recognition because they mainly focus on dynamic foreground objects. In this paper, a background-attentive camera-radar fusion-based method, named CRPlace, is proposed to generate background-attentive global descriptors from multi-view images and radar point clouds for accurate place recognition. To extract stationary background features effectively, we design an adaptive module that generates the background-attentive mask by utilizing the camera BEV feature and radar dynamic points. With the guidance of a background mask, we devise a bidirectional cross-attention-based spatial fusion strategy to facilitate comprehensive spatial interaction between the background information of the camera BEV feature and the radar BEV feature. As the first camera-radar fusion-based place recognition network, CRPlace has been evaluated thoroughly on the nuScenes dataset. The results show that our algorithm outperforms a variety of baseline methods across a comprehensive set of metrics (recall@1 reaches 91.2%). Shaowei Fu, Yifan Duan, Yao Li 0016, Chengzhen Meng, Jianmin Ji, Yanyong Zhang |
IROS | 7 |
| 2024 | MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded ScenesabstractLocalization and mapping are critical tasks for various applications such as autonomous vehicles and robotics. The challenges posed by outdoor environments present particular complexities due to their unbounded characteristics. In this work, we present MM-Gaussian, a LiDAR-camera multimodal fusion system for localization and mapping in unbounded scenes. Our approach is inspired by the recently developed 3D Gaussians, which demonstrate remarkable capabilities in achieving high rendering quality and fast rendering speed. Specifically, our system fully utilizes the geometric structure information provided by solid-state LiDAR to address the problem of inaccurate depth encountered when relying solely on visual solutions in unbounded, outdoor scenarios. Additionally, we utilize 3D Gaussian point clouds, with the assistance of pixel-level gradient descent, to fully exploit the color information in photos, thereby achieving realistic rendering effects. To further bolster the robustness of our system, we designed a relocalization module, which assists in returning to the correct trajectory in the event of a localization failure. Experiments conducted in multiple scenarios demonstrate the effectiveness of our method. Yifan Duan, Yu Sheng, Jianmin Ji, Yanyong Zhang |
IROS | 6 |
| 2024 | RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric StrategiesabstractThe recent advances in query-based multi-camera 3D object detection are featured by initializing object queries in the 3D space, and then sampling features from perspective-view images to perform multi-round query refinement. In such a framework, query points near the same camera ray are likely to sample similar features from very close pixels, resulting in ambiguous query features and degraded detection accuracy. To this end, we introduce RayFormer, a camera-ray-inspired query-based 3D object detector that aligns the initialization and feature extraction of object queries with the optical characteristics of cameras. Specifically, RayFormer transforms perspective-view image features into bird's eye view (BEV) via the lift-splat-shoot method and segments the BEV map to sectors based on the camera rays. Object queries are uniformly and sparsely initialized along each camera ray, facilitating the projection of different queries onto different areas in the image to extract distinct features. Besides, we leverage the instance information of images to supplement the uniformly initialized object queries by further involving additional queries along the ray from 2D object detection boxes. To extract unique object-level features that cater to distinct queries, we design a ray sampling method that suitably organizes the distribution of feature sampling points on both images and bird's eye view. Extensive experiments are conducted on the nuScenes dataset to validate our proposed ray-inspired model design. The proposed RayFormer achieves 55.5% mAP and 63.3% NDS, respectively. Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Yao Li 0016, Yanyong Zhang |
ACM Multimedia | 6 |
| 2024 | FARFusion V2: A Geometry-based Radar-Camera Fusion Method on the Ground for Roadside Far-Range 3D Object DetectionabstractFusing the data of millimeter-wave Radar sensors and high-definition cameras has emerged as a viable approach to achieving precise 3D object detection for roadside traffic surveillance. For roadside perception systems, earlier studies have pointed out that it is better to perform the fusion on the 2D image plane than on the BEV plane (which is popular for on-car perception systems), especially when the perception range is large (e.g., >150m). Image-plane fusion requires critical transformations, like perspective projection from the Radar's BEV to the camera's 2D plane and reverse IPM. However, real-world issues like uneven terrain and sensor movement degrade these transformations' precision, impacting fusion effectiveness. To alleviate these issues, we propose a geometry-based Radar-camera fusion method on the ground, namely FARFusion V2. Specifically, we extend the ground-plane assumption in FARFusion[20] to support arbitrary shapes by formulating the ground height as an implicit representation based on geometric transformations. By incorporating the ground information, we can enhance Radar data with target height measurements. Consequently, we can thus project the enhanced Radar data onto the 2D plane to obtain more accurate depth information, thereby assisting the IPM process. A real-time parameterized transformation parameters estimation module is further introduced to refine the view transformation processes. Moreover, considering various measurement noises across these two sensors, we introduce an uncertainty-based depth fusion strategy into the 2D fusion process to maximize the probability of obtaining the optimal depth value. Extensive experiments are conducted on our collected roadside OWL benchmark, demonstrating the excellent localization capacity of FARFusion V2 in far-range scenarios. Our method achieves an average location accuracy of 0.771m when we extend the detection range up to 500m. Yao Li 0016, Jiajun Deng, Yingjie Wang 0004, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang |
ACM Multimedia | 7 |
| 2024 | See Through Vehicles: Fully Occluded Vehicle Detection with Millimeter Wave RadarabstractA crucial task in autonomous driving is to continuously detect nearby vehicles. Problems thus arise when a vehicle is occluded and becomes "unseeable", which may lead to accidents. In this study, we develop mmOVD, a system that can detect fully occluded vehicles by involving millimeter-wave radars to capture the ground-reflected signals passing beneath the blocking vehicle's chassis. The foremost challenge here is coping with ghost points caused by frequent multi-path reflections, which highly resemble the true points. We devise a set of features that can efficiently distinguish the ghost points by exploiting the neighbor points' spatial and velocity distributions. We also design a cumulative clustering algorithm to effectively aggregate the unstable ground-reflected radar points over consecutive frames to derive the bounding boxes of the vehicles. Chenming He, Chengzhen Meng, Chunwang He, Xiaoran Fan, Yubo Yan, Yanyong Zhang |
MobiCom | 7 |
| 2024 | Map++: Towards User-Participatory Visual SLAM Systems with Efficient Map Expansion and SharingabstractConstructing precise 3D maps is crucial for the development of future map-based systems such as self-driving and navigation. However, generating these maps in complex environments, such as multi-level parking garages or shopping malls, remains a formidable challenge. In this paper, we introduce a participatory sensing approach that delegates map-building tasks to map users, thereby enabling cost-effective and continuous data collection. The proposed method harnesses the collective efforts of users, facilitating the expansion and ongoing update of the maps as the environment evolves. Hanqi Zhu, Yifan Duan, Wuyang Zhang, Longfei Shangguan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
MobiCom | 8 |
| 2024 | Enabling Tensor Language Model to Assist in Generating High-Performance Tensor Programs for Deep Learning
Yi Zhai 0005, Keyu Pan, Renwei Zhang, Shuo Liu 0019, Zichun Ye, Jianmin Ji, Jie Zhao 0002, Yu Zhang 0086, Yanyong Zhang |
OSDI | 11 |
| 2024 | Dynamic facial expression recognition with pseudo-label guided multi-modal pre-trainingabstractAbstract Due to the huge cost of manual annotations, the labelled data may not be sufficient to train a dynamic facial expression (DFR) recogniser with good performance. To address this, the authors propose a multi‐modal pre‐training method with a pseudo‐label guidance mechanism to make full use of unlabelled video data for learning informative representations of facial expressions. First, the authors build a pre‐training dataset of videos with aligned vision and audio modals. Second, the vision and audio feature encoders are trained through an instance discrimination strategy and a cross‐modal alignment strategy on the pre‐training data. Third, the vision feature encoder is extended as a dynamic expression recogniser and is fine‐tuned on the labelled training data. Fourth, the fine‐tuned expression recogniser is adopted to predict pseudo‐labels for the pre‐training data, and then start a new pre‐training phase with the guidance of pseudo‐labels to alleviate the long‐tail distribution problem and the instance‐class confliction. Fifth, since the representations learnt with the guidance of pseudo‐labels are more informative, a new fine‐tuning phase is added to further boost the generalisation performance on the DFR recognition task. Experimental results on the Dynamic Facial Expression in the Wild dataset demonstrate the superiority of the proposed method. Cong Liu 0006, Yanyong Zhang, Changfeng Xi, Zhen-Hua Ling |
IET Comput. Vis. | 4 |
| 2024 | HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
Sha Zhang 0002, Jiajun Deng, Lei Bai 0001, Houqiang Li, Wanli Ouyang, Yanyong Zhang |
Int. J. Comput. Vis. | 6 |
| 2023 | TLP: A Deep Learning-Based Cost Model for Tensor Program TuningabstractTensor program tuning is a non-convex objective optimization problem, to which search-based approaches have proven to be effective. At the core of the search-based approaches lies the design of the cost model. Though deep learning-based cost models perform significantly better than other methods, they still fall short and suffer from the following problems. First, their feature extraction heavily relies on expert-level domain knowledge in hardware architectures. Even so, the extracted features are often unsatisfactory and require separate considerations for CPUs and GPUs. Second, a cost model trained on one hardware platform usually performs poorly on another, a problem we call cross-hardware unavailability. Yi Zhai 0005, Yu Zhang 0086, Shuo Liu 0019, Xiaomeng Chu, Jie Peng 0002, Jianmin Ji, Yanyong Zhang |
ASPLOS (2) | 7 |
| 2023 | Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object DetectionabstractLiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear. The main challenge arises from that Radar data are extremely sparse and lack height information. Therefore, directly integrating Radar features into LiDAR-centric detection networks is not optimal. In this work, we introduce a bi-directional LiDAR-Radar fusion framework, termed Bi-LRFusion, to tackle the challenges and improve 3D detection for dynamic objects. Technically, Bi-LRFusion involves two steps: first, it enriches Radar's local features by learning important details from the LiDAR branch to alleviate the problems caused by the absence of height information and extreme sparsity; second, it combines LiDAR features with the enhanced Radar features in a unified bird's-eye-view representation. We conduct extensive experiments on nuScenes and ORR datasets, and show that our Bi-LRFusion achieves state-of-the-art performance for detecting dynamic objects. Notably, Radar data in these two datasets have different formats, which demonstrates the generalizability of our method. Codes will be published. Yingjie Wang 0005, Jiajun Deng, Yao Li 0016, Jinshui Hu, Cong Liu 0006, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang |
CVPR | 9 |
| 2023 | P3O: Transferring Visual Representations for Reinforcement Learning via PromptingabstractIt is important for deep reinforcement learning (DRL) algorithms to transfer their learned policies to new environments that have different visual inputs. In this paper, we introduce Prompt based Proximal Policy Optimization (P3O), a three-stage DRL algorithm that transfers visual representations from a target to a source environment by applying prompting. The process of P3O consists of three stages: pre-training, prompting, and predicting. In particular, we specify a prompt-transformer for representation conversion and propose a two-step training process to train the prompt-transformer for the target environment, while the rest of the DRL pipeline remains unchanged. We implement P3O and evaluate it on the OpenAI CarRacing video game. The experimental results show that P3O outperforms the state-of-the-art visual transferring schemes. In particular, P3O allows the learned policies to perform well in environments with different visual inputs, which is much more effective than retraining the policies in these environments. Guoliang You, Xiaomeng Chu, Yifan Duan, Jie Peng 0002, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang |
ICME | 7 |
| 2023 | Reinforcement Learning for Robot Navigation with Adaptive Forward Simulation Time (AFST) in a Semi-Markov ModelabstractDeep reinforcement learning (DRL) algorithms have proven effective in robot navigation, especially in unknown environments, by directly mapping perception inputs into robot control commands. However, most existing methods ignore the local minimum problem in navigation and thereby cannot handle complex unknown environments. In this paper, we propose the first DRL-based navigation method modeled by a semi-Markov decision process (SMDP) with continuous action space, named Adaptive Forward Simulation Time (AFST), to overcome this problem. Specifically, we reduce the dimensions of the action space and improve the distributed proximal policy optimization (DPPO) algorithm for the specified SMDP problem by modifying its GAE to better estimate the policy gradient in SMDPs. Experiments in various unknown environments demonstrate the effectiveness of AFST. Yu'an Chen, Ruosong Ye, Ziyang Tao, Hongjian Liu, Guangda Chen, Jie Peng 0002, Jun Ma 0034, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
IROS | 10 |
| 2023 | NN-Stretch: Automatic Neural Network Branching for Parallel Inference on Heterogeneous Multi-ProcessorsabstractMobile devices are increasingly equipped with heterogeneous multiprocessors, e.g., CPU + GPU + DSP. Yet existing Neural Network (NN) inference fails to fully utilize the computing power of the heterogeneous multi-processors due to the sequential structures of NN models. Towards this end, this paper proposes NN-Stretch, a new model adaption strategy, as well as the supporting system. It automatically branches a given model according to the processor architecture characteristics. Compared to other popular model adaption techniques such as model pruning that often sacrifices accuracy, NN-Stretch accelerates inference while preserving accuracy. Jianyu Wei, Ting Cao 0003, Shijie Cao, Shiqi Jiang 0002, Shaowei Fu, Mao Yang 0004, Yanyong Zhang, Yunxin Liu 0001 |
MobiSys | 7 |
| 2023 | CluB: Cluster Meets BEV for LiDAR-Based 3D Object DetectionabstractCurrently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors.
BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution layers, which, however, weakens the capability of presenting an object with the center point.
On the other hand, cluster-based detectors exploit the voting mechanism and aggregate the foreground points into object-centric clusters for further prediction.
In this paper, we explore how to effectively combine these two complementary representations into a unified framework.
Specifically, we propose a new 3D object detection framework, referred to as CluB, which incorporates an auxiliary cluster-based branch into the BEV-based detector by enriching the object representation at both feature and query levels.
Technically, CluB is comprised of two steps.
First, we construct a cluster feature diffusion module to establish the association between cluster features and BEV features in a subtle and adaptive fashion.
Based on that, an imitation loss is introduced to distill object-centric knowledge from the cluster features to the BEV features.
Second, we design a cluster query generation module to leverage the voting centers directly from the cluster branch, thus enriching the diversity of object queries.
Meanwhile, a direction loss is employed to encourage a more accurate voting center for each cluster.
Extensive experiments are conducted on Waymo and nuScenes datasets, and our CluB achieves state-of-the-art performance on both benchmarks. Yingjie Wang 0005, Jiajun Deng, Yuenan Hou, Yao Li 0016, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang |
NeurIPS | 8 |
| 2023 | Multi-Modal 3D Object Detection in Autonomous Driving: A Survey
Yingjie Wang 0005, Qiuyu Mao, Hanqi Zhu, Jiajun Deng, Yu Zhang 0086, Jianmin Ji, Houqiang Li, Yanyong Zhang |
Int. J. Comput. Vis. | 8 |
| 2023 | TransVG++: End-to-End Visual Grounding With Language Conditioned Vision TransformerabstractIn this work, we explore neat yet effective Transformer-based frameworks for visual grounding. The previous methods generally address the core problem of visual grounding, i.e., multi-modal fusion and reasoning, with manually-designed mechanisms. Such heuristic designs are not only complicated but also make models easily overfit specific data distributions. To avoid this, we first propose TransVG, which establishes multi-modal correspondences by Transformers and localizes referred regions by directly regressing box coordinates. We empirically show that complicated fusion modules can be replaced by a simple stack of Transformer encoder layers with higher performance. However, the core fusion Transformer in TransVG is stand-alone against uni-modal encoders, and thus should be trained from scratch on limited visual grounding data, which makes it hard to be optimized and leads to sub-optimal performance. To this end, we further introduce TransVG++ to make two-fold improvements. For one thing, we upgrade our framework to a purely Transformer-based one by leveraging Vision Transformer (ViT) for vision feature encoding. For another, we devise Language Conditioned Vision Transformer that removes external fusion modules and reuses the uni-modal ViT for vision-language fusion at the intermediate layers. We conduct extensive experiments on five prevalent datasets, and report a series of state-of-the-art records. Jiajun Deng, Zhengyuan Yang, Daqing Liu, Wengang Zhou 0001, Yanyong Zhang, Houqiang Li, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Reinforcement Learning Based Energy-Efficient Collaborative Inference for Mobile Edge ComputingabstractCollaborative inference in mobile edge computing (MEC) enables mobile devices to offload the computation tasks for the computation-intensive perception services, and the inference policy determines the inference latency and energy consumption. The optimal inference policy depends on the inference performance model of deep learning, the data generation model and the network model that are rarely known by mobile devices in time. In this paper, we propose a multi-agent reinforcement learning (RL) based energy-efficient MEC collaborative inference scheme, which enables each mobile device to choose both the partition point of deep learning and the collaborative edge of each mobile device based on the image quantity, the channel conditions and the previous inference performance. A learning experience exchange mechanism exploits the Q-values of the neighboring mobile devices to accelerate the inference policy optimization with less energy consumption. We also provide a deep multi-agent RL based inference scheme to accelerate learning for large-scale MEC networks, in which an actor network yields the collaborative inference policy probability distribution and a critic network guides the weight update of the actor network to enhance sample efficiency. We provide the inference performance bound and analyze the computational complexity. Both simulation and experimental results show that our proposed schemes reduce the inference latency and save the MEC energy consumption. Yilin Xiao 0001, Liang Xiao 0003, Kunpeng Wan, Helin Yang, Yi Zhang 0035, Yi Wu 0010, Yanyong Zhang |
IEEE Trans. Commun. | 7 |
| 2023 | TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs Through Trajectory MatchingabstractRecently, deploying sensors such as LiDARs on the roadside to monitor the passing traffic and assist autonomous vehicle perception has become popular. However, unlike autonomous vehicle systems, roadside sensor systems involve sensors from different subsystems, resulting in a lack of synchronization in both time and space between the sensors. Calibration is a critical technology that enables the central server to fuse data generated by different location infrastructures, which vastly improves sensing range and detection robustness. Regrettably, existing calibration algorithms frequently assume that LiDARs have significant overlap or that temporal calibration has already been achieved. However, since these assumptions do not always hold in real-world scenarios, the calibration results obtained from existing algorithms are frequently unsatisfactory. In this paper, we propose TrajMatch - the first system that can automatically calibrate roadside LiDARs in both time and space. The main idea is to automatically calibrate the sensors based on the result of the detection/tracking task, rather than relying on extracting special features. Furthermore, we propose a novel mechanism for evaluating calibration parameters that align with our algorithm, and we demonstrate its effectiveness through experiments. This mechanism can also guide parameter iterations for multiple calibrations, further enhancing the accuracy and efficiency of our calibration method. Finally, to evaluate the performance of TrajMatch, we collected two datasets, one simulated dataset LiDARnet-sim 1.0 and one real-world dataset. The experimental results show that TrajMatch can achieve a spatial calibration error of less than$10cm$and a temporal calibration error of less than$1.5ms$. Haojie Ren, Sha Zhang 0002, Sugang Li, Yao Li 0016, Xinchen Li, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2023 | Back-Guard: Wireless Backscattering Based User Sensing With Parallel Attention ModelabstractWith the rapid advance of wireless sensing techniques, it becomes possible to provide a fine-grained user activity tracking service at home and office. Such a technique is of broad applications in various domains such as personal activity diary, elderly care, and customized services. For example, several radio frequency (RF) based sensing systems were recently proposed for human activity recognition. However, most of them focused on specific scenarios and suffered from interference caused by other users and wireless devices. In this work, we propose Back-Guard, a backscattering-based sensing system that achieves accurate and non-intrusive user activity recognition and further user identification/authentication. Back-Guard carefully examines the backscatter spectrogram data and extracts high-level features from both spatial and temporal domains. Leveraging the parallel attention based deep learning model, our system can discriminate different motions and users accurately and robustly in various situations. We implemented a prototype system and collected data from 25 users for more than 2 months. Extensive experiments demonstrate that Back-Guard achieves 93.4$\%$activity recognition accuracy and 91.5$\%$user identification accuracy, respectively. In particular, Back-Guard can also tackle multiple user scenarios, which has little accuracy reduction when the users are separated, e.g., by around 2 meters. Xiang-Yang Li 0001, Manjiang Yin, Yanyong Zhang, Panlong Yang, Chengchen Wan, Haisheng Tan |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | VPFNet: Improving 3D Object Detection With Virtual Point Based LiDAR and Stereo Data FusionabstractIt has been well recognized that fusing the complementary information from depth-aware LiDAR point clouds and semantic-rich stereo images would benefit 3D object detection. Nevertheless, it is non-trivial to explore the inherently unnatural interaction between sparse 3D points and dense 2D pixels. To ease this difficulty, the recent approaches generally project the 3D points onto the 2D image plane to sample the image data and then aggregate the data at the points. However, these approaches often suffer from the mismatch between the resolution of point clouds and RGB images, leading to sub-optimal performance. Specifically, taking the sparse points as the multi-modal data aggregation locations causes severe information loss for high-resolution images, which in turn undermines the effectiveness of multi-sensor fusion. In this paper, we presentVPFNet—a new architecture that cleverly aligns and aggregates the point cloud and image data at the “virtual” points. Particularly, with their density lying between that of the 3D points and 2D pixels, the virtual points can nicely bridge the resolution gap between the two sensors, and thus preserve more information for processing. Moreover, we also investigate the data augmentation techniques that can be applied to both point clouds and RGB images, as the data augmentation has made non-negligible contribution towards 3D object detectors to date. We have conducted extensive experiments on KITTI dataset, and have observed good performance compared to the state-of-the-art methods. Remarkably, ourVPFNetachieves 83.21% moderate$AP_{3D}$and 91.86% moderate$AP_{BEV}$on the KITTI test set. The network design also takes computation efficiency into consideration – we can achieve a FPS of 15 on a single NVIDIA RTX 2080Ti GPU. Hanqi Zhu, Jiajun Deng, Yu Zhang 0086, Jianmin Ji, Qiuyu Mao, Houqiang Li, Yanyong Zhang |
IEEE Trans. Multim. | 7 |
| 2022 | PF-MOT: Probability Fusion Based 3D Multi-Object Tracking for Autonomous Vehiclesabstract3D Multi-Object Tracking (MOT) plays a crucial role in efficient and safe operation of automatic driving, especially in scenarios of occlusion or poor visibility. Most 3D MOT methods leverage only positional distance, which is insufficient for scenes with high density of objects or drastic changes in the motion state. In order to address this, we propose a new 3D MOT model which fuses information pertaining to positional distance and geometric similarity. Our proposed solution comprises of four parts: a) a feature extraction mechanism integrated into a commonly used detector to extract individual features for each detection, b) computation of two distance matrices based on Euclidean distance and feature similarity, c) conversion of the distance matrices to probability matrices by a cluster based Earth-Mover Distance (EMD) algorithm, and d) a data association method that fuses both sources to boost the tracking accuracy. Our proposed model demonstrates state-of-the-art performance on the nuScenes tracking dataset, with extensive experiments attesting to an improved tracking accuracy over baselines that operate solely on positional distance. Tao Wen 0018, Yanyong Zhang, Nikolaos M. Freris |
ICRA | 2 |
| 2022 | PFilter: Building Persistent Maps through Feature Filtering for Fast and Accurate LiDAR-based SLAMabstractSimultaneous localization and mapping (SLAM) based on laser sensors has been widely adopted by mobile robots and autonomous vehicles. These SLAM systems are required to support accurate localization with limited computational resources. In particular, point cloud registration, i.e., the process of matching and aligning multiple LiDAR scans collected at multiple locations in a global coordinate framework, has been deemed as the bottleneck step in SLAM. In this paper, we propose a feature filtering algorithm, PFilter, that can filter out invalid features and can thus greatly alleviate this bottleneck. Meanwhile, the overall registration accuracy is also improved due to the carefully curated feature points. We integrate PFilter into the well-established scan-to-map LiDAR odometry framework, F-LOAM, and evaluate its performance on the KITTI dataset. The experimental results show that PFilter can remove about 48.4% of the points in the local feature map and reduce feature points in scan by 19.3% on average, which save 20.9% processing time per frame. In the mean time, we improve the accuracy by 9.4%. Yifan Duan, Jie Peng 0002, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
IROS | 5 |
| 2022 | Reinforcement learning based energy efficient robot relay for unmanned aerial vehicles against smart jamming
Xiaozhen Lu, Jingfang Jie, Liang Xiao 0003, Jin Li 0002, Yanyong Zhang |
Sci. China Inf. Sci. | 6 |
| 2022 | Special Issue on Device-Free Sensing for Human Behavior Recognition II
Zhu Wang 0001, Bin Guo 0001, Yanyong Zhang, Daqing Zhang 0001 |
Pers. Ubiquitous Comput. | 3 |
| 2021 | Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionabstractRecent advances on 3D object detection heavily rely on how the 3D data are represented, i.e., voxel-based or point-based representation. Many existing high performance 3D detectors are point-based because this structure can better retain precise point positions. Nevertheless, point-level features lead to high computation overheads due to unordered storage. In contrast, the voxel-based structure is better suited for feature extraction but often yields lower accuracy because the input data are divided into grids. In this paper, we take a slightly different viewpoint --- we find that precise positioning of raw points is not essential for high performance 3D object detection and that the coarse voxel granularity can also offer sufficient detection accuracy. Bearing this view in mind, we devise a simple but effective voxel-based framework, named Voxel R-CNN. By taking full advantage of voxel features in a two-stage approach, our method achieves comparable detection accuracy with state-of-the-art point-based models, but at a fraction of the computation cost. Voxel R-CNN consists of a 3D backbone network, a 2D bird-eye-view (BEV) Region Proposal Network, and a detect head. A voxel RoI pooling is devised to extract RoI features directly from voxel features for further refinement. Extensive experiments are conducted on the widely used KITTI Dataset and the more recent Waymo Open Dataset. Our results show that compared to existing voxel-based methods, Voxel R-CNN delivers a higher detection accuracy while maintaining a real-time frame processing rate, i.e., at a speed of 25 FPS on an NVIDIA RTX 2080 Ti GPU. The code is available at https://github.com/djiajunustc/Voxel-R-CNN. Jiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 0001, Yanyong Zhang, Houqiang Li |
AAAI | 5 |
| 2021 | Towards an Online RRT-based Path Planning Algorithm for Ackermann-steering VehiclesabstractIt is challenging to develop an online path planning algorithm for Ackermann-steering vehicles to find collision-free and kinematically-feasible paths, that is efficient for dense environments, adaptable to various environments, and suitable for environments with narrow passages. In this paper, we propose a kinematically constrained RRT-based path planning algorithm integrating with a trajectory parameter space (TP-space) with three novel improvements to meet the above requirements. In specific, we introduce a new way to choose candidate nodes to expand the tree for an RRT-based algorithm, which can significantly increase the success rate of the expansion and improve the efficiency of the algorithm. We also introduce a procedure to incrementally adjust the step size for the expansion, which enables the algorithm to automatically adapt to various environments. At last, we integrate rapidly-exploring random vines (RRV) with a TP-space to handle kinematic constraints and improve the performance of the algorithm to expand the tree through a narrow passage. We also prove that the algorithm is probabilistic complete and asymptotically near-optimal. An ablation study shows that all three improvements can notably improve the performance of the RRT-based path planning algorithm. We also evaluate the algorithm in various environments. The experimental results show that our algorithm achieves competitive performance compared with the state-of-the-art. The source code is available at https://github.com/PengJieb/fastbkrrt. Jie Peng 0002, Yu'an Chen, Yifan Duan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
ICRA | 6 |
| 2021 | DRQN-based 3D Obstacle Avoidance with a Limited Field of ViewabstractIn this paper, we propose a map-based end-to-end DRL approach for three-dimensional (3D) obstacle avoidance in a partially observed environment, which is applied to achieve autonomous navigation for an indoor mobile robot using a depth camera with a narrow field of view. We first train a neural network with LSTM units in a 3D simulator of mobile robots to approximate the Q-value function in double DRQN. We also use a curriculum learning strategy to accelerate and stabilize the training process. Then we deploy the trained model to a real robot to perform 3D obstacle avoidance in its navigation. We evaluate the proposed approach both in the simulated environment and on a robot in the real world. The experimental results show that the approach is efficient and easy to be deployed, and it performs well for 3D obstacle avoidance with a narrow observation angle, which outperforms other existing DRL-based models by 15.5% on success rate. Yu'an Chen, Guangda Chen, Lifan Pan, Jun Ma 0034, Yu Zhang 0086, Yanyong Zhang, Jianmin Ji |
IROS | 6 |
| 2021 | Neighbor-Vote: Improving Monocular 3D Object Detection through Neighbor Distance VotingabstractAs cameras are increasingly deployed in new application domains such as autonomous driving, performing 3D object detection on monocular images becomes an important task for visual scene understanding. Recent advances on monocular 3D object detection mainly rely on the "pseudo-LiDAR'' generation, which performs monocular depth estimation and lifts the 2D pixels to pseudo 3D points. However, depth estimation from monocular images, due to its poor accuracy, leads to inevitable position shift of pseudo-LiDAR points within the object. Therefore, the predicted bounding boxes may suffer from inaccurate location and deformed shape. In this paper, we present a novel neighbor-voting method that incorporates neighbor predictions to ameliorate object detection from severely deformed pseudo-LiDAR point clouds. Specifically, each feature point around the object forms their own predictions, and then the "consensus'' is achieved through voting. In this way, we can effectively combine the neighbors' predictions with local prediction and achieve more accurate 3D detection. To further enlarge the difference between the foreground region of interest (ROI) pseudo-LiDAR points and the background points, we also encode the ROI prediction scores of 2D foreground pixels into the corresponding pseudo-LiDAR points. We conduct extensive experiments on the KITTI benchmark to validate the merits of our proposed method. Our results on the bird's eye view detection outperform the state-of-the-art performance, especially for the "hard" level detection. The code is available at https://github.com/cxmomo/Neighbor-Vote. Xiaomeng Chu, Jiajun Deng, Yao Li 0016, Zhenxun Yuan, Yanyong Zhang, Jianmin Ji, Yu Zhang 0086 |
ACM Multimedia | 5 |
| 2021 | HeadFi: bringing intelligence to all headphonesabstractHeadphones continue to become more intelligent as new functions (e.g., touch-based gesture control) appear. These functions usually rely on auxiliary sensors (e.g., accelerometer and gyroscope) that are available in smart headphones. However, for those headphones that do not have such sensors, supporting these functions becomes a daunting task. This paper presents HeadFi, a new design paradigm for bringing intelligence to headphones. Instead of adding auxiliary sensors into headphones, HeadFi turns the pair of drivers that are readily available inside all headphones into a versatile sensor to enable new applications spanning across mobile health, user-interface, and context-awareness. HeadFi works as a plug-in peripheral connecting the headphones and the pairing device (e.g., a smartphone). The simplicity (can be as simple as only two resistors) and small form factor of this design lend itself to be embedded into the pairing device as an integrated circuit. We envision HeadFi can serve as a vital supplementary solution to existing smart headphone design by directly transforming large amounts of existing "dumb" headphones into intelligent ones. We prototype HeadFi on PCB and conduct extensive experiments with 53 volunteers using 54 pairs of non-smart headphones under the institutional review board (IRB) protocols. The results show that HeadFi can achieve 97.2%--99.5% accuracy on user identification, 96.8%--99.2% accuracy on heart rate monitoring, and 97.7%--99.3% accuracy on gesture recognition. Xiaoran Fan, Longfei Shangguan, Siddharth Rupavatharam, Yanyong Zhang, Jie Xiong 0001, Richard E. Howard |
MobiCom | 4 |
| 2021 | Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingabstractAs mobile devices continuously generate streams of images and videos, a new class of mobile deep vision applications are rapidly emerging, which usually involve running deep neural networks on these multimedia data in real-time. To support such applications, having mobile devices offload the computation, especially the neural network inference, to edge clouds has proved effective. Existing solutions often assume there exists a dedicated and powerful server, to which the entire inference can be offloaded. In reality, however, we may not be able to find such a server but need to make do with less powerful ones. To address these more practical situations, we propose to partition the video frame and offload the partial inference tasks to multiple servers for parallel processing. This paper presents the design of Elf, a framework to accelerate the mobile deep vision applications with any server provisioning through the parallel offloading. Elf employs a recurrent region proposal prediction algorithm, a region proposal centric frame partitioning, and a resource-aware multi-offloading scheme. We implement and evaluate Elf upon Linux and Android platforms using four commercial mobile devices and three deep vision applications with ten state-of-the-art models. The comprehensive experiments show that Elf can speed up the applications by 4.85× with saving bandwidth usage by 52.6%, while with <1% application accuracy sacrifice. Wuyang Zhang, Zhezhi He, Zhenhua Jia, Yunxin Liu 0001, Marco Gruteser, Dipankar Raychaudhuri, Yanyong Zhang |
MobiCom | 8 |
| 2021 | Feasibility study of practical vital sign detection using millimeter-wave radios
Zhenhua Jia, Chenren Xu, Guojie Luo, Daqing Zhang 0001, Ning An 0001, Yanyong Zhang |
CCF Trans. Pervasive Comput. Interact. | 7 |
| 2021 | From Multi-View to Hollow-3D: Hallucinated Hollow-3D R-CNN for 3D Object DetectionabstractAs an emerging data modal with precise distance sensing, LiDAR point clouds have been placed great expectations on 3D scene understanding. However, point clouds are always sparsely distributed in the 3D space, and with unstructured storage, which makes it difficult to represent them for effective 3D object detection. To this end, in this work, we regard point clouds as hollow-3D data and propose a new architecture, namely Hallucinated Hollow-3D R-CNN (H23D R-CNN), to address the problem of 3D object detection. In our approach, we first extract the multi-view features by sequentially projecting the point clouds into the perspective view and the bird-eye view. Then, we hallucinate the 3D representation by a novel bilaterally guided multi-view fusion block. Finally, the 3D objects are detected via a box refinement module with a novel Hierarchical Voxel RoI Pooling operation. The proposed H23D R-CNN provides a new angle to take full advantage of complementary information in the perspective view and the bird-eye view with an efficient framework. We evaluate our approach on the public KITTI Dataset and Waymo Open Dataset. Extensive experiments demonstrate the superiority of our method over the state-of-the-art algorithms with respect to both effectiveness and efficiency. The code is available athttps://github.com/djiajunustc/H-23D_R-CNN. Jiajun Deng, Wengang Zhou 0001, Yanyong Zhang, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Lightweight Map-Enhanced 3D Object Detection and Tracking for Autonomous Drivingabstract3D object detection and tracking are crucial to the real-time and accurate perception of the surrounding environment for autonomous driving. Recent approaches on 3D object detection and tracking have made great progress, thanks to the rapid development of deep learning models. Even though these models have achieved superior performance on specific datasets, the actual self-driving systems still cannot deal with real-world driving situations properly, especially in complicated scenarios like road intersections. With the development of vehicle-infrastructure cooperation technology, scene information such as map is considered to have great potential in alleviating these problems. In this paper, we explore the potential of solving corner cases in real driving scenarios through the cooperation between autonomous vehicles and map information. We propose a holistic approach that integrates and utilizes the map information in system following the tracking-by-detection paradigm. In order to ensure that the use of map information does not bring much overhead to detection and tracking, we propose a representation method for concise information extracted from rich map. We show that our framework can improve the detection and tracking accuracy with mild or no increase of latency. Specifically, in some cases, our results demonstrate a MOTA improvement of nearly 2% . Shunhong Wang, Yu Zhang 0086, Yanyong Zhang, Jianmin Ji |
Internetware | 4 |
| 2020 | Back-Guard: Wireless Backscattering based User Activity Recognition and Identification with Parallel Attention ModelabstractWith the rapid advance of smart home and office systems, it becomes possible to provide a fine-grained user activity tracking service accurately recognizing user activities and identities in a seamless and non-invasive manner. Such a system can find applications in various domains, such as elder safeguard, customized services, and simply personal activity diary. Recently, several radio frequency (RF) based sensing systems were proposed for human sensing, most of which focus on limited scenarios and suffer from interference caused by other users or wireless devices. To tackle this challenge, we propose Back-Guard, which achieves accurate and non-intrusive user activity recognition and then user identification through battery-free wireless backscattering. Back-Guard carefully examines the backscatter spectrogram data and extracts high-level features from both spatial and temporal domains that can characterize the user behaviors. Leveraging the parallel attention based deep learning models, our system can discriminate different motions and users accurately and robustly in various situations. We implement a prototype system and collect data in actual scenarios from 25 users for over 2 months. Extensive experiments demonstrate the promising performance of our system. In particular, Back-Guard achieves 93.4% activity recognition accuracy and 91.5% user identification accuracy, respectively. Our experiments also demonstrate little accuracy reduction when multiple users are separated, e.g., by around 2 meters. Manjiang Yin, Xiang-Yang Li 0001, Yanyong Zhang, Panlong Yang, Chengchen Wan |
IWQoS | 3 |
| 2020 | Towards flexible wireless charging for medical implants using distributed antenna systemabstractThis paper presents the design, implementation and evaluation of In-N-Out, a software-hardware solution for far-field wireless power transfer. In-N-Out can continuously charge a medical implant residing in deep tissues at near-optimal beamforming power, even when the implant moves around inside the human body. To accomplish this, we exploit the unique energy ball pattern of distributed antenna array and devise a backscatter-assisted beamforming algorithm that can concentrate RF energy on a tiny spot surrounding the medical implant. Meanwhile, the power levels on other body parts stay in low level, reducing the risk of overheating. We proto-type In-N-Out on 21 software-defined radios and a printed circuit board (PCB). Extensive experiments demonstrate that In-N-Out achieves 0.37 mW average charging power inside a 10 cm-thick pork belly, which is sufficient to wirelessly power a range of commercial medical devices. Our head-to-head comparison with the state-of-the-art approach shows that In-N-Out achieves 5.4X-18.1X power gain when the implant is stationary, and 5.3X-7.4X power gain when the implant is in motion. Xiaoran Fan, Longfei Shangguan, Richard E. Howard, Yanyong Zhang, Yao Peng 0002, Jie Xiong 0001, Xiang-Yang Li 0001 |
MobiCom | 4 |
| 2020 | A Close Look at Multi-tenant Parallel CNN Inference for Autonomous Driving
Yitong Huang, Yu Zhang 0086, Boyuan Feng, Yanyong Zhang, Yufei Ding 0001 |
NPC | 5 |
| 2020 | Reinforcement Learning-Based Mobile Offloading for Edge Computing Against Jamming and InterferenceabstractMobile edge computing systems help improve the performance of computational-intensive applications on mobile devices and have to resist jamming attacks and heavy interference. In this paper, we present a reinforcement learning based mobile offloading scheme for edge computing against jamming attacks and interference, which uses safe reinforcement learning to avoid choosing the risky offloading policy that fails to meet the computational latency requirements of the tasks. This scheme enables the mobile device to choose the edge device, the transmit power and the offloading rate to improve its utility including the sharing gain, the computational latency, the energy consumption and the signal-to-interference-plus-noise ratio of the offloading signals without knowing the task generation model, the edge computing model, and the jamming/interference model. We also design a deep reinforcement learning based mobile offloading for edge computing that uses an actor network to choose the offloading policy and a critic network to update the actor network weights to improve the computational performance. We discuss the computational complexity and provide the performance bound that consists of the computational latency and the energy consumption based on the Nash equilibrium of the mobile offloading game. Simulation results show that this scheme can reduce the computational latency and save energy consumption. Liang Xiao 0003, Xiaozhen Lu, Tangwei Xu, Xiaoyue Wan, Wen Ji 0003, Yanyong Zhang |
IEEE Trans. Commun. | 6 |
| 2020 | In-Bed Body Motion Detection and Classification SystemabstractIn-bed motion detection and classification are important techniques that can enable an array of applications, among which are sleep monitoring and abnormal movement detection. In this article, we present a low-cost, low-overhead, and highly robust system for in-bed movement detection and classification that uses low-end load cells. To detect movements, we have designed a feature that we refer to as Log-Peak, which can be extracted from load cell data that is collected through wireless links in an energy-efficient manner. After detection, we set out to achieve a precise body motion classification. Toward this goal, we define nine classes of movements, and design a machine learning algorithm using Support Vector Machine, Random Forest, and XGBoost techniques to classify a movement into one of nine classes. For every movement, we have extracted 24 features and used them in our model. This movement detection/classification system was evaluated on data collected from 40 subjects who performed 35 predefined movements in each experiment. We have applied multiple tree topologies for each technique to reach their best results. After examining various combinations, we have achieved a final classification accuracy of 91.5%. This system can be used conveniently for long-term home monitoring. Musaab Alaziz, Zhenhua Jia, Richard E. Howard, Xiaodong Lin 0004, Yanyong Zhang |
ACM Trans. Sens. Networks | 5 |
| 2019 | Hetero-Edge: Orchestration of Real-time Vision Applications on Heterogeneous Edge CloudsabstractRunning computer vision algorithms on images or videos collected by mobile devices represent a new class of latency-sensitive applications that expect to benefit from edge cloud computing. These applications often demand real-time responses (e.g., <;100 ms), which can not be satisfied by traditional cloud computing. However, the edge cloud architecture is inherently distributed and heterogeneous, requiring new approaches to resource allocation and orchestration. This paper presents the design and evaluation of a latency-aware edge computing platform, aiming to minimize the end-to-end latency for edge applications. The proposed platform is built on Apache Storm, and consists of multiple edge servers with heterogeneous computation (including both GPUs and CPUs) and networking resources. Central to our platform is an orchestration framework that breaks down an edge application into Storm tasks as defined by a directed acyclic graph (DAG) and then maps these tasks onto heterogeneous edge servers for efficient execution. An experimental proof-of-concept testbed is used to demonstrate that the proposed platform can indeed achieve low end-to-end latency: considering a real-time 3D scene reconstruction application, it is shown that the testbed can support up to 30 concurrent streams with an average perframe latency of 32ms, and can achieve 40% latency reduction relative to the baseline Storm scheduling approach. Wuyang Zhang, Sugang Li, Zhenhua Jia, Yanyong Zhang, Dipankar Raychaudhuri |
INFOCOM | 5 |
| 2019 | Demo: The RFID Can Hear Your Music PlayabstractIn this work, we devise RF-DJ, a contactless music recognition system with the help of COTS RFID device. Since the music is caused by vibration and the vibration can influence the RF signal, our system could accurately recover the frequency of every tone, especially string instruments. Specifically, RF-DJ is immune to noises from the player/instrument motions and the ambient environment. Further more, it can recover the high frequency signal from the relatively low sampling rate data. In our demonstration, we put one tag on the surface of ukulele (not the string) and achieve the overall recognition accuracy of $93%, 90%, 87%, 81%$ when using 1,2,3,4 strings, respectively. Compared to typical machine learning based RF sensing systems, our system is model driven instead of data driven, which requires little training effort and could be applicable across different locations. Last but not the least, our system can also be used for other instruments such as zither, violin and kalimba and shows similarly good performances. Yuanhao Feng, Panlong Yang, Yanyong Zhang, Xiang-Yang Li 0001 |
MobiCom | 3 |
| 2019 | Special issue on device-free sensing for human behavior recognition
Bin Guo 0001, Yanyong Zhang, Daqing Zhang 0001, Zhu Wang 0001 |
Pers. Ubiquitous Comput. | 2 |
| 2019 | Dynamic Resource Allocation for Streaming Scalable Videos in SDN-Aided Dense Small-Cell NetworksabstractBoth wireless small-cell communications and software-defined networking (SDN) in wired systems continue to evolve rapidly, aiming for improving the quality of experience (QoE) of users. Against this emerging landscape, we conceive scalable video streaming over SDN-aided dense smell-cell networks by jointly optimizing the video layer selection, the wireless resource allocation, and the dynamic routing of video streams. In the light of this ambitious objective, we conceive a dense software-defined small-cell network architecture for the fine-grained manipulation of the video streams relying on the cooperation of small-cell base stations. Based on this framework, we formulate the scalable video streaming problem as maximizing the time-averaged QoE subject to a specific time-averaged rate constraint as well as to a resource constraint. By employing the classic Lyapunov optimization method, the problem is further decomposed into the twin sub-problems of video layer selection and wireless resource allocation. Via solving these sub-problems, we derive a video layer selection strategy and a wireless resource allocation algorithm. Furthermore, we propose a beneficial routing policy for scalable video streams with the aid of the so-called segment routing technique in the context of SDN, which additionally exploits the collaboration of small-cell base stations. Our results demonstrate compelling performance improvements compared with the classic PID control theory-based method. Jian Yang 0014, Shuangwu Chen, Yongdong Zhang 0001, Yanyong Zhang, Lajos Hanzo |
IEEE Trans. Commun. | 5 |
| 2018 | Preventing Unauthorized Access on Passive TagsabstractAs the Ultra High Frequency (UHF) passive Radio Frequency IDentification (RFID) technology becomes increasingly deployed, it faces an array of new security attacks. In this paper, we consider a type of attack in which a malicious RFID reader could arbitrarily modify the tags via standard commands, e.g., IDs or other data in the memory. To deal with this type of attack, we propose a physical-layer RF signal based reader authentication solution, namely Arbitrator, that involves passively listening on RF channels, analyzing the communication signals, identifying unauthorized readers and jamming the commands from such readers. Our solution does not need to modify RFID devices or the underlying communication standards, hence fully compatible with the existing RFID infrastructure. In this study, we have implemented a prototype Arbitrator over the Universal Software Radio Peripheral (USRP) platform, and conducted extensive experiments to evaluate its performance. Our results show that Arbitrator can detect unauthorized RFID readers with high accuracy, and thus effectively diminish the unauthorized access attacks. Han Ding 0002, Jinsong Han, Yanyong Zhang, Fu Xiao 0001, Wei Xi 0003, Ge Wang 0003, Zhiping Jiang |
INFOCOM | 3 |
| 2018 | Secret-Focus: A Practical Physical Layer Secret Communication System by Perturbing Focused Phases in Distributed BeamformingabstractEnsuring confidentiality of communication is fundamental to securing the operation of a wireless system, where eavesdropping is easily facilitated by the broadcast nature of the wireless medium. By applying distributed beamforming among a coalition, we show that a new approach for assuring physical layer secrecy, without requiring any knowledge about the eavesdropper or injecting any additional cover noise, is possible if the transmitters frequently perturb their phases around the proper alignment phase while transmitting messages. This approach is readily applied to amplitude-based modulation schemes, such as PAM or QAM. We present our secrecy mechanisms, prove several important secrecy properties, and develop a practical secret communication system design. We further implement and deploy a prototype that consists of 16 distributed transmitters using USRP N210s in a 20×20×3m3area. By sending more than 160M bits over our system to the receiver, depending on system parameter settings, we measure that the eavesdroppers failed to decode 30%-60% of the bits cross multiple locations while the intended receiver has an estimated bit error ratio of 3×10-6. Xiaoran Fan, Wade Trappe, Yanyong Zhang, Richard E. Howard, Zhu Han 0001 |
INFOCOM | 4 |
| 2018 | Enabling Concurrent IoT Transmissions in Distributed C-RANabstractAs rapid expansion of the low-cost next billion devices, wireless sensor networks (WSN) undertake much denser low-end internet of things (IoT) nodes nowadays. In the meantime, the future next 5 generation (5G) radio base stations (BS) are granted more capabilities. Distributed cloud radio access network (C-RAN) is becoming available for the future massive WSN. However, real-world distributed C-RAN is less explored for low-end IoT based WSN due to its difficulties in implementation. In this paper, we built a distributed C-RAN which has tens of distributed radio frontends using USRP N210s in a 20 × 20 × 3 m3 area. By exploiting the inherent hardware properties of low-end IoT devices and the spatial diversity of distributed C-RAN system, we show the distributed C-RAN can potentially decode collided signals from low-end IoT devices with all signal processing been done on the cloud. Xiaoran Fan, Zhenzhou Qi, Zhenhua Jia, Yanyong Zhang |
SenSys | 4 |
| 2018 | Continuous Low-Power Ammonia Monitoring Using Long Short-Term Memory Neural NetworksabstractAccurate and continuous ammonia monitoring is important for laboratory animal studies and many other applications. Existing solutions are often expensive, inaccurate, or unsuitable for long-term monitoring. In this work, we propose a new ammonia monitoring approach that is low-power, automatic, accurate, and wireless. Zhenhua Jia, Xinmeng Lyu, Wuyang Zhang, Richard P. Martin, Richard E. Howard, Yanyong Zhang |
SenSys | 6 |
| 2017 | HB-phone: a bed-mounted geophone-based heartbeat monitoring system: demo abstractabstractMonitoring heartbeats takes an important role to ensure a person's health and well-being. Few of the existing systems are accurate, unobtrusive, robust and easy to install at the same time. Thus, we propose a completely unobtrusive system which can detect heartbeats during sleep by sensing the weak ballistic vibrations caused by heartbeats on any bed. The system, HB-Phone, is centered around the off-the-shelf geophone sensor and can be easily installed on an existing bed. In this demo, we demonstrate that our system can detect and extract heartbeats accurately and in real time, even with the presence of noise from the environment and gross body movements during sleep. Zhenhua Jia, Richard E. Howard, Yanyong Zhang, Pei Zhang 0001 |
IPSN | 3 |
| 2017 | Monitoring a Person's Heart Rate and Respiratory Rate on a Shared Bed Using GeophonesabstractUsing geophones to sense bed vibrations caused by ballistic force has shown great potential in monitoring a person's heart rate during sleep. It does not require a special mattress or sheets, and the user is free to move around and change position during sleep. Earlier work has studied how to process the geophone signal to detect heartbeats when a single subject occupies the entire bed. In this study, we develop a system called VitalMon, aiming to monitor a person's respiratory rate as well as heart rate, even when she is sharing a bed with another person. In such situations, the vibrations from both persons are mixed together. VitalMon first separates the two heartbeat signals, and then distinguishes the respiration signal from the heartbeat signal for each person. Our heartbeat separation algorithm relies on the spatial difference between two signal sources with respect to each vibration sensor, and our respiration extraction algorithm deciphers the breathing rate embedded in amplitude fluctuation of the heartbeat signal. Zhenhua Jia, Amelie Bonde, Sugang Li, Chenren Xu, Yanyong Zhang, Richard E. Howard, Pei Zhang 0001 |
SenSys | 6 |
| 2017 | Transmit Only: An Ultra Low Overhead MAC Protocol for Dense Wireless SystemsabstractThe number of small wireless devices is rapidly increasing, making the radio channel efficiency in limited geographic areas (individual rooms or buildings) an important metric for MAC protocols. Many of these emerging devices have use-cases that are difficult to satisfy with current hardware solutions and channel access methods; for instance device mobility, small energy reserves, and requirements for low cost and small form factors. However, for most of these applications, such as health care monitoring or sensing, feedback to the radio device is unnecessary and unidirectional communication techniques are not only sufficient, but can also be advantageous. We propose an efficient, reliable technique for unidirectional communication, called Transmit Only (TO), that satisfies these requirements while maintaining packet throughput guarantees and reducing energy consumption. In this paper we will demonstrate the feasibility and performance of this kind of highly asymmetric, transmit-only protocol through theoretical, simulated, and experimental results. Yanyong Zhang, Bernhard Firner, Richard E. Howard, Richard P. Martin, Narayan B. Mandayam, Junichiro Fukuyama, Chenren Xu |
SMARTCOMP | 1 |
| 2017 | Special issue on big data computing, analytics and applications
Chenren Xu, Zhu Han 0001, Yanyong Zhang |
Pers. Ubiquitous Comput. | 3 |
| 2016 | SEGUE: Quality of Service Aware Edge Cloud Service MigrationabstractEdge cloud computing moves cloud services to the edge of the network, thereby allowing clients to access services with a significantly reduced network delay. This service migration is intended to enable a range of latency sensitive mobile applications. In this paper, we propose to manage user QoS by actively migrating services to different edge clouds in response to degraded server or network performance. Previous studies have proposed a distance-based Markov Decision Process (MDP) for optimizing migration decisions. These models provide the feasibility of applying MDP to edge cloud service migration decisions. However, these models fail to consider dynamic network and server states in migration decisions. In this work, we address these limitations by designing a comprehensive edge cloud migration decision system, which we call SEGUE. SEGUE achieves optimal migration decisions by providing a long-term optimal QoS to mobile users in the presence of link quality and server load variation. The basis of SEGUE is in its QoS-aware service migration and its state based MDP model which effectively incorporates the two dominant factors in making migration decisions: 1) network state, and 2) server state. An evaluation of SEGUE performance is given through an augmented reality application. Our results demonstrate that SEGUE reduces the response time of this application by 27.21% and 53.70% compared to the lowest load migration model and the least hop migration model, respectively. Wuyang Zhang, Yanyong Zhang, Dipankar Raychaudhuri |
CloudCom | 3 |
| 2016 | HB-Phone: A Bed-Mounted Geophone-Based Heartbeat Monitoring SystemabstractHeartbeat monitoring during sleep is critically important to ensuring the well-being of many people, ranging from patients to elderly. Technologies that support heartbeat monitoring should be unobtrusive, and thus solutions that are accurate and can be easily applied to existing beds is an important need that has been unfulfilled. We tackle the challenge of accurate, low-cost and easy to deploy heartbeat monitoring by investigating whether off-the- shelf analog geophone sensors can be used to detect heartbeats when installed under a bed. Geophones have the desirable property of being insensitive to lower-frequency movements, which lends itself to heartbeat monitoring as the heartbeat signal has harmonic frequencies that are easily captured by the geophone. At the same time, lower-frequency movements such as respiration, can be naturally filtered out by the geophone. With carefully-designed signal processing algorithms, we show it is possible to detect and extract heartbeats in the presence of environmental noise and other body movements a person may have during sleep. We have built a prototype sensor and conducted detailed experiments that involve 43 subjects (with IRB approval), which demonstrate that the geophone sensor is a compelling solution to long-term at-home heartbeat monitoring. We compared the average heartbeat rate estimated by our prototype and that reported by a pulse oximeter. The results revealed that the average error rate is around 1.30% over 500 data samples when the subjects were still on the bed, and 3.87% over 300 data samples when the subjects had different types of body movements while lying on the bed. We also deployed the prototype in the homes of 9 subjects for a total of 25 nights, and found that the average estimation error rate was 8.25% over more than 181 hours' data. Zhenhua Jia, Musaab Alaziz, Xiang Chi, Richard E. Howard, Yanyong Zhang, Pei Zhang 0001, Wade Trappe, Anand Sivasubramaniam, Ning An 0001 |
IPSN | 5 |
| 2016 | Poster Abstract: Node Deployment Mechanism for Quick, Indoor, and Device-Free LocalizationabstractRadio frequency based device-free passive localization has been actively studied. An important issue with this technique is that a large amount of sensor nodes are usually deployed in localization environments, causing high deployment cost, high computation complexity and other drawbacks. To reduce the number of nodes while maintaining high localization accuracy, this paper proposes a fingerprint-based node deployment mechanism, which saves hardware cost, speeds up node deployment process as well as maintains a similar accuracy. The experiments are conducted in two environments with considerably different sizes of 150 m2and 25 m2. Maintaining the accuracies above 95%, the proposed mechanism can efficiently reduce the total number of nodes in large room by 54.55% and in small room by 68.75%. Especially, the localization response time sharply reduces by 93.08% in the large room and 79.87% in the small room. Jinjun Liu, Ning An 0001, Md. Tanbir Hassan, Guilin Chen, Yanyong Zhang |
IPSN | 5 |
| 2016 | Whose move is it anyway? Authenticating smart wearable devices using unique head movement patternsabstractIn this paper, we present the design, implementation and evaluation of a user authentication system, Headbanger, for smart head-worn devices, through monitoring the user's unique head-movement patterns in response to an external audio stimulus. Compared to today's solutions, which primarily rely on indirect authentication mechanisms via the user's smartphone, thus cumbersome and susceptible to adversary intrusions, the proposed head-movement based authentication provides an accurate, robust, light-weight and convenient solution. Through extensive experimental evaluation with 95 participants, we show that our mechanism can accurately authenticate users with an average true acceptance rate of 95.57% while keeping the average false acceptance rate of 4.43%. We also show that even simple head-movement patterns are robust against imitation attacks. Finally, we demonstrate our authentication algorithm is rather light-weight: the overall processing latency on Google Glass is around 1.9 seconds. Sugang Li, Ashwin Ashok, Yanyong Zhang, Chenren Xu, Janne Lindqvist, Marco Gruteser |
PerCom | 3 |
| 2016 | What Am I Looking At? Low-Power Radio-Optical Beacons for In-View Recognition on Smart-GlassabstractApplications on wearable personal imaging devices, or Smart-glasses as they are called, can largely benefit from accurate and energy-efficient recognition of objects that are within the user's view. Existing solutions such as optical or computer vision approaches are too energy intensive, while low-power active radio tags suffer from imprecise orientation estimates. To address this challenge, this paper presents the design, implementation, and evaluation of a radio-optical hybrid system where a radio-optical transmitter, or tag, whose radio-optical beacons are used for accurate relative orientation tracking of tagged objects by a wearable radio-optical receiver. A low-power radio link that conveys identity is used to reduce the battery drain by synchronizing the radio-optical transmitter and receiver so that extremely short optical (infrared) pulses are sufficient for orientation (angle and distance) estimation. Through extensive experiments with our prototype we show that our system can achieve orientation estimates with 1-to-2 degree accuracy and within 40 cm ranging error, with a maximum range of 9 m in typical indoor use cases. With a tag and receiver battery power consumption of 81 μW and 90 mW, respectively, our radio-optical tags and receiver are at least 1.5x energy efficient than prior works in this space. Ashwin Ashok, Chenren Xu, Tam Vu 0001, Marco Gruteser, Richard E. Howard, Yanyong Zhang, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana |
IEEE Trans. Mob. Comput. | 6 |
| 2016 | Improving Access Point Association Protocols Through Channel Utilization and Adaptive ProbingabstractWe propose a distributed access point selection scheme by which nodes select an appropriate access point to associate with based upon each individual device's channel utilization. In this paper, we define channel utilization as the ratio of required bandwidth to estimated available bandwidth. By incorporating channel utilization into the access point selection protocol, we can effectively reduce unnecessary reassociations and improve upper layer performance such as throughput and packet delivery delay. We have further enhanced our association protocol by using reinforcement learning to dynamically schedule the probing of neighboring access points (APs), ultimately bringing down the probing overhead by learning from past experience. When channel utilization is combined with adaptive probing, we observe a significant performance improvement compared to traditional association approaches. Yanyong Zhang, Wade Trappe |
IEEE Trans. Mob. Comput. | 2 |
| 2016 | The Case for Efficient and Robust RF-Based Device-Free LocalizationabstractRadio frequency based device-free localization has been proposed as an alternative localization technique. Unlike its active localization counterpart, it does not require subjects to wear any radio device, but tries to determine the subject's location by observing how much the subject disturbs the radio propagation patterns. This problem is very challenging due to the well known multipath effect, especially in a complex indoor environment where it is impractical to accurately model the effects of a subject on the surrounding radio links. In this article, we formulate the device-free localization problem using probabilistic classification approaches that are based on discriminant analysis.To boost the localization accuracies, we adopt methods to mitigate errors caused by the multipath effect, as well as methods to automatically recalibrate training data so that accuracy can be maintained as the environment evolves. We validate our method in a one-bedroom apartment that consists of 32 cells, using eight fixed transmitters and eight fixed receivers. When the space has a single occupant, our method can correctly estimate the occupied cell with a likelihood as high as 97.2 percent. Further, we show that we can maintain a high localization accuracy, while substantially reducing the deployment overhead, which is an important concern for device-free localization methods. To achieve this goal, we have improved our training and testing procedures to reduce the overhead, studied the radio device placement to optimize the device cost, devised algorithms to extend the lifetime of the training data, and designed a set of auxiliary sensors and incorporate them into the system to achieve automatic re-calibration. Chenren Xu, Bernhard Firner, Yanyong Zhang, Richard E. Howard |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | QA-share: Towards efficient QoS-aware dispatching approach for urban taxi-sharingabstractTaxi-sharing allows occupied taxis to pick up new passengers on the fly, promising to reduce waiting time for taxi riders and increase productivity for drivers. However, if not carefully designed, taxi-sharing may cause more harm than benefit - it becomes harder to strike the balance between driver's profit and passenger's quality of service (e.g. travel time, number of strangers that share a taxi, etc.). In this paper, we propose a QoS-aware taxi-sharing system design - QA-Share - by addressing two important challenges. First, QA-Share aims to maximize driver profit and user experience at the same time. Second, QA-Share continuously optimizes these two metrics by dynamically adapting its schedule as new requests arrive, without entering an oscillation state. To address these two challenges, we have formulated the optimization problem using integer linear programming, and derived the optimal solution under a small system scale. When the number of requests and taxis becomes large, we have devised a heuristic algorithm that has a much faster execution time. We have also studied how to minimize oscillations caused by schedule re-calculations by dynamically tuning the update threshold. We have evaluated our approach with real-world dataset in a Chinese city - ZhenJiang - which contains the GPS traces recorded by over 3,000 taxis during a period of three months in 2013. Our results show that the QoS and profit is increased by 38% compared to earlier schemes. Shanfeng Zhang, Qiang Ma 0007, Yanyong Zhang, Kebin Liu 0001, Tong Zhu 0001, Yunhao Liu 0001 |
SECON | 3 |
| 2015 | Low-Power Radio-Optical Beacons for In-View RecognitionabstractObject recognition on wearable devices using computer vision is too energy intensive and challenging when objects are similar looking, while low-power active radio frequency identification (RFID) systems suffer from imprecise orientation (angle and distance) estimates. To address this challenge, this paper presents a novel radio-optical based recognition system where a radio-optical transmitter, or tag, that emits a beacon whose infra-red (IR) signal strength is used for accurate relative orientation tracking of tagged objects at a wearable radio-optical receiver. A low-power radio link that conveys identity is used to reduce the battery drain by synchronizing the radio- optical transmitter and receiver so that extremely short optical pulses are sufficient for precise orientation estimation. Through extensive experiments with our prototype we show that our system can achieve orientation estimates with 1-2° accuracy and within 40cm ranging error, with a maximum range of 9m in typical indoor use cases. With a tag battery power consumption of 86μW, the radio-optical tags show potential to achieve about half a decade lifetimes. Ashwin Ashok, Chenren Xu, Tam Vu 0001, Marco Gruteser, Richard E. Howard, Yanyong Zhang, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana |
VTC Fall | 6 |
| 2015 | EdgeBuffer: Caching and prefetching content at the edge in the MobilityFirst future Internet architectureabstractThe prevalence of mobile devices especially smartphones has attracted research on mobile content delivery techniques. In this paper, we propose to take advantage of the storage available at wireless access points to bring content closer to mobile devices, hence improving the downloading performance. Specifically, we propose to have a separate popularity based cache and a prefetch buffer at the network edge to capture both long-term and short-term content access patterns. Further, we point out that it is insufficient to rely on a device's past history to predict when and where to prefetch, especially in urban settings; instead, we propose to derive a prediction model based on the aggregated network-level statistics. We discuss the proposed mobile content caching/prefetching method in the context of the MobilityFirst future Internet architecture. In MobilityFirst, when mobile clients move between network attachment points (e.g., Wi-Fi access points), their network association records are logged by the network, which then naturally facilitates the network-level mobility prediction. Through detailed simulations with real taxi mobility traces, we show that such a strategy is more effective than earlier schemes in satisfying content requests at the edge (higher cache hit ratios), leading to shorter content download latencies. Specifically, the fraction of requests satisfied at the edge increases by a factor of 2.9 compared to a caching only approach, and by 45% compared to individual user-based prediction and prefetching. Feixiong Zhang, Chenren Xu, Yanyong Zhang, K. K. Ramakrishnan, Shreyasee Mukherjee, Roy D. Yates, Thu D. Nguyen |
WOWMOM | 3 |
| 2015 | Providing explicit congestion control and multi-homing support for content-centric networking transport
Feixiong Zhang, Yanyong Zhang, Alex Reznik, Hang Liu 0003, Chen Qian 0001, Chenren Xu |
Comput. Commun. | 2 |
| 2015 | Speech Emotion Recognition Using Fourier ParametersabstractRecently, studies have been performed on harmony features for speech emotion recognition. It is found in our study that the first- and second-order differences of harmony features also play an important role in speech emotion recognition. Therefore, we propose a new Fourier parameter model using the perceptual content of voice quality and the first- and second-order differences for speaker-independent speech emotion recognition. Experimental results show that the proposed Fourier parameter (FP) features are effective in identifying various emotional states in speech signals. They improve the recognition rates over the methods using Mel frequency cepstral coefficient (MFCC) features by 16.2, 6.8 and 16.6 points on the German database (EMODB), Chinese language database (CASIA) and Chinese elderly emotion database (EESDB). In particular, when combining FP with MFCC, the recognition rates can be further improved on the aforementioned databases by 17.5, 10 and 10.5 points, respectively. Kunxia Wang, Ning An 0001, Bing Nan Li, Yanyong Zhang, Lian Li 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2014 | A transport protocol for content-centric networking with explicit congestion controlabstractContent-centric networking (CCN) adopts a receiver-driven, hop-by-hop transport approach that facilitates in-network caching, which in turn leads to multiple sources and multiple paths for transferring content. In such a case, keeping a single round trip time (RTT) estimator for a multi-path flow is insufficient as each path may experience different round trip times. To solve this problem, it has been proposed to use multiple RTT estimators to predict network condition. In this paper, we examine an alternative approach to this problem, CHoPCoP, which utilizes explicit congestion control to cope with the multiple-source, multiple-path situation. Protocol design innovations of CHoPCoP include a random early marking (REM) scheme that explicitly signals network congestion, and a per-hop fair share Interest shaping algorithm (FISP) and a receiver Interest control method (RIC) that regulate the Interest rates at routers and the receiver respectively. We have implemented CHoPCoP on the ORBIT testbed and conducted experiments under various network and traffic settings. The evaluation shows that CHoPCoP is a viable approach that can effectively deal with congestion in the multipath environment. Feixiong Zhang, Yanyong Zhang, Alex Reznik, Hang Liu 0003, Chen Qian 0001, Chenren Xu |
ICCCN | 2 |
| 2014 | Boe: Context-Aware Global Power Management for Mobile Devices Balancing Battery Outage and User ExperienceabstractEnergy conservation on mobile devices is now more important than ever due to the increasing benefits that smartphones and tablets provide to our daily life. However, most existing power management approaches either focus narrowly on a particular sub-system of the mobile device such as the sensor system, the LCD display, or the communication system, or use heuristic approaches to maximize energy efficiency at the cost of user experience. In this paper, we present Boe, a context-aware global power management scheme for mobile devices Balancing battery outage and user experience. To meet the mobile device's expected battery life while sacrificing end user experience as little as possible. Boe takes into account the users' phone usage patterns and activities to dynamically adjust the device's global power management policy to minimize outage time and maximize user experience. We demonstrate our proposed technique by controlling display brightness level and GPS sampling rate on smartphones. We evaluate our approach through real world smartphone data from 10 users over two months. Compared to the best fixed user experience policies, we show that: (i) Boe eliminates all frustrating battery outage events for light, moderate, and heavy phone users, and (ii) Boe improves user experience by 20% for light users, maintains the same user experience for moderate users, and degrades user experience by 23% for heavy smartphone users. Chenren Xu, Vijay Srinivasan, Yoshiya Hirase, Emmanuel Munguia Tapia, Yanyong Zhang |
MASS | 6 |
| 2014 | Poster: enabling mobile content-oriented networking in the mobilityfirst future internet architectureabstractThe prevalence of mobile devices has attracted research on mobile content delivery techniques. The MobilityFirst(MF) project, discussed in this paper, proposes a clean-slate Internet architecture that enables mobile content-oriented operations at the network level. We describe the design details of the architecture in realizing this and provides a preliminary evaluation on scalability and performance. Feixiong Zhang, Yanyong Zhang, Dipankar Raychaudhuri |
MobiHoc | 3 |
| 2014 | A comparative study of MobilityFirst and NDN based ICN-IoT architecturesabstractTo develop unified IoT platforms where objects can be made accessible to applications across organizations and domains, popular solutions are based on client-server overlays on today's Internet. These solutions, however, inherit the inefficiencies of the current Internet - especially in terms of mobility, scalability, and communication reliability. To address this problem, we propose to build the unified IoT platform leveraging the salient feats of Information-Centric Network (ICN) architectures, which we call ICN-IoT. Specifically, we explore two ICN architectures - MobilityFirst and NDN - to support IoT, and refer to them as MF-IoT and NDN-IoT, respectively. Through detailed simulations, we find that though these two architectures fare comparably, MF-IoT incurs lower control overheads. Sugang Li, Yanyong Zhang, Dipankar Raychaudhuri, Ravishankar Ravindran |
QSHINE | 2 |
| 2013 | Crowd++: unsupervised speaker count with smartphonesabstractSmartphones are excellent mobile sensing platforms, with the microphone in particular being exercised in several audio inference applications. We take smartphone audio inference a step further and demonstrate for the first time that it's possible to accurately estimate the number of people talking in a certain place -- with an average error distance of 1.5 speakers -- through unsupervised machine learning analysis on audio segments captured by the smartphones. Inference occurs transparently to the user and no human intervention is needed to derive the classification model. Our results are based on the design, implementation, and evaluation of a system called Crowd++, involving 120 participants in 10 very different environments. We show that no dedicated external hardware or cumbersome supervised learning approaches are needed but only off-the-shelf smartphones used in a transparent manner. We believe our findings have profound implications in many research fields, including social sensing and personal wellbeing assessment. Chenren Xu, Sugang Li, Gang Liu 0001, Yanyong Zhang, Emiliano Miluzzo, Yih-Farn Robin Chen, Jun Li 0034, Bernhard Firner |
UbiComp | 4 |
| 2013 | Secure Name Resolution for Identifier-to-Locator Mappings in the Global InternetabstractA recent trend in clean-slate network design has been to separate the role of identifiers from network locators. An essential component to such a separation is the ability to resolve names into network addresses. One challenge facing name resolution is securing the name resolution service. This paper examines the security of a clean-slate name resolution service suitable for mobile networking. We begin with a high- level threat analysis, and identify several types of attacks that may be used against name resolution services. We then present secure protocols that together form a secure global name resolution service. Specifically, we present a secure update protocol that allows users to update their network addresses as they migrate and that includes several checkpoints that prevents spoofing, collusion, stale identifiers and false identifier announcements. Since the primary function behind a name resolution service is to respond to address-lookup queries, we also present a secure query protocol. Finally, we address the security risks associated with IP holes that can arise in a global name resolution service. Xiruo Liu, Wade Trappe, Yanyong Zhang |
ICCCN | 3 |
| 2013 | SCPL: indoor device-free multi-subject counting and localization using radio signal strengthabstractRadio frequency based device-free passive (DfP) localization techniques have shown great potentials in localizing individual human subjects, without requiring them to carry any radio devices. In this study, we extend the DfP technique to count and localize multiple subjects in indoor environments. To address the impact of multipath on indoor radio signals, we adopt a fingerprinting based approach to infer subject locations from observed signal strengths through profiling the environment. When multiple subjects are present, our objective is to use the profiling data collected by a single subject to count and localize multiple subjects without any extra effort. In order to address the non-linearity of the impact of multiple subjects, we propose a successive cancellation based algorithm to iteratively determine the number of subjects. We model indoor human trajectories as a state transition process, exploit indoor human mobility constraints and integrate all information into a conditional random field (CRF) to simultaneously localize multiple subjects. As a result, we call the proposed algorithm SCPL -- sequential counting, parallel localizing. We test SCPL with two different indoor settings, one with size 150 m2 and the other 400 m2. In each setting, we have four different subjects, walking around in the deployed areas, sometimes with overlapping trajectories. Through extensive experimental results, we show that SCPL can count the present subjects with 86% counting percentage when their trajectories are not completely overlapping. Our localization algorithms are also highly accurate, with an average localization error distance of 1.3 m. Chenren Xu, Bernhard Firner, Robert S. Moore, Yanyong Zhang, Wade Trappe, Richard E. Howard, Feixiong Zhang, Ning An 0001 |
IPSN | 4 |
| 2013 | BiFocus: using radio-optical beacons for an augmented reality search applicationabstractAugmented Reality (AR) applications benefit from accurate detection of the objects that are within a person's view. Typically, it is not only desirable to identify what is currently within view, but also to navigate the users view to the item of interest - for example, finding a misplaced object. In this paper we demonstrate a low-power hybrid radio-optical beaconing system, where objects of interest are tagged with battery-powered RFID-like tags equipped with infrared light emitting diodes (LED) that emit periodic infrared beacons. These beacons are used for accurately estimating the angle and distance from the object to the receiver so as to locate it. The beacons are synchronized using the radio link that is also used to convey the object's unique ID. Ashwin Ashok, Chenren Xu, Tam Vu 0001, Marco Gruteser, Richard E. Howard, Yanyong Zhang, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana |
MobiSys | 6 |
| 2012 | Popularity-driven coordinated caching in named data networkingabstractThe built-in caching capability of future Named Data Networking (NDN) promises to enable effective content distribution at a global scale without requiring special infrastructure. The aim of this work is to design efficient caching schemes in NDN to achieve better performance at both the network layer and application layer. With the specific objective of minimizing the inter-ISP (Internet Service Provider) traffic and average access latency, we first formulate the optimization problems for different objectives and then solve them to obtain the optimal replica placement. Then we develop popularity-driven caching schemes which dynamically place the replicas in the caches on the en-route path in a coordination fashion. Simulation results show that the performances of our caching algorithms are much closer to the optimum and outperform the widely used schemes in terms of the inter-ISP traffic and the average number of access hops. Finally, we thoroughly evaluate the impact of several important design issues such as network topology, cache size, access pattern and content popularity on the caching performance and demonstrate that the proposed schemes are effective, stable, scalable and with reasonably light overhead. Jun Li 0003, Hao Wu 0023, Bin Liu 0001, Jianyuan Lu, Yi Wang 0004, Xin Wang 0001, Yanyong Zhang, Lijun Dong |
ANCS | 7 |
| 2012 | DMap: A Shared Hosting Scheme for Dynamic Identifier to Locator Mappings in the Global InternetabstractThis paper presents the design and evaluation of a novel distributed shared hosting approach, DMap, for managing dynamic identifier to locator mappings in the global Internet. DMap is the foundation for a fast global name resolution service necessary to enable emerging Internet services such as seamless mobility support, content delivery and cloud computing. Our approach distributes identifier to locator mappings among Autonomous Systems (ASs) by directly applying K>1 consistent hash functions on the identifier to produce network addresses of the AS gateway routers at which the mapping will be stored. This direct mapping technique leverages the reach ability information of the underlying routing mechanism that is already available at the network layer, and achieves low lookup latencies through a single overlay hop without additional maintenance overheads. The proposed DMap technique is described in detail and specific design problems such as address space fragmentation, reducing latency through replication, taking advantage of spatial locality, as well as coping with inconsistent entries are addressed. Evaluation results are presented from a large-scale discrete event simulation of the Internet with ~26,000 ASs using real-world traffic traces from the DIMES repository. The results show that the proposed method evenly balances storage load across the global network while achieving lookup latencies with a mean value of ~50 ms and 95th percentile value of ~100 ms, considered adequate for support of dynamic mobility across the global Internet. Tam Vu 0001, Akash Baid, Yanyong Zhang, Thu D. Nguyen, Junichiro Fukuyama, Richard P. Martin, Dipankar Raychaudhuri |
ICDCS | 3 |
| 2012 | Exploiting human mobility trajectory information in indoor device-free passive trackingabstractDevice-free passive (DfP) localization is proposed to localize human subjects indoors by observing how the subject disturbs the pattern of the radio signals without having the subject wear a tag. In our previous work, we have proposed a probabilistic classification based DfP technique, which we call PC-DfP in short, and demonstrated that PC-DfP can classify which cell (32 cells in total) is occupied by the stationary subject with an accuracy as high as 97.2% in a one-bedroom apartment. In this poster, we focus on extending PC-DfP to track a mobile subject in indoor environments by taking into consideration that a human subject's locations should form a continuous trajectory. Through experiments in a 10 × 15 meters open plan office, we show that we can achieve better accuracies by exploiting the property of continuous mobility trajectories. Chenren Xu, Bernhard Firner, Yanyong Zhang, Richard E. Howard, Jun Li 0034 |
IPSN | 3 |
| 2012 | Improving RF-based device-free passive localization in cluttered indoor environments through probabilistic classification methodsabstractRadio frequency based device-free passive localization has been proposed as an alternative to indoor localization because it does not require subjects to wear a radio device. This technique observes how people disturb the pattern of radio waves in an indoor space and derives their positions accordingly. The well-known multipath effect makes this problem very challenging, because in a complex environment it is impractical to have enough knowledge to be able to accurately model the effects of a subject on the surrounding radio links. In addition, even minor changes in the environment over time change radio propagation sufficiently to invalidate the datasets needed by simple fingerprint-based methods. In this paper, we develop a fingerprinting-based method using probabilistic classification approaches based on discriminant analysis. We also devise ways to mitigate the error caused by multipath effect in data collection, further boosting the classification likelihood. Chenren Xu, Bernhard Firner, Yanyong Zhang, Richard E. Howard, Jun Li 0034, Xiaodong Lin 0004 |
IPSN | 3 |
| 2012 | Towards robust device-free passive localization through automatic camera-assisted recalibrationabstractDevice-free passive localization (DfP) techniques can localize human subjects without wearing a radio tag. Being convenient and private, DfP can find many applications in ubiquitous/pervasive computing. Unfortunately, DfP techniques need frequent manual recalibration of the radio signal values, which can be cumbersome and costly. We present SenCam, a sensor-camera collaboration solution that conducts automatic recalibration by leveraging existing surveillance camera(s). When the camera detects a subject, it can periodically trigger recalibration and update the radio signal data accordingly. This technique requires camera access occasionally each month, minimizing computational costs and reducing privacy concerns when compared to localization techniques solely based on cameras. Through experiments in an open indoor space, we show that this scheme can retain good localization results while avoiding manual recalibration. Chenren Xu, Mingchen Gao, Bernhard Firner, Yanyong Zhang, Richard E. Howard, Jun Li 0034 |
SenSys | 4 |
| 2012 | Association attacks: Identifying association protocolsabstractIn this paper, we examine the problem of identifying different association protocols based on client probing patterns. We take the view point of an attacker, who aims to trick certain clients to switch their association to a compromised AP, so that the attacker can easily perform various attacks, such as passing false management frames and stealing client information. In order to do that, the attacker must know what association protocol the client is using since it determines the clients switching criteria. Therefore, the attacker must be able to identify the association protocol by monitoring the network traffic. We investigated methods to identify four association protocols and propose an approach which combines k-means clustering and Gaussian fitting to classify the association protocols based on probing patterns. We tested the designed scheme on traffic traces for a variety of network scenarios. We also designed a method to quantify the likelihood of the identification using confidence intervals. Results show that the proposed method can correctly identify association protocols. Further interpretation of the results also reveals information regarding important metrics of the clients chosen association protocol. Yanyong Zhang, Wade Trappe |
WOWMOM | 2 |
| 2012 | The Boomerang Protocol: Tying Data to Geographic Locations in Mobile Disconnected NetworksabstractWe present the boomerang protocol to efficiently retain information at a particular geographic location in a sparse network of highly mobile nodes without using infrastructure networks. To retain information around certain physical location, each mobile device passing that location will carry the information for a short while. This approach can become challenging for remote locations around which only few nodes pass by. To address this challenge, the boomerang protocol, similar to delay-tolerant communication, first allows a mobile node to carry packets away from their location of origin and periodically returns them to the anchor location. A unique feature of this protocol is that it records the geographical trajectory while moving away from the origin and exploits the recorded trajectory to optimize the return path. Simulations using automotive traffic traces for a southern New Jersey region show that the boomerang protocol improves packet return rate by 70 percent compared to a baseline shortest path routing protocol. This performance gain can become even more significant when the road map is less connected. Finally, we look at adaptive protocols that can return information within specified time limits. Bin Zan, Yanyong Zhang, Marco Gruteser |
IEEE Trans. Mob. Comput. | 3 |
| 2011 | Optimal Caching with Content Broadcast in Cache-and-Forward NetworksabstractWith the rapid advance in the technology area of data storage, storage capacities have increased substantially while the price has been dropping fast. Motivated by this trend, it has been proposed in the Cache-and-Forward architecture that storage is incorporated into each intermediate CNF router. Content can be cached at CNF routers when they flow through the network, and therefore, routers can serve the subsequent requests later on, without forwarding the requests to the host server, we refer to this caching paradigm as In-Network Caching. In this paper, the content caching is enhanced by Content Broadcast(CB), by which a CNF router broadcasts the information of cached content to its neighboring nodes. In order to solve the problem that with limited storage, how an intermediate CNF router optimally decides which passing content should be cached, we develop a mathematical model for CB to minimize the average content retrieval latency, and propose the Independent Allocation algorithm. We compare the average content retrieval latencies of the proposed caching scheme with two other commonly used cache replacement policies. We study the impact of cache size and locality parameter. The proposed scheme is shown to provide significant performance improvement under various settings by as large as 65%. Lijun Dong, Dan Zhang 0010, Yanyong Zhang, Dipankar Raychaudhuri |
ICC | 3 |
| 2011 | Enhance content broadcast efficiency in routers with integrated cachingabstractWith integrated in-network caching diagram, each router advertises the cached content to its immediate neighbors. Although the baseline content broadcast strategy can significantly improve the performance, it leads to low overall cache utilization while each router makes independent caching decisions. In this paper, we enhance the efficiency of content broadcast by providing implicit coordination among neighboring routers. Through detailed simulations, we show that the proposed technique has dramatic performance improvement over the baseline content broadcast scheme. More importantly, this performance gain can be achieved with a minimal communication overhead. Lijun Dong, Yanyong Zhang, Dipankar Raychaudhuri |
ISCC | 2 |
| 2011 | Improving Access Point Association Protocols through Channel Utilization and Adaptive SwitchingabstractIn this paper, we propose a distributed access point selection scheme by which nodes select an appropriate access point to associate with based on each individual device's channel utilization. We define channel utilization as the ratio of required bandwidth to estimated available bandwidth. By incorporating channel utilization into the access point association protocol, we can effectively reduce unnecessary reassociations and improve upper layer performance such as throughput and packet delivery delay. We have further enhanced our association protocol by using adaptive switching to schedule the switching to neighboring access points (APs), ultimately bringing down the association overhead. When channel utilization is combined with adaptive switching, we observe a significant performance improvement compared to traditional association approaches. Yanyong Zhang, Wade Trappe |
MASS | 2 |
| 2011 | Virtual wireless network mapping: An approach to housing MVNOs on wireless meshesabstractVirtual network (VN) mapping is a useful tool for mapping VNs to physical mesh networks. This study extends the idea of mapping VNs from the wired world to the wireless domain by showing its potential applications. Since the generic VN mapping problem is NP-Hard, this study shows how the wireless VN mapping problem can be simplified and be used instead as a mechanism for provisioning wireless points of presences (POPs) as additions to conventional cellular voice and data services. Two heuristic algorithms GSA and GDR are proposed for producing a 2-phase solution to the mapping problem, which corresponds to conventional network deployment process. The results obtained from the VN mapping algorithms proposed here can be used for comparison of overall performance achieved by deploying a particular type of physical network. Further, using this setup, the network operator can determine the costs and benefits associated with setting wired or wireless links on the physical network. Performance is determined based on perceived revenue, and substrate utilization. Gautam D. Bhanage, Yanyong Zhang, Dipankar Raychaudhuri |
PIMRC | 2 |
| 2011 | Smart buildings, sensor networks, and the Internet of ThingsabstractIn contrast to traditional sensor networks, the "Internet of Things" focuses on interactions between humans and physical objects rather than on sensing and reporting low level information. While several middle-ware systems have been created to simplify the task of managing and aggregating data from multiple sensor networks that use different hardware and software, management of the data is not sufficient to build an Internet of Things. Bernhard Firner, Robert S. Moore, Richard E. Howard, Richard P. Martin, Yanyong Zhang |
SenSys | 5 |
| 2011 | Statistical learning strategies for RF-based indoor device-free passive localizationabstractIn this paper, we present the design, implementation and evaluation of a RF-based device-free passive localization strategy using active RFID nodes. Patterns of the measured power on multiple radio links are used to determine the location of a person in a room in a home environment. We develop an adaptive algorithm and training technique to minimize multi-path effects. With experimental deployment in a 5 x 8 meters room, we demonstrate that our system can successfully localize an individual to a 30-inch grid square with an 97.2% accuracy and 0.36 meters average error distance. Chenren Xu, Bernhard Firner, Yanyong Zhang, Richard E. Howard, Jun Li 0034 |
SenSys | 3 |
| 2010 | The Boomerang Protocol: Tieing Data to Geographic Locations in Mobile Disconnected NetworksabstractWe present the novel boomerang protocol to efficiently retain information at a particular geographic location in a sparse network of highly mobile nodes without use of infrastructure networks. Our proof-of-concept implementation revealed the main challenge in implementing the Boomerang protocol is to accurately detect whether a node is divergent from a recorded trajectory and then followed up with a detailed study to address the challenge. Simulation with automotive traffic traces for a southern New Jersey region shows that the protocol improves packet return rate by 70% compared to a baseline implementation using shortest path geographic routing. Bin Zan, Marco Gruteser, Yanyong Zhang |
Mobile Data Management | 4 |
| 2010 | Multiple receiver strategies for minimizing packet loss in dense sensor networksabstractA typical wireless sensor network consists of many small sensors that collect instrument data around their locations and forward it to a central location for data processing. These networks can be deployed to monitor livestock and agricultural assets, products in a store, patients in a hospital, and so on. In many cases sensors have to be densely deployed, and collisions or overhead due to collision avoidance will considerably degrade the system performance below an application's required levels. With the decreasing cost of radio devices the obvious solution to this problem is the use of multiple receivers on different radio channels. However, we show that if receivers can be placed in different locations then increasing the number of receivers on a single channel will increase the rate of the capture effect and decrease collision losses, while also increasing the fairness of the transmitters' radio links. Not only can this single channel approach be more effective than using multiple channels, it is also required for some techniques, such as localization, where each receiver must be able to detect a transmission from any transmitter. We also show that the optimal choice between these two solutions is influenced by the radio attenuation rate and the number of receivers in the system. Bernhard Firner, Chenren Xu, Richard E. Howard, Yanyong Zhang |
MobiHoc | 4 |
| 2010 | Detecting intra-room mobility with signal strength descriptorsabstractWe explore the problem of detecting whether a device has moved within a room. Our approach relies on comparing summaries of received signal strength measurements over time, which we call descriptors. We consider descriptors based on the differences in the mean, standard deviation, and histogram comparison. In close to 1000 mobility events we conducted, our approach delivers perfect recall and near perfect precision for detecting mobility at a granularity of a few seconds. It is robust to the movement of dummy objects near the transmitter as well as people moving within the room. The detection is successful because true mobility causes fast fading, while environmental mobility causes shadow fading, which exhibit considerable difference in signal distributions. The ability to produce good detection accuracy throughout the experiments also demonstrates that our approach can be applied to varying room environments and radio technologies, thus enabling novel security, health care, and inventory control applications. Konstantinos Kleisouris, Bernhard Firner, Richard E. Howard, Yanyong Zhang, Richard P. Martin |
MobiHoc | 4 |
| 2009 | On the Cache-and-Forward Network ArchitectureabstractIn order to meet the increasing demands of content dissemination in Internet, we propose a novel architecture for the future Internet called cache-and-forward (CNF), which transports content as "packages" in a hop-by-hop manner towards the destination, instead of transporting a stream of fragmented packets along an established TCP/IP connection. In this paper, we discuss how the CNF network architecture can be designed for efficient content retrieval. We first introduce several specific services provided in CNF network which are centered around content handling and mobile access. We then give an overview of the CNF protocol stack, which is built on top of IP, and consists of a data plane and a control plane. We provide detailed descriptions of each protocol within both planes. Then we present two caching algorithms, where one involves each CNF router making independent decisions on content caching while the other coordinates node caching within an autonomous system (AS) through hashing. Finally, we gave the initial simulation results to show the performance benefits of hop-by-hop transport and content caching. Lijun Dong, Hongbo Liu 0005, Yanyong Zhang, Sanjoy Paul, Dipankar Raychaudhuri |
ICC | 3 |
| 2009 | Demo abstract: Towards continuous tracking: Low-power communication and fail-safe presence assurance
Bernhard Firner, Prashant Jadhav, Yanyong Zhang, Richard E. Howard, Wade Trappe |
IPSN | 3 |
| 2009 | Towards Continuous Asset Tracking: Low-Power Communication and Fail-Safe Presence AssuranceabstractAsset tracking is an important application domain for wireless sensor networks. However, continuous tracking of a large number of items at the individual item level over a significant period of time is still not feasible. There are two main obstacles. The first is the need for efficient, low-power communication protocols. Many current protocols employ energy-expensive methods to achieve reliable communication for arbitrary traffic situations. Such protocols are not suitable for continuous asset tracking applications. The second challenge is the lack of a robust presence detection algorithm that can differentiate packet losses caused by a missing item from packet losses caused by the ambient radio environment. In this paper, we designed a simple communication protocol, Uni-HB, and demonstrated it can lead to longer system lifetime and higher communication reliability than several popular protocols. We also devised two robust detection algorithms that can yield low false alarm rates while achieving timely loss notification. We took an experimental approach, and evaluated protocols on a generic embedded hardware platform that has an similar architecture to motes. We also derived analytical models to validate our experimental measurements. Bernhard Firner, Prashant Jadhav, Yanyong Zhang, Richard E. Howard, Wade Trappe, Eitan Fenson |
SECON | 3 |
| 2009 | Temporal privacy in wireless sensor networks: Theory and practiceabstractAlthough the content of sensor messages describing “events of interest” may be encrypted to provide confidentiality, the context surrounding these events may also be sensitive and therefore should be protected from eavesdroppers. An adversary armed with knowledge of the network deployment, routing algorithms, and the base-station (data sink) location can infer the temporal patterns of interesting events by merely monitoring the arrival of packets at the sink, thereby allowing the adversary to remotely track the spatio-temporal evolution of a sensed event. In this paper we introduce the problem of temporal privacy for delay-tolerant sensor networks, and propose adaptive buffering at intermediate nodes on the source-sink routing path to obfuscate temporal information from the adversary. We first present the effect of buffering on temporal privacy using an information-theoretic formulation, and then examine the effect that delaying packets has on buffer occupancy. We observe that temporal privacy and efficient buffer utilization are contrary objectives, and then present an adaptive buffering strategy that effectively manages these tradeoffs. Finally, we evaluate our privacy enhancement strategies using simulations, where privacy is quantified in terms of the adversary's mean square error. Pandurang Kamat, Wenyuan Xu 0001, Wade Trappe, Yanyong Zhang |
ACM Trans. Sens. Networks | 4 |
| 2009 | An optimal resource control scheme under fidelity and energy constraints in sensor networks
Jaewon Kang, Yanyong Zhang, B. R. Badrinath |
Wirel. Networks | 2 |
| 2008 | Mining joules and bits: towards a long-life pervasive systemabstractIn this paper, we investigate one of the major challenges in pervasive systems: energy efficiency, by exploring the design of an RFID system intended to support the simultaneous and real time monitoring of thousands of entities. These entities, which may be individuals or inventory items, each carry a lowpower transmit-only tag and are monitored by a collection of networked base-stations reporting to a central database. We have built a customized transmit-only tag with a small formfactor, and have implemented a real-time monitoring application intended to verify the presence of each tag in order to detect potential disappearance of a tag (perhaps due to item theft). Throughout the construction of our pervasive system, we have carefully engineered it for extended tag lifetime and reliable monitoring capabilities in the presence of packet collisions, while keeping the tags small and inexpensive. The major challenge in this architecture (called Roll-Call™) is to supply the energy needed for long range continuous tracking for a year or more while keeping the tags (called PIPs) small and inexpensive. We have used this as a model problem for optimizing cost, size and lifetime across the entire pervasive, persistent system from firmware to protocol. Shweta Medhekar, Richard E. Howard, Wade Trappe, Yanyong Zhang, Peter Wolniansky |
IPDPS | 4 |
| 2008 | Failure prediction in IBM BlueGene/L event logsabstractIn this paper, we present our effort in developing a failure prediction model based on event logs collected from IBM BlueGene/L. We first show how the event records can be converted into a data set that is appropriate for running classification techniques. Then we apply classifiers on the data, including RIPPER (a rule-based classifier), support vector machines (SVMs), a traditional Nearest Neighbor method, and a customized nearest neighbor method. We show that the customized nearest neighbor approach can outperform RIPPER and SVMs in terms of both coverage and precision. The results suggest that the customized nearest neighbor approach can be used to alleviate the impact of failures. Yanyong Zhang, Anand Sivasubramaniam |
IPDPS | 1 |
| 2008 | Anti-jamming timing channels for wireless networksabstractWireless communication is susceptible to radio interference, which prevents the reception of communications. Although evasion strategies have been proposed, such strategies are costly or ineffective against broadband jammers. In this paper, we explore an alternative to evasion strategies that involves the establishment of a timing channel that exists in spite of the presence of jamming. The timing channel is built using failed packet reception times. We first show that it is possible to detect failed packet events inspite of jamming. We then explore single sender and multisender timing channel constructions that may be used to build a low-rate overlay link-layer. We discuss implementation issues that we have overcome in constructing such jamming-resistant timing channel, and present the results of validation efforts using the MICA2 platform. Finally, we examine additional error correction and authentication mechanisms that may be used to cope with adversaries that both jam and seek to corrupt our timing channel. Wenyuan Xu 0001, Wade Trappe, Yanyong Zhang |
WISEC | 3 |
| 2008 | Defending wireless sensor networks from radio interference through channel adaptationabstractRadio interference, whether intentional or otherwise, represents a serious threat to assuring the availability of sensor network services. As such, techniques that enhance the reliability of sensor communications in the presence of radio interference are critical. In this article, we propose to cope with this threat through a technique called channel surfing, whereby the sensor nodes in the network adapt their channel assignments to restore network connectivity in the presence of interference. We explore two different approaches to channel surfing: coordinated channel switching, in which the entire sensor network adjusts its channel; and spectral multiplexing, in which nodes in a jammed region switch channels and nodes on the boundary of a jammed region act as radio relays between different spectral zones. For coordinated channel switching, we examine an autonomous strategy where each node detects the loss of its neighbors in order to initiate channel switching. To cope with latency issues in the autonomous strategy, we propose a broadcast-assisted channel switching strategy to more rapidly coordinate channel switching. For spectral multiplexing, we have devised both synchronous and asynchronous strategies to facilitate the scheduling of nodes in order to improve network fidelity when sensor nodes operate on multiple channels. In designing these algorithms, we have taken a system-oriented approach that has focused on exploring actual implementation issues under realistic network settings. We have implemented these proposed methods on a testbed of 30 Mica2 sensor nodes, and the experimental results show that channel surfing, in its various forms, is an effective technique for repairing network connectivity in the presence of radio interference, while not introducing significant performance-overhead. Wenyuan Xu 0001, Wade Trappe, Yanyong Zhang |
ACM Trans. Sens. Networks | 3 |
| 2008 | Managing the Mobility of a Mobile Sensor Network Using Network DynamicsabstractIt has been discussed in the literature that the mobility of a mobile sensor network (MSN) can be used to improve its sensing coverage. How the mobility can efficiently be managed toward a better coverage, however, remains unanswered. In this paper, motivated by classical dynamics that study the movement of objects, we propose the concept of network dynamics and define the associated potential functions that capture the operational goals, as well as the environment of an MSN. We find that in managing the mobility of an MSN, Newton's laws of motion in classical dynamics are insufficient, for they introduce oscillations into the movement of sensor nodes. Instead, in network dynamics, the laws of motion are formulated using the steepest descent method in optimization. Based on the network dynamics model, we first devise a parallel and distributed algorithm (parallel and distributed network dynamics (PDND)) that runs on each sensor node to guide its movement. PDND then turns sensor nodes into autonomous entities that are capable of adjusting their locations according to the operational goals and environmental changes. After that, we formally prove the convergence of PDND. Finally, we apply PDND in three applications to demonstrate its effectiveness. Ke Ma 0004, Yanyong Zhang, Wade Trappe |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2007 | R-Sentry: Providing Continuous Sensor Services against Random Node FailuresabstractThe success of sensor-driven applications is reliant on whether a steady stream of data can be provided by the underlying system. This need, however, poses great challenges to sensor systems, mainly because the sensor nodes from which these systems are built have extremely short lifetimes. In order to extend the lifetime of the networked system beyond the lifetime of an individual sensor node, a common practice is to deploy a large array of sensor nodes and, at any time, have only a minimal set of nodes active performing duties while others stay in sleep mode to conserve energy. With this rationale, random node failures, either from active nodes or from redundant nodes, can seriously disrupt system operations. To address this need, we propose R-Sentry, which attempts to bound the service loss duration due to node failures, by coordinating the schedules among redundant nodes. Our simulation results show that compared to PEAS, a popular node scheduling algorithm, R-Sentry can provide a continuous 95% coverage through bounded recoveries from frequent node failures, while prolonging the lifetime of a sensor network by roughly 30%. Shengchao Yu, Yanyong Zhang |
DSN | 2 |
| 2007 | Temporal Privacy in Wireless Sensor NetworksabstractAlthough the content of sensor messages describing "events of interest" may be encrypted to provide confidentiality, the context surrounding these events may also be sensitive and therefore should be protected from eavesdroppers. An adversary armed with knowledge of the network deployment, routing algorithms, and the base-station (data sink) location can infer the temporal patterns of interesting events by merely monitoring the arrival of packets at the sink, thereby allowing the adversary to remotely track the spatio-temporal evolution of a sensed event. In this paper, we introduce the problem of temporal privacy for delay- tolerant sensor networks and propose adaptive buffering at intermediate nodes on the source-sink routing path to obfuscate temporal information from an adversary. We first present the effect of buffering on temporal privacy using an information-theoretic formulation and then examine the effect that delaying packets has on buffer occupancy. We evaluate our privacy enhancement strategies using simulations, where privacy is quantified in terms of the adversary's estimation error. Pandurang Kamat, Wenyuan Xu 0001, Wade Trappe, Yanyong Zhang |
ICDCS | 4 |
| 2007 | Failure Prediction in IBM BlueGene/L Event LogsabstractFrequent failures are becoming a serious concern to the community of high-end computing, especially when the applications and the underlying systems rapidly grow in size and complexity. In order to develop effective fault-tolerant strategies, there is a critical need to predict failure events. To this end, we have collected detailed event logs from IBM BlueGene/L, which has 128 K processors, and is currently the fastest supercomputer in the world. In this study, we first show how the event records can be converted into a data set that is appropriate for running classification techniques. Then we apply classifiers on the data, including RIPPER (a rule-based classifier), Support Vector Machines (SVMs), a traditional Nearest Neighbor method, and a customized Nearest Neighbor method. We show that the customized nearest neighbor approach can outperform RIPPER and SVMs in terms of both coverage and precision. The results suggest that the customized nearest neighbor approach can be used to alleviate the impact of failures. Yinglung Liang, Yanyong Zhang, Hui Xiong 0001, Ramendra K. Sahoo |
ICDM | 2 |
| 2007 | An Adaptive Semantic Filter for Blue Gene/L Failure Log AnalysisabstractFrequent failure occurrences are becoming a serious concern to the community of high-end computing, especially when the applications and the underlying systems rapidly grow in size and complexity. In order to better understand the failure behavior of such systems and further develop effective fault-tolerant strategies, we have collected detailed event logs from IBM Blue Gene/L, which has as many as 128K processors, and is currently the fastest supercomputer in the world. Due to the scale of such machines and the granularity of the logging mechanisms, the logs can get voluminous and usually contain records which may not all be distinct. Consequently, it is crucial to filter these logs towards isolating the specific failures, which can then be useful for subsequent analysis. However, existing filtering methods either require too much domain expertise, or produce erroneous results. This paper thus fills this crucial void by designing and developing an adaptive semantic filtering (ASF) method, which is accurate, light-weight, and more importantly, easy to automate. Specifically, ASF exploits the semantic correlation between two events, and dynamically adapts the correlation threshold based on the temporal gap between the events. We have validated the ASF method using the failure logs collected from Blue Gene/L over a period of 98 days. Our experimental results show that ASF can effectively remove redundant entries in the logs, and the filtering results can serve as a good base for future failure analysis studies. Yinglung Liang, Yanyong Zhang, Hui Xiong 0001, Ramendra K. Sahoo |
IPDPS | 2 |
| 2007 | Channel surfing: defending wireless sensor networks from interferenceabstractWireless sensor networks are susceptible to interference that can disrupt sensor communication. In order to cope with this disruption, we explore channel surfing, whereby the sensor nodes adapt their channel assignments to restore network connectivity in the presence of interference. We explore two different approaches to channel surfing: coordinated channel switching, where the entire sensor network adjusts its channel; and spectral multiplexing, where nodes in a jammed region switch channels while nodes on the boundary of a jammed region act as radio relays between different spectral zones. For spectral multiplexing, we have devised both synchronous and asynchronous strategies to facilitate the spectral scheduling needed to improve network fidelity when sensor nodes operate on multiple channels. In designing these algorithms, we have taken a system-oriented approach that has focused on exploring actual implementation issues under realistic network settings. We have implemented these proposed methods on a testbed of 30 Mica2 sensor nodes, and the experimental results show that these strategies can each repair network connectivity in the presence of interference without introducing significant overhead. Wenyuan Xu 0001, Wade Trappe, Yanyong Zhang |
IPSN | 3 |
| 2007 | TARA: Topology-Aware Resource Adaptation to Alleviate Congestion in Sensor NetworksabstractNetwork congestion can be alleviated either by reducing demand (traffic control) or by increasing capacity (resource control). Unlike in traditional wired or other wireless counterparts, sensor network deployments provide elastic resource availability for satisfying the fidelity level required by applications. In many cases, using traffic control can violate fidelity requirements. Hence, we propose the use of resource control: increasing capacity by enabling more nodes to become active during periods of congestion. However, a naive approach to increase resources without a careful consideration of the type of congestion, traffic pattern, and network topology make the situation worse. In this paper, we present TARA, a topology-aware resource adaptation strategy to alleviate congestion. The core of TARA is our capacity analysis model, which can be used to estimate capacity of various topologies. Detailed performance results show that TARA can achieve data delivery rate and energy consumption that is close to an ideal offline resource control algorithm. Jaewon Kang, Yanyong Zhang, B. R. Badrinath |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | BlueGene/L Failure Analysis and Prediction ModelsabstractThe growing computational and storage needs of several scientific applications mandate the deployment of extreme-scale parallel machines, such as IBM’s BlueGene/L which can accommodate as many as 128K processors. One of the challenges when designing and deploying these systems in a production setting is the need to take failure occurrences, whether it be in the hardware or in the software, into account. Ear- lier work has shown that conventional runtime fault- tolerant techniques such as periodic checkpointing are not effective to the emerging systems. Instead, the ability to predict failure occurrences can help develop more effective checkpointing strategies. Failure predic- tion has long been regarded as a challenging research problem, mainly due to the lack of realistic failure data from actual production systems. In this study, we have collected RAS event logs from BlueGene/L over a pe- riod of more than 100 days. We have investigated the characteristics of fatal failure events, as well as the correlation between fatal events and non-fatal events. Based on the observations, we have developed three simple yet effective failure prediction methods, which can predict around 80% of the memory and network failures, and 47% of the application I/O failures. Yinglung Liang, Yanyong Zhang, Anand Sivasubramaniam, Morris Jette, Ramendra K. Sahoo |
DSN | 2 |
| 2006 | Relay MAC: A Collision Free And Power Efficient Reading Protocol For Active RFID TagsabstractA promising application for RFID tags is to trace valuable assets in an inventory. In such systems, the key challenge is to achieve reliable and energy-efficient tag reads. This paper proposes a novel tag reading protocol, Relay-MAC, which aims at reducing the information sent over the network and the energy spent in collision detection and handling by introducing deliberate sequencing at runtime. This paper provides an in-depth study of the design issues one may face in implementing such a protocol on RFID tags, and validates its feasibility using simulation studies. These studies clearly demonstrate that Relay- MAC can yield much better throughput and energy conservation when compared to a conventional select-and-read protocol. Gautam D. Bhanage, Yanyong Zhang |
ICCCN | 2 |
| 2006 | Analysis of Resource Increase and Decrease Algorithm in Wireless Sensor NetworksabstractIn this paper, we first attempt to formally define the resource control framework that adjusts the resource provisioning at the hotspot during congestion. In an effort to find the optimal resource control under the fidelity and energy constraints, we present a resource increase and decrease algorithm called Early Increase/Early Decrease (EIED) that tries to adjust the effective channel capacity quickly to the incoming traffic volume in an energy-efficient manner, thereby increasing the fidelity (or accuracy) level observed by the application during congestion. Under the framework of energy-constrained optimization, we prove this algorithm incurs the lowest overhead of energy consumption for the given fidelity level that is required by the application. Jaewon Kang, Yanyong Zhang, B. R. Badrinath |
ISCC | 2 |
| 2006 | Channel surfing: defending wireless sensor networks from jamming and interference
Wenyuan Xu 0005, Wade Trappe, Yanyong Zhang |
SenSys | 3 |
| 2005 | Filtering Failure Logs for a BlueGene/L PrototypeabstractThe growing computational and storage needs of several scientific applications mandate the deployment of extreme-scale parallel machines, such as IBM's BlueGene/L, which can accommodate as many as 128K processors. In this paper, we present our experiences in collecting and filtering error event logs from a 8192 processor BlueGene/L prototype at IBM Rochester, which is currently ranked #8 in the Top-500 list. We analyze the logs collected from this machine over a period of 84 days starting from August 26, 2004. We perform a three-step filtering algorithm on these logs: extracting and categorizing failure events; temporal filtering to remove duplicate reports from the same location; and finally coalescing failure reports of the same error across different locations. Using this approach, we can substantially compress these logs, removing over 99.96% of the 828,387 original entries, and more accurately portray the failure occurrences on this system. Yinglung Liang, Yanyong Zhang, Anand Sivasubramaniam, Ramendra K. Sahoo, José E. Moreira, Manish Gupta 0002 |
DSN | 2 |
| 2005 | Enhancing Source-Location Privacy in Sensor Network RoutingabstractOne of the most notable challenges threatening the successful deployment of sensor systems is privacy. Although many privacy-related issues can be addressed by security mechanisms, one sensor network privacy issue that cannot be adequately addressed by network security is source-location privacy. Adversaries may use RF localization techniques to perform hop-by-hop traceback to the source sensor’s location. This paper provides a formal model for the source-location privacy problem in sensor networks and examines the privacy characteristics of different sensor routing protocols. We examine two popular classes of routing protocols: the class of flooding protocols, and the class of routing protocols involving only a single path from the source to the sink. While investigating the privacy performance of routing protocols, we considered the tradeoffs between location-privacy and energy consumption. We found that most of the current protocols cannot provide efficient source-location privacy while maintaining desirable system performance. In order to provide efficient and private sensor communications, we devised new techniques to enhance source-location privacy that augment these routing protocols. One of our strategies, a technique we have called phantom routing, has proven flexible and capable of protecting the source’s location, while not incurring a noticeable increase in energy overhead. Further, we examined the effect of source mobility on location privacy. We showed that, even with the natural privacy amplification resulting from source mobility, our phantom routing techniques yield improved source-location privacy relative to other routing methods. Pandurang Kamat, Yanyong Zhang, Wade Trappe, Celal Öztürk |
ICDCS | 2 |
| 2005 | Robust statistical methods for securing wireless localization in sensor networksabstractMany sensor applications are being developed that require the location of wireless devices, and localization schemes have been developed to meet this need. However, as location-based services become more prevalent, the localization infrastructure will become the target of malicious attacks. These attacks will not be conventional security threats, but rather threats that adversely affect the ability of localization schemes to provide trustworthy location information. This paper identifies a list of attacks that are unique to localization algorithms. Since these attacks are diverse in nature, and there may be many unforeseen attacks that can bypass traditional security countermeasures, it is desirable to alter the underlying localization algorithms to be robust to intentionally corrupted measurements. In this paper, we develop robust statistical methods to make localization attack-tolerant. We examine two broad classes of localization: triangulation and RF-based fingerprinting methods. For triangulation-based localization, we propose an adaptive least squares and least median squares position estimator that has the computational advantages of least squares in the absence of attacks and is capable of switching to a robust mode when being attacked. We introduce robustness to fingerprinting localization through the use of a median-based distance metric. Finally, we evaluate our robust localization schemes under different threat conditions. Zang Li, Wade Trappe, Yanyong Zhang, B. R. Badrinath |
IPSN | 3 |
| 2005 | Mobile network management and robust spatial retreats via network dynamicsabstractThe mobility provided by mobile ad hoc and sensor networks will facilitate new mobility-oriented services. Recent work has demonstrated that, for many issues, mobility is advantageous to network operations. This paper proposes that the need for mobility may be captured by formulating the movement of nodes as a classical dynamical system. Motivated by classical mechanics, we propose the notion of network dynamics, where the position and movement of mobile devices evolve according to forces arising from system potential functions that capture the operational goals of the network. We argue that, in the context of moving communicating nodes, the equations of motion should be formulated as a steepest descent minimization of the system potential energy. Further, since global information is not practical in sensor networks, we introduce distributed algorithms that yield more practical implementations of network dynamics. The resulting algorithms are generic, and may be applied to produce balanced network configurations for different initial network deployments. As a second application of network dynamics, we examine the problem of adapting a mobile sensor network to the threat of a jammer. We show that the combination of spatial escape strategies with network dynamics prevents network partitioning that might arise from a mobile jammer Ke Ma 0004, Yanyong Zhang, Wade Trappe |
MASS | 2 |
| 2005 | The feasibility of launching and detecting jamming attacks in wireless networksabstractWireless networks are built upon a shared medium that makes it easy for adversaries to launch jamming-style attacks. These attacks can be easily accomplished by an adversary emitting radio frequency signals that do not follow an underlying MAC protocol. Jamming attacks can severely interfere with the normal operation of wireless networks and, consequently, mechanisms are needed that can cope with jamming attacks. In this paper, we examine radio interference attacks from both sides of the issue: first, we study the problem of conducting radio interference attacks on wireless networks, and second we examine the critical issue of diagnosing the presence of jamming attacks. Specifically, we propose four different jamming attack models that can be used by an adversary to disable the operation of a wireless network, and evaluate their effectiveness in terms of how each method affects the ability of a wireless node to send and receive packets. We then discuss different measurements that serve as the basis for detecting a jamming attack, and explore scenarios where each measurement by itself is not enough to reliably classify the presence of a jamming attack. In particular, we observe that signal strength and carrier sensing time are unable to conclusively detect the presence of a jammer. Further, we observe that although by using packet delivery ratio we may differentiate between congested and jammed scenarios, we are nonetheless unable to conclude whether poor link utility is due to jamming or the mobility of nodes. The fact that no single measurement is sufficient for reliably classifying the presence of a jammer is an important observation, and necessitates the development of enhanced detection schemes that can remove ambiguity when detecting a jammer. To address this need, we propose two enhanced detection protocols that employ consistency checking. The first scheme employs signal strength measurements as a reactive consistency check for poor packet delivery ratios, while the second scheme employs location information to serve as the consistency check. Throughout our discussions, we examine the feasibility and effectiveness of jamming attacks and detection schemes using the MICA2 Mote platform. Wenyuan Xu 0001, Wade Trappe, Yanyong Zhang, Timothy Wood 0001 |
MobiHoc | 3 |
| 2005 | Accurate and energy-efficient congestion level measurement in ad hoc networksabstractCongestion in ad hoc networks not only degrades throughput, but also wastes scarce energy due to a large number of retransmissions and packet drops. For efficient congestion control, an accurate and timely estimation of resource demands by measuring the network congestion level is necessary. Congestion level measurement in ad hoc networks is more difficult than in wired networks due to time-variant channel capacity, contention among neighboring nodes, and non-deterministic node scheduling. We propose a new congestion detection mechanism that quantifies the congestion level accurately and energy-efficiently at both a node-level (implemented at the MAC layer) and a flow-level (implemented at the routing layer) in ad hoc networks. For accurate congestion measurement, a set of metrics that decouple the measurement from various MAC protocol characteristics is defined. For energy-efficient congestion measurement, an asynchronous channel loading measurement scheme, called lazy measurement, which emulates synchronous measurement by using virtual channel sampling, is incorporated into the proposed scheme. Simulation results show that the proposed mechanism significantly cuts down the energy needed to measure congestion accurately, while maintaining the high level of accuracy needed for timely congestion control. Jaewon Kang, Yanyong Zhang, B. R. Badrinath |
WCNC | 2 |
| 2004 | Failure Data Analysis of a Large-Scale Heterogeneous Server EnvironmentabstractThe growing complexity of hardware and software mandates the recognition of fault occurrence in system deployment and management. While there are several techniques to prevent and/or handle faults, there continues to be a growing need for an in-depth understanding of system errors and failures and their empirical and statistical properties. This understanding can help evaluate the effectiveness of different techniques for improving system availability, in addition to developing new solutions. In this paper, we analyze the empirical and statistical properties of system errors and failures from a network of nearly 400 heterogeneous servers running a diverse workload over a year. While improvements in system robustness continue to limit the number of actual failures to a very small fraction of the recorded errors, the failure rates are significant and highly variable. Our results also show that the system error and failure patterns are comprised of time-varying behavior containing long stationary intervals. These stationary intervals exhibit various strong correlation structures and periodic patterns, which impact performance but also can be exploited to address such performance issues. Ramendra K. Sahoo, Anand Sivasubramaniam, Mark S. Squillante, Yanyong Zhang |
DSN | 4 |
| 2004 | Synthesizing Representative I/O Workloads for TPC-HabstractSynthesizing I/O requests that can accurately capture workload behavior is extremely valuable for the design, implementation and optimization of disk subsystems. This paper presents a synthetic workload generator for TPC-H, an important decision-support commercial workload, by completely characterizing the arrival and access patterns of its queries. We present a novel approach for parameterizing the behavior of inter-mingling streams of sequential requests, and exploit correlations between multiple attributes of these requests, to generate disk block-level traces that are shown to accurately mimic the behavior of a real trace in terms of response time characteristics for each TPC-H query. Jianyong Zhang, Anand Sivasubramaniam, Hubertus Franke, Natarajan Gautam, Yanyong Zhang, Shailabh Nagar |
HPCA | 5 |
| 2004 | Performance Implications of Failures in Large-Scale Cluster Scheduling
Yanyong Zhang, Mark S. Squillante, Anand Sivasubramaniam, Ramendra K. Sahoo |
JSSPP | 1 |
| 2003 | Decision-Support Workload Characteristics on a Clustered Database Server from the OS PerspectiveabstractA range of database services are being offered on clusters of workstations today to meet the demanding needs of applications with voluminous datasets, high computational and I/O requirements and a large number of users. The underlying database engine runs on cost-effective off-the-shelf hardware and software components that may not really be tailored/tuned for these applications. At the same time, many of these databases have legacy codes that may not be easy to modulate based on the evolving capabilities and limitations of clusters. An indepth understanding of the interaction between these database engines and the underlying operating system (OS) can identify a set of characteristics that would be extremely valuable for future research on systems support for these environments. To our knowledge, there is no prior work that has embarked on such a characterization for a clustered database server. Using IBM DB2 Universal Database (UDB) Extended Enterprise Edition (EEE) V7.2 Trial version and TPC-H like/sup 1/ decision support queries, this paper studies numerous issues by evaluating performance on an off-the-shelf Pentium/Linux cluster connected by Myrinet. These include detailed performance profiles of all kernel activities, as well as qualitative and quantitative insights on the interaction between the database engine and the operating system. Yanyong Zhang, Jianyong Zhang, Anand Sivasubramaniam, Chun Liu 0001, Hubertus Franke |
ICDCS | 1 |
| 2003 | Gang Scheduling Extensions for I/O Intensive Workloads
Yanyong Zhang, Antony Yang, Anand Sivasubramaniam, José E. Moreira |
JSSPP | 1 |
| 2003 | An Integrated Approach to Parallel Scheduling Using Gang-Scheduling, Backfilling, and MigrationabstractEffective scheduling strategies to improve response times, throughput, and utilization are an important consideration in large supercomputing environments. Parallel machines in these environments have traditionally used space-sharing strategies to accommodate multiple jobs at the same time by dedicating the nodes to a single job until it completes. This approach, however, can result in low system utilization and large job wait times. This paper discusses three techniques that can be used beyond simple space-sharing to improve the performance of large parallel systems. The first technique we analyze is backfilling, the second is gang-scheduling, and the third is migration. The main contribution of this paper is an analysis of the effects of combining the above techniques. Using extensive simulations based on detailed models of realistic workloads, the benefits of combining the various techniques are shown over a spectrum of performance criteria. Yanyong Zhang, Hubertus Franke, José E. Moreira, Anand Sivasubramaniam |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2002 | Characterizing the Scalability of Decision-Support Workloads on Clusters and SMP SystemsabstractUsing a public domain version of a commercial clustered database server and TPC-H like decision support queries, this paper studies the performance and scalability issues of a Pentium/Linux cluster and an 8-way Linux SMP. The execution profile demonstrates the dominance of the I/O subsystem in the execution, and the importance of the communication subsystem for cluster scalability. In addition to quantifying their importance, this paper provides further details on how these subsystems are exercised by the database engine. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Yanyong Zhang, Anand Sivasubramaniam, Jianyong Zhang, Shailabh Nagar, Hubertus Franke |
Euro-Par | 1 |
| 2002 | Modeling and analysis of dynamic coscheduling in parallel and distributed environmentsabstractScheduling in large-scale parallel systems has been and continues to be an important and challenging research problem. Several key factors, including the increasing use of off-the-shelf clusters of workstations to build such parallel systems, have resulted in the emergence of a new class of scheduling strategies, broadly referred to as dynamic coscheduling. Unfortunately, the size of both the design and performance spaces of these emerging scheduling strategies is quite large, due in part to the numerous dynamic interactions among the different components of the parallel computing environment as well as the wide range of applications and systems that can comprise the parallel environment. This in turn makes it difficult to fully explore the benefits and limitations of the various proposed dynamic coscheduling approaches for large-scale systems solely with the use of simulation and/or experimentation.To gain a better understanding of the fundamental properties of different dynamic coscheduling methods, we formulate a general mathematical model of this class of scheduling strategies within a unified framework that allows us to investigate a wide range of parallel environments. We derive a matrix-analytic analysis based on a stochastic decomposition and a fixed-point iteration. A large number of numerical experiments are performed in part to examine the accuracy of our approach. These numerical results are in excellent agreement with detailed simulation results. Our mathematical model and analysis is then used to explore several fundamental design and performance tradeoffs associated with the class of dynamic coscheduling policies across a broad spectrum of parallel computing environments. Mark S. Squillante, Yanyong Zhang, Anand Sivasubramaniam, Natarajan Gautam, Hubertus Franke, José E. Moreira |
SIGMETRICS | 2 |
| 2001 | An Integrated Approach to Parallel Scheduling Using Gang-Scheduling, Backfilling, and Migration
Yanyong Zhang, Hubertus Franke, José E. Moreira, Anand Sivasubramaniam |
JSSPP | 1 |
| 2001 | Scheduling best-effort and real-time pipelined applications on time-shared clustersabstractTwo important emerging trends are influencing the design, implementation and deployment of high performance parallel systems. The first is on the architectural end, where both economic and technological factors are compelling the use of off-the-shelf computing elements (workstations/PCs and networks) to put together high performance systems called clusters. The second is from the user community that is finding an increasing number of applications to benefit from such high performance systems. Apart from the scientific applications that have traditionally needed supercomputing power, a large number of graphics, visualization, database, web service and e-commerce applications have started using clusters because of their high processing and storage requirements. These applications have diverse characteristics and can place different Quality-of-Service (QoS) requirements on the underlying system (low response time, high throughput, high I/O demands, guaranteed response/throughput etc.). Further, clusters running such applications need to cater to potentially a large number of users (or other applications) in a time-shared manner. The underlying system needs to accommodate the requirements of each application, while ensuring that they do not interfere with each other. Yanyong Zhang, Anand Sivasubramaniam |
SPAA | 1 |
| 2001 | Impact of Workload and System Parameters on Next Generation Cluster Scheduling MechanismsabstractScheduling of processes onto processors of a parallel machine has always been an important and challenging area of research. The issue becomes even more crucial and difficult as we gradually progress to the use of off-the-shelf workstations, operating systems, and high bandwidth networks to build cost-effective clusters for demanding applications. Clusters are gaining acceptance not just in scientific applications that need supercomputing power, but also in domains such as databases, web service, and multimedia which place diverse Quality-of-Service (QoS) demands on the underlying system. Further, these applications have diverse characteristics in terms of their computation, communication, and I/O requirements, making conventional parallel scheduling solutions, such as space sharing or gang scheduling, unattractive. At the same time, leaving it to the native operating system of each node to make decisions independently can lead to ineffective use of system resources whenever there is communication. Instead, an emerging class of dynamic coscheduling mechanisms that attempt to take remedial actions to guide the system toward coscheduled execution without requiring explicit synchronization offers a lot of promise for cluster scheduling. Using a detailed simulator, this paper evaluates the pros and cons of different dynamic coscheduling alternatives while comparing their advantages over traditional gang scheduling (and not performing any coordinated scheduling at all). The impact of dynamic job arrivals, job characteristics, and different system parameters on these alternatives is evaluated in terms of several performance criteria. In addition, heuristics to enhance one of the alternatives even further are identified, classified, and evaluated. It is shown that these heuristics can significantly outperform the other alternatives over a spectrum of workload and system parameters and is thus a much better option for clusters than conventional gang scheduling. Yanyong Zhang, Anand Sivasubramaniam, José E. Moreira, Hubertus Franke |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2000 | The Impact of Migration on Parallel Job Scheduling for Distributed Systems
Yanyong Zhang, Hubertus Franke, José E. Moreira, Anand Sivasubramaniam |
Euro-Par | 1 |
| 2000 | A simulation-based study of scheduling mechanisms for a dynamic cluster environmentabstractScheduling of processes onto processors of a parallel machine has always been an important and challenging area of research. The issue becomes even more crucial and difficult as we gradually progress to the use of off-the-shelf workstations, operating systems, and high bandwidth networks to build cost-effective clusters for demanding applications. Clusters are gaining acceptance not just in scientific applications that need supercomputing power, but also in domains such as databases, web service and multimedia, which place diverse Quality-of-Service (QoS) demands on the underlying system. Further, these applications have diverse characteristics in terms of their computation, communication and I/O requirements, making conventional parallel scheduling solutions, such as space sharing or coscheduling, an unattractive option. At the same time, leaving it to the native operating system of each node to make decisions independently can lead to ineffective use of system resources whenever there is communication. Instead, an emerging class of dynamic coscheduling mechanisms, that attempt to take remedial actions to guide the system towards coscheduled execution without requiring explicit synchronization, offer a lot of promise for cluster scheduling. Using a detailed simulator, this paper evaluates the pros and cons of different dynamic coscheduling alternatives, while comparing their advantages over traditional coscheduling (and not performing any coordinated scheduling at all). The impact of dynamic job arrivals, job characteristics and different system parameters on these alternatives are evaluated in terms of several performance criteria. Yanyong Zhang, Anand Sivasubramaniam, José E. Moreira, Hubertus Franke |
ICS | 1 |
| 2000 | Improving Parallel Job Scheduling by Combining Gang Scheduling and Backfilling TechniquesabstractTwo different approaches have been commonly used to address problems associated with space sharing scheduling strategies: (a) augmenting space sharing with backfilling, which performs out of order job scheduling; and (b) augmenting space sharing with time sharing, using a technique called coscheduling or gang scheduling. With three important experimental results-impact of priority queue order on backfilling, impact of overestimation of job execution times, and comparison of scheduling techniques-this paper presents an integrated strategy that combines backfilling with gang scheduling. Using extensive simulations based on detailed models of realistic workloads, the benefits of combining backfilling and gang scheduling are clearly demonstrated over a spectrum of performance criteria. Yanyong Zhang, Anand Sivasubramaniam, Hubertus Franke, José E. Moreira |
IPDPS | 1 |